$stack/claude

AI that survives production

We build on Claude where the work is reasoning over a client's own material — support, research, operations. Retrieval, evaluation and guardrails first; the demo is the easy part.

$claude services

Three reasons clients build with us on Claude

Why Claude

Long context, strong reasoning and tool use make it the right engine for work over dense internal material. Where a smaller or cheaper model is enough, we use that instead.

Long context for real documents, not excerptsReasoning that holds up on domain materialTool use and structured output you can validateSafety behaviour appropriate for regulated dataPredictable cost once usage patterns are known

Retrieval and evaluation

Almost every AI failure we are called in to fix is a retrieval or evaluation failure. We build the harness before the feature, so quality is measured rather than asserted.

Retrieval designed against your real corpusAn eval set built with your domain expertsRegression testing on every prompt or model changeHuman review loops where stakes require themCost and latency tracked per use case

Shipping and operating

We put it in front of users with the guardrails, observability and fallbacks a production system needs — and keep it honest about what it does not know.

Guardrails, refusal behaviour and escalation pathsTraceable answers with citations to source materialObservability: prompts, tokens, failures, driftFallbacks for outage and rate limitsHandover so your team can extend it safelyBook a call
$evaluation

Measured, not asserted

Retrieval quality is a number. We report it before and after every change, and we separate it from how well the model writes.

100%Features shipped with an owned eval set
2-6 wksPrototype to evaluated production system
Every answerTraceable to its source material
0Claims we cannot reproduce in a test

What we build around it

A short, deliberate list — each in production on a system we maintain. See the full partner stack.

Claude (Anthropic)AI platform partner
OpenAIPlatform — production use
LangSmithEvals and tracing
PineconeVector search
SupabaseData platform partner
PostgreSQLCore data engine
SnowflakeWarehouse — healthcare data
Google CloudCertified — data & AI

Clients we work with

Bayt Travel
Travel Secrets
DAX — Doha Express
Orangetheory Fitness
Nova Fertility
Octillion Global
Ornamint
Shubhra Krishan
IKISAKI
$case studies

Claude case studies

Research assistant UI
ResearchNova FertilityReasoning over a decade of clinical research with citations to source, evaluated with the research team before launch.
Ops console
OperationsDAX — Doha ExpressFreight documentation triage with human review where the stakes require it.

What we will not do

Most AI briefs we receive should be smaller than they arrive. These are the jobs we hand back.

$./ai-readiness-review
01Ship an AI feature with no eval set and no owner.02Use a language model where a query or a rule would do.03Put a chatbot in front of a process problem.04Train on client data without an explicit, written basis.05Demo something we could not run under real load.