AI that survives production
We build on Claude where the work is reasoning over a client's own material — support, research, operations. Retrieval, evaluation and guardrails first; the demo is the easy part.
Three reasons clients build with us on Claude
Why Claude
Long context, strong reasoning and tool use make it the right engine for work over dense internal material. Where a smaller or cheaper model is enough, we use that instead.
→Long context for real documents, not excerpts→Reasoning that holds up on domain material→Tool use and structured output you can validate→Safety behaviour appropriate for regulated data→Predictable cost once usage patterns are knownRetrieval and evaluation
Almost every AI failure we are called in to fix is a retrieval or evaluation failure. We build the harness before the feature, so quality is measured rather than asserted.
→Retrieval designed against your real corpus→An eval set built with your domain experts→Regression testing on every prompt or model change→Human review loops where stakes require them→Cost and latency tracked per use caseShipping and operating
We put it in front of users with the guardrails, observability and fallbacks a production system needs — and keep it honest about what it does not know.
→Guardrails, refusal behaviour and escalation paths→Traceable answers with citations to source material→Observability: prompts, tokens, failures, drift→Fallbacks for outage and rate limits→Handover so your team can extend it safelyBook a call→Measured, not asserted
Retrieval quality is a number. We report it before and after every change, and we separate it from how well the model writes.
What we build around it
A short, deliberate list — each in production on a system we maintain. See the full partner stack.
Clients we work with








Claude case studies
What we will not do
Most AI briefs we receive should be smaller than they arrive. These are the jobs we hand back.
$./ai-readiness-review→