Software crashes. AI systems just start being wrong
Conventional monitoring watches for things that stop. An AI system rarely stops — it drifts, retrieves worse, answers plausibly and is quietly wrong for three weeks before somebody notices. Running these systems is a different discipline from running software, and it is the part most projects have no plan for.
- Quality watched, not just uptime. Answer quality, retrieval relevance and refusal rates tracked continuously, with alerts on movement rather than on outage.
- Response times you can hold us to. By severity, with escalation paths and a named engineer who knows your deployment rather than a queue.
- Cost under control. Per feature and per user, with anomalies caught in days rather than on the invoice.
- A shrinking engagement. Documentation and pairing so your team takes over the routine parts, with the support tier declining on a written schedule.
What is delivered
Quality monitoring
A continuously-scored evaluation set, retrieval quality tracked separately from answer quality, and alerting on drift. The core of the service and what makes it different from hosting.
Incident process for quiet failure
Severity defined in business terms, a first response that narrows the system's autonomy while investigating, and a written record afterwards.
Model lifecycle
Provider deprecations, version upgrades and prompt changes tested against your evaluation set before they reach you, not after a user reports something odd.
Cost management
Per-feature attribution, routing and caching reviewed on a cadence, and a monthly note explaining any movement in plain terms.
Data and retention operations
Re-indexing, embedding refresh, retention enforcement and the checks that keep the corpus honest as your documents change.
Reporting somebody reads
One page a month: what changed, what broke, what it cost, and what we recommend. Not a dashboard nobody opens.
How it runs
- 01
Take over properly
A documented handover with the evaluation set, runbooks and access. If those do not exist we build them first and say so up front.
- 02
Baseline
Current quality, cost and latency measured before we change anything, so improvement is demonstrable rather than asserted.
- 03
Run and report
Monitoring, incidents, upgrades and a monthly note. Boring by design, which is the whole objective.
- 04
Reduce
Transfer routine operations to your team on an agreed schedule. If the tier does not decline, one of us is not doing the job.
A good fit when
- An AI system is live and nobody is watching its quality.
- The team that built it has moved on or was external and is gone.
- Costs are unpredictable and nobody can attribute them.
- You need an SLA on an AI system for internal or customer reasons.
Not the right service when
- You want infrastructure hosting only. A cloud provider does that better and cheaper.
- The system has no evaluation set and you do not want one built. Without it, monitoring quality is not possible and we would be selling you the appearance of it.
- You want us to guarantee the model is never wrong. Nobody can promise that, and a supplier who does is describing something else.
Frequently asked questions
What is actually in the SLA?
How do you detect a model getting worse?
Do you need access to our production data?
Can you run a system you did not build?
What does it cost?
Other services
AI Strategy
Which decisions are worth changing, what each would cost, and what would measurably be different. Including the ones where the answer is not AI.
Read articleAI Transformation
The part after the strategy: sequencing against real capacity, changing how work is done, and making adoption somebody's job rather than a hope.
Read articleAI Agents Engineering
Agents that do work rather than answer questions — with bounds, tools, approval gates and a decision log that survives the first incident.
Read article
