Software crashes. AI systems just start being wrong

Conventional monitoring watches for things that stop. An AI system rarely stops — it drifts, retrieves worse, answers plausibly and is quietly wrong for three weeks before somebody notices. Running these systems is a different discipline from running software, and it is the part most projects have no plan for.

  • Quality watched, not just uptime. Answer quality, retrieval relevance and refusal rates tracked continuously, with alerts on movement rather than on outage.
  • Response times you can hold us to. By severity, with escalation paths and a named engineer who knows your deployment rather than a queue.
  • Cost under control. Per feature and per user, with anomalies caught in days rather than on the invoice.
  • A shrinking engagement. Documentation and pairing so your team takes over the routine parts, with the support tier declining on a written schedule.
Talk to us about this

What is delivered

  • Quality monitoring

    A continuously-scored evaluation set, retrieval quality tracked separately from answer quality, and alerting on drift. The core of the service and what makes it different from hosting.

  • Incident process for quiet failure

    Severity defined in business terms, a first response that narrows the system's autonomy while investigating, and a written record afterwards.

  • Model lifecycle

    Provider deprecations, version upgrades and prompt changes tested against your evaluation set before they reach you, not after a user reports something odd.

  • Cost management

    Per-feature attribution, routing and caching reviewed on a cadence, and a monthly note explaining any movement in plain terms.

  • Data and retention operations

    Re-indexing, embedding refresh, retention enforcement and the checks that keep the corpus honest as your documents change.

  • Reporting somebody reads

    One page a month: what changed, what broke, what it cost, and what we recommend. Not a dashboard nobody opens.

How it runs

  1. 01

    Take over properly

    A documented handover with the evaluation set, runbooks and access. If those do not exist we build them first and say so up front.

  2. 02

    Baseline

    Current quality, cost and latency measured before we change anything, so improvement is demonstrable rather than asserted.

  3. 03

    Run and report

    Monitoring, incidents, upgrades and a monthly note. Boring by design, which is the whole objective.

  4. 04

    Reduce

    Transfer routine operations to your team on an agreed schedule. If the tier does not decline, one of us is not doing the job.

A good fit when

  • An AI system is live and nobody is watching its quality.
  • The team that built it has moved on or was external and is gone.
  • Costs are unpredictable and nobody can attribute them.
  • You need an SLA on an AI system for internal or customer reasons.

Not the right service when

  • You want infrastructure hosting only. A cloud provider does that better and cheaper.
  • The system has no evaluation set and you do not want one built. Without it, monitoring quality is not possible and we would be selling you the appearance of it.
  • You want us to guarantee the model is never wrong. Nobody can promise that, and a supplier who does is describing something else.

Frequently asked questions

What is actually in the SLA?
Response times by severity, escalation paths, a named engineer, maintenance windows and defined quality-monitoring coverage. What cannot honestly be promised is a fix by a fixed time for any possible defect, or that the model will never be wrong.
How do you detect a model getting worse?
A held-out evaluation set scored continuously, retrieval relevance tracked separately, plus signals from real usage: rising refusals, rising corrections, changing output length. Any one alone is noisy; together they catch drift before a user reports it.
Do you need access to our production data?
As little as possible. Usually metrics, logs and a sampled evaluation set rather than the corpus. Where more is needed, the access is scoped, time-boxed and logged, and the arrangement is written into the contract.
Can you run a system you did not build?
Yes, and it is common. The first weeks go into understanding it and building an evaluation set if none exists. That step usually finds a problem that was already there, which is uncomfortable and better known.
What does it cost?
Priced per system and scope rather than per user or per request, so the fee does not rise because the system succeeded. The tier is expected to decline as your team takes over the routine parts.