Think Build Implement Repeat
London, UK +44 7367 067226
WhatsApp FOLLOW f in X
  1. Home
  2. Blog
  3. Seeing What Your AI Service Is Actually Doing
Python & Django

Seeing What Your AI Service Is Actually Doing

Observability for AI services that call language models: logging context, output, model version, latency and cost on every call, and the metrics to watch.

Updated 2 min readBy SpiderHunts Technologies

Free estimateNo obligation

Get a free estimate

Tell us what you need. A senior engineer reads every enquiry.

Takes under a minute. We never share your details.

  • Free consultation
  • No commitment
  • NDA on request

Prefer to talk? Book a free 30-minute call →

Quick answer — TL;DR

Record every call with its input, context, output, model version, latency and cost. Without that, diagnosing a bad output is guesswork and explaining one is impossible.

You cannot debug what you did not record

When someone reports a wrong answer, the questions are what was asked, what was retrieved, which prompt and model were used, and what came back. All four have to have been recorded at the time.

Reproducing an AI failure after the fact is nearly impossible without the trace. Recording it costs almost nothing and it is the difference between diagnosing and guessing.

What to record per call

  1. The input, or a reference to it
  2. The retrieved context — which passages, from which documents
  3. The prompt version and model version
  4. The raw output, before any post-processing
  5. Latency, tokens and cost
  6. Any validation failure or retry

Metrics that matter

MetricWatch for
Latency percentilesDegradation, not just averages
Error and retry rateProvider or prompt problems
Validation failure rateOutput drifting from the schema
Refusal rateRetrieval broken
Cost per requestContext growing
Correction rateQuality degrading

Correlate with the originating request

A single user action may trigger retrieval, several model calls and a queued job. Without a shared identifier across all of them, tracing what happened means guessing.

Generate it once at the boundary and carry it everywhere, including into queued work.

Mind what you log

  • Retrieved passages may contain personal or sensitive data
  • Log references rather than content where the material is sensitive
  • Apply a retention period automatically
  • Remember these logs are in scope for data requests
  • Restrict who can read them

FAQ

Frequently asked questions

The questions readers ask us after this guide.

Still have a question?

Ask us directly — a senior engineer will get back to you.

Ask about your project

Do we need a dedicated observability tool?

Useful at scale. Structured logging with a correlation identifier covers most needs for a single service.

How long should traces be kept?

Thirty to ninety days for diagnosis. Longer where they serve as an audit record, subject to your data obligations.

What about the storage cost?

Modest for text. Storing references rather than full content keeps it small even at volume.

Should we sample rather than record everything?

Record everything at business volumes. Sampling means the failure you need to diagnose is the one you did not record.

Keep reading

More on Python & Django

Python & Django

An API Other Systems Can Depend On

Designing a Python API service others can depend on: validation at the boundary, consistent errors and status codes, early versioning and documentation.

Python & Django

Moving and Transforming Data Reliably

Building data pipelines in Python that cope with malformed input: restartable stages, quarantining failures, reconciling counts and alerting on absence.

Start here

Cannot explain why your AI gave that answer?

That is a tracing gap rather than a model problem. Happy to add proper observability to an existing service.

  1. You tell us what you needTwo minutes on the form, or a message on WhatsApp.
  2. A senior engineer reviews itAnd comes back with questions, a realistic range and an honest view on fit.
  3. Free 30-minute scoping callWe talk through scope, options and a realistic estimate — with no obligation.
Free estimateNo obligation

Talk to someone who builds this

Send a short brief and we will come back with an honest view and a realistic range.

Takes under a minute. We never share your details.

  • Free consultation
  • No commitment
  • NDA on request

Prefer to talk? Book a free 30-minute call →