Think Build Implement Repeat
London, UK +44 7367 067226
WhatsApp FOLLOW f in X
  1. Home
  2. Blog
  3. Wrapping a Model in Something Dependable
Python & Django

Wrapping a Model in Something Dependable

Architecture for a dependable AI service in Python: queue the work, force structured output, log every call, cap costs and keep prompts in configuration.

Updated 2 min readBy SpiderHunts Technologies

Free estimateNo obligation

Get a free estimate

Tell us what you need. A senior engineer reads every enquiry.

Takes under a minute. We never share your details.

  • Free consultation
  • No commitment
  • NDA on request

Prefer to talk? Book a free 30-minute call →

Quick answer — TL;DR

Queue the work, validate the output against a schema, record everything, cap the cost, and always leave a path that does not involve the model.

The service is more than the model call

Calling a model is a few lines. Making that call dependable — retried, validated, logged, capped and observable — is the actual work, and it is what separates a demonstration from a system.

The model call is perhaps five per cent of an AI service. The rest is everything that makes it safe to depend on.

The shape that works

  1. Receive the request and validate it
  2. Queue it, unless a person is genuinely waiting
  3. Build the prompt from configuration, not from code
  4. Call with a timeout and bounded retries
  5. Validate the output against a schema before accepting it
  6. Record input, context, output, model version and confidence

Force structured output

Ask for a defined structure and validate what comes back. If it does not match, retry once with a stricter instruction and then quarantine rather than accepting something malformed.

  • A schema the output must satisfy
  • Validation before anything downstream sees it
  • One retry with a stricter prompt
  • Quarantine with the raw response for diagnosis

Cost control is not optional

ControlWhy
Hard daily capA loop must not spend the budget
Per-caller capOne consumer cannot exhaust it
Cost per request trackedRising cost is the early warning
Token limits on outputGeneration is the expensive direction
Decided behaviour at the capDegrade or queue, not crash

Keep the prompts in configuration

Prompts change more often than code and are frequently adjusted by someone who is not a developer. They belong in configuration, versioned, with the ability to roll back.

That also makes it possible to run an evaluation set against a prompt change before deploying it, which is what stops quality regressions.

FAQ

Frequently asked questions

The questions readers ask us after this guide.

Still have a question?

Ask us directly — a senior engineer will get back to you.

Ask about your project

Should the service call the model synchronously?

Only where a person is waiting. Everything else should be queued, which gives you retries and isolates provider slowness.

How do we handle provider outages?

Queue and retry for background work. For interactive work, degrade to a smaller model or fail with a clear message and a manual route.

Do we need to record everything?

Input, retrieved context, model version, output and confidence. Without those you cannot explain a decision later.

Which Python framework?

Any competent one. The architecture matters considerably more than the framework choice.

Keep reading

More on Python & Django

Python & Django

An API Other Systems Can Depend On

Designing a Python API service others can depend on: validation at the boundary, consistent errors and status codes, early versioning and documentation.

Python & Django

Moving and Transforming Data Reliably

Building data pipelines in Python that cope with malformed input: restartable stages, quarantining failures, reconciling counts and alerting on absence.

Start here

Building a service around a model?

The model call is the easy part. Happy to review the architecture around it.

  1. You tell us what you needTwo minutes on the form, or a message on WhatsApp.
  2. A senior engineer reviews itAnd comes back with questions, a realistic range and an honest view on fit.
  3. Free 30-minute scoping callWe talk through scope, options and a realistic estimate — with no obligation.
Free estimateNo obligation

Talk to someone who builds this

Send a short brief and we will come back with an honest view and a realistic range.

Takes under a minute. We never share your details.

  • Free consultation
  • No commitment
  • NDA on request

Prefer to talk? Book a free 30-minute call →