Think Build Implement Repeat
London, UK +44 7367 067226
WhatsApp FOLLOW f in X
  1. Home
  2. Blog
  3. Making Sure the Bill Is Predictable
Python & Django

Making Sure the Bill Is Predictable

Why AI features in Python services get expensive and how to control it: context size over request volume, cost per request, hard caps and better retrieval.

Updated 2 min readBy SpiderHunts Technologies

Free estimateNo obligation

Get a free estimate

Tell us what you need. A senior engineer reads every enquiry.

Takes under a minute. We never share your details.

  • Free consultation
  • No commitment
  • NDA on request

Prefer to talk? Book a free 30-minute call →

Quick answer — TL;DR

Context size drives cost more than request volume. Track cost per request, cap hard, and improve retrieval — which usually reduces cost and improves accuracy at once.

Context, not volume

Most expensive AI services are expensive because they send far more context than the task needs, not because they handle many requests.

Sending an entire document when three paragraphs would do multiplies your bill and slows every response. Better retrieval is a cost lever, not only an accuracy one.

Five levers

  1. Retrieve fewer, better passages
  2. Summarise conversation history rather than resending everything
  3. Cache stable instructions and repeated context where the provider supports it
  4. Route by difficulty — cheap models for simple steps
  5. Cap output length, because generation costs more than input

Track cost per request, not total

Total cost rising with volume is expected. Cost per request rising means something has changed — longer prompts, more retries, or retrieval returning more than it should.

  • Record tokens in and out on every request
  • Aggregate by endpoint and by caller
  • Alert when the per-request figure moves
  • Review monthly against the previous month

Hard caps with decided behaviour

CapBehaviour at the cap
Daily totalQueue non-urgent work
Per user or callerReject with a clear message
Per requestTruncate context, not silently
Monthly budgetAlert well before, then degrade

A cap with no decided behaviour becomes a crash at the worst moment. Decide what happens before you need it.

Conversations get expensive quickly

Resending full history each turn means cost grows with the square of conversation length. A twenty-turn conversation can cost more than the first nineteen combined.

Summarise older turns, keep recent ones verbatim. Users notice nothing and the cost halves.

FAQ

Frequently asked questions

The questions readers ask us after this guide.

Still have a question?

Ask us directly — a senior engineer will get back to you.

Ask about your project

What does a typical service cost to run?

£80–£600 a month for typical business volumes. Higher for heavy document processing.

Is a cheaper model always cheaper?

Not if it needs retries or a second pass. Compare cost per completed task rather than per token.

Does caching help much?

Substantially where a large stable instruction or document is resent every call. One of the easiest savings available.

Should we self-host to save money?

Rarely at small business volumes. Hosting and operational cost usually exceeds the API spend until volumes are large.

Keep reading

More on Python & Django

Python & Django

An API Other Systems Can Depend On

Designing a Python API service others can depend on: validation at the boundary, consistent errors and status codes, early versioning and documentation.

Python & Django

Moving and Transforming Data Reliably

Building data pipelines in Python that cope with malformed input: restartable stages, quarantining failures, reconciling counts and alerting on absence.

Start here

AI bill higher than expected?

It is usually context size rather than volume. Happy to look at where the spend goes.

  1. You tell us what you needTwo minutes on the form, or a message on WhatsApp.
  2. A senior engineer reviews itAnd comes back with questions, a realistic range and an honest view on fit.
  3. Free 30-minute scoping callWe talk through scope, options and a realistic estimate — with no obligation.
Free estimateNo obligation

Talk to someone who builds this

Send a short brief and we will come back with an honest view and a realistic range.

Takes under a minute. We never share your details.

  • Free consultation
  • No commitment
  • NDA on request

Prefer to talk? Book a free 30-minute call →