Think Build Implement Repeat
London, UK +44 7367 067226
WhatsApp FOLLOW f in X
  1. Home
  2. Blog
  3. What You Are Actually Paying For in an AI Integration
AI Integration

What You Are Actually Paying For in an AI Integration

Why prompts get expensive: how context size drives AI cost and latency, five levers to cut it, the cost of long conversations and why caps are essential.

Updated 2 min readBy SpiderHunts Technologies

Free estimateNo obligation

Get a free estimate

Tell us what you need. A senior engineer reads every enquiry.

Takes under a minute. We never share your details.

  • Free consultation
  • No commitment
  • NDA on request

Prefer to talk? Book a free 30-minute call →

Quick answer — TL;DR

You pay per unit of text in and out. Sending an entire document when three paragraphs would do multiplies your bill and slows every response. Retrieval quality is a cost lever, not just an accuracy one.

Where the money goes

Every request carries the instruction, the retrieved context, the conversation history and the question. Output is charged too, usually at a higher rate.

Most expensive integrations are expensive because they send far more context than the task needs, not because the model is costly.

Five levers

  1. Retrieve fewer, better passages — five relevant beats twenty adequate
  2. Summarise long history rather than resending every turn
  3. Cache repeated instructions and stable context where the provider supports it
  4. Route by difficulty so simple steps use cheaper models
  5. Cap output length, because generation is the expensive direction
Improving retrieval usually cuts cost and improves accuracy at the same time. It is the rare optimisation with no trade-off.

Long conversations get expensive quickly

Resending the full history every turn means cost grows with the square of the conversation length. A twenty-turn conversation can cost more than the first nineteen combined.

Summarise older turns and keep the recent ones verbatim. Users notice nothing; the bill halves.

Model the cost before you build

DriverQuestion to answer
VolumeHow many requests per day, at peak?
Context sizeHow much text per request, realistically?
Output sizeHow long is a typical answer?
Retry rateHow often does a request need a second attempt?

Those four give a monthly figure within about twenty per cent, which is enough to decide.

Caps are not optional

A hard daily cap, a per-user cap and an alert well below both. A loop or an abusive user should cost you an alert, not an invoice.

Decide the behaviour at the cap too — degrade or queue rather than crash.

FAQ

Frequently asked questions

The questions readers ask us after this guide.

Still have a question?

Ask us directly — a senior engineer will get back to you.

Ask about your project

What does a typical integration cost to run?

£80–£600 a month for most business volumes. Higher for heavy document processing, lower for internal tools with modest use.

Is a cheaper model always cheaper?

Not if it needs retries or produces answers that require a second pass. Compare cost per completed task, not per token.

Does caching really help?

Substantially, where a large stable instruction or document is resent on every call. It is one of the easiest savings available.

Should we self-host to save money?

Rarely, at small business volumes. The hosting and operational cost usually exceeds the API spend until volumes are large.

Keep reading

More on AI Integration

Start here

AI bill higher than you expected?

It is usually context size rather than volume. We can look at where the spend is going and what would cut it.

  1. You tell us what you needTwo minutes on the form, or a message on WhatsApp.
  2. A senior engineer reviews itAnd comes back with questions, a realistic range and an honest view on fit.
  3. Free 30-minute scoping callWe talk through scope, options and a realistic estimate — with no obligation.
Free estimateNo obligation

Talk to someone who builds this

Send a short brief and we will come back with an honest view and a realistic range.

Takes under a minute. We never share your details.

  • Free consultation
  • No commitment
  • NDA on request

Prefer to talk? Book a free 30-minute call →