CreativeCape

Cutting LLM Costs Without Cutting Quality

Model routing, prompt caching and output control — the levers that cut LLM spend without users noticing a difference.

July 23, 2026·4 min read

LLM costs rarely become a problem gradually. They are negligible in development, then a feature ships, usage grows, and finance asks what the line item is.

The good news: most LLM spend has obvious waste in it, and removing it rarely costs quality.

Measure cost per feature first

You cannot optimise an aggregate. A single monthly total tells you nothing about which of eight features is responsible.

Log, for every call: feature name, model, input tokens, output tokens, cached tokens, latency, and whether the user kept the result. Aggregate daily by feature.

The result is almost always lopsided — a single feature dominating, often one nobody considered expensive. A background summarisation job re-running over unchanged content is a classic. Find that before tuning prompts.

Also track cost per successful outcome, not cost per call. A cheap model that needs three attempts is more expensive than an accurate one that needs one, and only the outcome metric shows it.

Model routing

The largest single lever. Frontier models are priced for frontier work; most requests are not frontier work.

Classify by difficulty and route:

  • Simple — classification, extraction, formatting, short rewrites → small fast model

  • Moderate — summarisation, routine drafting → mid-tier

  • Hard — multi-step reasoning, code generation, nuanced analysis → largest model

Routing can be a heuristic (input length, feature, user tier) before it is anything clever. Even a crude split moves the number substantially, because the simple bucket is usually the biggest by volume.

Two patterns worth adding:

Escalate on failure. Try the small model; if output fails validation, retry on the larger one. Most requests settle at the cheap tier.

Draft then refine. A small model drafts, a large one polishes. Cheaper than generating entirely at the top tier.

Prompt caching

If you have a long stable prefix — a system prompt, a schema, a document being discussed — prompt caching is close to free money. Cached input tokens are billed at a large discount.

To benefit, the prefix must be byte-identical. Practical implications:


[ system prompt ][ schema ][ examples ][ document ]  ← stable, cacheable

[ conversation turns ][ current question ]          ← varies

Put everything stable first and everything variable last. A timestamp or a user id injected near the top of the prompt invalidates the cache on every call — and this is a genuinely common mistake, because it looks harmless.

Controlling output length

Output tokens typically cost several times more than input tokens, so verbosity is expensive.

  • Set max_tokens deliberately per feature rather than leaving a large default

  • Ask for the format you want: "Answer in at most three sentences", "Return JSON only, no prose"

  • Use structured output / tool schemas — they eliminate the explanatory padding around the data you actually wanted

  • Stop generating as soon as the UI has what it needs

A tone instruction like "be concise" is not decoration; it is a cost control.

Batching

For anything not user-facing — nightly summarisation, bulk classification, backfills — batch APIs offer a substantial discount in exchange for latency.

Ask which jobs genuinely need a synchronous answer. Usually only the interactive ones do, and moving the rest to batch is a large saving for very little work.

Caching at the application layer

Before the model call, ask whether you need it at all.

Exact-match cache. Hash the normalised input; serve identical requests from cache. Content-generation workloads repeat more than you expect.

Semantic cache. Embed the query and serve a cached answer above a similarity threshold. Effective for support-style questions asked many ways, but set the threshold conservatively — a wrong cache hit is worse than a cache miss.

Cache derived artefacts. Embeddings keyed by content hash, summaries keyed by document version. Recomputing unchanged content is the most common waste of all.

Do not call the model for deterministic work. Parsing dates, validating formats, simple arithmetic and lookups do not need an LLM. This sounds obvious; it appears in production code constantly.

When a smaller model is genuinely enough

The real question is not "which model is best" but "what is the smallest model that passes my evals".

Build a test set from real inputs with acceptable outputs. Run each candidate. Compare accuracy and cost. Often a small model matches on the actual distribution of work, and the frontier model was bought for a hard case that represents 2% of traffic — which routing handles.

Re-run that comparison periodically. Model pricing and capability move quickly, and a routing decision from a year ago is likely no longer optimal.

The order that works: measure per feature → cache what repeats → route by difficulty → cache the prompt prefix → constrain output → batch the asynchronous work. Each step is independent, and the first two usually pay for the rest.


→ AI Development

Tagged with
#llm#cost optimization#prompt caching#token usage#ai engineering

Found this useful? Share it.

Keep Reading

Related articles

Booking Q2 2026 Projects

Ready to Build Something Great?

From idea to launch — let our senior engineers build, ship and scale your next product. No commitment, just a conversation.

Senior Engineers
On-Time Delivery
Enterprise-Grade
Free Consultation

Free 30-min discovery call

Talk to a senior engineer — not a salesperson.

We'll review your goals, suggest the leanest path forward, and send a clear proposal within 24 hours.

24h

Response Time

100+

Projects Delivered

No commitment · No automated bots · Fully transparent