Here's a question worth asking your finance team: how much of this year's AI budget is left? For a surprising number of companies, the honest answer is "not much". Some are running dry by April. The annual budget, gone in a quarter. The bigger reason is how we're using it. We've moved from asking chatbots questions to running agents. And agents don't just answer. They call tools, read files, check results and call more tools. Every one of those steps burns tokens. One agentic request can cost many times what a plain chat message does.
It doesn't help that the pricing itself keeps moving. There are subscriptions. There's pay-as-you-go API access. There are promos that appear for a few months to stop people switching, and new top-tier models landing every few months with their own price points. So before we get into how to cut the bill, let's make sure we all understand what we're actually paying for. Because once you see how the meter works, the savings become obvious.
An LLM predicts one token at a time. For each token, it runs a stack of maths. To predict the next token, it needs the maths for every single token before it. Round and round, until it produces a special "I'm done" token. Picture every API call as a toll road. You pay to drive in (input tokens), and you pay to drive back out (output tokens). Output is the pricier lane. Take Anthropic's lineup as an example. Claude Fable 5 sits at $10 per million tokens in and $50 per million out. That's a 10x spread between the cheapest and priciest model from the same company.
Now look for the smallest line on that price sheet. Cached tokens cost one tenth of normal tokens. Keep that number in your pocket. It's worth more than everything else in this article combined. Without caching, the model needs the whole conversation every single time. You send message one. It sends back message one plus its reply. You send the whole lot plus message three. It sends all of that back plus reply two.
Source: Towards Data Science · Summarized by HeadlinesBriefing