← Back to The Mixing Bowl
building in publicAI-Nativeengineering

The Spend Report Was Right About One AI Cache and Wrong About the Other

By Chad Holdorf·October 5, 2026·5 min read

Bakery Buddy runs a weekly report against its own AI spend: every Claude call, broken down by feature, with a "what to do about it" section at the bottom. Two weeks ago that section told me to turn off caching on a feature that was burning money for no reason, and to turn it off on a second feature for the same stated reason. The first call was right. The second one, if I'd followed it, would have made the bill bigger, not smaller.

What a cache actually costs

Anthropic's prompt caching lets a repeated block of text, a system prompt, a set of style rules, a long order history, get written once and read back cheaply on a later call instead of being billed at full price every time. Writing a cache entry costs 1.25 times the normal input price. Reading one back costs 0.1 times. So caching only pays for itself once enough reads land on a write to make up for that 25% premium paid up front.

The question "how many reads is enough" has one real answer, and it wasn't the one sitting in this codebase's own comments: "break-even is a single read." That's the kind of number that sounds right (a write costs more, one cheap read should cancel it out) without anyone actually doing the algebra behind it.

Doing the algebra

A write costs 1.25x. A read costs 0.1x. Without caching at all, that same write plus however many reads (R) would each cost 1.0x, for a total of R+1. Set the two totals equal and solve for R, and the real break-even isn't 1.0 reads per write. It's 0.25 divided by 0.9, about 0.278.

That's not a rounding difference. The number this codebase had been using was almost 3.6 times too conservative. A feature that was genuinely saving money could sit comfortably on the wrong side of the wrong threshold and still get flagged as waste.

Two features, one wrong number, two different outcomes

The report measured two features over the same 30 days and flagged both as cache waste.

`classify_occasion` reads an order's free text and tags what kind of occasion it's for, a classification each order only ever needs once. Over 1,796 calls it wrote 64,019 cache tokens and read back exactly zero. 0.000 reads per write. That isn't close to either threshold, the real one or the wrong one. It's a feature that structurally can't read its own cache back, because nothing ever asks Claude to classify the same order twice, and it had been paying the 25% write premium 1,796 times for nothing. The report was right about this one, and it would have been right under either version of the math.

`social_compose_post_and_email`, which drafts a social caption and a short thank-you email together whenever a finished cake photo goes up, told a different story. 213 calls, 1,114,868 cache tokens written, 382,589 read back. 0.343 reads per write. That number lands in the one place where the two thresholds disagree: below the wrong one (1.0), above the real one (0.278). Judged against "you need one read to break even," it looks like a loser, and the report called it one. Judged against the real arithmetic, it was already a net saving, and turning it off on the report's advice would have raised the cost, not lowered it.

Why the mistake was invisible on the first feature and decisive on the second

Both features were measured with the same wrong formula, and only one of them had its verdict flipped by it. That's the part worth sitting with. A number that's off by 3.6x will still give you the right answer, as long as the real number you're checking it against is far enough from both thresholds that it doesn't matter which one you use. `classify_occasion` at 0.000 is nowhere near either line. `social_compose_post_and_email` at 0.343 sits exactly in the gap between them, which is the one range where a wrong formula stops being a rounding error and starts being a wrong decision.

A wrong constant doesn't announce itself by being wrong on everything it touches. If it announces itself at all, it's on the one case that happens to land between where the wrong answer and the right answer disagree. Every other case just quietly confirms a formula that was never actually correct.

The fix wasn't the one feature, it was the formula

`classify_occasion` got added to the list of features that skip caching entirely, which was the easy part. The harder part was making sure the next borderline case doesn't get decided by the same wrong number again. "Break-even is one read" was a sentence in a comment, the kind of thing that's easy to type once and never re-derive. It's now an exported, tested constant, computed from the two real prices (0.25 divided by 0.9) instead of restated from memory, with a test asserting that `social_compose_post_and_email` stays cached specifically because it clears that number.

That's the actual repair. Not catching the one feature that was wasting money. Making sure the rule that decides the next one can't quietly drift back to a number nobody ever checked.

What generalizes past caching

A rule of thumb that happens to produce the right answer on every case you've already looked at isn't the same as a correct rule. It's a rule that hasn't yet been tested on a case landing in its blind spot. You don't find that case by staring at the formula harder. You find it by measuring enough real instances that one of them eventually lands exactly where a wrong constant and a right constant disagree, and then doing the algebra instead of trusting the sentence that's been sitting in a comment since before anyone checked it.

If your own system has a threshold anywhere, a break-even point, a timeout, a retry count, a "good enough" cutoff, it's worth asking where that number actually came from. If the honest answer is "it sounded about right," that's worth five minutes to rederive from the real prices or the real constraints, before it quietly tells you to turn off the one thing that was already saving you money.

Frequently Asked Questions

What was the actual mistake?

Bakery Buddy's own AI spend report used the rule that a prompt cache breaks even after one read per write to flag two features as wasteful. The real break-even, from the actual prices (a cache write costs 1.25x normal input, a read costs 0.1x), is about 0.278 reads per write, nearly 3.6 times lower. One flagged feature, classify_occasion, genuinely was wasteful either way, at 0.000 reads per write. The other, social_compose_post_and_email, was sitting at 0.343 reads per write: already a net saving under the real number, and incorrectly flagged as a loss under the wrong one.

What would have happened if the report's advice had been followed on both?

Turning off caching for social_compose_post_and_email would have raised its cost, not lowered it, because it was already saving money at 0.343 reads per write against a true break-even of 0.278. Only classify_occasion, at 0.000 reads per write, actually needed to stop caching.

How was this caught?

By re-deriving the break-even formula from the two real per-token prices instead of trusting the one-read rule already sitting in the code's own comments, then checking both features' real measured numbers against the correct threshold instead of the approximate one.

What changed in the code?

classify_occasion was added to the list of features that skip caching entirely. The break-even number itself was turned from a sentence in a comment into an exported, tested constant derived from the real prices, with a test that locks in social_compose_post_and_email staying cached because it clears that exact number.

Written by
Chad Holdorf
Founder, Bakery Buddy

I build Bakery Buddy for my wife Lindsay's cake studio, Marin Cake Studio, and this is the week I learned a break-even number is only as good as the arithmetic nobody checked behind it.


Ready to put this into practice?

Join the Waitlist