Claim first: most of what you pay a model is not the clever last sentence. It is rereading your own rules.
If you run AI on real work, you already have a packet that repeats: how we talk to customers, what tools exist, what "done" means, what not to invent. Every job starts with that packet, then a new ask. Cost vs quality is not a forced trade if you stop paying full price for the packet, clean the leftover instructions, and pick how hard the model should think on that job.
Prove it on one repeating kind of work. Not a vendor leaderboard.
Prompt caching: the handbook on the desk
This section is prompt caching. Vendors will say prefixes and cache hits. Same idea: stop making the model reread your standing rules on every question.
Picture a new hire. You give them a 40-page employee handbook, then a one-line customer question. They answer. Next customer: you take the handbook back, hand them a fresh copy, and ask again.
That is how a lot of AI use still bills. The model has to "reread" the long setup every time. The setup is your system prompt, your tools, your examples, your policies. The question is short. The reread is not.
That leave-it-on-the-desk move is prompt caching.
1/ First time through, they read it and keep it handy. You pay to set the desk.
2/ Next question, if the handbook is still the same book, they do not start at page one. They look up. That reread is much cheaper.
3/ Only the new question is new work.
You do not need the vendor jargon. You need the rule: same front of the packet, same model, still on the desk.
What breaks prompt caching
The handbook falls off the desk. The cache is picky on purpose. It is not "close enough." If the front of the packet changed, it is a new book.
Three ordinary ways teams knock it off:
1/ You stamp today's date on page 1. A clock, a ticket ID, or "generated at 9:02" in the standing rules makes every call a new handbook. Put changing bits in the question, not in the rules.
2/ You reshuffle the chapters. Tools and policies have an order. Shuffle them "to be tidy" and the desk is gone.
3/ You change how hard they should think mid-shift, or you swap which person is at the desk. Different model, or a different "think hard" setting, is a different reader. The old desk does not transfer.
There is also time. If they sit idle long enough, they put the book back on the shelf. Long waits for a tool or a sub-task can do that. Then the next turn pays to set the desk again, and that rewrite costs more than a cheap reread.
None of this is exotic. It is the same discipline as a shared drive folder: do not rename the parent directory every hour.
Leftover sticky notes (prompt debt)
Handbooks collect sticky notes. "Double-check everything." "Be as thorough as possible." "Always think in this six-step pad." Those notes were often written when the intern was new.
A stronger model will follow them literally. Double-check becomes two full lookups. Thorough becomes a tour of every file. The pad stacks on top of the thinking the model already does.
That is prompt debt. Old instructions that patched a weaker model now tax a stronger one.
You do not need a special command to start. Read the standing rules as if a sharp person will obey every line. Cut:
1/ Verify-twice rituals
2/ "Maximally thorough" and all-caps MUST language
3/ Mandatory scratchpads
4/ Example chats from a model that no longer works here
5/ Two rules that fight (always refund / never refund without a manager)
6/ Settings written for last year's model
Labs have shown that cleaning those notes can cut cost and raise accuracy on their own support tests. Treat that as a pattern, not as your number. Your number comes from your repeating job.
How hard to think
"Effort" is the vendor word for how hard you ask the model to work. Low is "give me the obvious good answer." High is "slow down, look again, consider the other path."
Higher is not always better.
1/ A sharp person thinking lightly can beat a junior grinding, and cost less.
2/ If extra thinking does not change whether the work would have passed, you are paying for rumination.
3/ If the curve is flat, the job is not "think harder." The job is a cleaner packet, a cache that actually hits, or a smaller ask.
The miss in the other direction is real too. Too light, and it answers from the first search instead of the third. The writeup looks finished. The evidence is thin.
Calibrate on the work, not on a vibe that "this is important so max it."
A benchmark with no lab required
Vendor charts are not your work shape. You already have the only eval that matters: jobs you already accepted.
Do this once. Keep it boring.
1/ Pick one kind of repeating work. Ten to twenty past items. Freeze the inputs. Write down what "good enough" meant when a human signed it. Do not edit the gold after you see scores.
2/ Run the same items a few ways: a stronger model thinking lightly, a mid model thinking normally, a small model thinking hard. If you cleaned the sticky notes, run that as a second pass, not a silent overwrite.
3/ For each way, record only what you can see: would a human have accepted it, roughly what it cost, how long it took. No invented dollars. If you only have tokens, price them from a dated public table and keep the date.
4/ Write one line: effort stopped paying after ___ on this job.
That line is the blog figure and the ops habit. Re-run when a new frontier ships, or monthly. Keep old tables. Do not chase a score by changing the answers.
Do not publish the tickets. Do not put customer names, internal cost programs, or your org's private packet in the public essay. The method is public. The cases stay inside the building.
What to do Monday
1/ Find the repeating packet. If the standing rules change every call, caching cannot help yet.
2/ Move dates, IDs, and "now" out of the rules and into the question.
3/ Peel the sticky notes. One pass. Read them as orders, not vibes.
4/ On that one job, run the small matrix. Keep the one-line finding.
5/ Only then raise effort, or spend on a bigger model, where the table still moves.
The model should not reread the handbook every time. Leave it on the desk. Ask a new question. Measure where extra thinking stops earning its keep.
Related: Personal AI is in Uber time (the sticker vs the invoice). This piece is the operator sequel inside a single workflow: you can waste production money on rereads and leftover instructions long before you ever self-host a GPU.