Preheader: Cheap tokens still need boring proof.
Before you switch your agents to a cheaper model, run the stupid test.
The dangerous part of a new model release is not the marketing page. It is the Slack thread after it.
Someone sees a lower token price. Someone else drops a benchmark screenshot. Then a team changes the default model for a bunch of agent workflows and calls it cost control.
That is not cost control.
That is guessing with a nicer chart.
Sonnet 5 may deserve a test. The temporary price window is real. The tool-use claims are relevant if you run agents that touch browsers, terminals, docs, tickets, or code.
Fine. Test it.
Do not crown it.
The thing to measure is boring: did the job finish, and how much cleanup did it create?
Use a one-page sheet. Five rows is enough.
Columns:
workflow
current model
test model
effort setting
token cost
wall time
retries
tool-call mistakes
human cleanup minutes
final artifact
pass or fail
Pick real work. A research brief. A Google Doc handoff. A browser/source triage job. A Kanban debugging task. A content transformation that usually needs taste.
Run each task once with the current default and once with the cheaper model.
If the cheaper model saves tokens but adds review time, mark it honestly.
If it finishes faster but lies about the artifact, fail it.
If it needs two retries to produce the same usable output, include the retries in the cost.
After five tasks, make the routing decision.
If it wins on at least three workflows without lowering verification quality, use it for those workflows.
If it only wins on the easy ones, route only the easy ones.
If it loses on the ugly jobs, keep the expensive model where the ugly jobs live.
The mistake is treating model choice like a brand preference.
It is closer to hiring a contractor. Start with small jobs. Check the work. Track the rework. Stop sending it the jobs it keeps messing up.
If nobody is willing to fill in the sheet, do not change the default.