Multi-agent designs cost 4x to 8x more tokens. When does that pay off?
Published papers put a single agent at 1.9K tokens against 15.5K for the best multi-agent method. The overhead is real. So is the accuracy gain.
Dana Reyes edits Handoff and writes its orchestration and cost coverage. The brief is narrow on purpose: what it takes to keep autonomous software running once it is past a demo, and what that costs per completed task.
Her working rule is that a claim without a number is not reporting. Every figure in a Handoff article carries a source, a sample size and a date, and where a figure is illustrative rather than measured, the table caption says so.
She reads every incident report the labs publish and treats vendor announcements as claims to be checked, not news to be relayed.
Tips, corrections and data sets go to dana@readhandoff.com. If you are reporting an error in a published piece, quote the sentence. It gets fixed faster.
Published papers put a single agent at 1.9K tokens against 15.5K for the best multi-agent method. The overhead is real. So is the accuracy gain.
Headlines say inference is getting cheaper. Trackers say the market split in two, and agents run on the half that got more expensive since January.
Caching is the cheapest lever on an agent bill, and agents keep breaking it. What invalidates a prefix, and how to order context so it survives.
Published numbers are not wrong, they answer different questions. Utilisation is the hidden variable, and nobody puts it in the model.