The question that has no single answer
Say your team runs agent tasks all day, and every one of them gets logged with a token count. At the end of the month you divide total spend by number of runs and get "average cost per task." That number is confidently wrong, and it's wrong in a specific, well-documented way.
What does one unit of output actually cost?
Kostenrechnung — the German cost-accounting tradition, built up across the twentieth century by practitioners and academics working on how manufacturing firms should account for cost internally, distinct from the external financial-reporting rules a public company files under — established that this is not one question but a family of them, and that most costing disasters come from confidently answering the wrong member.
What did this absorb of total cost? is a different question from what does one more unit add? — which is different again from what would disappear if we stopped doing this altogether?
Steering on the wrong one produces prices that chase away good business, product lines killed for phantom losses, and the fixed-cost death spiral.
That was all worked out for factories. It applies now, urgently, to something new: for the first time, an organisation exists whose workers are metered by the unit. Your labour bill arrives in tokens. Your supervision bill arrives in minutes. Every task leaves a logged trail.
The accountant's century-old dream — costs traceable by cause rather than smeared by allocation — is suddenly the default data condition. And the discourse around agents is busily re-deriving, badly, every mistake the tradition already catalogued.
The finding that reorganises everything
The steering figure was never cost. It is contribution margin — what a unit earns above the costs that unit alone causes.
Apply that honestly to agent work and you get a conclusion that follows directly from the accounting logic, even though it rarely shows up stated this plainly in discussion of agent costs:
It is not a low-quality unit. It consumed the proportional cost — tokens — and it consumed the scarce complementary factor, your review minutes, and it returned nothing you can bank.
Both inputs spent. No output. That is below zero, not merely below expectations.
Which means the unit that earns anything is not "an output." It is a verified output. Everything else on this site — gates, checks, evidence over assertion — stops being engineering hygiene at that point and becomes the thing that decides whether the arithmetic works.
Why your cost-per-task dashboard is lying to you
Current agent tooling reports average cost per task by dividing total spend by number of runs — silently folding in orchestrator turns, retries, system prompts and retrieved context. Shared costs allocated by volume.
That is full costing with a volume surcharge, which is the exact construction the tradition spent decades building alternatives to.
Two consequences, both textbook material in cost-accounting courses long before anyone had a token bill — this is not a novel finding, it's the standard critique of full-cost allocation:
High-volume simple tasks subsidise low-volume complex ones. The average tells you nothing about either.
And the number moves inversely with utilisation. Do fewer tasks, and each one "costs more" — which invites cuts, which raises the average further. The fixed-cost death spiral is model-agnostic. It will run perfectly happily on a token dashboard.
The streetlight fallacy, in cost-accounting form
Tokens are the visible, metered, proportional block. They are usually the smallest one.
Review minutes, orchestration, standing context and rework dominate the real cost structure of an operated harness. And cutting tokens often inflates them — a cheaper model produces more failures, which produces more review, which is the expensive input.
Minimising the metered factor because it is the metered factor is looking for your keys under the streetlight.
Here's a hypothetical that illustrates the mechanism, not a measured result: a cheap model that ships work only 80% of which passes verification can be strictly dominated in margin terms by a model at three times the token price, once you price in the review minutes the failing 20% costs.
The arithmetic: the expensive model's extra tokens are trivial next to the review minutes the cheap model's failures consume. But routing objectives optimise benchmark accuracy per dollar of inference — not contribution per minute of your review capacity.
That computation is routinely absent, and it is the one that decides whether the operation is profitable.
Cost attaches to a decision, not to a thing
Public discussion of agent economics slides freely between marginal token cost, subscription cost, orchestration cost and development amortisation — attributing all of them to whatever object is being argued about. "An agent costs X per task."
The missing discipline is attaching each cost to the level of decision that causes it:
| Level | What it carries | The question it answers |
|---|---|---|
| The task | tokens, the review minutes it consumed | should I run this one? |
| The stream | standing context, memory maintenance, monitoring | should this workstream exist? |
| The org / the period | subscriptions, orchestration, the tooling | what capacity should we hold at all? |
And the practical consequence is the layered margin statement: a task can be margin-positive while the stream it belongs to is not covering its own readiness costs.
Conflating the layers gives you both failure modes at once — killing tasks that were actually paying, while feeling good about cheap tasks in a stream that has never covered its standing cost.
A ZEO that never looks at the second and third layers quietly accumulates streams that are individually cheap and collectively insolvent.
Isn't this just Taylorism with better sensors?
The question deserves a straight answer, because the objection is serious and the answer is not obvious.
Applied to people, per-minute work metering is surveillance. Frederick Taylor's own scientific-management program, and the piecework and time-and-motion systems it inspired through the early twentieth century, are the namechecked historical case: they measured effort rather than judgement, and workers gamed the meter in ways that corrupted both the work and the data — a critique made at the time and repeated by labor historians since, not a claim unique to this piece.
Applied to agents, it is simply correct engineering practice. They are owned instruments with no privacy interest, whose runs are legitimately fully logged, and refusing to instrument them would be negligent rather than principled.
And the genuinely human-facing number here — review minutes — measures the organisation's consumption of the operator's attention.
Taylorism metered the human as the resource to be economised on.
Metering review minutes accounts the human as the scarce resource to be economised for.
Same instrument, opposite direction. That difference is the whole ethical position, and it is worth being explicit about rather than assuming people will grant it.
What smart people get wrong
What to actually track
Four numbers per work product, and none of them require a system:
Already computed. Read it, write it down.
Time spent reviewing, correcting and re-running. This is the binding constraint and the one nobody records, which is why nobody knows their real cost.
The single best proxy for whether the work was verifiable at all. A stream with a falling first-pass rate is a stream getting more expensive while its token line looks flat.
Task, stream, or standing capacity. One word. It's what stops you conflating them later.
Four numbers, thirty seconds, at the moment you finish. That is a cost record most small firms never achieve about their human work, available for free about your agent work — and it is the only thing that turns "this feels faster" into something you can price against.
The harness articles on this site carry a dated list of claims likely to expire, because product mechanics change weekly.
This piece doesn't need one. Contribution margin was worked out in the 1950s and the volume- surcharge trap in the 1930s. When the tokens are billed differently next year, the arithmetic above will be unchanged — which is rather the point of reading the old books.
The mechanics of the token side: what drives the bill, which dials move it, and the ledger line to keep per work product.

