Deterministic agent benchmarks

How many tokens does your framework really cost an AI agent?

Agent Tokenomics measures token usage across development frameworks with the rigor a benchmark deserves: fixed tasks, pinned model versions, multiple trials, and raw, independently reproducible transcripts behind every number. No self-reported estimates, no cherry-picked runs.

02
published comparisons
26
measured trials
SHA-256
verified transcripts
API
provider-reported usage

Replace anecdote with a fixed methodology.

AI coding agents now write a meaningful share of application code, and every token they spend has a real cost — in latency, in dollars, and in context budget. Claims about which framework is more “agent-friendly” circulate constantly, almost always backed by a single anecdote rather than a controlled run. Agent Tokenomics holds the model, the task, and the tooling constant, varies only the framework, and publishes the raw evidence alongside the number.

01

Provider-reported usage only

Every token count is read from the model provider’s own usage field on the raw API response (never a tokenizer estimate applied after the fact).

02

Multiple trials, full disclosure

Minimum five independent runs per task. Every nondeterministic input (tool output, timestamps, network calls) is disclosed, not papered over.

03

Hash-verified transcripts

Every raw request/response log ships with a SHA-256 manifest, so a results table can never silently drift from the evidence behind it.

04

One variable at a time

Model, prompt, tool access, and trial count are held constant across frameworks. Only the framework under test changes.

Get involved

Add the next verified submission.

Every comparison here ships with a fully measured, fully published run. Replicate one on a different model, a new framework version, or bring your own framework pair — and submit the results.

Read the submission format