Benchmarks
Put the claims against the results.
90.0% airline. 99.1% telecom. Self-reported runs from 2 September 2026, using GLM-5.3-Flash with provider routing pinned. τ²-bench tests an agent on customer-service tasks with tools, policies and a simulated customer. Run date: 2 September 2026.
τ²-bench · airline
50 tasks covering flight changes, cancellations and refunds. Success depends on making the correct changes to the database.
- 90.0
- 77.3
- 76.5
- 67.5
- 67.0
- Cinderpaw
- τ² reference
- LongCat 2601
- LongCat
- GPT-5.1
| Cinderpaw · GLM 5.3 Flash (my test setup) | 90.0% |
| τ² reference agent · same model, same test setup, same day | 82.0% |
| τ² reference agent · as published (their test setup) | 77.3% |
An observed eight-point difference under the same evaluation conditions. The chart also includes published results from different setups; those are context, not a controlled ranking. Why public leaderboard comparisons need caution.
τ²-bench · telecom
113/114 tasks passed. Of the checked actions, 72.6% are performed by the simulated customer. This measures diagnosis and guidance, not just agent tool execution. Published results already approach the ceiling, so 99.1% does not establish superiority.
ARC-AGI-3
Interactive reasoning tasks where the agent learns how an environment works through its actions. (ARC Prize)
No Cinderpaw result yet. ARC-AGI-3 is a planned evaluation, subject to time and budget. An empty chart means no run, not a score of zero.
Agents’ Last Exam
Professional tasks evaluated through the files an agent produces. Completing the artifact matters, not how persuasive the explanation sounds.
Not run yet. I’m prioritizing the alpha and reproducible evaluations within the project’s budget. Any future result will identify the task set, model, setup and cost.
Task definitions and evaluation details: their own task lists from Berkeley RDI.
Cinderpaw results are self-reported. Read the model, provider, costs and comparison limits before drawing conclusions. How these were run in full. The exact model. The test setup. What each part truly checks. What it cost. And where my setup is not the same as the public boards.