Benchmarks

Put the claims against the results.

90.0% airline. 99.1% telecom. Self-reported runs from 2 September 2026, using GLM-5.3-Flash with provider routing pinned. τ²-bench tests an agent on customer-service tasks with tools, policies and a simulated customer. Run date: 2 September 2026.

τ²-bench · airline

50 tasks covering flight changes, cancellations and refunds. Success depends on making the correct changes to the database.

  • 90.0
  • 77.3
  • 76.5
  • 67.5
  • 67.0
  • Cinderpaw
  • τ² reference
  • LongCat 2601
  • LongCat
  • GPT-5.1
τ²-bench airline. Share of tasks done. Tap a bar to see more.
Cinderpaw · GLM 5.3 Flash (my test setup)90.0%
τ² reference agent · same model, same test setup, same day82.0%
τ² reference agent · as published (their test setup)77.3%

An observed eight-point difference under the same evaluation conditions. The chart also includes published results from different setups; those are context, not a controlled ranking. Why public leaderboard comparisons need caution.

τ²-bench · telecom

113/114 tasks passed. Of the checked actions, 72.6% are performed by the simulated customer. This measures diagnosis and guidance, not just agent tool execution. Published results already approach the ceiling, so 99.1% does not establish superiority.

99.1%Cinderpaw · GLM 5.3 Flash, 113 of 114 tasks

ARC-AGI-3

Interactive reasoning tasks where the agent learns how an environment works through its actions. (ARC Prize)

      ARC-AGI-3. Share of games won. Zero runs yet. So zero dots.

      No Cinderpaw result yet. ARC-AGI-3 is a planned evaluation, subject to time and budget. An empty chart means no run, not a score of zero.

      Agents’ Last Exam

      Professional tasks evaluated through the files an agent produces. Completing the artifact matters, not how persuasive the explanation sounds.

          Agents' Last Exam. Share of tasks done. Zero runs yet. So zero dots.

          Not run yet. I’m prioritizing the alpha and reproducible evaluations within the project’s budget. Any future result will identify the task set, model, setup and cost.

          Task definitions and evaluation details: their own task lists from Berkeley RDI.

          Cinderpaw results are self-reported. Read the model, provider, costs and comparison limits before drawing conclusions. How these were run in full. The exact model. The test setup. What each part truly checks. What it cost. And where my setup is not the same as the public boards.