Methodology

Good scores need good context.

These are self-reported runs, not an independent evaluation. The airline comparison is 90.0% for Cinderpaw versus 82.0% for the reference agent under the same setup. Telecom reached 99.1%, in a domain where published scores are already close to the ceiling.

The setup

Sierra Research’s τ²-bench, run on 2 September 2026. Agent model: z-ai/glm-5.3-flash. Simulated customer: gemini-2.5-flash. Maximum 200 steps per task. No failed runs were excluded from the reported totals.

Provider routing matters

In an exploratory check without pinned routing, the same 21 telecom tasks scored 11/21 and 20/21 across two runs with the same seed, model name and code. Nine outcomes changed. Those conversations went from 201 messages to between 34 and 72. This observation does not isolate the cause, but it makes unreported routing a serious comparison problem.

The reported runs pin the provider to Z.AI, with no provider fallback. A model name alone is not enough to reproduce a result.

Airline: 90.0% versus 82.0%

Cinderpaw passed 45 of 50 tasks. Database state matched on 46/50; write actions matched on 45/49. The reference agent scored 82.0% with the same model, provider, day and evaluation setup: an observed eight-percentage-point difference.

The published reference result is 77.3%, versus 82.0% in this setup. That 4.7-point discrepancy limits comparison with the public leaderboard. It is not a universal correction factor: subtracting it from 90.0% would not produce a validated adjusted score. These runs also do not establish statistical significance or performance on unrelated tasks.

Telecom: strong score, limited ranking value

Cinderpaw passed 113 of 114 tasks (99.1%). Published high-performing results are around 97.9–99.3%. This score does not establish a lead over those systems.

Telecom evaluates diagnosis and guiding a simulated customer through a fix. Of 351 checked actions, 255 (72.6%) are customer actions. The customer has 30 tools; the agent has 13. This is a different skill mix from airline, so the two scores are not combined.

Across 1,110 customer tool calls in 114 conversations, none came before the agent’s first message. The median was eight agent messages before the first customer tool call, with a minimum of four. That supports reading the transcripts as guided interactions; it does not quantify how much of the score comes from Cinderpaw itself.

Cost

The two airline runs totalled 100 task runs and about $0.55. Telecom cost about $1.45 for 114 tasks, or $0.013 per telecom task. Combined inference cost was about $2. These are costs for the reported runs, not a promise of current pricing or typical production cost.

What is still missing

A telecom ablation with a minimally helpful agent has not been run. Without it, the contribution of the agent versus the simulated customer is not isolated. Public leaderboard rows use different setups. Repeated trials and independent reproduction would strengthen the evidence.

Reproduce the runs

The harness is in scripts/tau2/ in the Cinderpaw repository. It reports the provider configuration and domain checks before running. Record the code revision, model, provider, simulator and evaluation settings when comparing results.