Fireworks reports model routing scores 97.6% on DeepSWE at $1.88 per task
Fireworks reports that an oracle choosing among 18 models per DeepSWE task reached a 97.6% pass rate at $1.88 per task, versus 74.1% at $6.52 for GPT-6 Astra. The result is hindsight-based and evaluated on the same 113 tasks used to choose winners, so it is not evidence of a production router achieving that performance.
SessionWatcher editorial · Published · Updated · Source announcement: 2026-09-21
What Fireworks measured
Fireworks analyzed DeepSWE v1.1, an agentic coding benchmark of engineering tasks involving repository inspection, tool use, code changes, execution, and tests. Its fixed-model policy chose one model at the start of each task and did not switch mid-session.
The best fixed model listed was GPT-6 Astra, with a 74.1% pass rate at $6.52 per task. Fireworks' oracle router selected the winning model separately for each task across 18 models, reaching 97.6% at $1.88 per task.
Why the headline result is a ceiling, not a production forecast
The oracle result uses hindsight: Fireworks ran all 18 models on every task before selecting the winner. It scored on the same 113 tasks it selected from, with four rollouts per model-task pair; the post notes that choosing among 18 noisy estimates biases the score upward.
Fireworks says an actual router must predict which model to use before seeing the outcome. It cites model recall as a challenge and says routing approaches may fail to reliably beat simple baselines. The benchmark therefore shows potential within this model pool, not demonstrated performance from a deployable router.
Open-model result and cost context
Restricting the oracle to the six named open-weight models, Fireworks reports a 90.3% pass rate at $1.45 per task. These are benchmark results, not a guarantee that a user can achieve the same accuracy or cost in their coding workflow.
For the benchmark costs, Fireworks says it scaled each model's task costs so its mean matched the DeepSWE leaderboard's published figure, because raw trial costs did not match the board. The analysis used trials refreshed September 17, 2026, covering 113 tasks and 18 models.
AI assisted reporting, checked against the linked official sources. Source pages checked 2026-09-30. Editorial process and corrections.