All AI updates

OpenRouter compares Jev 1.13 and Claude Opus 5 on Banking77 classification

OpenRouter reports Jev 1.13 scored 81.0% accuracy versus Claude Opus 5 at 84.4% on 3,080 Banking77 utterances. With prompt caching, Jev cost $0.11 per 1,000 requests versus $2.42 for Opus; the post also tests confidence-based routing and notes important limits to the comparison.

SessionWatcher editorial · Published · Updated · Source announcement: 2026-09-22

Test results

OpenRouter tested Jev 1.13 and Claude Opus 5 on the 3,080-example Banking77 test split, classifying utterances into 77 intents. Jev reached 81.0% accuracy and 80.5% macro-F1; Opus reached 84.4% and 83.6%. Both returned valid responses for all examples.

Median client-observed latency was 175 ms for Jev and 2,266 ms for Opus. The full run cost $0.34 for Jev and $7.44 for Opus, or $0.11 and $2.42 per 1,000 requests respectively. The Opus run used prompt caching for its repeated system prompt.

Sources: Is Jev as Accurate as Frontier Models at Classification?

Confidence-based routing

The report tested sending lower-confidence Jev classifications to Opus. At a 0.90 threshold, 75.9% of requests were handled by Jev, combined accuracy was 84.0%, and cost was $0.69 per 1,000 requests. The report says Jev's confidence score is not a calibrated probability, so teams should select and test a threshold on their own traffic.

Sources: Is Jev as Accurate as Frontier Models at Classification?

Limits of the comparison

This was one dataset, one prompt design, and one short test window. The cascade results were evaluated on the same examples used to measure accuracy, which OpenRouter describes as an upper bound. Claude Opus 5 was tested with reasoning off and structured outputs on, not with other settings or few-shot examples. The results are evidence about this classification setup, not a general ranking of the models.

Sources: Is Jev as Accurate as Frontier Models at Classification?

AI assisted reporting, checked against the linked official sources. Source pages checked 2026-09-30. Editorial process and corrections.