OpenRouter explains confidence-based model escalation routing
OpenRouter describes using structured outputs to return a numeric confidence score, measuring score bands against your own error rates, and routing low-confidence answers to a stronger model. The score is a self-report, not a calibrated probability, and escalation requires application code; OpenRouter model fallbacks address errors instead.
SessionWatcher editorial · Published · Updated · Source announcement: 2026-10-01
Build a score you can route on
OpenRouter’s guide recommends asking the model for an answer and numeric confidence field in a schema-validated structured response, rather than trying to infer uncertainty from prose. It cautions that a reported score is a self-report, not a calibrated probability: a score of 0.85 does not mean an 85% chance the answer is correct.
Before relying on the score, run representative requests through the cheaper model, grade the answers, and check whether lower score bands have higher error rates. Use that observed ordering to choose a cutoff for your own traffic, rather than copying a threshold from an example.
Balance accuracy, cost, and latency
A higher cutoff sends more requests for a second model call, which can reduce errors among answers you keep but adds cost and latency. OpenRouter says the cost depends on the price difference between the models and how often escalation happens; check current per-model pricing when estimating it.
The guide’s example uses 200 requests and illustrates that a 0.7 cutoff would escalate 14% of traffic. OpenRouter explicitly labels the numbers illustrative, not a target, so use measured error rates from your workload instead.
Implement and keep monitoring
Your application must make the confidence-based escalation decision: read the score and send low-scoring requests to a stronger model, either in the same function or as a separate first pass. OpenRouter says its model fallbacks trigger on errors such as provider downtime or rate limits, not on a valid answer with a low confidence score.
Log score distributions, escalation rates, and errors on answers that were not escalated. Recheck the cutoff when models or traffic change. Structured-output support can vary by provider endpoint, so validate parsed JSON and use require_parameters: true if you need routing restricted to endpoints that support the requested parameters.
AI assisted reporting, checked against the linked official sources. Source pages checked 2026-10-01. Editorial process and corrections.