OpenRouter guide shows how to gate pull requests on LLM evals
OpenRouter’s guide explains how to run a fixed eval set in GitHub Actions, set a regression threshold, and make the result a required pull-request check. It also covers repeated sampling, provider routing, and evals for agents that call tools.
SessionWatcher editorial · Published · Updated · Source announcement: 2026-10-01
A fixed eval set in CI
OpenRouter’s October 1 guide walks through a pull-request gate for prompt and agent changes. It recommends keeping test cases in the repository, running them in a CI job, and making the script fail when the pass rate falls below a threshold.
The example uses a path-filter job to trigger evals only for relevant changes. The guide cautions that skipped jobs and skipped workflows behave differently in GitHub, and says to require both the filter job and eval job as status checks so failures cannot silently bypass evaluation.
Measure noise and choose a threshold
The tutorial recommends repeating runs against an unchanged branch to measure variation before setting the threshold. Its sample script runs each case several times and uses the majority verdict; it also routes requests to one provider to avoid provider differences appearing as prompt regressions.
The guide says temperature and seed only help on models that list them in supported_parameters. It also notes that repeated sampling increases cost, which the example measures using the response usage cost field.
When the agent calls tools
For tool-calling agents, OpenRouter presents Ori Eval as an option for assertions about tool use as well as model responses. The guide contrasts this with a plain script that checks answer text, and advises calibrating a judge model against human assessments before letting its grades fail builds.
AI assisted reporting, checked against the linked official sources. Source pages checked 2026-10-01. Editorial process and corrections.