Independent, reproducible measurements of the knobs you can actually set on an OpenRouter request: models, providers, search engines, and tool budgets. Every score links to the configuration, costs, and telemetry behind it.
11 benchmarks2,454,759 task evaluationslast run Sep 18, 2026
Generated images, clips and speech, graded against the request and priced per output.
Image prompts built to fail: depth, direction, counting, and text in the frame.
44 models

Six seconds of video, held to one duration and resolution across every model.
24 models
Edit a meme as an image or animate it as a clip, graded on whether the brief landed.
57 models

Text models take on unconventional tasks like creating drawings and playable games.
Text models create drawings from prompts. The results are graded as images.
204 models


Text models create playable games from a single brief in one attempt.
29 models



Hard-to-locate facts on the live web, scored on persistent multi-step research.
4 modelslast run Aug 18, 2026
Questions whose answers are lists, scored for exhaustive retrieval with no padding.
4 modelslast run Aug 18, 2026
Humanity's Last Exam as a search benchmark: expert questions answered with live search.
2 modelslast run Aug 17, 2026
Fill an entire table; answer-item accuracy scores partial matches.
4 modelslast run Aug 18, 2026
For usage-based views of the same models, see the model rankings and the full model list.