Anthropic: Claude Opus 4.1 Benchmarks
Review independent benchmark scores and TrustedRouter route measurements for Claude Opus 4.1, with cited sources and links to current evaluation results.
anthropic/claude-opus-4.1
Published benchmark scores
Benchmark scores for Anthropic: Claude Opus 4.1 — every row links to its source, and a score is only ever attached to the exact checkpoint it was measured on. Vendor model-card and open-leaderboard numbers are cited, not run by us. Rows marked TrustedRouter · replays published are our own runs of this model through the gateway, with the full per-item replay published in trustedrouter-benchmarks so anyone can re-grade them.
| Benchmark | Category | Score | Source |
|---|---|---|---|
| SWE-bench Verified | Coding | 74.5% | Anthropic — Claude Opus 4.1 2025-08-05 |
TrustedRouter measurements
TrustedRouter publishes route and status measurements without storing prompt or output content. Provider latency and uptime are exposed through the model performance and uptime pages.
External benchmark references
- TrustedRouter performance pageTrustedRouter measurement
- TrustedRouter uptime pageTrustedRouter measurement
- AI IQ profile · IQ 97Independent model IQ score
- Anthropic model docsOfficial model information
- LMArena leaderboardIndependent benchmark index
- LiveBenchIndependent benchmark index
- Artificial Analysis modelsIndependent benchmark index
- HELMIndependent benchmark index