LLM Provider Latency Benchmarks
Measured time-to-first-token, time-to-first-byte, throughput, and success rate for LLM providers routed through TrustedRouter.
Provider speed data from real routed requests.
TrustedRouter publishes metadata-only measurements for time-to-first-token, time-to-first-byte, throughput, uptime, and excluded probe-configuration rows. The goal is to show what the router actually sees, not what a provider claims in a launch post.
- ✓ Provider and model leaderboards
- ✓ Per-provider performance pages when enough samples exist
- ✓ Per-model performance pages when enough samples exist
- ✓ Prompt and output content never stored for these rollups
{
"provider": "tinfoil",
"model": "moonshotai/kimi-k2.6",
"p50_ttft_ms": 1192,
"uptime": 0.999,
"sample_count": 42
}
Provider pages
- Tinfoil performanceConfidential and E2EE route samples
- Anthropic performanceClaude route samples
- Google Vertex performanceManaged Google Cloud route samples
- Google AI Studio performanceGemini Developer API route samples
Model pages
- Kimi K2.6 performanceProvider-specific route metrics
- Gemini Flash performanceFast multimodal route metrics
- GPT Nano performanceSmall-model latency metrics
Current routes, prices, privacy, and measured performance.
Catalog facts come from the routes currently configured in TrustedRouter. Performance uses the same cached metadata snapshot as the public leaderboard. Prompts and outputs are not part of these measurements.
| Model | Providers | Context | Input | Output | Privacy | Measured route |
|---|---|---|---|---|---|---|
Anthropic: Claude Opus 4.8anthropic/claude-opus-4.8 |
3 routes | 1,000,000 | $5.25/1M | $26.25/1M | varies 5 cited scores | 1803 ms TTFT anthropic · 58 tok/s · 100.00% available · n=7 |
OpenAI: GPT-5.5openai/gpt-5.5 |
4 routes | 1,050,000 | $5.25/1M | $31.5/1M | ZDR 3 cited scores | 2180 ms TTFT openai · 100.00% available · n=5 |
Google: Gemini 3.5 Flashgoogle/gemini-3.5-flash |
7 routes | 1,048,576 | $1.575/1M | $9.45/1M | ZDR | 1997 ms TTFT google-ai-studio · 100.00% available · n=17 |
MoonshotAI: Kimi K2.7 Codemoonshotai/kimi-k2.7-code |
19 routes | 262,144 | $0.735/1M to $0.9975/1M | $3.675/1M to $4.2/1M | ZDR 5 cited scores | 1618 ms TTFT inceptron · 100.00% available · n=42 |
Z.ai: GLM 5.2z-ai/glm-5.2 |
45 routes | 1,048,576 | $0.714/1M to $1.575/1M | $1.575/1M to $5.5125/1M | E2EE 4 cited scores | 4152 ms TTFT parasail · 29 tok/s · 97.44% available · n=228 |
MiniMax: MiniMax M3minimax/minimax-m3 |
21 routes | 1,048,576 | $0.2835/1M to $0.63/1M | $1.155/1M to $2.52/1M | ZDR 4 cited scores | 1785 ms TTFT minimax · 69 tok/s · 100.00% available · n=42 |
Browse every modelReview provider policiesOpen the full leaderboardSnapshot 2026-08-07T03:00:56Z
Questions
Are these vendor claims?
No. The leaderboard is generated from TrustedRouter synthetic probes and runtime metadata, not provider marketing claims.
Do latency probes store prompts or outputs?
No. Status and leaderboard records store provider, model, latency, token, route, cost, and outcome metadata only.