OpenAI compatible API · Attested · Public status
Nebius Token Factory performance
Review measured TTFT, TTFB, effective throughput, uptime, and sampled model routes for Nebius Token Factory on TrustedRouter using metadata-only production probes.
Onebase URL to migrate
100sof models and routes
0prompt or output logs. Always.
Nebius Token Factorynebius
48 samplesContinuously sampled provider performance. TrustedRouter reports unsupported route and probe-configuration rows separately from provider downtime. Prompt and output content is not stored.
| p50 TTFT | 1859 ms |
|---|---|
| p95 TTFT | 5211 ms |
| p50 TTFB | 2078 ms |
| Effective throughput | 80 tok/s n=9 |
| Uptime | 100.00% |
Measured model routes
| Model | p50 TTFT | p50 TTFB | Effective throughput | Uptime | Config excluded | Availability samples |
|---|---|---|---|---|---|---|
| google/gemma-3-27b-it | 819 ms | 819 ms | — | 100.00% | — | 1 |
| openai/gpt-oss-120b | 1212 ms | 1212 ms | 57 tok/s n=2 | 100.00% | — | 2 |
| Qwen/Qwen2.5-VL-72B-Instruct | 1256 ms | 1256 ms | — | 100.00% | — | 3 |
| nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B | 1304 ms | 1304 ms | — | 100.00% | — | 2 |
| Qwen/Qwen3-32B | 1405 ms | 1404 ms | — | 100.00% | — | 3 |
| deepseek-ai/DeepSeek-V4-Pro | 1423 ms | 1423 ms | — | 100.00% | — | 2 |
| moonshotai/Kimi-K2.7-Code | 1665 ms | 1665 ms | — | 100.00% | — | 4 |
| moonshotai/kimi-k3 | 1689 ms | 1689 ms | 62 tok/s n=1 | 100.00% | — | 2 |
| moonshotai/Kimi-K2.6 | 1720 ms | 1720 ms | — | 100.00% | — | 4 |
| nvidia/nemotron-3-ultra-550b-a55b | 1859 ms | 1859 ms | 105 tok/s n=2 | 100.00% | 1 unsupported_route |
3 |
| MiniMaxAI/MiniMax-M3 | 2078 ms | 2078 ms | 80 tok/s n=2 | 100.00% | — | 2 |
| nvidia/Llama-3_1-Nemotron-Ultra-253B-v1 | 2118 ms | 2118 ms | — | 100.00% | — | 1 |
| meta-llama/Llama-3.3-70B-Instruct | 2186 ms | 2185 ms | — | 100.00% | — | 4 |
| NousResearch/Hermes-4-405B | 2224 ms | 2224 ms | — | 100.00% | — | 2 |
| nvidia/Cosmos3-Super-Reasoner | 2236 ms | 2236 ms | 54 tok/s n=1 | 100.00% | — | 1 |
| Qwen/Qwen3-30B-A3B-Instruct-2507 | 2281 ms | 2280 ms | — | 100.00% | — | 2 |
| openbmb/MiniCPM-V-4_5 | 2410 ms | 2409 ms | 87 tok/s n=1 | 100.00% | — | 1 |
| Qwen/Qwen3-235B-A22B-Instruct-2507 | 2490 ms | 2490 ms | — | 100.00% | — | 3 |
| nvidia/Nemotron-3-Nano-Omni | 2769 ms | 2769 ms | — | 100.00% | — | 1 |
| Qwen/Qwen3-Next-80B-A3B-Thinking | 2792 ms | 2792 ms | — | 100.00% | — | 3 |
| NousResearch/Hermes-4-70B | 3437 ms | 3436 ms | — | 100.00% | — | 1 |
| zai-org/GLM-5.1 | 4248 ms | 4247 ms | — | 100.00% | — | 1 |