OpenAI compatible API · Attested · Public status
Nebius Token Factory
Explore Nebius Token Factory models on TrustedRouter with current routes, token pricing, policy sources, privacy posture, regional availability, and API support.
Onebase URL to migrate
100sof models and routes
0prompt or output logs. Always.
Nebius Token Factorynebius
No logs| Provider | Nebius Token Factory |
|---|---|
| Provider website | https://nebius.com/ai-studio |
| Models | 28 public models |
| Prepaid routes | 26 |
| BYOK routes | 28 |
| Zero data retention | yes |
| Confidential compute | not claimed |
| Provider E2EE | not claimed |
| Policy note | Marked ZDR via TrustedRouter's arrangement — Nebius RETAINS inputs/outputs by default (for speculative decoding); zero retention is an opt-in control, which the deployed Nebius account has enabled. Nebius does not train on customer data. Policy source |
Measured performance
50 samplesContinuously sampled across Nebius Token Factory's routed models: p50 TTFT, effective throughput, and success rate. Effective throughput uses provider-reported output tokens over complete request time. Unsupported route and probe-configuration rows are separated from provider downtime. No prompt or output content stored.
| p50 TTFT | 1720 ms |
|---|---|
| Effective throughput | 74 tok/s n=8 |
| Uptime | 100.00% |
| Model | p50 TTFT | p50 TTFB | Effective throughput | Uptime | Config excluded | Availability samples |
|---|---|---|---|---|---|---|
| google/gemma-3-27b-it | 819 ms | 819 ms | — | 100.00% | — | 1 |
| Qwen/Qwen2.5-VL-72B-Instruct | 1053 ms | 1053 ms | — | 100.00% | — | 4 |
| openai/gpt-oss-120b | 1212 ms | 1212 ms | 57 tok/s n=2 | 100.00% | — | 2 |
| nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B | 1304 ms | 1304 ms | — | 100.00% | — | 2 |
| Qwen/Qwen3-32B | 1401 ms | 1401 ms | — | 100.00% | — | 4 |
| deepseek-ai/DeepSeek-V4-Pro | 1423 ms | 1423 ms | — | 100.00% | — | 2 |
| moonshotai/Kimi-K2.7-Code | 1665 ms | 1665 ms | — | 100.00% | — | 4 |
| moonshotai/kimi-k3 | 1689 ms | 1689 ms | 62 tok/s n=1 | 100.00% | — | 2 |
| moonshotai/Kimi-K2.6 | 1720 ms | 1720 ms | — | 100.00% | — | 4 |
| nvidia/nemotron-3-ultra-550b-a55b | 1859 ms | 1859 ms | 105 tok/s n=2 | 100.00% | 1 unsupported_route |
3 |
| MiniMaxAI/MiniMax-M3 | 2078 ms | 2078 ms | 86 tok/s n=1 | 100.00% | — | 2 |
| nvidia/Llama-3_1-Nemotron-Ultra-253B-v1 | 2118 ms | 2118 ms | — | 100.00% | — | 1 |
| meta-llama/Llama-3.3-70B-Instruct | 2186 ms | 2185 ms | — | 100.00% | — | 4 |
| NousResearch/Hermes-4-405B | 2224 ms | 2224 ms | — | 100.00% | — | 2 |
| nvidia/Cosmos3-Super-Reasoner | 2236 ms | 2236 ms | 54 tok/s n=1 | 100.00% | — | 1 |
| Qwen/Qwen3-30B-A3B-Instruct-2507 | 2281 ms | 2280 ms | — | 100.00% | — | 2 |
| openbmb/MiniCPM-V-4_5 | 2410 ms | 2409 ms | 87 tok/s n=1 | 100.00% | — | 1 |
| Qwen/Qwen3-235B-A22B-Instruct-2507 | 2490 ms | 2490 ms | — | 100.00% | — | 3 |
| nvidia/Nemotron-3-Nano-Omni | 2769 ms | 2769 ms | — | 100.00% | — | 1 |
| Qwen/Qwen3-Next-80B-A3B-Thinking | 2792 ms | 2792 ms | — | 100.00% | — | 3 |
| NousResearch/Hermes-4-70B | 3437 ms | 3436 ms | — | 100.00% | — | 1 |
| zai-org/GLM-5.1 | 4248 ms | 4247 ms | — | 100.00% | — | 1 |
Nebius Token Factory performance history · Full provider & model leaderboard.
Provider models
Models served by Nebius Token Factory.
Each row links to pricing, provider, benchmark, and API pages for the model.
| Model | AI IQ | Context | Endpoints | Prompt | Completion | Routes |
|---|---|---|---|---|---|---|
MiniMaxAI/MiniMax-M2.5MiniMax M2.5 |
IQ 105#56 | 204,800 | 2 | $0.315/1M | $1.26/1M | prepaid BYOK |
MiniMaxAI/MiniMax-M3MiniMax-M3 |
IQ 114#33 | 1,048,576 | 2 | $0.315/1M | $1.26/1M | prepaid BYOK |
NousResearch/Hermes-4-405BHermes 4 405B |
— | 131,072 | 2 | $1.05/1M | $3.15/1M | prepaid BYOK |
NousResearch/Hermes-4-70BHermes 4 70B |
— | 131,072 | 2 | $0.1365/1M | $0.42/1M | prepaid BYOK |
Qwen/Qwen2.5-VL-72B-InstructQwen2.5 VL 72B Instruct |
— | 32,768 | 2 | $0.2625/1M | $0.7875/1M | prepaid BYOK |
Qwen/Qwen3-235B-A22B-Instruct-2507Qwen3 235B A22B Instruct 2507 |
— | 131,072 | 2 | $0.21/1M | $0.63/1M | prepaid BYOK |
Qwen/Qwen3-30B-A3B-Instruct-2507Qwen3 30B A3B Instruct 2507 |
— | 131,072 | 2 | $0.105/1M | $0.315/1M | prepaid BYOK |
Qwen/Qwen3-32BQwen3 32B |
— | 131,072 | 2 | $0.105/1M | $0.315/1M | prepaid BYOK |
Qwen/Qwen3-Next-80B-A3B-ThinkingQwen3 Next 80B A3B Thinking |
— | 131,072 | 2 | $0.1575/1M | $1.26/1M | prepaid BYOK |
Qwen/Qwen3.5-397B-A17BQwen3.5 397B A17B |
— | 262,144 | 2 | $0.63/1M | $3.78/1M | prepaid BYOK |
deepseek-ai/DeepSeek-V4-ProDeepSeek V4 Pro |
IQ 115#30 | 1,048,576 | 2 | $1.8375/1M | $3.675/1M | prepaid BYOK |
google/gemma-2-2b-itgemma 2 2b it |
— | 8,192 | 1 | $0.021/1M | $0.063/1M | BYOK |
google/gemma-3-27b-itGoogle: Gemma 3 27B |
— | 262,144 | 2 | $0.105/1M | $0.315/1M | prepaid BYOK |
meta-llama/Llama-3.3-70B-InstructLlama 3.3 70B Instruct |
— | 131,072 | 2 | $0.1365/1M | $0.42/1M | prepaid BYOK |
meta-llama/Meta-Llama-3.1-8B-InstructMeta Llama 3.1 8B Instruct |
— | 128,000 | 1 | $0.021/1M | $0.063/1M | BYOK |
moonshotai/Kimi-K2.6Kimi-K2.6 |
IQ 119#19 | 8,000 | 2 | $0.9975/1M | $4.2/1M | prepaid BYOK |
moonshotai/Kimi-K2.7-CodeKimi-K2.7-Code |
IQ 118#22 | 262,144 | 2 | $0.9975/1M | $4.2/1M | prepaid BYOK |
moonshotai/kimi-k3MoonshotAI: Kimi K3 |
IQ 123#14 | 1,048,576 | 2 | $3.15/1M | $15.75/1M | prepaid BYOK |
nvidia/Cosmos3-Super-ReasonerCosmos3-Super-Reasoner |
— | 8,000 | 2 | $0.105/1M | $0.315/1M | prepaid BYOK |
nvidia/Llama-3_1-Nemotron-Ultra-253B-v1Llama 3_1 Nemotron Ultra 253B v1 |
— | 128,000 | 2 | $0.63/1M | $1.89/1M | prepaid BYOK |
nvidia/NVIDIA-Nemotron-3-Nano-30B-A3BNVIDIA Nemotron 3 Nano 30B A3B |
— | 131,072 | 2 | $0.063/1M | $0.252/1M | prepaid BYOK |
nvidia/Nemotron-3-Nano-OmniNemotron 3 Nano Omni |
— | 131,072 | 2 | $0.063/1M | $0.252/1M | prepaid BYOK |
nvidia/nemotron-3-super-120b-a12bNVIDIA: Nemotron 3 Super |
— | 1,000,000 | 2 | $0.315/1M | $0.945/1M | prepaid BYOK |
nvidia/nemotron-3-ultra-550b-a55bNVIDIA: Nemotron 3 Ultra |
— | 512,288 | 2 | $1.05/1M | $3.15/1M | prepaid BYOK |
openai/gpt-oss-120bOpenAI: gpt-oss-120b |
IQ 105#53 | 131,072 | 2 | $0.1575/1M | $0.63/1M | prepaid BYOK |
openbmb/MiniCPM-V-4_5openbmb/MiniCPM-V-4_5 |
— | 8,000 | 2 | $0.6909/1M | $1.1655/1M | prepaid BYOK |
zai-org/GLM-5.1GLM 5.1 |
IQ 114#32 | 204,800 | 2 | $1.47/1M | $4.62/1M | prepaid BYOK |
zai-org/GLM-5.2GLM-5.2 |
IQ 120#16 | 1,048,576 | 2 | $1.47/1M | $4.62/1M | prepaid BYOK |