OpenAI compatible API · Attested · Public status
DeepInfra
Explore DeepInfra models on TrustedRouter with current routes, token pricing, policy sources, privacy posture, regional availability, and API support.
Onebase URL to migrate
100sof models and routes
0prompt or output logs. Always.
DeepInfradeepinfra
No provider claim| Provider | DeepInfra |
|---|---|
| Provider website | https://deepinfra.com/ |
| Models | 50 public models |
| Prepaid routes | 49 |
| BYOK routes | 50 |
| Zero data retention | no |
| Confidential compute | not claimed |
| Provider E2EE | not claimed |
| Policy note | Tracked as no-store, not strict ZDR. DeepInfra documents memory-only handling and no training for ordinary inference, but reserves the right to log a small portion of requests for debugging or security. Google- and Anthropic-backed routes also inherit those vendors' terms. Policy source |
Measured performance
47 samplesContinuously sampled across DeepInfra's routed models: p50 TTFT, effective throughput, and success rate. Effective throughput uses provider-reported output tokens over complete request time. Unsupported route and probe-configuration rows are separated from provider downtime. No prompt or output content stored.
| p50 TTFT | 1511 ms |
|---|---|
| Effective throughput | 38 tok/s n=17 |
| Uptime | 97.87% |
| Model | p50 TTFT | p50 TTFB | Effective throughput | Uptime | Config excluded | Availability samples |
|---|---|---|---|---|---|---|
| qwen/qwen3-14b | 395 ms | 395 ms | — | 100.00% | — | 1 |
| microsoft/phi-4 | 589 ms | 589 ms | — | 100.00% | — | 1 |
| nousresearch/hermes-3-llama-3.1-405b | 757 ms | 757 ms | — | 100.00% | — | 1 |
| meta-llama/llama-3.1-70b-instruct | 850 ms | 850 ms | — | 100.00% | — | 1 |
| qwen/qwen3-235b-a22b-thinking-2507 | 929 ms | 929 ms | — | 100.00% | — | 3 |
| moonshotai/kimi-k2.5 | 993 ms | 993 ms | — | 100.00% | — | 2 |
| minimax/minimax-m3 | 1022 ms | 1022 ms | 22 tok/s n=1 | 100.00% | — | 3 |
| thinkingmachines/inkling | 1314 ms | 1314 ms | 38 tok/s n=2 | 100.00% | — | 1 |
| z-ai/glm-4.7 | 1327 ms | 1327 ms | — | 100.00% | — | 1 |
| qwen/qwen3.5-397b-a17b | 1358 ms | 1358 ms | — | 100.00% | — | 1 |
| mistralai/mistral-small-24b-instruct-2501 | 1370 ms | 1370 ms | — | 100.00% | — | 1 |
| qwen/qwen3.5-35b-a3b | 1390 ms | 1390 ms | — | 100.00% | — | 1 |
| google/gemma-3-4b-it | 1391 ms | 1391 ms | — | 100.00% | — | 1 |
| google/gemma-4-26b-a4b-it | 1411 ms | 1411 ms | — | 100.00% | — | 3 |
| qwen/qwen3-30b-a3b | 1447 ms | 1446 ms | — | 100.00% | — | 1 |
| z-ai/glm-4.6 | 1511 ms | 1511 ms | — | 100.00% | — | 2 |
| google/gemma-3-27b-it | 1681 ms | 1681 ms | — | 100.00% | — | 1 |
| deepseek/deepseek-v4-pro | 1757 ms | 1757 ms | 97 tok/s n=1 | 100.00% | — | 1 |
| deepseek/deepseek-v3.1-terminus | 1987 ms | 1987 ms | — | 100.00% | — | 1 |
| openai/gpt-oss-20b | 2578 ms | 2578 ms | — | 100.00% | — | 1 |
| nvidia/nemotron-3-nano-30b-a3b | 2625 ms | 2625 ms | — | 100.00% | — | 2 |
| deepseek/deepseek-v4-flash-0731 | 2659 ms | 2659 ms | 73 tok/s n=2 | 100.00% | — | 2 |
| qwen/qwen3-next-80b-a3b-instruct | 2760 ms | 2759 ms | — | 100.00% | — | 1 |
| qwen/qwen3.6-27b | 2866 ms | 2865 ms | — | 100.00% | — | 1 |
| google/gemini-3.1-flash-lite | 2969 ms | 2969 ms | — | 100.00% | — | 2 |
| qwen/qwen3.6-35b-a3b | 3776 ms | 3775 ms | — | 100.00% | — | 4 |
| google/gemma-4-31b-it | 3807 ms | 3807 ms | 14 tok/s n=1 | 100.00% | — | 2 |
| z-ai/glm-4.7-flash | 4655 ms | 4655 ms | — | 100.00% | — | 1 |
| tencent/hy3 | 5790 ms | 5790 ms | 40 tok/s n=2 | 100.00% | — | 1 |
| z-ai/glm-5.2 | 27546 ms | 27546 ms | 34 tok/s n=1 | 100.00% | — | 1 |
| deepseek/deepseek-v3.2 | — | — | — | 100.00% | — | 1 |
| deepseek/deepseek-v4-flash | — | — | 41 tok/s n=1 | — | — | 0 |
| moonshotai/kimi-k2.6 | — | — | 27 tok/s n=2 | — | — | 0 |
| openai/gpt-oss-120b | — | — | 37 tok/s n=2 | 0.00% | — | 1 |
| thinkingmachines/inkling-small | — | — | 74 tok/s n=2 | — | — | 0 |
DeepInfra performance history · Full provider & model leaderboard.
Provider models
Models served by DeepInfra.
Each row links to pricing, provider, benchmark, and API pages for the model.
| Model | AI IQ | Context | Endpoints | Prompt | Completion | Routes |
|---|---|---|---|---|---|---|
Qwen/Qwen3-Embedding-8BQwen3 Embedding 8B |
— | 32,000 | 2 | $0.0105/1M | selected route | prepaid BYOK |
anthropic/claude-opus-5Claude Opus 5 |
IQ 134#3 | 1,000,000 | 1 | $5.25/1M | $26.25/1M | BYOK |
deepseek/deepseek-r1-0528DeepSeek: R1 0528 |
— | 163,840 | 2 | $0.525/1M | $2.2575/1M | prepaid BYOK |
deepseek/deepseek-v3.1-terminusDeepSeek: DeepSeek V3.1 Terminus |
— | 163,840 | 2 | $0.2835/1M | $0.9975/1M | prepaid BYOK |
deepseek/deepseek-v3.2DeepSeek: DeepSeek V3.2 |
IQ 103#57 | 163,840 | 2 | $0.273/1M | $0.399/1M | prepaid BYOK |
deepseek/deepseek-v4-flashDeepSeek: DeepSeek V4 Flash 0423 |
IQ 108#46 | 1,048,576 | 2 | $0.0945/1M | $0.189/1M | prepaid BYOK |
deepseek/deepseek-v4-flash-0731DeepSeek: DeepSeek V4 Flash 0731 |
— | 1,048,576 | 2 | $0.0945/1M | $0.189/1M | prepaid BYOK |
deepseek/deepseek-v4-proDeepSeek: DeepSeek V4 Pro |
IQ 115#30 | 1,048,576 | 2 | $1.365/1M | $2.73/1M | prepaid BYOK |
google/gemini-2.5-flashGoogle: Gemini 2.5 Flash |
— | 1,048,576 | 2 | $0.315/1M | $2.625/1M | prepaid BYOK |
google/gemini-2.5-proGoogle: Gemini 2.5 Pro |
IQ 103#58 | 1,048,576 | 2 | $1.3125/1M | $10.5/1M | prepaid BYOK |
google/gemini-3.1-flash-liteGoogle: Gemini 3.1 Flash Lite |
IQ 101#65 | 1,048,576 | 2 | $0.2625/1M | $1.575/1M | prepaid BYOK |
google/gemma-3-12b-itGoogle: Gemma 3 12B |
— | 131,072 | 2 | $0.0525/1M | $0.1575/1M | prepaid BYOK |
google/gemma-3-27b-itGoogle: Gemma 3 27B |
— | 262,144 | 2 | $0.084/1M | $0.168/1M | prepaid BYOK |
google/gemma-3-4b-itGoogle: Gemma 3 4B |
— | 131,072 | 2 | $0.0525/1M | $0.105/1M | prepaid BYOK |
google/gemma-4-26b-a4b-itGoogle: Gemma 4 26B A4B |
IQ 96#81 | 262,144 | 2 | $0.0735/1M | $0.357/1M | prepaid BYOK |
google/gemma-4-31b-itGoogle: Gemma 4 31B |
IQ 101#66 | 262,144 | 2 | $0.1365/1M | $0.399/1M | prepaid BYOK |
gryphe/mythomax-l2-13bMythoMax 13B |
— | 8,192 | 2 | $0.42/1M | $0.42/1M | prepaid BYOK |
meta-llama/llama-3.1-70b-instructMeta: Llama 3.1 70B Instruct |
— | 131,072 | 2 | $0.42/1M | $0.42/1M | prepaid BYOK |
meta-llama/llama-guard-4-12bMeta: Llama Guard 4 12B |
— | 1,048,576 | 2 | $0.189/1M | $0.189/1M | prepaid BYOK |
microsoft/phi-4Microsoft: Phi 4 |
— | 16,384 | 2 | $0.0735/1M | $0.147/1M | prepaid BYOK |
minimax/minimax-m2.7MiniMax: MiniMax M2.7 |
IQ 109#44 | 204,800 | 2 | $0.2625/1M | $1.05/1M | prepaid BYOK |
minimax/minimax-m3MiniMax: MiniMax M3 |
IQ 114#33 | 1,048,576 | 2 | $0.315/1M | $1.26/1M | prepaid BYOK |
mistralai/mistral-small-24b-instruct-2501Mistral: Mistral Small 3 |
— | 32,768 | 2 | $0.0525/1M | $0.084/1M | prepaid BYOK |
moonshotai/kimi-k2.5MoonshotAI: Kimi K2.5 |
IQ 111#39 | 262,144 | 2 | $0.4725/1M | $2.3625/1M | prepaid BYOK |
moonshotai/kimi-k2.6MoonshotAI: Kimi K2.6 |
IQ 119#19 | 262,144 | 2 | $0.7875/1M | $3.675/1M | prepaid BYOK |
nousresearch/hermes-3-llama-3.1-405bNous: Hermes 3 405B Instruct |
— | 131,072 | 2 | $1.05/1M | $1.05/1M | prepaid BYOK |
nvidia/nemotron-3-nano-30b-a3bNVIDIA: Nemotron 3 Nano 30B A3B |
— | 262,144 | 2 | $0.0525/1M | $0.21/1M | prepaid BYOK |
openai/gpt-oss-120bOpenAI: gpt-oss-120b |
IQ 105#53 | 131,072 | 2 | $0.03885/1M | $0.1785/1M | prepaid BYOK |
openai/gpt-oss-20bOpenAI: gpt-oss-20b |
IQ 100#70 | 131,072 | 2 | $0.0315/1M | $0.147/1M | prepaid BYOK |
qwen/qwen3-14bQwen: Qwen3 14B |
— | 131,072 | 2 | $0.126/1M | $0.252/1M | prepaid BYOK |
qwen/qwen3-235b-a22b-thinking-2507Qwen: Qwen3 235B A22B Thinking 2507 |
— | 262,144 | 2 | $0.2415/1M | $2.415/1M | prepaid BYOK |
qwen/qwen3-30b-a3bQwen: Qwen3 30B A3B |
— | 131,072 | 2 | $0.126/1M | $0.525/1M | prepaid BYOK |
qwen/qwen3-next-80b-a3b-instructQwen: Qwen3 Next 80B A3B Instruct |
— | 262,144 | 2 | $0.0945/1M | $1.155/1M | prepaid BYOK |
qwen/qwen3-vl-235b-a22b-instructQwen: Qwen3 VL 235B A22B Instruct |
— | 262,144 | 2 | $0.21/1M | $0.924/1M | prepaid BYOK |
qwen/qwen3-vl-30b-a3b-instructQwen: Qwen3 VL 30B A3B Instruct |
— | 262,144 | 2 | $0.1575/1M | $0.63/1M | prepaid BYOK |
qwen/qwen3.5-122b-a10bQwen: Qwen3.5-122B-A10B |
— | 262,144 | 2 | $0.3045/1M | $2.52/1M | prepaid BYOK |
qwen/qwen3.5-27bQwen: Qwen3.5-27B |
— | 262,144 | 2 | $0.273/1M | $2.73/1M | prepaid BYOK |
qwen/qwen3.5-35b-a3bQwen: Qwen3.5-35B-A3B |
— | 262,144 | 2 | $0.147/1M | $1.05/1M | prepaid BYOK |
qwen/qwen3.5-397b-a17bQwen: Qwen3.5 397B A17B |
— | 262,144 | 2 | $0.4725/1M | $3.15/1M | prepaid BYOK |
qwen/qwen3.5-9bQwen: Qwen3.5-9B |
IQ 93#90 | 262,144 | 2 | $0.105/1M | $0.1575/1M | prepaid BYOK |
qwen/qwen3.6-27bQwen: Qwen3.6 27B |
IQ 111#40 | 262,144 | 2 | $0.336/1M | $3.36/1M | prepaid BYOK |
qwen/qwen3.6-35b-a3bQwen: Qwen3.6 35B A3B |
IQ 100#71 | 262,144 | 2 | $0.105/1M | $0.9975/1M | prepaid BYOK |
tencent/hy3Tencent: Hy3 |
IQ 103#60 | 262,144 | 2 | $0.147/1M | $0.609/1M | prepaid BYOK |
thinkingmachines/inklingThinking Machines: Inkling |
IQ 107#48 | 1,048,576 | 2 | $0.9975/1M | $4.2525/1M | prepaid BYOK |
thinkingmachines/inkling-smallThinking Machines: Inkling Small |
IQ 105#54 | 524,288 | 2 | $0.4725/1M | $1.26/1M | prepaid BYOK |
z-ai/glm-4.6Z.ai: GLM 4.6 |
— | 204,800 | 2 | $0.525/1M | $2.1/1M | prepaid BYOK |
z-ai/glm-4.7Z.ai: GLM 4.7 |
IQ 103#59 | 204,800 | 2 | $0.42/1M | $1.8375/1M | prepaid BYOK |
z-ai/glm-4.7-flashZ.ai: GLM 4.7 Flash |
— | 202,752 | 2 | $0.063/1M | $0.42/1M | prepaid BYOK |
z-ai/glm-5Z.ai: GLM 5 |
IQ 105#52 | 204,800 | 2 | $0.63/1M | $2.184/1M | prepaid BYOK |
z-ai/glm-5.2Z.ai: GLM 5.2 |
IQ 120#16 | 1,048,576 | 2 | $1.26/1M | $4.41/1M | prepaid BYOK |