OpenAI compatible API · Attested · Public status

Novita AI performance

Review measured TTFT, TTFB, effective throughput, uptime, and sampled model routes for Novita AI on TrustedRouter using metadata-only production probes.

Verify gateway
Onebase URL to migrate
100sof models and routes
0prompt or output logs. Always.

Novita AInovita

47 samples

Provider overview

Continuously sampled provider performance. TrustedRouter reports unsupported route and probe-configuration rows separately from provider downtime. Prompt and output content is not stored.

p50 TTFT1859 ms
p95 TTFT4327 ms
p50 TTFB1858 ms
Effective throughput62 tok/s n=20
Uptime89.36%

Measured model routes

Modelp50 TTFTp50 TTFBEffective throughputUptimeConfig excludedAvailability samples
meta-llama/llama-3.3-70b-instruct 833 ms 833 ms 100.00% 3
google/gemma-4-31b-it 949 ms 949 ms 12 tok/s n=1 100.00% 2
deepseek/deepseek-ocr-2 966 ms 966 ms 100.00% 1
moonshotai/kimi-k2-instruct 1070 ms 1070 ms 100.00% 1
deepseek/deepseek-v4-flash-0731 1222 ms 1222 ms 86 tok/s n=2 100.00% 2
zai-org/glm-5 1256 ms 1256 ms 100.00% 1
zai-org/glm-4.5v 1338 ms 1338 ms 100.00% 1
deepseek/deepseek-v4-pro 1432 ms 1432 ms 60 tok/s n=1 100.00% 1
mistralai/mistral-nemo 1521 ms 1520 ms 100.00% 1
minimax/minimax-m2.1 1562 ms 1562 ms 100.00% 2
xiaomimimo/mimo-v2.5-pro 1661 ms 1661 ms 100.00% 1
sao10k/l3-8b-lunaris 1708 ms 1707 ms 100.00% 1
zai-org/glm-4.5-air 1714 ms 1714 ms 100.00% 1
zai-org/glm-4.7-flash 1787 ms 1787 ms 100.00% 1
minimax/minimax-m2 1832 ms 1831 ms 100.00% 1
qwen/qwen-mt-plus 1859 ms 1858 ms 100.00% 2
minimax/minimax-m2.7 1872 ms 1872 ms 100.00% 1
deepseek/deepseek-v3.1 2073 ms 2073 ms 100.00% 1
deepseek/deepseek-v3-0324 2081 ms 2081 ms 100.00% 2
minimax/minimax-m2.5-highspeed 2340 ms 2340 ms 100.00% 1
qwen/qwen3-coder-30b-a3b-instruct 2455 ms 2455 ms 100.00% 1
qwen/qwen3-vl-30b-a3b-instruct 2665 ms 2665 ms 100.00% 1
qwen/qwen3.7-max 2757 ms 2757 ms 100.00% 3
deepseek/deepseek-v3.2 2803 ms 2803 ms 100.00% 1
moonshotai/kimi-k2.6 2869 ms 2869 ms 35 tok/s n=2 100.00% 1
qwen/qwen3.8-max 2920 ms 2920 ms 68 tok/s n=1 100.00% 2
zai-org/glm-4.6v 3000 ms 3000 ms 100.00% 1
deepseek/deepseek-r1-0528 3141 ms 3141 ms 100.00% 1
kwaipilot/kat-coder-pro 3224 ms 3224 ms 100.00% 1
qwen/qwen3-omni-30b-a3b-instruct 4449 ms 4449 ms 100.00% 1
google/gemma-3-27b-it 6693 ms 6693 ms 100.00% 1
qwen/qwen3.5-122b-a10b 4327 ms 4327 ms 50.00% 2
deepseek/deepseek-v4-flash 77 tok/s n=1 0
google/gemma-3-12b-it 0.00% 2
inclusionai/ling-3.0-flash 141 tok/s n=2 0
mindai/macaron-v1-tall 0.00% 2
mindai/macaron-v1-venti 13 tok/s n=2 0
minimax/minimax-m3 56 tok/s n=1 0
moonshotai/kimi-k2.7-code 34 tok/s n=1 0
moonshotai/kimi-k3 14 tok/s n=1 0
openai/gpt-oss-120b 64 tok/s n=2 0
tencent/hy3 93 tok/s n=2 0
z-ai/glm-5.2 22 tok/s n=1 0
Workspace access

Sign in

Choose a sign in method to access your TrustedRouter workspace.

By signing in you agree to the terms of service and privacy policy.