OpenAI compatible API · Attested · Public status

Baseten

Explore Baseten models on TrustedRouter with current routes, token pricing, policy sources, privacy posture, regional availability, and API support.

Verify gateway
Onebase URL to migrate
100sof models and routes
0prompt or output logs. Always.

Basetenbaseten

No provider claim

All providers

ProviderBaseten
Provider websitehttps://www.baseten.co/
Models12 public models
Prepaid routes12
BYOK routes12
Zero data retentionnot claimed
Confidential computenot claimed
Provider E2EEnot claimed
Policy noteNo provider-ZDR claim is tracked here. Baseten's inference and security documentation are linked for users who need to review API data handling.
Policy source

Measured performance

57 samples

Continuously sampled across Baseten's routed models: p50 TTFT, effective throughput, and success rate. Effective throughput uses provider-reported output tokens over complete request time. Unsupported route and probe-configuration rows are separated from provider downtime. No prompt or output content stored.

p50 TTFT1767 ms
Effective throughput83 tok/s n=17
Uptime100.00%
Modelp50 TTFTp50 TTFBEffective throughputUptimeConfig excludedAvailability samples
moonshotai/kimi-k2.7-code 631 ms 630 ms 178 tok/s n=2 100.00% 4
thinkingmachines/inkling-small 725 ms 725 ms 223 tok/s n=1 100.00% 3
deepseek/deepseek-v4-pro 1267 ms 1266 ms 59 tok/s n=1 100.00% 7
moonshotai/kimi-k3 1385 ms 1385 ms 33 tok/s n=2 100.00% 5
openai/gpt-oss-120b 1553 ms 1553 ms 82 tok/s n=2 100.00% 6
z-ai/glm-4.7 1573 ms 1573 ms 100.00% 3
deepseek/deepseek-v4-flash-0731 1767 ms 1766 ms 145 tok/s n=1 100.00% 5
moonshotai/kimi-k2.6 2165 ms 2165 ms 92 tok/s n=2 100.00% 6
nvidia/nemotron-3-ultra-550b-a55b 2186 ms 2186 ms 83 tok/s n=2 100.00% 7
thinkingmachines/inkling-1m 2831 ms 2831 ms 71 tok/s n=1 100.00% 3
z-ai/glm-5.2 2918 ms 2917 ms 49 tok/s n=1 100.00% 7
z-ai/glm-5.2-fast 3069 ms 3069 ms 86 tok/s n=2 100.00% 1

Baseten performance history · Full provider & model leaderboard.

Provider models

Models served by Baseten.

Each row links to pricing, provider, benchmark, and API pages for the model.

Model AI IQ Context Endpoints Prompt Completion Routes
deepseek/deepseek-v4-flash-0731
DeepSeek: DeepSeek V4 Flash 0731
1,048,576 2 $0.1365/1M $0.273/1M prepaid BYOK
deepseek/deepseek-v4-pro
DeepSeek: DeepSeek V4 Pro
IQ 115#30 1,048,576 2 $1.827/1M $3.654/1M prepaid BYOK
moonshotai/kimi-k2.6
MoonshotAI: Kimi K2.6
IQ 119#19 262,144 2 $0.9975/1M $4.2/1M prepaid BYOK
moonshotai/kimi-k2.7-code
MoonshotAI: Kimi K2.7 Code
IQ 118#22 262,144 2 $0.9975/1M $4.2/1M prepaid BYOK
moonshotai/kimi-k3
MoonshotAI: Kimi K3
IQ 123#14 1,048,576 2 $3.15/1M $15.75/1M prepaid BYOK
nvidia/nemotron-3-ultra-550b-a55b
NVIDIA: Nemotron 3 Ultra
512,288 2 $0.63/1M $2.52/1M prepaid BYOK
openai/gpt-oss-120b
OpenAI: gpt-oss-120b
IQ 105#53 131,072 2 $0.105/1M $0.525/1M prepaid BYOK
thinkingmachines/inkling-1m
Inkling
1,048,576 2 $1.05/1M $4.2525/1M prepaid BYOK
thinkingmachines/inkling-small
Thinking Machines: Inkling Small
IQ 105#54 524,288 2 $0.525/1M $1.26/1M prepaid BYOK
z-ai/glm-4.7
Z.ai: GLM 4.7
IQ 103#59 204,800 2 $0.63/1M $2.31/1M prepaid BYOK
z-ai/glm-5.2
Z.ai: GLM 5.2
IQ 120#16 1,048,576 2 $1.47/1M $4.62/1M prepaid BYOK
z-ai/glm-5.2-fast
GLM 5.2 Fast on Fireworks
1,048,576 2 $2.205/1M $6.93/1M prepaid BYOK
Workspace access

Sign in

Choose a sign in method to access your TrustedRouter workspace.

By signing in you agree to the terms of service and privacy policy.