OpenAI compatible API · Attested · Public status
Cloudflare Workers AI
Explore Cloudflare Workers AI models on TrustedRouter with current routes, token pricing, policy sources, privacy posture, regional availability, and API support.
Onebase URL to migrate
100sof models and routes
0prompt or output logs. Always.
Cloudflare Workers AIcloudflare-workers-ai
No provider claim| Provider | Cloudflare Workers AI |
|---|---|
| Provider website | https://www.cloudflare.com/developer-platform/products/workers-ai/ |
| Models | 23 public models |
| Prepaid routes | 23 |
| BYOK routes | 0 |
| Zero data retention | not claimed |
| Confidential compute | not claimed |
| Provider E2EE | not claimed |
| Policy note | No provider-ZDR claim is tracked here. Cloudflare's Workers AI documentation is linked for model and data-handling review. Policy source |
Measured performance
33 samplesContinuously sampled across Cloudflare Workers AI's routed models: p50 TTFT, effective throughput, and success rate. Effective throughput uses provider-reported output tokens over complete request time. Unsupported route and probe-configuration rows are separated from provider downtime. No prompt or output content stored.
| p50 TTFT | 1441 ms |
|---|---|
| Effective throughput | 33 tok/s n=4 |
| Uptime | 100.00% |
| Model | p50 TTFT | p50 TTFB | Effective throughput | Uptime | Config excluded | Availability samples |
|---|---|---|---|---|---|---|
| meta-llama/llama-4-scout-17b-16e-instruct | 598 ms | 598 ms | — | 100.00% | — | 1 |
| meta-llama/llama-3.1-8b-instruct-fp8 | 713 ms | 713 ms | — | 100.00% | — | 1 |
| openai/gpt-oss-120b | 766 ms | 766 ms | 42 tok/s n=2 | 100.00% | — | 2 |
| meta-llama/llama-3.3-70b-instruct-fp8-fast | 1247 ms | 1247 ms | — | 100.00% | — | 1 |
| openai/gpt-oss-20b | 1305 ms | 1305 ms | — | 100.00% | — | 5 |
| meta-llama/llama-3.2-3b-instruct | 1402 ms | 1402 ms | — | 100.00% | — | 3 |
| ibm-granite/granite-4.0-h-micro | 1407 ms | 1406 ms | — | 100.00% | — | 3 |
| nvidia/nemotron-3-120b-a12b | 1441 ms | 1441 ms | — | 100.00% | — | 3 |
| qwen/qwen3-30b-a3b-fp8 | 1466 ms | 1465 ms | — | 100.00% | — | 1 |
| z-ai/glm-4.7-flash | 1490 ms | 1490 ms | — | 100.00% | — | 2 |
| google/gemma-4-26b-a4b-it | 1576 ms | 1576 ms | — | 100.00% | 2 probe_config_error |
4 |
| qwen/qwen2.5-coder-32b-instruct | 2448 ms | 2447 ms | — | 100.00% | — | 4 |
| moonshotai/kimi-k3 | 2923 ms | 2923 ms | 25 tok/s n=2 | 100.00% | — | 2 |
| mistralai/mistral-small-3.1-24b-instruct | 7361 ms | 7361 ms | — | 100.00% | — | 1 |
Cloudflare Workers AI performance history · Full provider & model leaderboard.
Provider models
Models served by Cloudflare Workers AI.
Each row links to pricing, provider, benchmark, and API pages for the model.
| Model | AI IQ | Context | Endpoints | Prompt | Completion | Routes |
|---|---|---|---|---|---|---|
aisingapore/gemma-sea-lion-v4-27b-it@cf/aisingapore/gemma-sea-lion-v4-27b-it |
— | 128,000 | 1 | $0.36855/1M | $0.58275/1M | prepaid |
deepseek/deepseek-r1-distill-qwen-32b@cf/deepseek-ai/deepseek-r1-distill-qwen-32b |
— | 80,000 | 1 | $0.52185/1M | $5.12505/1M | prepaid |
google/gemma-4-26b-a4b-itGoogle: Gemma 4 26B A4B |
IQ 96#81 | 262,144 | 1 | $0.105/1M | $0.315/1M | prepaid |
ibm-granite/granite-4.0-h-micro@cf/ibm-granite/granite-4.0-h-micro |
— | 131,000 | 1 | $0.01785/1M | $0.1176/1M | prepaid |
meta-llama/llama-3.1-8b-instruct-fp8@cf/meta/llama-3.1-8b-instruct-fp8 |
— | 32,000 | 1 | $0.1596/1M | $0.30135/1M | prepaid |
meta-llama/llama-3.2-11b-vision-instruct@cf/meta/llama-3.2-11b-vision-instruct |
— | 128,000 | 1 | $0.050925/1M | $0.7098/1M | prepaid |
meta-llama/llama-3.2-1b-instruct@cf/meta/llama-3.2-1b-instruct |
— | 60,000 | 1 | $0.02835/1M | $0.21105/1M | prepaid |
meta-llama/llama-3.2-3b-instructMeta: Llama 3.2 3B Instruct |
— | 131,072 | 1 | $0.053445/1M | $0.35175/1M | prepaid |
meta-llama/llama-3.3-70b-instruct-fp8-fast@cf/meta/llama-3.3-70b-instruct-fp8-fast |
— | 24,000 | 1 | $0.30765/1M | $2.36565/1M | prepaid |
meta-llama/llama-4-scout-17b-16e-instructLlama 4 Scout Instruct |
— | 131,072 | 1 | $0.2835/1M | $0.8925/1M | prepaid |
meta-llama/llama-guard-3-8b@cf/meta/llama-guard-3-8b |
— | 131,072 | 1 | $0.5082/1M | $0.0315/1M | prepaid |
mistralai/mistral-small-3.1-24b-instruct@cf/mistralai/mistral-small-3.1-24b-instruct |
— | 128,000 | 1 | $0.36855/1M | $0.58275/1M | prepaid |
moonshotai/kimi-k2.6MoonshotAI: Kimi K2.6 |
IQ 119#19 | 262,144 | 1 | $0.9975/1M | $4.2/1M | prepaid |
moonshotai/kimi-k2.7-codeMoonshotAI: Kimi K2.7 Code |
IQ 118#22 | 262,144 | 1 | $0.9975/1M | $4.2/1M | prepaid |
moonshotai/kimi-k3MoonshotAI: Kimi K3 |
IQ 123#14 | 1,048,576 | 1 | $3.15/1M | $15.75/1M | prepaid |
nvidia/nemotron-3-120b-a12b@cf/nvidia/nemotron-3-120b-a12b |
— | 256,000 | 1 | $0.525/1M | $1.575/1M | prepaid |
openai/gpt-oss-120bOpenAI: gpt-oss-120b |
IQ 105#53 | 131,072 | 1 | $0.3675/1M | $0.7875/1M | prepaid |
openai/gpt-oss-20bOpenAI: gpt-oss-20b |
IQ 100#70 | 131,072 | 1 | $0.21/1M | $0.315/1M | prepaid |
qwen/qwen2.5-coder-32b-instruct@cf/qwen/qwen2.5-coder-32b-instruct |
— | 32,768 | 1 | $0.693/1M | $1.05/1M | prepaid |
qwen/qwen3-30b-a3b-fp8Qwen3 30B A3B |
— | 40,960 | 1 | $0.053445/1M | $0.35175/1M | prepaid |
qwen/qwq-32b@cf/qwen/qwq-32b |
— | 24,000 | 1 | $0.693/1M | $1.05/1M | prepaid |
z-ai/glm-4.7-flashZ.ai: GLM 4.7 Flash |
— | 202,752 | 1 | $0.063525/1M | $0.42/1M | prepaid |
z-ai/glm-5.2Z.ai: GLM 5.2 |
IQ 120#16 | 1,048,576 | 1 | $1.47/1M | $4.62/1M | prepaid |