OpenAI compatible API · Attested · Public status
Cloudflare Workers AI
Cloudflare Workers AI models on TrustedRouter with prices, routes, policy notes, and source links.
Onebase URL to migrate
100sof models and routes
0prompt or output logs. Always.
cloudflare-workers-ai
No provider claim
| Provider | Cloudflare Workers AI |
|---|---|
| Models | 23 public models |
| Prepaid routes | 23 |
| BYOK routes | 0 |
| Zero data retention | not claimed |
| Confidential compute | not claimed |
| Provider E2EE | not claimed |
| Policy note | No provider-ZDR claim is tracked here. Cloudflare's Workers AI documentation is linked for model and data-handling review. Policy source |
Measured performance
53 samplesContinuously sampled across Cloudflare Workers AI's routed models: p50 TTFT, effective throughput, and success rate. Effective throughput uses provider-reported output tokens over complete request time. Unsupported route and probe-configuration rows are separated from provider downtime. No prompt or output content stored.
| p50 TTFT | 1462 ms |
|---|---|
| Effective throughput | 37 tok/s n=2 |
| Uptime | 98.11% |
| Model | p50 TTFT | p50 TTFB | Effective throughput | Uptime | Config excluded | Availability samples |
|---|---|---|---|---|---|---|
| meta-llama/llama-3.3-70b-instruct-fp8-fast | 816 ms | 816 ms | — | 100.00% | — | 3 |
| openai/gpt-oss-20b | 830 ms | 829 ms | — | 100.00% | — | 4 |
| z-ai/glm-4.7-flash | 1055 ms | 1055 ms | — | 100.00% | — | 2 |
| meta-llama/llama-3.2-3b-instruct | 1100 ms | 1100 ms | — | 100.00% | — | 2 |
| openai/gpt-oss-120b | 1240 ms | 1240 ms | 48 tok/s n=1 | 100.00% | — | 3 |
| qwen/qwen2.5-coder-32b-instruct | 1260 ms | 1260 ms | — | 100.00% | — | 1 |
| aisingapore/gemma-sea-lion-v4-27b-it | 1456 ms | 1456 ms | — | 100.00% | — | 5 |
| meta-llama/llama-3.2-1b-instruct | 1459 ms | 1459 ms | — | 100.00% | — | 4 |
| qwen/qwq-32b | 1462 ms | 1461 ms | — | 100.00% | — | 3 |
| ibm-granite/granite-4.0-h-micro | 1477 ms | 1476 ms | — | 100.00% | — | 2 |
| meta-llama/llama-4-scout-17b-16e-instruct | 1502 ms | 1502 ms | — | 100.00% | — | 3 |
| google/gemma-4-26b-a4b-it | 1756 ms | 1756 ms | — | 100.00% | — | 1 |
| qwen/qwen3-30b-a3b-fp8 | 1903 ms | 1903 ms | — | 100.00% | — | 5 |
| meta-llama/llama-3.1-8b-instruct-fp8 | 2145 ms | 2145 ms | — | 100.00% | — | 3 |
| moonshotai/kimi-k3 | 2269 ms | 2269 ms | 25 tok/s n=1 | 100.00% | — | 2 |
| deepseek/deepseek-r1-distill-qwen-32b | 2715 ms | 2715 ms | — | 100.00% | — | 1 |
| mistralai/mistral-small-3.1-24b-instruct | 2862 ms | 2862 ms | — | 100.00% | — | 5 |
| nvidia/nemotron-3-120b-a12b | 2787 ms | 2786 ms | — | 75.00% | — | 4 |
Cloudflare Workers AI performance history · Full provider & model leaderboard.
Provider models
Models served by Cloudflare Workers AI.
Each row links to pricing, provider, benchmark, and API pages for the model.
| Model | AI IQ | Context | Endpoints | Prompt | Completion | Routes |
|---|---|---|---|---|---|---|
aisingapore/gemma-sea-lion-v4-27b-it@cf/aisingapore/gemma-sea-lion-v4-27b-it |
— | 128,000 | 1 | $0.36855/1M | $0.58275/1M | prepaid |
deepseek/deepseek-r1-distill-qwen-32b@cf/deepseek-ai/deepseek-r1-distill-qwen-32b |
— | 80,000 | 1 | $0.52185/1M | $5.12505/1M | prepaid |
google/gemma-4-26b-a4b-itGoogle: Gemma 4 26B A4B |
IQ 96#79 | 262,144 | 1 | $0.105/1M | $0.315/1M | prepaid |
ibm-granite/granite-4.0-h-micro@cf/ibm-granite/granite-4.0-h-micro |
— | 131,000 | 1 | $0.01785/1M | $0.1176/1M | prepaid |
meta-llama/llama-3.1-8b-instruct-fp8@cf/meta/llama-3.1-8b-instruct-fp8 |
— | 32,000 | 1 | $0.1596/1M | $0.30135/1M | prepaid |
meta-llama/llama-3.2-11b-vision-instruct@cf/meta/llama-3.2-11b-vision-instruct |
— | 128,000 | 1 | $0.050925/1M | $0.7098/1M | prepaid |
meta-llama/llama-3.2-1b-instruct@cf/meta/llama-3.2-1b-instruct |
— | 60,000 | 1 | $0.02835/1M | $0.21105/1M | prepaid |
meta-llama/llama-3.2-3b-instructMeta: Llama 3.2 3B Instruct |
— | 131,072 | 1 | $0.053445/1M | $0.35175/1M | prepaid |
meta-llama/llama-3.3-70b-instruct-fp8-fast@cf/meta/llama-3.3-70b-instruct-fp8-fast |
— | 24,000 | 1 | $0.30765/1M | $2.36565/1M | prepaid |
meta-llama/llama-4-scout-17b-16e-instructLlama 4 Scout Instruct |
— | 131,072 | 1 | $0.2835/1M | $0.8925/1M | prepaid |
meta-llama/llama-guard-3-8b@cf/meta/llama-guard-3-8b |
— | 131,072 | 1 | $0.5082/1M | $0.0315/1M | prepaid |
mistralai/mistral-small-3.1-24b-instruct@cf/mistralai/mistral-small-3.1-24b-instruct |
— | 128,000 | 1 | $0.36855/1M | $0.58275/1M | prepaid |
moonshotai/kimi-k2.6MoonshotAI: Kimi K2.6 |
IQ 119#19 | 262,144 | 1 | $0.9975/1M | $4.2/1M | prepaid |
moonshotai/kimi-k2.7-codeMoonshotAI: Kimi K2.7 Code |
IQ 118#22 | 262,144 | 1 | $0.9975/1M | $4.2/1M | prepaid |
moonshotai/kimi-k3MoonshotAI: Kimi K3 |
IQ 122#15 | 1,048,576 | 1 | $3.15/1M | $15.75/1M | prepaid |
nvidia/nemotron-3-120b-a12b@cf/nvidia/nemotron-3-120b-a12b |
— | 256,000 | 1 | $0.525/1M | $1.575/1M | prepaid |
openai/gpt-oss-120bOpenAI: gpt-oss-120b |
IQ 105#52 | 131,072 | 1 | $0.3675/1M | $0.7875/1M | prepaid |
openai/gpt-oss-20bOpenAI: gpt-oss-20b |
IQ 100#69 | 131,072 | 1 | $0.21/1M | $0.315/1M | prepaid |
qwen/qwen2.5-coder-32b-instruct@cf/qwen/qwen2.5-coder-32b-instruct |
— | 32,768 | 1 | $0.693/1M | $1.05/1M | prepaid |
qwen/qwen3-30b-a3b-fp8Qwen3 30B A3B |
— | 40,960 | 1 | $0.053445/1M | $0.35175/1M | prepaid |
qwen/qwq-32b@cf/qwen/qwq-32b |
— | 24,000 | 1 | $0.693/1M | $1.05/1M | prepaid |
z-ai/glm-4.7-flashZ.ai: GLM 4.7 Flash |
— | 202,752 | 1 | $0.063525/1M | $0.42/1M | prepaid |
z-ai/glm-5.2Z.ai: GLM 5.2 |
IQ 120#16 | 1,048,576 | 1 | $1.47/1M | $4.62/1M | prepaid |