OpenAI compatible API · Attested · Public status

Cloudflare Workers AI

Cloudflare Workers AI models on TrustedRouter with prices, routes, policy notes, and source links.

Verify gateway
Onebase URL to migrate
100sof models and routes
0prompt or output logs. Always.

cloudflare-workers-ai

No provider claim

All providers

ProviderCloudflare Workers AI
Models23 public models
Prepaid routes23
BYOK routes0
Zero data retentionnot claimed
Confidential computenot claimed
Provider E2EEnot claimed
Policy noteNo provider-ZDR claim is tracked here. Cloudflare's Workers AI documentation is linked for model and data-handling review.
Policy source

Measured performance

53 samples

Continuously sampled across Cloudflare Workers AI's routed models: p50 TTFT, effective throughput, and success rate. Effective throughput uses provider-reported output tokens over complete request time. Unsupported route and probe-configuration rows are separated from provider downtime. No prompt or output content stored.

p50 TTFT1462 ms
Effective throughput37 tok/s n=2
Uptime98.11%
Modelp50 TTFTp50 TTFBEffective throughputUptimeConfig excludedAvailability samples
meta-llama/llama-3.3-70b-instruct-fp8-fast 816 ms 816 ms 100.00% 3
openai/gpt-oss-20b 830 ms 829 ms 100.00% 4
z-ai/glm-4.7-flash 1055 ms 1055 ms 100.00% 2
meta-llama/llama-3.2-3b-instruct 1100 ms 1100 ms 100.00% 2
openai/gpt-oss-120b 1240 ms 1240 ms 48 tok/s n=1 100.00% 3
qwen/qwen2.5-coder-32b-instruct 1260 ms 1260 ms 100.00% 1
aisingapore/gemma-sea-lion-v4-27b-it 1456 ms 1456 ms 100.00% 5
meta-llama/llama-3.2-1b-instruct 1459 ms 1459 ms 100.00% 4
qwen/qwq-32b 1462 ms 1461 ms 100.00% 3
ibm-granite/granite-4.0-h-micro 1477 ms 1476 ms 100.00% 2
meta-llama/llama-4-scout-17b-16e-instruct 1502 ms 1502 ms 100.00% 3
google/gemma-4-26b-a4b-it 1756 ms 1756 ms 100.00% 1
qwen/qwen3-30b-a3b-fp8 1903 ms 1903 ms 100.00% 5
meta-llama/llama-3.1-8b-instruct-fp8 2145 ms 2145 ms 100.00% 3
moonshotai/kimi-k3 2269 ms 2269 ms 25 tok/s n=1 100.00% 2
deepseek/deepseek-r1-distill-qwen-32b 2715 ms 2715 ms 100.00% 1
mistralai/mistral-small-3.1-24b-instruct 2862 ms 2862 ms 100.00% 5
nvidia/nemotron-3-120b-a12b 2787 ms 2786 ms 75.00% 4

Cloudflare Workers AI performance history · Full provider & model leaderboard.

Provider models

Models served by Cloudflare Workers AI.

Each row links to pricing, provider, benchmark, and API pages for the model.

Model AI IQ Context Endpoints Prompt Completion Routes
aisingapore/gemma-sea-lion-v4-27b-it
@cf/aisingapore/gemma-sea-lion-v4-27b-it
128,000 1 $0.36855/1M $0.58275/1M prepaid
deepseek/deepseek-r1-distill-qwen-32b
@cf/deepseek-ai/deepseek-r1-distill-qwen-32b
80,000 1 $0.52185/1M $5.12505/1M prepaid
google/gemma-4-26b-a4b-it
Google: Gemma 4 26B A4B
IQ 96#79 262,144 1 $0.105/1M $0.315/1M prepaid
ibm-granite/granite-4.0-h-micro
@cf/ibm-granite/granite-4.0-h-micro
131,000 1 $0.01785/1M $0.1176/1M prepaid
meta-llama/llama-3.1-8b-instruct-fp8
@cf/meta/llama-3.1-8b-instruct-fp8
32,000 1 $0.1596/1M $0.30135/1M prepaid
meta-llama/llama-3.2-11b-vision-instruct
@cf/meta/llama-3.2-11b-vision-instruct
128,000 1 $0.050925/1M $0.7098/1M prepaid
meta-llama/llama-3.2-1b-instruct
@cf/meta/llama-3.2-1b-instruct
60,000 1 $0.02835/1M $0.21105/1M prepaid
meta-llama/llama-3.2-3b-instruct
Meta: Llama 3.2 3B Instruct
131,072 1 $0.053445/1M $0.35175/1M prepaid
meta-llama/llama-3.3-70b-instruct-fp8-fast
@cf/meta/llama-3.3-70b-instruct-fp8-fast
24,000 1 $0.30765/1M $2.36565/1M prepaid
meta-llama/llama-4-scout-17b-16e-instruct
Llama 4 Scout Instruct
131,072 1 $0.2835/1M $0.8925/1M prepaid
meta-llama/llama-guard-3-8b
@cf/meta/llama-guard-3-8b
131,072 1 $0.5082/1M $0.0315/1M prepaid
mistralai/mistral-small-3.1-24b-instruct
@cf/mistralai/mistral-small-3.1-24b-instruct
128,000 1 $0.36855/1M $0.58275/1M prepaid
moonshotai/kimi-k2.6
MoonshotAI: Kimi K2.6
IQ 119#19 262,144 1 $0.9975/1M $4.2/1M prepaid
moonshotai/kimi-k2.7-code
MoonshotAI: Kimi K2.7 Code
IQ 118#22 262,144 1 $0.9975/1M $4.2/1M prepaid
moonshotai/kimi-k3
MoonshotAI: Kimi K3
IQ 122#15 1,048,576 1 $3.15/1M $15.75/1M prepaid
nvidia/nemotron-3-120b-a12b
@cf/nvidia/nemotron-3-120b-a12b
256,000 1 $0.525/1M $1.575/1M prepaid
openai/gpt-oss-120b
OpenAI: gpt-oss-120b
IQ 105#52 131,072 1 $0.3675/1M $0.7875/1M prepaid
openai/gpt-oss-20b
OpenAI: gpt-oss-20b
IQ 100#69 131,072 1 $0.21/1M $0.315/1M prepaid
qwen/qwen2.5-coder-32b-instruct
@cf/qwen/qwen2.5-coder-32b-instruct
32,768 1 $0.693/1M $1.05/1M prepaid
qwen/qwen3-30b-a3b-fp8
Qwen3 30B A3B
40,960 1 $0.053445/1M $0.35175/1M prepaid
qwen/qwq-32b
@cf/qwen/qwq-32b
24,000 1 $0.693/1M $1.05/1M prepaid
z-ai/glm-4.7-flash
Z.ai: GLM 4.7 Flash
202,752 1 $0.063525/1M $0.42/1M prepaid
z-ai/glm-5.2
Z.ai: GLM 5.2
IQ 120#16 1,048,576 1 $1.47/1M $4.62/1M prepaid
Workspace access

Sign in

Choose a sign in method. New email and OAuth accounts include $0.10 in starter credit; wallet-only accounts start at $0.

By signing in you agree to the terms of service and privacy policy.