OpenAI compatible API · Attested · Public status
Cloudflare Workers AI performance
Measured TTFT, TTFB, effective throughput, uptime, and sampled model routes for Cloudflare Workers AI.
Onebase URL to migrate
100sof models and routes
0prompt or output logs. Always.
cloudflare-workers-ai
54 samples
Continuously sampled provider performance. TrustedRouter reports unsupported route and probe-configuration rows separately from provider downtime. Prompt and output content is not stored.
| p50 TTFT | 1462 ms |
|---|---|
| p95 TTFT | 4089 ms |
| p50 TTFB | 1647 ms |
| Effective throughput | 37 tok/s n=2 |
| Uptime | 98.15% |
Measured model routes
| Model | p50 TTFT | p50 TTFB | Effective throughput | Uptime | Config excluded | Availability samples |
|---|---|---|---|---|---|---|
| meta-llama/llama-3.3-70b-instruct-fp8-fast | 816 ms | 816 ms | — | 100.00% | — | 3 |
| z-ai/glm-4.7-flash | 1055 ms | 1055 ms | — | 100.00% | — | 2 |
| meta-llama/llama-3.2-3b-instruct | 1100 ms | 1100 ms | — | 100.00% | — | 2 |
| openai/gpt-oss-120b | 1240 ms | 1240 ms | 48 tok/s n=1 | 100.00% | — | 3 |
| qwen/qwen2.5-coder-32b-instruct | 1260 ms | 1260 ms | — | 100.00% | — | 1 |
| openai/gpt-oss-20b | 1286 ms | 1286 ms | — | 100.00% | — | 5 |
| aisingapore/gemma-sea-lion-v4-27b-it | 1456 ms | 1456 ms | — | 100.00% | — | 5 |
| meta-llama/llama-3.2-1b-instruct | 1459 ms | 1459 ms | — | 100.00% | — | 4 |
| qwen/qwq-32b | 1462 ms | 1461 ms | — | 100.00% | — | 3 |
| ibm-granite/granite-4.0-h-micro | 1477 ms | 1476 ms | — | 100.00% | — | 2 |
| meta-llama/llama-4-scout-17b-16e-instruct | 1502 ms | 1502 ms | — | 100.00% | — | 3 |
| google/gemma-4-26b-a4b-it | 1756 ms | 1756 ms | — | 100.00% | — | 1 |
| qwen/qwen3-30b-a3b-fp8 | 1903 ms | 1903 ms | — | 100.00% | — | 5 |
| meta-llama/llama-3.1-8b-instruct-fp8 | 2145 ms | 2145 ms | — | 100.00% | — | 3 |
| moonshotai/kimi-k3 | 2269 ms | 2269 ms | 25 tok/s n=1 | 100.00% | — | 2 |
| deepseek/deepseek-r1-distill-qwen-32b | 2715 ms | 2715 ms | — | 100.00% | — | 1 |
| mistralai/mistral-small-3.1-24b-instruct | 2862 ms | 2862 ms | — | 100.00% | — | 5 |
| nvidia/nemotron-3-120b-a12b | 2787 ms | 2786 ms | — | 75.00% | — | 4 |