OpenAI compatible API · Attested · Public status
DeepInfra
DeepInfra models on TrustedRouter with prices, routes, policy notes, and source links.
Onebase URL to migrate
100sof models and routes
0prompt or output logs. Always.
deepinfra
No provider claim
| Provider | DeepInfra |
|---|---|
| Models | 50 public models |
| Prepaid routes | 49 |
| BYOK routes | 50 |
| Zero data retention | no |
| Confidential compute | not claimed |
| Provider E2EE | not claimed |
| Policy note | Tracked as no-store, not strict ZDR. DeepInfra documents memory-only handling and no training for ordinary inference, but reserves the right to log a small portion of requests for debugging or security. Google- and Anthropic-backed routes also inherit those vendors' terms. Policy source |
Measured performance
54 samplesContinuously sampled across DeepInfra's routed models: p50 TTFT, effective throughput, and success rate. Effective throughput uses provider-reported output tokens over complete request time. Unsupported route and probe-configuration rows are separated from provider downtime. No prompt or output content stored.
| p50 TTFT | 1874 ms |
|---|---|
| Effective throughput | 51 tok/s n=14 |
| Uptime | 98.15% |
| Model | p50 TTFT | p50 TTFB | Effective throughput | Uptime | Config excluded | Availability samples |
|---|---|---|---|---|---|---|
| qwen/qwen3.5-27b | 369 ms | 368 ms | — | 100.00% | — | 1 |
| qwen/qwen3.5-9b | 476 ms | 476 ms | — | 100.00% | — | 2 |
| deepseek/deepseek-r1-0528 | 517 ms | 516 ms | — | 100.00% | — | 2 |
| qwen/qwen3-30b-a3b | 548 ms | 548 ms | — | 100.00% | — | 1 |
| qwen/qwen3.5-35b-a3b | 698 ms | 698 ms | — | 100.00% | — | 2 |
| qwen/qwen3-235b-a22b-thinking-2507 | 764 ms | 764 ms | — | 100.00% | — | 1 |
| deepseek/deepseek-v4-flash | 882 ms | 882 ms | 37 tok/s n=1 | 100.00% | — | 1 |
| openai/gpt-oss-120b | 923 ms | 923 ms | 28 tok/s n=1 | 100.00% | — | 2 |
| google/gemma-3-12b-it | 938 ms | 938 ms | — | 100.00% | — | 2 |
| tencent/hy3 | 1187 ms | 1187 ms | 53 tok/s n=1 | 100.00% | — | 2 |
| qwen/qwen3.5-397b-a17b | 1338 ms | 1338 ms | — | 100.00% | — | 2 |
| nousresearch/hermes-3-llama-3.1-405b | 1498 ms | 1498 ms | — | 100.00% | — | 2 |
| qwen/qwen3-14b | 1595 ms | 1594 ms | — | 100.00% | — | 4 |
| google/gemma-3-27b-it | 1645 ms | 1645 ms | — | 100.00% | — | 1 |
| meta-llama/llama-3.1-70b-instruct | 1653 ms | 1653 ms | — | 100.00% | — | 1 |
| z-ai/glm-4.7 | 1874 ms | 1873 ms | — | 100.00% | — | 1 |
| google/gemma-3-4b-it | 1944 ms | 1943 ms | — | 100.00% | — | 2 |
| moonshotai/kimi-k2.5 | 2274 ms | 2274 ms | — | 100.00% | — | 4 |
| google/gemini-2.5-flash | 2584 ms | 2583 ms | — | 100.00% | — | 1 |
| openai/gpt-oss-20b | 2819 ms | 2819 ms | — | 100.00% | — | 3 |
| google/gemini-3.1-flash-lite | 2872 ms | 2872 ms | — | 100.00% | — | 2 |
| gryphe/mythomax-l2-13b | 3012 ms | 3012 ms | — | 100.00% | — | 1 |
| qwen/qwen3.6-35b-a3b | 3189 ms | 3189 ms | — | 100.00% | — | 1 |
| deepseek/deepseek-v3.1-terminus | 3346 ms | 3345 ms | — | 100.00% | — | 3 |
| z-ai/glm-5 | 3512 ms | 3512 ms | — | 100.00% | — | 3 |
| thinkingmachines/inkling-small | 3641 ms | 3641 ms | 63 tok/s n=1 | 100.00% | — | 1 |
| google/gemini-2.5-pro | 4535 ms | 4535 ms | — | 100.00% | — | 1 |
| z-ai/glm-4.7-flash | 7660 ms | 7659 ms | — | 100.00% | — | 1 |
| minimax/minimax-m2.7 | 8778 ms | 8778 ms | — | 100.00% | — | 3 |
| deepseek/deepseek-v4-flash-0731 | — | — | 63 tok/s n=1 | — | — | 0 |
| deepseek/deepseek-v4-pro | — | — | 73 tok/s n=2 | — | — | 0 |
| google/gemma-4-31b-it | — | — | 67 tok/s n=1 | — | — | 0 |
| minimax/minimax-m3 | — | — | 19 tok/s n=2 | — | — | 0 |
| moonshotai/kimi-k2.6 | — | — | 37 tok/s n=1 | — | — | 0 |
| qwen/qwen3.5-122b-a10b | — | — | — | 0.00% | — | 1 |
| thinkingmachines/inkling | — | — | 67 tok/s n=1 | — | — | 0 |
| z-ai/glm-5.2 | — | — | 49 tok/s n=2 | — | — | 0 |
DeepInfra performance history · Full provider & model leaderboard.
Provider models
Models served by DeepInfra.
Each row links to pricing, provider, benchmark, and API pages for the model.
| Model | AI IQ | Context | Endpoints | Prompt | Completion | Routes |
|---|---|---|---|---|---|---|
Qwen/Qwen3-Embedding-8BQwen3 Embedding 8B |
— | 32,000 | 2 | $0.0105/1M | selected route | prepaid BYOK |
anthropic/claude-opus-5Claude Opus 5 |
IQ 134#3 | 1,000,000 | 1 | $5.25/1M | $26.25/1M | BYOK |
deepseek/deepseek-r1-0528DeepSeek: R1 0528 |
— | 163,840 | 2 | $0.525/1M | $2.2575/1M | prepaid BYOK |
deepseek/deepseek-v3.1-terminusDeepSeek: DeepSeek V3.1 Terminus |
— | 163,840 | 2 | $0.2835/1M | $0.9975/1M | prepaid BYOK |
deepseek/deepseek-v3.2DeepSeek: DeepSeek V3.2 |
IQ 103#56 | 163,840 | 2 | $0.273/1M | $0.399/1M | prepaid BYOK |
deepseek/deepseek-v4-flashDeepSeek: DeepSeek V4 Flash |
IQ 108#46 | 1,048,576 | 2 | $0.0945/1M | $0.189/1M | prepaid BYOK |
deepseek/deepseek-v4-flash-0731DeepSeek: DeepSeek V4 Flash 0731 |
— | 1,048,576 | 2 | $0.0945/1M | $0.189/1M | prepaid BYOK |
deepseek/deepseek-v4-proDeepSeek: DeepSeek V4 Pro |
IQ 115#30 | 1,048,576 | 2 | $1.365/1M | $2.73/1M | prepaid BYOK |
google/gemini-2.5-flashGoogle: Gemini 2.5 Flash |
— | 1,048,576 | 2 | $0.315/1M | $2.625/1M | prepaid BYOK |
google/gemini-2.5-proGoogle: Gemini 2.5 Pro |
IQ 103#57 | 1,048,576 | 2 | $1.3125/1M | $10.5/1M | prepaid BYOK |
google/gemini-3.1-flash-liteGoogle: Gemini 3.1 Flash Lite |
IQ 101#64 | 1,048,576 | 2 | $0.2625/1M | $1.575/1M | prepaid BYOK |
google/gemma-3-12b-itGoogle: Gemma 3 12B |
— | 131,072 | 2 | $0.0525/1M | $0.1575/1M | prepaid BYOK |
google/gemma-3-27b-itGoogle: Gemma 3 27B |
— | 262,144 | 2 | $0.084/1M | $0.168/1M | prepaid BYOK |
google/gemma-3-4b-itGoogle: Gemma 3 4B |
— | 131,072 | 2 | $0.0525/1M | $0.105/1M | prepaid BYOK |
google/gemma-4-26b-a4b-itGoogle: Gemma 4 26B A4B |
IQ 96#79 | 262,144 | 2 | $0.0735/1M | $0.357/1M | prepaid BYOK |
google/gemma-4-31b-itGoogle: Gemma 4 31B |
IQ 101#65 | 262,144 | 2 | $0.1365/1M | $0.399/1M | prepaid BYOK |
gryphe/mythomax-l2-13bMythoMax 13B |
— | 8,192 | 2 | $0.42/1M | $0.42/1M | prepaid BYOK |
meta-llama/llama-3.1-70b-instructMeta: Llama 3.1 70B Instruct |
— | 131,072 | 2 | $0.42/1M | $0.42/1M | prepaid BYOK |
meta-llama/llama-guard-4-12bMeta: Llama Guard 4 12B |
— | 1,048,576 | 2 | $0.189/1M | $0.189/1M | prepaid BYOK |
microsoft/phi-4Microsoft: Phi 4 |
— | 16,384 | 2 | $0.0735/1M | $0.147/1M | prepaid BYOK |
minimax/minimax-m2.7MiniMax: MiniMax M2.7 |
IQ 109#44 | 204,800 | 2 | $0.2625/1M | $1.05/1M | prepaid BYOK |
minimax/minimax-m3MiniMax: MiniMax M3 |
IQ 114#33 | 1,048,576 | 2 | $0.315/1M | $1.26/1M | prepaid BYOK |
mistralai/mistral-small-24b-instruct-2501Mistral: Mistral Small 3 |
— | 32,768 | 2 | $0.0525/1M | $0.084/1M | prepaid BYOK |
moonshotai/kimi-k2.5MoonshotAI: Kimi K2.5 |
IQ 111#39 | 262,144 | 2 | $0.4725/1M | $2.3625/1M | prepaid BYOK |
moonshotai/kimi-k2.6MoonshotAI: Kimi K2.6 |
IQ 119#19 | 262,144 | 2 | $0.7875/1M | $3.675/1M | prepaid BYOK |
nousresearch/hermes-3-llama-3.1-405bNous: Hermes 3 405B Instruct |
— | 131,072 | 2 | $1.05/1M | $1.05/1M | prepaid BYOK |
nvidia/nemotron-3-nano-30b-a3bNVIDIA: Nemotron 3 Nano 30B A3B |
— | 262,144 | 2 | $0.0525/1M | $0.21/1M | prepaid BYOK |
openai/gpt-oss-120bOpenAI: gpt-oss-120b |
IQ 105#52 | 131,072 | 2 | $0.03885/1M | $0.1785/1M | prepaid BYOK |
openai/gpt-oss-20bOpenAI: gpt-oss-20b |
IQ 100#69 | 131,072 | 2 | $0.0315/1M | $0.147/1M | prepaid BYOK |
qwen/qwen3-14bQwen: Qwen3 14B |
— | 131,072 | 2 | $0.126/1M | $0.252/1M | prepaid BYOK |
qwen/qwen3-235b-a22b-thinking-2507Qwen: Qwen3 235B A22B Thinking 2507 |
— | 262,144 | 2 | $0.2415/1M | $2.415/1M | prepaid BYOK |
qwen/qwen3-30b-a3bQwen: Qwen3 30B A3B |
— | 131,072 | 2 | $0.126/1M | $0.525/1M | prepaid BYOK |
qwen/qwen3-next-80b-a3b-instructQwen: Qwen3 Next 80B A3B Instruct |
— | 262,144 | 2 | $0.0945/1M | $1.155/1M | prepaid BYOK |
qwen/qwen3-vl-235b-a22b-instructQwen: Qwen3 VL 235B A22B Instruct |
— | 262,144 | 2 | $0.21/1M | $0.924/1M | prepaid BYOK |
qwen/qwen3-vl-30b-a3b-instructQwen: Qwen3 VL 30B A3B Instruct |
— | 262,144 | 2 | $0.1575/1M | $0.63/1M | prepaid BYOK |
qwen/qwen3.5-122b-a10bQwen: Qwen3.5-122B-A10B |
— | 262,144 | 2 | $0.3045/1M | $2.52/1M | prepaid BYOK |
qwen/qwen3.5-27bQwen: Qwen3.5-27B |
— | 262,144 | 2 | $0.273/1M | $2.73/1M | prepaid BYOK |
qwen/qwen3.5-35b-a3bQwen: Qwen3.5-35B-A3B |
— | 262,144 | 2 | $0.147/1M | $1.05/1M | prepaid BYOK |
qwen/qwen3.5-397b-a17bQwen: Qwen3.5 397B A17B |
— | 262,144 | 2 | $0.4725/1M | $3.15/1M | prepaid BYOK |
qwen/qwen3.5-9bQwen: Qwen3.5-9B |
IQ 93#89 | 262,144 | 2 | $0.105/1M | $0.1575/1M | prepaid BYOK |
qwen/qwen3.6-27bQwen: Qwen3.6 27B |
IQ 111#40 | 262,144 | 2 | $0.336/1M | $3.36/1M | prepaid BYOK |
qwen/qwen3.6-35b-a3bQwen: Qwen3.6 35B A3B |
IQ 100#70 | 262,144 | 2 | $0.105/1M | $0.9975/1M | prepaid BYOK |
tencent/hy3Tencent: Hy3 |
IQ 103#59 | 262,144 | 2 | $0.147/1M | $0.609/1M | prepaid BYOK |
thinkingmachines/inklingThinking Machines: Inkling |
IQ 104#55 | 1,048,576 | 2 | $1.05/1M | $4.2525/1M | prepaid BYOK |
thinkingmachines/inkling-smallThinking Machines: Inkling Small |
— | 524,288 | 2 | $0.525/1M | $1.26/1M | prepaid BYOK |
z-ai/glm-4.6Z.ai: GLM 4.6 |
— | 204,800 | 2 | $0.525/1M | $2.1/1M | prepaid BYOK |
z-ai/glm-4.7Z.ai: GLM 4.7 |
IQ 103#58 | 204,800 | 2 | $0.42/1M | $1.8375/1M | prepaid BYOK |
z-ai/glm-4.7-flashZ.ai: GLM 4.7 Flash |
— | 202,752 | 2 | $0.063/1M | $0.42/1M | prepaid BYOK |
z-ai/glm-5Z.ai: GLM 5 |
IQ 105#51 | 204,800 | 2 | $0.63/1M | $2.184/1M | prepaid BYOK |
z-ai/glm-5.2Z.ai: GLM 5.2 |
IQ 120#16 | 1,048,576 | 2 | $1.26/1M | $4.41/1M | prepaid BYOK |