DeepInfra
Explore DeepInfra models on TrustedRouter with current routes, token pricing, policy sources, privacy posture, regional availability, and API support.
DeepInfradeepinfra
No provider claimThese privacy labels describe DeepInfra, the upstream model provider. ZDR is a retention policy; verified confidential inference additionally requires attested provider compute and end-to-end encryption.
| Provider | DeepInfra |
|---|---|
| Routing status | Active |
| Provider website | https://deepinfra.com/ |
| Models | 96 public models |
| Credits routes | 96 |
| Zero data retention | no |
| Verified confidential inference | Not verified |
| Policy note | Tracked as no-store, not strict ZDR. DeepInfra documents memory-only handling and no training for ordinary inference, but reserves the right to log a small portion of requests for debugging or security. Google- and Anthropic-backed routes also inherit those vendors' terms. Policy source |
Measured performance
496 samplesContinuously sampled across DeepInfra's routed models: p50 TTFT, effective throughput, and success rate. Effective throughput uses provider-reported output tokens over complete request time. Unsupported route and probe-configuration rows are separated from provider downtime. No prompt or output content stored.
| p50 TTFT | 3142 ms |
|---|---|
| Effective throughput | 88 tok/s n=2 |
| Uptime | 99.19% |
| Model | p50 TTFT | Effective throughput | Uptime | Config excluded | Availability samples |
|---|---|---|---|---|---|
| deepseek/deepseek-v4-flash-0731 | 3142 ms | — | 99.52% | — | 419 |
| meta-models/muse-glimmer-30b | 860 ms | — | 100.00% | — | 1 |
| qwen/qwen3.7-max | 3762 ms | — | 100.00% | — | 1 |
| qwen/qwen3.5-27b | 3995 ms | — | 100.00% | — | 1 |
| nvidia/nemotron-content-safety-3.5 | 4240 ms | — | 100.00% | — | 1 |
| stepfun-ai/step-3.7-flash | 4614 ms | — | 100.00% | — | 1 |
| meta-llama/meta-llama-3.1-8b-instruct-turbo | — | — | 100.00% | — | 1 |
| mistralai/mistral-nemo-instruct-2407 | — | — | 100.00% | — | 1 |
| xiaomi/mimo-v2.5-pro | — | 23 tok/s n=1 | 100.00% | — | 68 |
| deepseek/deepseek-v4-pro-0423 | — | 152 tok/s n=1 | — | — | 0 |
| openai/gpt-oss-120b-ultra | — | — | 0.00% | — | 1 |
| z-ai/glm-5.3 | — | — | 0.00% | — | 1 |
DeepInfra performance history · Full provider & model leaderboard.
Models served by DeepInfra.
Each row links to pricing, provider, benchmark, and API pages for the model.
| Model | AI IQ | Context | Input | Cached input | Output |
|---|---|---|---|---|---|
Qwen/Qwen3-Embedding-8BQwen3 Embedding 8B |
— | 32,000 | $0.01055/1M | Not published | selected route |
bytedance/seed-1.8ByteDance/Seed-1.8 |
— | 256,000 | $0.26375/1M | $0.05275/1M | $2.11/1M |
bytedance/seed-2.0-codeByteDance/Seed-2.0-code |
— | 256,000 | $0.5275/1M | $0.1055/1M | $3.165/1M |
bytedance/seed-2.0-miniByteDance/Seed-2.0-mini |
— | 256,000 | $0.1055/1M | $0.0211/1M | $0.422/1M |
bytedance/seed-2.0-proByteDance/Seed-2.0-pro |
— | 256,000 | $0.5275/1M | $0.1055/1M | $3.165/1M |
deepseek/deepseek-r1-0528DeepSeek: R1 0528 |
— | 163,840 | $0.5275/1M | $0.36925/1M | $2.26825/1M |
deepseek/deepseek-v3deepseek-ai/DeepSeek-V3 |
— | 163,840 | $0.3376/1M | Not published | $0.93895/1M |
deepseek/deepseek-v3-0324deepseek-ai/DeepSeek-V3-0324 |
— | 163,840 | $0.2532/1M | $0.142425/1M | $0.9495/1M |
deepseek/deepseek-v3.1DeepSeek V3.1 |
IQ 95#105 | 131,072 | $0.26375/1M | $0.13715/1M | $1.00225/1M |
deepseek/deepseek-v3.2DeepSeek: DeepSeek V3.2 |
IQ 103#77 | 163,840 | $0.2743/1M | $0.13715/1M | $0.4009/1M |
deepseek/deepseek-v4-flashDeepSeek: DeepSeek V4 Flash 0423 |
IQ 116#37 | 1,048,576 | $0.09495/1M | $0.01899/1M | $0.1899/1M |
deepseek/deepseek-v4-flash-0731DeepSeek: DeepSeek V4 Flash 0731 |
— | 1,048,576 | $0.0633/1M | $0.015825/1M | $0.1899/1M |
deepseek/deepseek-v4-flash-vision-expDeepSeek: DeepSeek V4 Flash Vision Exp |
IQ 115#42 | 1,048,576 | $0.4642/1M | $0.01477/1M | $1.3926/1M |
deepseek/deepseek-v4-proDeepSeek: DeepSeek V4 Pro 0423 |
IQ 114#47 | 1,048,576 | $1.3715/1M | $0.1055/1M | $2.743/1M |
deepseek/deepseek-v4-pro-0423DeepSeek V4 Pro 0423 |
— | 1,048,576 | $1.3715/1M | $0.1055/1M | $2.743/1M |
deepseek/deepseek-v4.1-flashDeepSeek: DeepSeek V4.1 Flash |
IQ 116#38 | 1,048,576 | $0.211/1M | $0.01/1M | $0.633/1M |
google/gemini-2.5-flashGoogle: Gemini 2.5 Flash |
— | 1,048,576 | $0.3165/1M | Not published | $2.6375/1M |
google/gemini-2.5-proGoogle: Gemini 2.5 Pro |
IQ 100#88 | 1,048,576 | $1.31875/1M | Not published | $10.55/1M |
google/gemini-3.1-flash-liteGoogle: Gemini 3.1 Flash Lite |
IQ 103#78 | 1,048,576 | $0.26375/1M | Not published | $1.5825/1M |
google/gemini-3.1-progoogle/gemini-3.1-pro |
IQ 129#8 | 1,000,000 | $2.11/1M | Not published | $12.66/1M |
google/gemini-3.7-flashGoogle: Gemini 3.7 Flash |
IQ 124#15 | 1,048,576 | $0.79125/1M | Not published | $3.95625/1M |
google/gemma-3-12b-itGoogle: Gemma 3 12B |
— | 131,072 | $0.05275/1M | Not published | $0.15825/1M |
google/gemma-3-27b-itGoogle: Gemma 3 27B |
— | 131,072 | $0.0844/1M | Not published | $0.1688/1M |
google/gemma-3-4b-itGoogle: Gemma 3 4B |
— | 131,072 | $0.05275/1M | Not published | $0.1055/1M |
google/gemma-4-26b-a4b-itGoogle: Gemma 4 26B A4B |
IQ 100#89 | 262,144 | $0.07385/1M | Not published | $0.3587/1M |
google/gemma-4-31b-itGoogle: Gemma 4 31B |
IQ 103#80 | 262,144 | $0.13715/1M | Not published | $0.4009/1M |
google/gemma-4-31b-it-turbogoogle/gemma-4-31B-it-turbo |
— | 262,144 | $0.09495/1M | $0.05275/1M | $0.3587/1M |
google/gemma-4-31b-it-ultragoogle/gemma-4-31B-it-Ultra |
— | 131,072 | $0.28485/1M | Not published | $0.8018/1M |
google/gemma-4-e4b-itgoogle/gemma-4-E4B-it |
IQ 84#129 | 131,072 | $0.0211/1M | Not published | $0.1055/1M |
gryphe/mythomax-l2-13bMythoMax 13B |
— | 4,096 | $0.422/1M | Not published | $0.422/1M |
ibm-granite/granite-4.2-30bibm-granite/granite-4.2-30b |
— | 131,072 | $0.1688/1M | $0.0422/1M | $0.68575/1M |
ibm-granite/granite-4.2-3bibm-granite/granite-4.2-3b |
— | 131,072 | $0.03165/1M | $0.01/1M | $0.1266/1M |
ibm-granite/granite-4.2-8bIBM: Granite 4.2 8B |
— | 131,072 | $0.0633/1M | $0.015825/1M | $0.26375/1M |
inclusionai/ling-3.0-flashinclusionAI: Ling 3.0 Flash |
IQ 94#109 | 131,072 | $0.0633/1M | $0.01266/1M | $0.1899/1M |
inclusionai/ling-3.0-flash-fininclusionAI: Ling 3.0 Flash Fin |
— | 262,144 | $0.0633/1M | $0.01266/1M | $0.1899/1M |
inclusionai/ling-3.0-flash-vlinclusionAI: Ling 3.0 Flash VL |
— | 131,072 | $0.0633/1M | $0.01266/1M | $0.1899/1M |
meta-llama/llama-3.1-70b-instructMeta: Llama 3.1 70B Instruct |
— | 131,072 | $0.422/1M | Not published | $0.422/1M |
meta-llama/llama-3.3-70b-instruct-turbometa-llama/Llama-3.3-70B-Instruct-Turbo |
— | 131,072 | $0.1055/1M | Not published | $0.3376/1M |
meta-llama/llama-4-maverick-17b-128e-instruct-fp8Llama 4 Maverick Instruct |
— | 1,048,576 | $0.211/1M | Not published | $0.844/1M |
meta-llama/llama-4-scout-17b-16e-instructLlama 4 Scout Instruct |
— | 131,072 | $0.1055/1M | Not published | $0.3165/1M |
meta-llama/llama-guard-4-12bMeta: Llama Guard 4 12B |
— | 163,840 | $0.1899/1M | Not published | $0.1899/1M |
meta-llama/meta-llama-3.1-8b-instruct-turbometa-llama/Meta-Llama-3.1-8B-Instruct-Turbo |
— | 131,072 | $0.0211/1M | Not published | $0.0422/1M |
meta-models/muse-glimmer-30bMuse Glimmer 30B on Fireworks |
— | 131,072 | $0.3165/1M | $0.0422/1M | $1.266/1M |
microsoft/phi-4Microsoft: Phi 4 |
— | 16,384 | $0.07385/1M | Not published | $0.1477/1M |
minimax/minimax-m2.7-turboMiniMaxAI/MiniMax-M2.7-Turbo |
— | 196,608 | $0.4009/1M | $0.07385/1M | $1.7935/1M |
minimax/minimax-m3MiniMax: MiniMax M3 |
IQ 115#45 | 524,288 | $0.2954/1M | $0.05908/1M | $1.1605/1M |
mistralai/mistral-nemo-instruct-2407mistralai/Mistral-Nemo-Instruct-2407 |
— | 131,072 | $0.020045/1M | Not published | $0.03165/1M |
mistralai/mistral-small-24b-instruct-2501Mistral: Mistral Small 3 |
— | 32,768 | $0.05275/1M | Not published | $0.0844/1M |
mistralai/mistral-small-3.2-24b-instruct-2506mistralai/Mistral-Small-3.2-24B-Instruct-2506 |
— | 128,000 | $0.079125/1M | Not published | $0.211/1M |
moonshotai/kimi-k2.6MoonshotAI: Kimi K2.6 |
IQ 118#35 | 262,144 | $0.79125/1M | $0.15825/1M | $3.6925/1M |
moonshotai/kimi-k3MoonshotAI: Kimi K3 |
IQ 121#27 | 1,048,576 | $3.00675/1M | $0.300675/1M | $15.03375/1M |
nousresearch/hermes-3-llama-3.1-405bNous: Hermes 3 405B Instruct |
— | 131,072 | $1.055/1M | Not published | $1.055/1M |
nvidia/nemotron-3-nano-30b-a3bNVIDIA: Nemotron 3 Nano 30B A3B |
— | 262,144 | $0.05275/1M | $0.026375/1M | $0.211/1M |
nvidia/nemotron-3-super-120b-a12bNVIDIA: Nemotron 3 Super |
— | 262,144 | $0.089675/1M | Not published | $0.422/1M |
nvidia/nemotron-3-ultra-550b-a55bNVIDIA: Nemotron 3 Ultra |
— | 262,144 | $0.5275/1M | $0.1055/1M | $2.321/1M |
nvidia/nemotron-3.5-lightningNVIDIA: Nemotron 3.5 Lightning |
— | 262,144 | $0.0844/1M | $0.0422/1M | $0.211/1M |
nvidia/nemotron-content-safety-3.5nvidia/Nemotron-Content-Safety-3.5 |
— | 131,072 | $0.211/1M | Not published | $0.211/1M |
openai/gpt-oss-120bOpenAI: gpt-oss-120b |
IQ 98#99 | 131,072 | $0.039035/1M | Not published | $0.17935/1M |
openai/gpt-oss-120b-turboopenai/gpt-oss-120b-Turbo |
— | 131,072 | $0.15825/1M | Not published | $0.633/1M |
openai/gpt-oss-120b-ultraopenai/gpt-oss-120b-Ultra |
— | 131,072 | $0.211/1M | Not published | $1.00225/1M |
openai/gpt-oss-20bOpenAI: gpt-oss-20b |
IQ 93#113 | 131,072 | $0.03165/1M | Not published | $0.1477/1M |
qwen/qwen2.5-72b-instructQwen/Qwen2.5-72B-Instruct |
— | 32,768 | $0.3798/1M | Not published | $0.422/1M |
qwen/qwen3-14bQwen: Qwen3 14B |
— | 40,960 | $0.1266/1M | Not published | $0.2532/1M |
qwen/qwen3-235b-a22b-instruct-2507Qwen3 235B A22B Instruct 2507 |
— | 131,072 | $0.09495/1M | Not published | $0.58025/1M |
qwen/qwen3-30b-a3bQwen: Qwen3 30B A3B |
— | 40,960 | $0.1266/1M | Not published | $0.5275/1M |
qwen/qwen3-coder-480b-a35b-instruct-turboQwen/Qwen3-Coder-480B-A35B-Instruct-Turbo |
— | 262,144 | $0.3165/1M | $0.1055/1M | $1.055/1M |
qwen/qwen3-maxQwen3 Max |
— | 262,144 | $1.266/1M | $0.2532/1M | $6.33/1M |
qwen/qwen3-max-thinkingQwen/Qwen3-Max-Thinking |
— | 256,000 | $1.266/1M | $0.2532/1M | $6.33/1M |
qwen/qwen3-next-80b-a3b-instructQwen: Qwen3 Next 80B A3B Instruct |
— | 262,144 | $0.09495/1M | Not published | $1.1605/1M |
qwen/qwen3-vl-235b-a22b-instructQwen: Qwen3 VL 235B A22B Instruct |
— | 262,144 | $0.211/1M | $0.11605/1M | $0.9284/1M |
qwen/qwen3-vl-30b-a3b-instructQwen: Qwen3 VL 30B A3B Instruct |
— | 262,144 | $0.15825/1M | Not published | $0.633/1M |
qwen/qwen3.5-122b-a10bQwen: Qwen3.5-122B-A10B |
— | 262,144 | $0.30595/1M | Not published | $2.532/1M |
qwen/qwen3.5-27bQwen: Qwen3.5-27B |
— | 262,144 | $0.2743/1M | Not published | $2.743/1M |
qwen/qwen3.5-35b-a3bQwen: Qwen3.5-35B-A3B |
— | 262,144 | $0.1477/1M | $0.05275/1M | $1.055/1M |
qwen/qwen3.5-397b-a17bQwen: Qwen3.5 397B A17B |
— | 262,144 | $0.47475/1M | $0.2321/1M | $3.165/1M |
qwen/qwen3.5-9bQwen: Qwen3.5-9B |
IQ 95#108 | 262,144 | $0.1055/1M | Not published | $0.15825/1M |
qwen/qwen3.6-27bQwen: Qwen3.6 27B |
IQ 108#68 | 262,144 | $0.3376/1M | Not published | $3.376/1M |
qwen/qwen3.6-35b-a3bQwen: Qwen3.6 35B A3B |
IQ 100#94 | 262,144 | $0.1055/1M | Not published | $1.00225/1M |
qwen/qwen3.7-maxQwen3.7 Max |
IQ 119#31 | 1,000,000 | $2.6375/1M | $0.5275/1M | $7.9125/1M |
qwen/qwen3.8-2.4t-a95bQwen: Qwen3.8 2.4T A95B |
IQ 122#24 | 262,144 | $2.11/1M | $0.211/1M | $6.33/1M |
qwen/qwen3.8-27bQwen: Qwen3.8 27B |
IQ 110#60 | 262,144 | $0.211/1M | $0.05275/1M | $2.6375/1M |
qwen/qwen3.8-flashQwen3.8 Flash |
IQ 115#46 | 1,000,000 | $0.119215/1M | $0.014876/1M | $0.40301/1M |
qwen/qwen3.8-maxQwen3.8 Max |
IQ 122#25 | 1,000,000 | $1.74075/1M | $0.21733/1M | $5.223305/1M |
sao10k/l3-8b-lunaris-v1-turboSao10K/L3-8B-Lunaris-v1-Turbo |
— | 8,192 | $0.0422/1M | Not published | $0.05275/1M |
sao10k/l3.1-70b-euryale-v2.2Sao10K/L3.1-70B-Euryale-v2.2 |
— | 131,072 | $0.89675/1M | Not published | $0.89675/1M |
stepfun-ai/step-3.7-flashstepfun-ai/Step-3.7-Flash |
IQ 102#84 | 262,144 | $0.211/1M | $0.0422/1M | $1.21325/1M |
tencent/hy3Tencent: Hy3 |
IQ 100#90 | 262,144 | $0.1477/1M | $0.036925/1M | $0.6119/1M |
thinkingmachines/inklingThinking Machines: Inkling |
IQ 107#70 | 524,288 | $1.00225/1M | $0.1688/1M | $4.27275/1M |
thinkingmachines/inkling-smallThinking Machines: Inkling Small |
IQ 107#71 | 524,288 | $0.47475/1M | $0.1055/1M | $1.266/1M |
xiaomi/mimo-v2.5Xiaomi: MiMo-V2.5 |
IQ 108#67 | 1,050,000 | $0.1477/1M | $0.01/1M | $0.2954/1M |
xiaomi/mimo-v2.5-proXiaomi: MiMo-V2.5-Pro |
IQ 112#55 | 1,050,000 | $1.055/1M | $0.211/1M | $3.165/1M |
z-ai/glm-4.6Z.ai: GLM 4.6 |
— | 202,752 | $0.5275/1M | $0.1055/1M | $2.11/1M |
z-ai/glm-4.7Z.ai: GLM 4.7 |
IQ 103#81 | 202,752 | $0.422/1M | $0.0844/1M | $1.84625/1M |
z-ai/glm-5.2Z.ai: GLM 5.2 |
IQ 120#28 | 1,048,576 | $0.79125/1M | $0.1477/1M | $2.532/1M |
z-ai/glm-5.3Z.ai: GLM 5.3 |
IQ 123#18 | 1,048,576 | $1.266/1M | $0.211/1M | $4.22/1M |
z-ai/glm-5.3-flashZ.ai: GLM 5.3 Flash |
IQ 116#40 | 1,048,576 | $0.15825/1M | $0.03165/1M | $0.5275/1M |
Questions
Does DeepInfra have zero data retention?
TrustedRouter does not currently mark DeepInfra as provider-level zero data retention. Use trustedrouter/zdr or provider.min_privacy=zdr to select a different eligible route, and review the linked policy source for changes.
Is DeepInfra end-to-end encrypted?
TrustedRouter does not currently mark DeepInfra as end-to-end encrypted at the provider boundary. The TrustedRouter gateway is still attested, but the selected provider normally receives the request in order to run the model. Use trustedrouter/e2e for the stronger route requirement.
Which DeepInfra models are available through TrustedRouter?
This page currently lists 96 public DeepInfra models, with live pricing, route count, context length, measured performance when available, and links to each model's provider and benchmark pages.