OpenAI compatible API · Attested · Public status

DeepInfra

Explore DeepInfra models on TrustedRouter with current routes, token pricing, policy sources, privacy posture, regional availability, and API support.

Verify gateway
Onebase URL to migrate
100sof models and routes
0prompt or output logs. Always.

DeepInfradeepinfra

No provider claim

All providers

These privacy labels describe DeepInfra, the upstream model provider. ZDR is a retention policy; verified confidential inference additionally requires attested provider compute and end-to-end encryption.

ProviderDeepInfra
Routing statusActive
Provider websitehttps://deepinfra.com/
Models96 public models
Credits routes96
Zero data retentionno
Verified confidential inferenceNot verified
Policy noteTracked as no-store, not strict ZDR. DeepInfra documents memory-only handling and no training for ordinary inference, but reserves the right to log a small portion of requests for debugging or security. Google- and Anthropic-backed routes also inherit those vendors' terms.
Policy source

Measured performance

496 samples

Continuously sampled across DeepInfra's routed models: p50 TTFT, effective throughput, and success rate. Effective throughput uses provider-reported output tokens over complete request time. Unsupported route and probe-configuration rows are separated from provider downtime. No prompt or output content stored.

p50 TTFT3142 ms
Effective throughput88 tok/s n=2
Uptime99.19%
Modelp50 TTFTEffective throughputUptimeConfig excludedAvailability samples
deepseek/deepseek-v4-flash-0731 3142 ms 99.52% 419
meta-models/muse-glimmer-30b 860 ms 100.00% 1
qwen/qwen3.7-max 3762 ms 100.00% 1
qwen/qwen3.5-27b 3995 ms 100.00% 1
nvidia/nemotron-content-safety-3.5 4240 ms 100.00% 1
stepfun-ai/step-3.7-flash 4614 ms 100.00% 1
meta-llama/meta-llama-3.1-8b-instruct-turbo 100.00% 1
mistralai/mistral-nemo-instruct-2407 100.00% 1
xiaomi/mimo-v2.5-pro 23 tok/s n=1 100.00% 68
deepseek/deepseek-v4-pro-0423 152 tok/s n=1 0
openai/gpt-oss-120b-ultra 0.00% 1
z-ai/glm-5.3 0.00% 1

DeepInfra performance history · Full provider & model leaderboard.

Provider models

Models served by DeepInfra.

Each row links to pricing, provider, benchmark, and API pages for the model.

Model AI IQ Context Input Cached input Output
Qwen/Qwen3-Embedding-8B
Qwen3 Embedding 8B
32,000 $0.01055/1M Not published selected route
bytedance/seed-1.8
ByteDance/Seed-1.8
256,000 $0.26375/1M $0.05275/1M $2.11/1M
bytedance/seed-2.0-code
ByteDance/Seed-2.0-code
256,000 $0.5275/1M $0.1055/1M $3.165/1M
bytedance/seed-2.0-mini
ByteDance/Seed-2.0-mini
256,000 $0.1055/1M $0.0211/1M $0.422/1M
bytedance/seed-2.0-pro
ByteDance/Seed-2.0-pro
256,000 $0.5275/1M $0.1055/1M $3.165/1M
deepseek/deepseek-r1-0528
DeepSeek: R1 0528
163,840 $0.5275/1M $0.36925/1M $2.26825/1M
deepseek/deepseek-v3
deepseek-ai/DeepSeek-V3
163,840 $0.3376/1M Not published $0.93895/1M
deepseek/deepseek-v3-0324
deepseek-ai/DeepSeek-V3-0324
163,840 $0.2532/1M $0.142425/1M $0.9495/1M
deepseek/deepseek-v3.1
DeepSeek V3.1
IQ 95#105 131,072 $0.26375/1M $0.13715/1M $1.00225/1M
deepseek/deepseek-v3.2
DeepSeek: DeepSeek V3.2
IQ 103#77 163,840 $0.2743/1M $0.13715/1M $0.4009/1M
deepseek/deepseek-v4-flash
DeepSeek: DeepSeek V4 Flash 0423
IQ 116#37 1,048,576 $0.09495/1M $0.01899/1M $0.1899/1M
deepseek/deepseek-v4-flash-0731
DeepSeek: DeepSeek V4 Flash 0731
1,048,576 $0.0633/1M $0.015825/1M $0.1899/1M
deepseek/deepseek-v4-flash-vision-exp
DeepSeek: DeepSeek V4 Flash Vision Exp
IQ 115#42 1,048,576 $0.4642/1M $0.01477/1M $1.3926/1M
deepseek/deepseek-v4-pro
DeepSeek: DeepSeek V4 Pro 0423
IQ 114#47 1,048,576 $1.3715/1M $0.1055/1M $2.743/1M
deepseek/deepseek-v4-pro-0423
DeepSeek V4 Pro 0423
1,048,576 $1.3715/1M $0.1055/1M $2.743/1M
deepseek/deepseek-v4.1-flash
DeepSeek: DeepSeek V4.1 Flash
IQ 116#38 1,048,576 $0.211/1M $0.01/1M $0.633/1M
google/gemini-2.5-flash
Google: Gemini 2.5 Flash
1,048,576 $0.3165/1M Not published $2.6375/1M
google/gemini-2.5-pro
Google: Gemini 2.5 Pro
IQ 100#88 1,048,576 $1.31875/1M Not published $10.55/1M
google/gemini-3.1-flash-lite
Google: Gemini 3.1 Flash Lite
IQ 103#78 1,048,576 $0.26375/1M Not published $1.5825/1M
google/gemini-3.1-pro
google/gemini-3.1-pro
IQ 129#8 1,000,000 $2.11/1M Not published $12.66/1M
google/gemini-3.7-flash
Google: Gemini 3.7 Flash
IQ 124#15 1,048,576 $0.79125/1M Not published $3.95625/1M
google/gemma-3-12b-it
Google: Gemma 3 12B
131,072 $0.05275/1M Not published $0.15825/1M
google/gemma-3-27b-it
Google: Gemma 3 27B
131,072 $0.0844/1M Not published $0.1688/1M
google/gemma-3-4b-it
Google: Gemma 3 4B
131,072 $0.05275/1M Not published $0.1055/1M
google/gemma-4-26b-a4b-it
Google: Gemma 4 26B A4B
IQ 100#89 262,144 $0.07385/1M Not published $0.3587/1M
google/gemma-4-31b-it
Google: Gemma 4 31B
IQ 103#80 262,144 $0.13715/1M Not published $0.4009/1M
google/gemma-4-31b-it-turbo
google/gemma-4-31B-it-turbo
262,144 $0.09495/1M $0.05275/1M $0.3587/1M
google/gemma-4-31b-it-ultra
google/gemma-4-31B-it-Ultra
131,072 $0.28485/1M Not published $0.8018/1M
google/gemma-4-e4b-it
google/gemma-4-E4B-it
IQ 84#129 131,072 $0.0211/1M Not published $0.1055/1M
gryphe/mythomax-l2-13b
MythoMax 13B
4,096 $0.422/1M Not published $0.422/1M
ibm-granite/granite-4.2-30b
ibm-granite/granite-4.2-30b
131,072 $0.1688/1M $0.0422/1M $0.68575/1M
ibm-granite/granite-4.2-3b
ibm-granite/granite-4.2-3b
131,072 $0.03165/1M $0.01/1M $0.1266/1M
ibm-granite/granite-4.2-8b
IBM: Granite 4.2 8B
131,072 $0.0633/1M $0.015825/1M $0.26375/1M
inclusionai/ling-3.0-flash
inclusionAI: Ling 3.0 Flash
IQ 94#109 131,072 $0.0633/1M $0.01266/1M $0.1899/1M
inclusionai/ling-3.0-flash-fin
inclusionAI: Ling 3.0 Flash Fin
262,144 $0.0633/1M $0.01266/1M $0.1899/1M
inclusionai/ling-3.0-flash-vl
inclusionAI: Ling 3.0 Flash VL
131,072 $0.0633/1M $0.01266/1M $0.1899/1M
meta-llama/llama-3.1-70b-instruct
Meta: Llama 3.1 70B Instruct
131,072 $0.422/1M Not published $0.422/1M
meta-llama/llama-3.3-70b-instruct-turbo
meta-llama/Llama-3.3-70B-Instruct-Turbo
131,072 $0.1055/1M Not published $0.3376/1M
meta-llama/llama-4-maverick-17b-128e-instruct-fp8
Llama 4 Maverick Instruct
1,048,576 $0.211/1M Not published $0.844/1M
meta-llama/llama-4-scout-17b-16e-instruct
Llama 4 Scout Instruct
131,072 $0.1055/1M Not published $0.3165/1M
meta-llama/llama-guard-4-12b
Meta: Llama Guard 4 12B
163,840 $0.1899/1M Not published $0.1899/1M
meta-llama/meta-llama-3.1-8b-instruct-turbo
meta-llama/Meta-Llama-3.1-8B-Instruct-Turbo
131,072 $0.0211/1M Not published $0.0422/1M
meta-models/muse-glimmer-30b
Muse Glimmer 30B on Fireworks
131,072 $0.3165/1M $0.0422/1M $1.266/1M
microsoft/phi-4
Microsoft: Phi 4
16,384 $0.07385/1M Not published $0.1477/1M
minimax/minimax-m2.7-turbo
MiniMaxAI/MiniMax-M2.7-Turbo
196,608 $0.4009/1M $0.07385/1M $1.7935/1M
minimax/minimax-m3
MiniMax: MiniMax M3
IQ 115#45 524,288 $0.2954/1M $0.05908/1M $1.1605/1M
mistralai/mistral-nemo-instruct-2407
mistralai/Mistral-Nemo-Instruct-2407
131,072 $0.020045/1M Not published $0.03165/1M
mistralai/mistral-small-24b-instruct-2501
Mistral: Mistral Small 3
32,768 $0.05275/1M Not published $0.0844/1M
mistralai/mistral-small-3.2-24b-instruct-2506
mistralai/Mistral-Small-3.2-24B-Instruct-2506
128,000 $0.079125/1M Not published $0.211/1M
moonshotai/kimi-k2.6
MoonshotAI: Kimi K2.6
IQ 118#35 262,144 $0.79125/1M $0.15825/1M $3.6925/1M
moonshotai/kimi-k3
MoonshotAI: Kimi K3
IQ 121#27 1,048,576 $3.00675/1M $0.300675/1M $15.03375/1M
nousresearch/hermes-3-llama-3.1-405b
Nous: Hermes 3 405B Instruct
131,072 $1.055/1M Not published $1.055/1M
nvidia/nemotron-3-nano-30b-a3b
NVIDIA: Nemotron 3 Nano 30B A3B
262,144 $0.05275/1M $0.026375/1M $0.211/1M
nvidia/nemotron-3-super-120b-a12b
NVIDIA: Nemotron 3 Super
262,144 $0.089675/1M Not published $0.422/1M
nvidia/nemotron-3-ultra-550b-a55b
NVIDIA: Nemotron 3 Ultra
262,144 $0.5275/1M $0.1055/1M $2.321/1M
nvidia/nemotron-3.5-lightning
NVIDIA: Nemotron 3.5 Lightning
262,144 $0.0844/1M $0.0422/1M $0.211/1M
nvidia/nemotron-content-safety-3.5
nvidia/Nemotron-Content-Safety-3.5
131,072 $0.211/1M Not published $0.211/1M
openai/gpt-oss-120b
OpenAI: gpt-oss-120b
IQ 98#99 131,072 $0.039035/1M Not published $0.17935/1M
openai/gpt-oss-120b-turbo
openai/gpt-oss-120b-Turbo
131,072 $0.15825/1M Not published $0.633/1M
openai/gpt-oss-120b-ultra
openai/gpt-oss-120b-Ultra
131,072 $0.211/1M Not published $1.00225/1M
openai/gpt-oss-20b
OpenAI: gpt-oss-20b
IQ 93#113 131,072 $0.03165/1M Not published $0.1477/1M
qwen/qwen2.5-72b-instruct
Qwen/Qwen2.5-72B-Instruct
32,768 $0.3798/1M Not published $0.422/1M
qwen/qwen3-14b
Qwen: Qwen3 14B
40,960 $0.1266/1M Not published $0.2532/1M
qwen/qwen3-235b-a22b-instruct-2507
Qwen3 235B A22B Instruct 2507
131,072 $0.09495/1M Not published $0.58025/1M
qwen/qwen3-30b-a3b
Qwen: Qwen3 30B A3B
40,960 $0.1266/1M Not published $0.5275/1M
qwen/qwen3-coder-480b-a35b-instruct-turbo
Qwen/Qwen3-Coder-480B-A35B-Instruct-Turbo
262,144 $0.3165/1M $0.1055/1M $1.055/1M
qwen/qwen3-max
Qwen3 Max
262,144 $1.266/1M $0.2532/1M $6.33/1M
qwen/qwen3-max-thinking
Qwen/Qwen3-Max-Thinking
256,000 $1.266/1M $0.2532/1M $6.33/1M
qwen/qwen3-next-80b-a3b-instruct
Qwen: Qwen3 Next 80B A3B Instruct
262,144 $0.09495/1M Not published $1.1605/1M
qwen/qwen3-vl-235b-a22b-instruct
Qwen: Qwen3 VL 235B A22B Instruct
262,144 $0.211/1M $0.11605/1M $0.9284/1M
qwen/qwen3-vl-30b-a3b-instruct
Qwen: Qwen3 VL 30B A3B Instruct
262,144 $0.15825/1M Not published $0.633/1M
qwen/qwen3.5-122b-a10b
Qwen: Qwen3.5-122B-A10B
262,144 $0.30595/1M Not published $2.532/1M
qwen/qwen3.5-27b
Qwen: Qwen3.5-27B
262,144 $0.2743/1M Not published $2.743/1M
qwen/qwen3.5-35b-a3b
Qwen: Qwen3.5-35B-A3B
262,144 $0.1477/1M $0.05275/1M $1.055/1M
qwen/qwen3.5-397b-a17b
Qwen: Qwen3.5 397B A17B
262,144 $0.47475/1M $0.2321/1M $3.165/1M
qwen/qwen3.5-9b
Qwen: Qwen3.5-9B
IQ 95#108 262,144 $0.1055/1M Not published $0.15825/1M
qwen/qwen3.6-27b
Qwen: Qwen3.6 27B
IQ 108#68 262,144 $0.3376/1M Not published $3.376/1M
qwen/qwen3.6-35b-a3b
Qwen: Qwen3.6 35B A3B
IQ 100#94 262,144 $0.1055/1M Not published $1.00225/1M
qwen/qwen3.7-max
Qwen3.7 Max
IQ 119#31 1,000,000 $2.6375/1M $0.5275/1M $7.9125/1M
qwen/qwen3.8-2.4t-a95b
Qwen: Qwen3.8 2.4T A95B
IQ 122#24 262,144 $2.11/1M $0.211/1M $6.33/1M
qwen/qwen3.8-27b
Qwen: Qwen3.8 27B
IQ 110#60 262,144 $0.211/1M $0.05275/1M $2.6375/1M
qwen/qwen3.8-flash
Qwen3.8 Flash
IQ 115#46 1,000,000 $0.119215/1M $0.014876/1M $0.40301/1M
qwen/qwen3.8-max
Qwen3.8 Max
IQ 122#25 1,000,000 $1.74075/1M $0.21733/1M $5.223305/1M
sao10k/l3-8b-lunaris-v1-turbo
Sao10K/L3-8B-Lunaris-v1-Turbo
8,192 $0.0422/1M Not published $0.05275/1M
sao10k/l3.1-70b-euryale-v2.2
Sao10K/L3.1-70B-Euryale-v2.2
131,072 $0.89675/1M Not published $0.89675/1M
stepfun-ai/step-3.7-flash
stepfun-ai/Step-3.7-Flash
IQ 102#84 262,144 $0.211/1M $0.0422/1M $1.21325/1M
tencent/hy3
Tencent: Hy3
IQ 100#90 262,144 $0.1477/1M $0.036925/1M $0.6119/1M
thinkingmachines/inkling
Thinking Machines: Inkling
IQ 107#70 524,288 $1.00225/1M $0.1688/1M $4.27275/1M
thinkingmachines/inkling-small
Thinking Machines: Inkling Small
IQ 107#71 524,288 $0.47475/1M $0.1055/1M $1.266/1M
xiaomi/mimo-v2.5
Xiaomi: MiMo-V2.5
IQ 108#67 1,050,000 $0.1477/1M $0.01/1M $0.2954/1M
xiaomi/mimo-v2.5-pro
Xiaomi: MiMo-V2.5-Pro
IQ 112#55 1,050,000 $1.055/1M $0.211/1M $3.165/1M
z-ai/glm-4.6
Z.ai: GLM 4.6
202,752 $0.5275/1M $0.1055/1M $2.11/1M
z-ai/glm-4.7
Z.ai: GLM 4.7
IQ 103#81 202,752 $0.422/1M $0.0844/1M $1.84625/1M
z-ai/glm-5.2
Z.ai: GLM 5.2
IQ 120#28 1,048,576 $0.79125/1M $0.1477/1M $2.532/1M
z-ai/glm-5.3
Z.ai: GLM 5.3
IQ 123#18 1,048,576 $1.266/1M $0.211/1M $4.22/1M
z-ai/glm-5.3-flash
Z.ai: GLM 5.3 Flash
IQ 116#40 1,048,576 $0.15825/1M $0.03165/1M $0.5275/1M

Questions

Does DeepInfra have zero data retention?

TrustedRouter does not currently mark DeepInfra as provider-level zero data retention. Use trustedrouter/zdr or provider.min_privacy=zdr to select a different eligible route, and review the linked policy source for changes.

Is DeepInfra end-to-end encrypted?

TrustedRouter does not currently mark DeepInfra as end-to-end encrypted at the provider boundary. The TrustedRouter gateway is still attested, but the selected provider normally receives the request in order to run the model. Use trustedrouter/e2e for the stronger route requirement.

Which DeepInfra models are available through TrustedRouter?

This page currently lists 96 public DeepInfra models, with live pricing, route count, context length, measured performance when available, and links to each model's provider and benchmark pages.

Workspace access

Sign in

Choose a sign in method to access your TrustedRouter workspace.

By signing in you agree to the terms of service and privacy policy.