OpenAI compatible API · Attested · Public status

DeepInfra

DeepInfra models on TrustedRouter with prices, routes, policy notes, and source links.

Verify gateway
Onebase URL to migrate
100sof models and routes
0prompt or output logs. Always.

deepinfra

No provider claim

All providers

ProviderDeepInfra
Models50 public models
Prepaid routes49
BYOK routes50
Zero data retentionno
Confidential computenot claimed
Provider E2EEnot claimed
Policy noteTracked as no-store, not strict ZDR. DeepInfra documents memory-only handling and no training for ordinary inference, but reserves the right to log a small portion of requests for debugging or security. Google- and Anthropic-backed routes also inherit those vendors' terms.
Policy source

Measured performance

54 samples

Continuously sampled across DeepInfra's routed models: p50 TTFT, effective throughput, and success rate. Effective throughput uses provider-reported output tokens over complete request time. Unsupported route and probe-configuration rows are separated from provider downtime. No prompt or output content stored.

p50 TTFT1874 ms
Effective throughput51 tok/s n=14
Uptime98.15%
Modelp50 TTFTp50 TTFBEffective throughputUptimeConfig excludedAvailability samples
qwen/qwen3.5-27b 369 ms 368 ms 100.00% 1
qwen/qwen3.5-9b 476 ms 476 ms 100.00% 2
deepseek/deepseek-r1-0528 517 ms 516 ms 100.00% 2
qwen/qwen3-30b-a3b 548 ms 548 ms 100.00% 1
qwen/qwen3.5-35b-a3b 698 ms 698 ms 100.00% 2
qwen/qwen3-235b-a22b-thinking-2507 764 ms 764 ms 100.00% 1
deepseek/deepseek-v4-flash 882 ms 882 ms 37 tok/s n=1 100.00% 1
openai/gpt-oss-120b 923 ms 923 ms 28 tok/s n=1 100.00% 2
google/gemma-3-12b-it 938 ms 938 ms 100.00% 2
tencent/hy3 1187 ms 1187 ms 53 tok/s n=1 100.00% 2
qwen/qwen3.5-397b-a17b 1338 ms 1338 ms 100.00% 2
nousresearch/hermes-3-llama-3.1-405b 1498 ms 1498 ms 100.00% 2
qwen/qwen3-14b 1595 ms 1594 ms 100.00% 4
google/gemma-3-27b-it 1645 ms 1645 ms 100.00% 1
meta-llama/llama-3.1-70b-instruct 1653 ms 1653 ms 100.00% 1
z-ai/glm-4.7 1874 ms 1873 ms 100.00% 1
google/gemma-3-4b-it 1944 ms 1943 ms 100.00% 2
moonshotai/kimi-k2.5 2274 ms 2274 ms 100.00% 4
google/gemini-2.5-flash 2584 ms 2583 ms 100.00% 1
openai/gpt-oss-20b 2819 ms 2819 ms 100.00% 3
google/gemini-3.1-flash-lite 2872 ms 2872 ms 100.00% 2
gryphe/mythomax-l2-13b 3012 ms 3012 ms 100.00% 1
qwen/qwen3.6-35b-a3b 3189 ms 3189 ms 100.00% 1
deepseek/deepseek-v3.1-terminus 3346 ms 3345 ms 100.00% 3
z-ai/glm-5 3512 ms 3512 ms 100.00% 3
thinkingmachines/inkling-small 3641 ms 3641 ms 63 tok/s n=1 100.00% 1
google/gemini-2.5-pro 4535 ms 4535 ms 100.00% 1
z-ai/glm-4.7-flash 7660 ms 7659 ms 100.00% 1
minimax/minimax-m2.7 8778 ms 8778 ms 100.00% 3
deepseek/deepseek-v4-flash-0731 63 tok/s n=1 0
deepseek/deepseek-v4-pro 73 tok/s n=2 0
google/gemma-4-31b-it 67 tok/s n=1 0
minimax/minimax-m3 19 tok/s n=2 0
moonshotai/kimi-k2.6 37 tok/s n=1 0
qwen/qwen3.5-122b-a10b 0.00% 1
thinkingmachines/inkling 67 tok/s n=1 0
z-ai/glm-5.2 49 tok/s n=2 0

DeepInfra performance history · Full provider & model leaderboard.

Provider models

Models served by DeepInfra.

Each row links to pricing, provider, benchmark, and API pages for the model.

Model AI IQ Context Endpoints Prompt Completion Routes
Qwen/Qwen3-Embedding-8B
Qwen3 Embedding 8B
32,000 2 $0.0105/1M selected route prepaid BYOK
anthropic/claude-opus-5
Claude Opus 5
IQ 134#3 1,000,000 1 $5.25/1M $26.25/1M BYOK
deepseek/deepseek-r1-0528
DeepSeek: R1 0528
163,840 2 $0.525/1M $2.2575/1M prepaid BYOK
deepseek/deepseek-v3.1-terminus
DeepSeek: DeepSeek V3.1 Terminus
163,840 2 $0.2835/1M $0.9975/1M prepaid BYOK
deepseek/deepseek-v3.2
DeepSeek: DeepSeek V3.2
IQ 103#56 163,840 2 $0.273/1M $0.399/1M prepaid BYOK
deepseek/deepseek-v4-flash
DeepSeek: DeepSeek V4 Flash
IQ 108#46 1,048,576 2 $0.0945/1M $0.189/1M prepaid BYOK
deepseek/deepseek-v4-flash-0731
DeepSeek: DeepSeek V4 Flash 0731
1,048,576 2 $0.0945/1M $0.189/1M prepaid BYOK
deepseek/deepseek-v4-pro
DeepSeek: DeepSeek V4 Pro
IQ 115#30 1,048,576 2 $1.365/1M $2.73/1M prepaid BYOK
google/gemini-2.5-flash
Google: Gemini 2.5 Flash
1,048,576 2 $0.315/1M $2.625/1M prepaid BYOK
google/gemini-2.5-pro
Google: Gemini 2.5 Pro
IQ 103#57 1,048,576 2 $1.3125/1M $10.5/1M prepaid BYOK
google/gemini-3.1-flash-lite
Google: Gemini 3.1 Flash Lite
IQ 101#64 1,048,576 2 $0.2625/1M $1.575/1M prepaid BYOK
google/gemma-3-12b-it
Google: Gemma 3 12B
131,072 2 $0.0525/1M $0.1575/1M prepaid BYOK
google/gemma-3-27b-it
Google: Gemma 3 27B
262,144 2 $0.084/1M $0.168/1M prepaid BYOK
google/gemma-3-4b-it
Google: Gemma 3 4B
131,072 2 $0.0525/1M $0.105/1M prepaid BYOK
google/gemma-4-26b-a4b-it
Google: Gemma 4 26B A4B
IQ 96#79 262,144 2 $0.0735/1M $0.357/1M prepaid BYOK
google/gemma-4-31b-it
Google: Gemma 4 31B
IQ 101#65 262,144 2 $0.1365/1M $0.399/1M prepaid BYOK
gryphe/mythomax-l2-13b
MythoMax 13B
8,192 2 $0.42/1M $0.42/1M prepaid BYOK
meta-llama/llama-3.1-70b-instruct
Meta: Llama 3.1 70B Instruct
131,072 2 $0.42/1M $0.42/1M prepaid BYOK
meta-llama/llama-guard-4-12b
Meta: Llama Guard 4 12B
1,048,576 2 $0.189/1M $0.189/1M prepaid BYOK
microsoft/phi-4
Microsoft: Phi 4
16,384 2 $0.0735/1M $0.147/1M prepaid BYOK
minimax/minimax-m2.7
MiniMax: MiniMax M2.7
IQ 109#44 204,800 2 $0.2625/1M $1.05/1M prepaid BYOK
minimax/minimax-m3
MiniMax: MiniMax M3
IQ 114#33 1,048,576 2 $0.315/1M $1.26/1M prepaid BYOK
mistralai/mistral-small-24b-instruct-2501
Mistral: Mistral Small 3
32,768 2 $0.0525/1M $0.084/1M prepaid BYOK
moonshotai/kimi-k2.5
MoonshotAI: Kimi K2.5
IQ 111#39 262,144 2 $0.4725/1M $2.3625/1M prepaid BYOK
moonshotai/kimi-k2.6
MoonshotAI: Kimi K2.6
IQ 119#19 262,144 2 $0.7875/1M $3.675/1M prepaid BYOK
nousresearch/hermes-3-llama-3.1-405b
Nous: Hermes 3 405B Instruct
131,072 2 $1.05/1M $1.05/1M prepaid BYOK
nvidia/nemotron-3-nano-30b-a3b
NVIDIA: Nemotron 3 Nano 30B A3B
262,144 2 $0.0525/1M $0.21/1M prepaid BYOK
openai/gpt-oss-120b
OpenAI: gpt-oss-120b
IQ 105#52 131,072 2 $0.03885/1M $0.1785/1M prepaid BYOK
openai/gpt-oss-20b
OpenAI: gpt-oss-20b
IQ 100#69 131,072 2 $0.0315/1M $0.147/1M prepaid BYOK
qwen/qwen3-14b
Qwen: Qwen3 14B
131,072 2 $0.126/1M $0.252/1M prepaid BYOK
qwen/qwen3-235b-a22b-thinking-2507
Qwen: Qwen3 235B A22B Thinking 2507
262,144 2 $0.2415/1M $2.415/1M prepaid BYOK
qwen/qwen3-30b-a3b
Qwen: Qwen3 30B A3B
131,072 2 $0.126/1M $0.525/1M prepaid BYOK
qwen/qwen3-next-80b-a3b-instruct
Qwen: Qwen3 Next 80B A3B Instruct
262,144 2 $0.0945/1M $1.155/1M prepaid BYOK
qwen/qwen3-vl-235b-a22b-instruct
Qwen: Qwen3 VL 235B A22B Instruct
262,144 2 $0.21/1M $0.924/1M prepaid BYOK
qwen/qwen3-vl-30b-a3b-instruct
Qwen: Qwen3 VL 30B A3B Instruct
262,144 2 $0.1575/1M $0.63/1M prepaid BYOK
qwen/qwen3.5-122b-a10b
Qwen: Qwen3.5-122B-A10B
262,144 2 $0.3045/1M $2.52/1M prepaid BYOK
qwen/qwen3.5-27b
Qwen: Qwen3.5-27B
262,144 2 $0.273/1M $2.73/1M prepaid BYOK
qwen/qwen3.5-35b-a3b
Qwen: Qwen3.5-35B-A3B
262,144 2 $0.147/1M $1.05/1M prepaid BYOK
qwen/qwen3.5-397b-a17b
Qwen: Qwen3.5 397B A17B
262,144 2 $0.4725/1M $3.15/1M prepaid BYOK
qwen/qwen3.5-9b
Qwen: Qwen3.5-9B
IQ 93#89 262,144 2 $0.105/1M $0.1575/1M prepaid BYOK
qwen/qwen3.6-27b
Qwen: Qwen3.6 27B
IQ 111#40 262,144 2 $0.336/1M $3.36/1M prepaid BYOK
qwen/qwen3.6-35b-a3b
Qwen: Qwen3.6 35B A3B
IQ 100#70 262,144 2 $0.105/1M $0.9975/1M prepaid BYOK
tencent/hy3
Tencent: Hy3
IQ 103#59 262,144 2 $0.147/1M $0.609/1M prepaid BYOK
thinkingmachines/inkling
Thinking Machines: Inkling
IQ 104#55 1,048,576 2 $1.05/1M $4.2525/1M prepaid BYOK
thinkingmachines/inkling-small
Thinking Machines: Inkling Small
524,288 2 $0.525/1M $1.26/1M prepaid BYOK
z-ai/glm-4.6
Z.ai: GLM 4.6
204,800 2 $0.525/1M $2.1/1M prepaid BYOK
z-ai/glm-4.7
Z.ai: GLM 4.7
IQ 103#58 204,800 2 $0.42/1M $1.8375/1M prepaid BYOK
z-ai/glm-4.7-flash
Z.ai: GLM 4.7 Flash
202,752 2 $0.063/1M $0.42/1M prepaid BYOK
z-ai/glm-5
Z.ai: GLM 5
IQ 105#51 204,800 2 $0.63/1M $2.184/1M prepaid BYOK
z-ai/glm-5.2
Z.ai: GLM 5.2
IQ 120#16 1,048,576 2 $1.26/1M $4.41/1M prepaid BYOK
Workspace access

Sign in

Choose a sign in method. New email and OAuth accounts include $0.10 in starter credit; wallet-only accounts start at $0.

By signing in you agree to the terms of service and privacy policy.