OpenAI compatible API · Attested · Public status

Baseten

Baseten models on TrustedRouter with prices, routes, policy notes, and source links.

Verify gateway
Onebase URL to migrate
100sof models and routes
0prompt or output logs. Always.

baseten

No provider claim

All providers

ProviderBaseten
Models12 public models
Prepaid routes12
BYOK routes12
Zero data retentionnot claimed
Confidential computenot claimed
Provider E2EEnot claimed
Policy noteNo provider-ZDR claim is tracked here. Baseten's inference and security documentation are linked for users who need to review API data handling.
Policy source

Measured performance

64 samples

Continuously sampled across Baseten's routed models: p50 TTFT, effective throughput, and success rate. Effective throughput uses provider-reported output tokens over complete request time. Unsupported route and probe-configuration rows are separated from provider downtime. No prompt or output content stored.

p50 TTFT1668 ms
Effective throughput71 tok/s n=12
Uptime93.75%
Modelp50 TTFTp50 TTFBEffective throughputUptimeConfig excludedAvailability samples
moonshotai/kimi-k3 1252 ms 1252 ms 47 tok/s n=1 100.00% 8
thinkingmachines/inkling-1m 1284 ms 1284 ms 64 tok/s n=1 100.00% 5
thinkingmachines/inkling-small 1285 ms 1285 ms 69 tok/s n=1 100.00% 4
openai/gpt-oss-120b 1590 ms 1590 ms 121 tok/s n=1 100.00% 3
z-ai/glm-4.7 1668 ms 1668 ms 100.00% 7
deepseek/deepseek-v4-pro 2080 ms 2080 ms 72 tok/s n=2 100.00% 7
z-ai/glm-5.2 2912 ms 2912 ms 45 tok/s n=2 100.00% 5
nvidia/nemotron-3-ultra-550b-a55b 3087 ms 3087 ms 234 tok/s n=1 100.00% 4
z-ai/glm-5.2-fast 3111 ms 3111 ms 64 tok/s n=1 100.00% 6
moonshotai/kimi-k2.6 2966 ms 2966 ms 97 tok/s n=1 90.00% 10
moonshotai/kimi-k2.7-code 567 ms 567 ms 74 tok/s n=1 40.00% 5

Baseten performance history · Full provider & model leaderboard.

Provider models

Models served by Baseten.

Each row links to pricing, provider, benchmark, and API pages for the model.

Model AI IQ Context Endpoints Prompt Completion Routes
deepseek/deepseek-v4-flash-0731
DeepSeek: DeepSeek V4 Flash 0731
1,048,576 2 $0.1365/1M $0.273/1M prepaid BYOK
deepseek/deepseek-v4-pro
DeepSeek: DeepSeek V4 Pro
IQ 115#30 1,048,576 2 $1.827/1M $3.654/1M prepaid BYOK
moonshotai/kimi-k2.6
MoonshotAI: Kimi K2.6
IQ 119#19 262,144 2 $0.9975/1M $4.2/1M prepaid BYOK
moonshotai/kimi-k2.7-code
MoonshotAI: Kimi K2.7 Code
IQ 118#22 262,144 2 $0.9975/1M $4.2/1M prepaid BYOK
moonshotai/kimi-k3
MoonshotAI: Kimi K3
IQ 122#15 1,048,576 2 $3.15/1M $15.75/1M prepaid BYOK
nvidia/nemotron-3-ultra-550b-a55b
NVIDIA: Nemotron 3 Ultra
512,288 2 $0.63/1M $2.52/1M prepaid BYOK
openai/gpt-oss-120b
OpenAI: gpt-oss-120b
IQ 105#52 131,072 2 $0.105/1M $0.525/1M prepaid BYOK
thinkingmachines/inkling-1m
Inkling
1,048,576 2 $1.05/1M $4.2525/1M prepaid BYOK
thinkingmachines/inkling-small
Thinking Machines: Inkling Small
524,288 2 $0.525/1M $1.26/1M prepaid BYOK
z-ai/glm-4.7
Z.ai: GLM 4.7
IQ 103#58 204,800 2 $0.63/1M $2.31/1M prepaid BYOK
z-ai/glm-5.2
Z.ai: GLM 5.2
IQ 120#16 1,048,576 2 $1.47/1M $4.62/1M prepaid BYOK
z-ai/glm-5.2-fast
GLM 5.2 Fast on Fireworks
1,048,576 2 $2.205/1M $6.93/1M prepaid BYOK
Workspace access

Sign in

Choose a sign in method. New email and OAuth accounts include $0.10 in starter credit; wallet-only accounts start at $0.

By signing in you agree to the terms of service and privacy policy.