OpenAI compatible API · Attested · Public status
Baseten performance
Measured TTFT, TTFB, effective throughput, uptime, and sampled model routes for Baseten.
Onebase URL to migrate
100sof models and routes
0prompt or output logs. Always.
baseten
63 samples
Continuously sampled provider performance. TrustedRouter reports unsupported route and probe-configuration rows separately from provider downtime. Prompt and output content is not stored.
| p50 TTFT | 1668 ms |
|---|---|
| p95 TTFT | 3969 ms |
| p50 TTFB | 2168 ms |
| Effective throughput | 71 tok/s n=12 |
| Uptime | 93.65% |
Measured model routes
| Model | p50 TTFT | p50 TTFB | Effective throughput | Uptime | Config excluded | Availability samples |
|---|---|---|---|---|---|---|
| moonshotai/kimi-k3 | 1252 ms | 1252 ms | 47 tok/s n=1 | 100.00% | — | 8 |
| thinkingmachines/inkling-1m | 1284 ms | 1284 ms | 64 tok/s n=1 | 100.00% | — | 5 |
| thinkingmachines/inkling-small | 1484 ms | 1484 ms | 69 tok/s n=1 | 100.00% | — | 5 |
| openai/gpt-oss-120b | 1590 ms | 1590 ms | 121 tok/s n=1 | 100.00% | — | 3 |
| z-ai/glm-4.7 | 1668 ms | 1668 ms | — | 100.00% | — | 7 |
| deepseek/deepseek-v4-pro | 2016 ms | 2016 ms | 72 tok/s n=2 | 100.00% | — | 6 |
| z-ai/glm-5.2 | 2912 ms | 2912 ms | 45 tok/s n=2 | 100.00% | — | 5 |
| nvidia/nemotron-3-ultra-550b-a55b | 3087 ms | 3087 ms | 234 tok/s n=1 | 100.00% | — | 4 |
| z-ai/glm-5.2-fast | 3179 ms | 3179 ms | 64 tok/s n=1 | 100.00% | — | 7 |
| moonshotai/kimi-k2.6 | 3440 ms | 3440 ms | 97 tok/s n=1 | 87.50% | — | 8 |
| moonshotai/kimi-k2.7-code | 567 ms | 567 ms | 74 tok/s n=1 | 40.00% | — | 5 |