OpenAI compatible API · Attested · Public status
Together
Together models on TrustedRouter with prices, routes, policy notes, and source links.
Onebase URL to migrate
100sof models and routes
0prompt or output logs. Always.
together
No logs
| Provider | Together |
|---|---|
| Models | 17 public models |
| Prepaid routes | 17 |
| BYOK routes | 17 |
| Zero data retention | yes |
| Confidential compute | not claimed |
| Provider E2EE | not claimed |
| Policy note | Tracked as provider ZDR. Together documents that inference inputs and outputs are not stored by default; temporary prompt caching may be used for performance, and sharing content for training is opt-in. Policy source |
Measured performance
35 samplesContinuously sampled across Together's routed models: p50 TTFT, effective throughput, and success rate. Effective throughput uses provider-reported output tokens over complete request time. Unsupported route and probe-configuration rows are separated from provider downtime. No prompt or output content stored.
| p50 TTFT | 1854 ms |
|---|---|
| Effective throughput | 66 tok/s n=10 |
| Uptime | 100.00% |
| Model | p50 TTFT | p50 TTFB | Effective throughput | Uptime | Config excluded | Availability samples |
|---|---|---|---|---|---|---|
| openai/gpt-oss-120b | 697 ms | 697 ms | 129 tok/s n=1 | 100.00% | — | 2 |
| moonshotai/kimi-k2.6 | 1000 ms | 1000 ms | — | 100.00% | — | 3 |
| nvidia/nemotron-3-ultra-550b-a55b | 1515 ms | 1515 ms | 124 tok/s n=1 | 100.00% | — | 1 |
| moonshotai/kimi-k2.7-code | 1696 ms | 1696 ms | — | 100.00% | — | 5 |
| moonshotai/kimi-k3 | 1707 ms | 1707 ms | 31 tok/s n=1 | 100.00% | — | 3 |
| openai/gpt-oss-20b | 1854 ms | 1854 ms | — | 100.00% | — | 5 |
| qwen/qwen-2.5-7b-instruct | 1868 ms | 1868 ms | — | 100.00% | — | 3 |
| thinkingmachines/inkling | 1987 ms | 1987 ms | 24 tok/s n=1 | 100.00% | — | 2 |
| z-ai/glm-5.2 | 2263 ms | 2263 ms | — | 100.00% | — | 7 |
| google/gemma-3n-e4b-it | 2342 ms | 2342 ms | — | 100.00% | — | 2 |
| meta-llama/llama-3.3-70b-instruct | 2412 ms | 2412 ms | 68 tok/s n=1 | 100.00% | — | 2 |
| deepseek/deepseek-v4-pro | — | — | 69 tok/s n=2 | — | 6 probe_config_error |
0 |
| google/gemma-4-31b-it | — | — | 64 tok/s n=1 | — | 1 probe_config_error |
0 |
| minimax/minimax-m3 | — | — | 61 tok/s n=2 | — | 7 probe_config_error |
0 |
Together performance history · Full provider & model leaderboard.
Provider models
Models served by Together.
Each row links to pricing, provider, benchmark, and API pages for the model.
| Model | AI IQ | Context | Endpoints | Prompt | Completion | Routes |
|---|---|---|---|---|---|---|
deepseek/deepseek-v4-proDeepSeek: DeepSeek V4 Pro |
IQ 115#30 | 1,048,576 | 2 | $1.827/1M | $3.654/1M | prepaid BYOK |
google/gemma-3n-e4b-itGoogle: Gemma 3n 4B |
— | 32,768 | 2 | $0.063/1M | $0.126/1M | prepaid BYOK |
google/gemma-4-31b-itGoogle: Gemma 4 31B |
IQ 101#65 | 262,144 | 2 | $0.4095/1M | $1.0185/1M | prepaid BYOK |
intfloat/multilingual-e5-large-instructMultilingual E5 Large Instruct |
— | 512 | 2 | $0.021/1M | selected route | prepaid BYOK |
meta-llama/llama-3.3-70b-instructMeta: Llama 3.3 70B Instruct |
— | 131,072 | 2 | $1.092/1M | $1.092/1M | prepaid BYOK |
minimax/minimax-m3MiniMax: MiniMax M3 |
IQ 114#33 | 1,048,576 | 2 | $0.315/1M | $1.26/1M | prepaid BYOK |
moonshotai/kimi-k2.6MoonshotAI: Kimi K2.6 |
IQ 119#19 | 262,144 | 2 | $1.26/1M | $4.725/1M | prepaid BYOK |
moonshotai/kimi-k2.7-codeMoonshotAI: Kimi K2.7 Code |
IQ 118#22 | 262,144 | 2 | $0.9975/1M | $4.2/1M | prepaid BYOK |
moonshotai/kimi-k3MoonshotAI: Kimi K3 |
IQ 122#15 | 1,048,576 | 2 | $3.15/1M | $15.75/1M | prepaid BYOK |
nvidia/nemotron-3-ultra-550b-a55bNVIDIA: Nemotron 3 Ultra |
— | 512,288 | 2 | $0.63/1M | $3.78/1M | prepaid BYOK |
openai/gpt-oss-120bOpenAI: gpt-oss-120b |
IQ 105#52 | 131,072 | 2 | $0.1575/1M | $0.63/1M | prepaid BYOK |
openai/gpt-oss-20bOpenAI: gpt-oss-20b |
IQ 100#69 | 131,072 | 2 | $0.0525/1M | $0.21/1M | prepaid BYOK |
pearl-ai/gemma-4-31b-itPearl-ai Gemma-4-31B-it-pearl |
IQ 101#65 | 262,144 | 2 | $0.294/1M | $0.903/1M | prepaid BYOK |
qwen/qwen-2.5-7b-instructQwen: Qwen2.5 7B Instruct |
— | 32,768 | 2 | $0.315/1M | $0.315/1M | prepaid BYOK |
qwen/qwen3.5-9bQwen: Qwen3.5-9B |
IQ 93#89 | 262,144 | 2 | $0.1785/1M | $0.2625/1M | prepaid BYOK |
thinkingmachines/inklingThinking Machines: Inkling |
IQ 104#55 | 1,048,576 | 2 | $1.05/1M | $4.2525/1M | prepaid BYOK |
z-ai/glm-5.2Z.ai: GLM 5.2 |
IQ 120#16 | 1,048,576 | 2 | $1.47/1M | $4.62/1M | prepaid BYOK |