Weights & Biases
Explore Weights & Biases models on TrustedRouter with current routes, token pricing, policy sources, privacy posture, regional availability, and API support.
Weights & Biaseswandb
No provider claimThese privacy labels describe Weights & Biases, the upstream model provider. ZDR is a retention policy; verified confidential inference additionally requires attested provider compute and end-to-end encryption.
| Provider | Weights & Biases |
|---|---|
| Routing status | Active |
| Provider website | https://wandb.ai/site/inference/ |
| Models | 26 public models |
| Credits routes | 26 |
| Zero data retention | no |
| Verified confidential inference | Not verified |
| Policy note | W&B Serverless Inference runs on CoreWeave infrastructure. W&B does not publish a zero-retention, confidential-compute, or end-to-end-encryption commitment for this service, so these routes use the Standard privacy tier. TrustedRouter does not enable W&B Weave tracing for provider calls. Policy source |
Measured performance
32 samplesContinuously sampled across Weights & Biases's routed models: p50 TTFT, effective throughput, and success rate. Effective throughput uses provider-reported output tokens over complete request time. Unsupported route and probe-configuration rows are separated from provider downtime. No prompt or output content stored.
| p50 TTFT | 2298 ms |
|---|---|
| Effective throughput | 98 tok/s n=4 |
| Uptime | 100.00% |
| Model | p50 TTFT | Effective throughput | Uptime | Config excluded | Availability samples |
|---|---|---|---|---|---|
| openpipe/qwen3-14b-instruct | 854 ms | — | 100.00% | — | 2 |
| meta-llama/llama-3.3-70b-instruct | 1253 ms | — | 100.00% | — | 2 |
| meta-llama/llama-3.1-70b-instruct | 1256 ms | — | 100.00% | — | 1 |
| meta-llama/llama-3.1-8b-instruct | 1469 ms | — | 100.00% | — | 2 |
| deepseek/deepseek-v4-pro | 1515 ms | — | 100.00% | — | 3 |
| moonshotai/kimi-k2.7-code | 1628 ms | — | 100.00% | — | 2 |
| minimax/minimax-m3 | 1910 ms | 74 tok/s n=1 | 100.00% | — | 1 |
| qwen/qwen3.5-35b-a3b | 2298 ms | — | 100.00% | — | 3 |
| z-ai/glm-5.2 | 2448 ms | — | 100.00% | — | 3 |
| moonshotai/kimi-k2.6 | 2659 ms | — | 100.00% | — | 1 |
| qwen/qwen3-30b-a3b-instruct-2507 | 3062 ms | — | 100.00% | — | 4 |
| deepseek/deepseek-v4-flash | 3764 ms | 129 tok/s n=1 | 100.00% | — | 2 |
| qwen/qwen3.6-27b | 3857 ms | — | 100.00% | — | 4 |
| openai/gpt-oss-120b | 6322 ms | 32 tok/s n=1 | 100.00% | — | 1 |
| z-ai/glm-5.3-flash | — | — | 100.00% | — | 1 |
| google/gemma-4-31b-it | — | 122 tok/s n=1 | — | — | 0 |
Weights & Biases performance history · Full provider & model leaderboard.
Models served by Weights & Biases.
Each row links to pricing, provider, benchmark, and API pages for the model.
| Model | AI IQ | Context | Input | Cached input | Output |
|---|---|---|---|---|---|
deepseek/deepseek-v3.1DeepSeek V3.1 |
IQ 95#105 | 131,072 | $0.58025/1M | Not published | $1.74075/1M |
deepseek/deepseek-v4-flashDeepSeek: DeepSeek V4 Flash 0423 |
IQ 116#37 | 1,048,576 | $0.1477/1M | $0.07385/1M | $0.2954/1M |
deepseek/deepseek-v4-flash-0731DeepSeek: DeepSeek V4 Flash 0731 |
— | 1,048,576 | $0.13715/1M | $0.07385/1M | $0.2954/1M |
deepseek/deepseek-v4-proDeepSeek: DeepSeek V4 Pro 0423 |
IQ 114#47 | 1,048,576 | $1.21325/1M | $0.211/1M | $2.69025/1M |
google/gemma-4-31b-itGoogle: Gemma 4 31B |
IQ 103#80 | 262,144 | $0.1055/1M | Not published | $0.3587/1M |
ibm-granite/granite-4.1-8bgranite-4.1-8b |
— | 32,768 | $0.05275/1M | Not published | $0.1055/1M |
ibm-granite/granite-4.2-8bIBM: Granite 4.2 8B |
— | 131,072 | $0.1055/1M | $0.05275/1M | $0.15825/1M |
jetbrains/mellum2-12b-a2.5b-instructJetBrains Mellum2 12B A2.5B |
— | 131,072 | $0.05275/1M | Not published | $0.1055/1M |
meta-llama/llama-3.1-70b-instructMeta: Llama 3.1 70B Instruct |
— | 131,072 | $0.844/1M | Not published | $0.844/1M |
meta-llama/llama-3.1-8b-instructMeta: Llama 3.1 8B Instruct |
— | 131,072 | $0.2321/1M | Not published | $0.2321/1M |
meta-llama/llama-3.3-70b-instructMeta: Llama 3.3 70B Instruct |
— | 131,072 | $0.74905/1M | Not published | $0.74905/1M |
minimax/minimax-m3MiniMax: MiniMax M3 |
IQ 115#45 | 524,288 | $0.24265/1M | $0.05275/1M | $1.0128/1M |
moonshotai/kimi-k2.6MoonshotAI: Kimi K2.6 |
IQ 118#35 | 262,144 | $0.68575/1M | $0.15825/1M | $3.59755/1M |
moonshotai/kimi-k2.7-codeMoonshotAI: Kimi K2.7 Code |
IQ 108#66 | 262,144 | $0.74905/1M | $0.15825/1M | $3.6925/1M |
nvidia/nemotron-3-ultra-550b-a55bNVIDIA: Nemotron 3 Ultra |
— | 262,144 | $0.5275/1M | $0.1055/1M | $2.26825/1M |
nvidia/nemotron-3.5-lightning-30b-a3bnvidia/Nemotron-3.5-Lightning-30B-A3B |
— | 262,144 | $0.07385/1M | $0.0422/1M | $0.211/1M |
openai/gpt-oss-120bOpenAI: gpt-oss-120b |
IQ 98#99 | 131,072 | $0.03165/1M | Not published | $0.17935/1M |
openai/gpt-oss-20bOpenAI: gpt-oss-20b |
IQ 93#113 | 131,072 | $0.03165/1M | Not published | $0.13715/1M |
openpipe/qwen3-14b-instructOpenPipe Qwen3 14B Instruct |
— | 32,768 | $0.05275/1M | Not published | $0.2321/1M |
qwen/qwen3-30b-a3b-instruct-2507Qwen: Qwen3 30B A3B Instruct 2507 |
— | 262,144 | $0.1055/1M | Not published | $0.3165/1M |
qwen/qwen3.5-35b-a3bQwen: Qwen3.5-35B-A3B |
— | 262,144 | $0.26375/1M | Not published | $1.31875/1M |
qwen/qwen3.6-27bQwen: Qwen3.6 27B |
IQ 108#68 | 262,144 | $0.633/1M | $0.1266/1M | $3.798/1M |
qwen/qwen3.6-35b-a3bQwen: Qwen3.6 35B A3B |
IQ 100#94 | 262,144 | $0.26375/1M | Not published | $1.31875/1M |
qwen/qwen3.8-27bQwen: Qwen3.8 27B |
IQ 110#60 | 262,144 | $0.422/1M | $0.15825/1M | $3.165/1M |
z-ai/glm-5.2Z.ai: GLM 5.2 |
IQ 120#28 | 1,048,576 | $0.8018/1M | $0.1477/1M | $2.5531/1M |
z-ai/glm-5.3-flashZ.ai: GLM 5.3 Flash |
IQ 116#40 | 1,048,576 | $0.15825/1M | $0.05275/1M | $0.5275/1M |
Questions
Does Weights & Biases have zero data retention?
TrustedRouter does not currently mark Weights & Biases as provider-level zero data retention. Use trustedrouter/zdr or provider.min_privacy=zdr to select a different eligible route, and review the linked policy source for changes.
Is Weights & Biases end-to-end encrypted?
TrustedRouter does not currently mark Weights & Biases as end-to-end encrypted at the provider boundary. The TrustedRouter gateway is still attested, but the selected provider normally receives the request in order to run the model. Use trustedrouter/e2e for the stronger route requirement.
Which Weights & Biases models are available through TrustedRouter?
This page currently lists 26 public Weights & Biases models, with live pricing, route count, context length, measured performance when available, and links to each model's provider and benchmark pages.