Novita AI
Explore Novita AI models on TrustedRouter with current routes, token pricing, policy sources, privacy posture, regional availability, and API support.
Novita AInovita
No provider claimThese privacy labels describe Novita AI, the upstream model provider. ZDR is a retention policy; verified confidential inference additionally requires attested provider compute and end-to-end encryption.
| Provider | Novita AI |
|---|---|
| Routing status | Active |
| Provider website | https://novita.ai/ |
| Models | 93 public models |
| Credits routes | 93 |
| Zero data retention | not claimed |
| Verified confidential inference | Not verified |
| Policy note | No provider-ZDR claim is tracked here. Novita's privacy policy says personal information is not used for model training; customer-content processing is governed by customer agreements. Policy source |
Measured performance
30 samplesContinuously sampled across Novita AI's routed models: p50 TTFT, effective throughput, and success rate. Effective throughput uses provider-reported output tokens over complete request time. Unsupported route and probe-configuration rows are separated from provider downtime. No prompt or output content stored.
| p50 TTFT | 2004 ms |
|---|---|
| Effective throughput | 78 tok/s n=9 |
| Uptime | 83.33% |
| Model | p50 TTFT | Effective throughput | Uptime | Config excluded | Availability samples |
|---|---|---|---|---|---|
| deepseek/deepseek-ocr-2 | 703 ms | — | 100.00% | — | 4 |
| sao10k/l31-70b-euryale-v2.2 | 1207 ms | — | 100.00% | — | 1 |
| moonshotai/kimi-k2.6 | 1426 ms | 29 tok/s n=1 | 100.00% | — | 2 |
| zai-org/autoglm-phone-9b-multilingual | 1430 ms | — | 100.00% | — | 1 |
| qwen/qwen3-next-80b-a3b-instruct | 1573 ms | — | 100.00% | — | 1 |
| zai-org/glm-4.6v | 1584 ms | — | 100.00% | — | 1 |
| qwen/qwen3.8-27b | 1668 ms | — | 100.00% | — | 1 |
| deepseek/deepseek-v4.1-flash | 1863 ms | — | 100.00% | — | 1 |
| meta-llama/llama-3.3-70b-instruct | 2004 ms | — | 100.00% | — | 1 |
| nvidia/nemotron-3-nano-30b-a3b | 2022 ms | — | 100.00% | — | 1 |
| qwen/qwen3-vl-30b-a3b-instruct | 2117 ms | — | 100.00% | — | 1 |
| minimax/minimax-m3 | 2189 ms | 87 tok/s n=2 | 100.00% | — | 1 |
| qwen/qwen3.8-max | 2466 ms | — | 100.00% | — | 2 |
| z-ai/glm-5.2 | 2827 ms | — | 100.00% | — | 1 |
| meta-llama/llama-4-scout-17b-16e-instruct | 3040 ms | — | 100.00% | — | 1 |
| minimax/minimax-m2.7 | 3817 ms | — | 100.00% | — | 1 |
| zai-org/glm-4.7 | 4556 ms | — | 100.00% | — | 1 |
| baidu/ernie-4.5-vl-424b-a47b | 5075 ms | — | 100.00% | — | 1 |
| deepseek/deepseek-v4-pro | 5524 ms | — | 100.00% | — | 1 |
| openai/gpt-oss-120b | 22168 ms | 6 tok/s n=1 | 100.00% | — | 1 |
| baichuan/baichuan-m2-32b | — | — | 0.00% | — | 2 |
| baidu/ernie-4.5-21B-a3b | — | — | 0.00% | — | 1 |
| deepseek/deepseek-v4-flash | — | 95 tok/s n=1 | — | — | 0 |
| google/gemma-4-31b-it | — | 78 tok/s n=2 | — | — | 0 |
| mindai/macaron-v1-tall | — | — | 0.00% | — | 1 |
| moonshotai/kimi-k3 | — | 45 tok/s n=1 | — | — | 0 |
| qwen/qwen3-235b-a22b-thinking-2507 | — | 101 tok/s n=1 | — | — | 0 |
| zai-org/glm-4.7-flash | — | — | 0.00% | — | 1 |
Novita AI performance history · Full provider & model leaderboard.
Models served by Novita AI.
Each row links to pricing, provider, benchmark, and API pages for the model.
| Model | AI IQ | Context | Input | Cached input | Output |
|---|---|---|---|---|---|
Sao10K/L3-8B-Stheno-v3.2Llama 3 8B Stheno V3.2 |
— | 8,192 | $0.05275/1M | Not published | $0.05275/1M |
baichuan/baichuan-m2-32bBaichuan M2 32B |
— | 131,072 | $0.07385/1M | Not published | $0.07385/1M |
baidu/cobuddyCoBuddy |
— | 131,072 | $0.2954/1M | $0.07385/1M | $1.19215/1M |
baidu/ernie-4.5-21B-a3bERNIE 4.5 21B A3B |
— | 120,000 | $0.07385/1M | Not published | $0.2954/1M |
baidu/ernie-4.5-vl-424b-a47bBaidu: ERNIE 4.5 VL 424B A47B |
— | 123,000 | $0.4431/1M | Not published | $1.31875/1M |
deepseek/deepseek-ocr-2DeepSeek OCR 2 |
— | 8,192 | $0.03165/1M | Not published | $0.03165/1M |
deepseek/deepseek-r1-0528DeepSeek: R1 0528 |
— | 163,840 | $0.7385/1M | $0.36925/1M | $2.6375/1M |
deepseek/deepseek-r1-distill-llama-70bDeepSeek: R1 Distill Llama 70B |
— | 8,192 | $0.844/1M | Not published | $0.844/1M |
deepseek/deepseek-r1-turboDeepSeek R1 Turbo |
— | 64,000 | $0.7385/1M | Not published | $2.6375/1M |
deepseek/deepseek-v3.1DeepSeek V3.1 |
IQ 95#105 | 131,072 | $0.28485/1M | $0.142425/1M | $1.055/1M |
deepseek/deepseek-v3.1-terminusDeepSeek: DeepSeek V3.1 Terminus |
— | 131,072 | $0.28485/1M | $0.142425/1M | $1.055/1M |
deepseek/deepseek-v3.2DeepSeek: DeepSeek V3.2 |
IQ 103#77 | 163,840 | $0.283795/1M | $0.141898/1M | $0.422/1M |
deepseek/deepseek-v3.2-expDeepSeek: DeepSeek V3.2 Exp |
— | 163,840 | $0.28485/1M | Not published | $0.43255/1M |
deepseek/deepseek-v4-flashDeepSeek: DeepSeek V4 Flash 0423 |
IQ 116#37 | 1,048,576 | $0.1477/1M | $0.02954/1M | $0.2954/1M |
deepseek/deepseek-v4-flash-0731DeepSeek: DeepSeek V4 Flash 0731 |
— | 1,048,576 | $0.4642/1M | $0.02954/1M | $1.3926/1M |
deepseek/deepseek-v4-flash-vision-expDeepSeek: DeepSeek V4 Flash Vision Exp |
IQ 115#42 | 1,048,576 | $0.4642/1M | $0.02954/1M | $1.3926/1M |
deepseek/deepseek-v4-proDeepSeek: DeepSeek V4 Pro 0423 |
IQ 114#47 | 1,048,576 | $1.688/1M | $0.142425/1M | $3.376/1M |
deepseek/deepseek-v4.1-flashDeepSeek: DeepSeek V4.1 Flash |
IQ 116#38 | 1,048,576 | $0.3165/1M | $0.01/1M | $1.266/1M |
google/gemma-3-12b-itGoogle: Gemma 3 12B |
— | 131,072 | $0.05275/1M | Not published | $0.1055/1M |
google/gemma-3-27b-itGoogle: Gemma 3 27B |
— | 131,072 | $0.125545/1M | Not published | $0.211/1M |
google/gemma-4-26b-a4b-itGoogle: Gemma 4 26B A4B |
IQ 100#89 | 262,144 | $0.13715/1M | Not published | $0.422/1M |
google/gemma-4-31b-itGoogle: Gemma 4 31B |
IQ 103#80 | 262,144 | $0.1477/1M | Not published | $0.422/1M |
inclusionai/ling-3.0-flashinclusionAI: Ling 3.0 Flash |
IQ 94#109 | 131,072 | $0.0633/1M | $0.01266/1M | $0.1899/1M |
meta-llama/llama-3.1-8b-instructMeta: Llama 3.1 8B Instruct |
— | 131,072 | $0.0211/1M | Not published | $0.05275/1M |
meta-llama/llama-3.3-70b-instructMeta: Llama 3.3 70B Instruct |
— | 131,072 | $0.142425/1M | Not published | $0.422/1M |
meta-llama/llama-4-maverick-17b-128e-instruct-fp8Llama 4 Maverick Instruct |
— | 1,048,576 | $0.28485/1M | Not published | $0.89675/1M |
meta-llama/llama-4-scout-17b-16e-instructLlama 4 Scout Instruct |
— | 131,072 | $0.1899/1M | Not published | $0.62245/1M |
microsoft/wizardlm-2-8x22bWizardLM-2 8x22B |
— | 65,535 | $0.6541/1M | Not published | $0.6541/1M |
mindai/macaron-v1-tallMacaron V1 Tall |
— | 262,144 | $0.261113/1M | $0.04642/1M | $1.50865/1M |
mindai/macaron-v1-ventiMacaron V1 Venti |
— | 1,048,576 | $0.870375/1M | $0.174075/1M | $2.611125/1M |
minimax/minimax-m2MiniMax: MiniMax M2 |
— | 204,800 | $0.3165/1M | $0.03165/1M | $1.266/1M |
minimax/minimax-m2.1MiniMax: MiniMax M2.1 |
IQ 101#87 | 204,800 | $0.3165/1M | $0.03165/1M | $1.266/1M |
minimax/minimax-m2.5MiniMax: MiniMax M2.5 |
IQ 106#74 | 196,608 | $0.3165/1M | $0.03165/1M | $1.266/1M |
minimax/minimax-m2.5-highspeedMiniMax M2.5 Highspeed |
— | 204,800 | $0.633/1M | $0.03165/1M | $2.532/1M |
minimax/minimax-m2.7MiniMax: MiniMax M2.7 |
IQ 107#72 | 196,608 | $0.3165/1M | $0.0633/1M | $1.266/1M |
minimax/minimax-m3MiniMax: MiniMax M3 |
IQ 115#45 | 524,288 | $0.3165/1M | $0.0633/1M | $1.266/1M |
minimaxai/minimax-m1-80kMiniMax M1 |
— | 1,000,000 | $0.58025/1M | Not published | $2.321/1M |
mistralai/mistral-nemoMistral: Mistral Nemo |
— | 131,072 | $0.0422/1M | Not published | $0.17935/1M |
moonshotai/kimi-k2-0905MoonshotAI: Kimi K2 0905 |
— | 262,144 | $0.633/1M | Not published | $2.6375/1M |
moonshotai/kimi-k2-instructKimi K2 Instruct |
IQ 91#116 | 131,072 | $0.60135/1M | Not published | $2.4265/1M |
moonshotai/kimi-k2-thinkingMoonshotAI: Kimi K2 Thinking |
IQ 91#116 | 262,144 | $0.633/1M | $0.15825/1M | $2.6375/1M |
moonshotai/kimi-k2.5MoonshotAI: Kimi K2.5 |
IQ 111#57 | 262,144 | $0.633/1M | $0.1055/1M | $3.165/1M |
moonshotai/kimi-k2.6MoonshotAI: Kimi K2.6 |
IQ 118#35 | 262,144 | $0.844/1M | $0.1688/1M | $3.587/1M |
moonshotai/kimi-k2.7-codeMoonshotAI: Kimi K2.7 Code |
IQ 108#66 | 262,144 | $1.00225/1M | $0.20045/1M | $4.22/1M |
moonshotai/kimi-k3MoonshotAI: Kimi K3 |
IQ 121#27 | 1,048,576 | $3.165/1M | $0.3165/1M | $15.825/1M |
nvidia/nemotron-3-nano-30b-a3bNVIDIA: Nemotron 3 Nano 30B A3B |
— | 262,144 | $0.05275/1M | Not published | $0.211/1M |
openai/gpt-oss-120bOpenAI: gpt-oss-120b |
IQ 98#99 | 131,072 | $0.05275/1M | Not published | $0.26375/1M |
openai/gpt-oss-20bOpenAI: gpt-oss-20b |
IQ 93#113 | 131,072 | $0.0422/1M | Not published | $0.15825/1M |
qwen/qwen-2.5-72b-instructQwen2.5 72B Instruct |
— | 32,000 | $0.4009/1M | Not published | $0.422/1M |
qwen/qwen-mt-plusQwen MT Plus |
— | 16,384 | $0.26375/1M | Not published | $0.79125/1M |
qwen/qwen3-235b-a22b-fp8Qwen3 235B A22B |
— | 40,960 | $0.211/1M | Not published | $0.844/1M |
qwen/qwen3-235b-a22b-instruct-2507Qwen3 235B A22B Instruct 2507 |
— | 131,072 | $0.09495/1M | Not published | $0.6119/1M |
qwen/qwen3-235b-a22b-thinking-2507Qwen: Qwen3 235B A22B Thinking 2507 |
— | 131,072 | $0.3165/1M | Not published | $3.165/1M |
qwen/qwen3-coder-30b-a3b-instructQwen: Qwen3 Coder 30B A3B Instruct |
— | 160,000 | $0.07385/1M | Not published | $0.28485/1M |
qwen/qwen3-coder-480b-a35b-instructQwen3 Coder 480B A35B Instruct |
— | 262,144 | $0.4009/1M | Not published | $1.63525/1M |
qwen/qwen3-coder-nextQwen: Qwen3 Coder Next |
— | 262,144 | $0.211/1M | Not published | $1.5825/1M |
qwen/qwen3-maxQwen3 Max |
— | 262,144 | $2.22605/1M | Not published | $8.91475/1M |
qwen/qwen3-next-80b-a3b-instructQwen: Qwen3 Next 80B A3B Instruct |
— | 262,144 | $0.15825/1M | Not published | $1.5825/1M |
qwen/qwen3-omni-30b-a3b-instructQwen3 Omni 30B A3B Instruct |
— | 65,536 | $0.26375/1M | Not published | $1.02335/1M |
qwen/qwen3-omni-30b-a3b-thinkingQwen3 Omni 30B A3B Thinking |
— | 65,536 | $0.26375/1M | Not published | $1.02335/1M |
qwen/qwen3-vl-235b-a22b-instructQwen: Qwen3 VL 235B A22B Instruct |
— | 262,144 | $0.3165/1M | Not published | $1.5825/1M |
qwen/qwen3-vl-235b-a22b-thinkingQwen: Qwen3 VL 235B A22B Thinking |
— | 131,072 | $1.0339/1M | Not published | $4.16725/1M |
qwen/qwen3-vl-30b-a3b-instructQwen: Qwen3 VL 30B A3B Instruct |
— | 262,144 | $0.211/1M | Not published | $0.7385/1M |
qwen/qwen3.5-122b-a10bQwen: Qwen3.5-122B-A10B |
— | 262,144 | $0.422/1M | Not published | $3.376/1M |
qwen/qwen3.5-27bQwen: Qwen3.5-27B |
— | 262,144 | $0.3165/1M | Not published | $2.532/1M |
qwen/qwen3.5-35b-a3bQwen: Qwen3.5-35B-A3B |
— | 262,144 | $0.26375/1M | Not published | $2.11/1M |
qwen/qwen3.5-397b-a17bQwen: Qwen3.5 397B A17B |
— | 262,144 | $0.633/1M | Not published | $3.798/1M |
qwen/qwen3.6-27bQwen: Qwen3.6 27B |
IQ 108#68 | 262,144 | $0.633/1M | Not published | $3.798/1M |
qwen/qwen3.6-35b-a3bQwen: Qwen3.6 35B A3B |
IQ 100#94 | 262,144 | $0.26164/1M | Not published | $1.566675/1M |
qwen/qwen3.7-maxQwen3.7 Max |
IQ 119#31 | 1,000,000 | $1.31875/1M | $0.26375/1M | $3.95625/1M |
qwen/qwen3.8-2.4t-a95bQwen: Qwen3.8 2.4T A95B |
IQ 122#24 | 262,144 | $2.11/1M | $0.26375/1M | $6.33/1M |
qwen/qwen3.8-27bQwen: Qwen3.8 27B |
IQ 110#60 | 262,144 | $0.4431/1M | $0.089675/1M | $3.165/1M |
qwen/qwen3.8-flashQwen3.8 Flash |
IQ 115#46 | 1,000,000 | $0.15825/1M | $0.01688/1M | $0.49585/1M |
qwen/qwen3.8-maxQwen3.8 Max |
IQ 122#25 | 1,000,000 | $2.11/1M | $0.26375/1M | $6.33/1M |
sao10k/l3-8b-lunarisLlama 3 8B Lunaris |
— | 8,192 | $0.05275/1M | Not published | $0.05275/1M |
sao10k/l31-70b-euryale-v2.2Llama 3.1 70B Euryale V2.2 |
— | 8,192 | $1.5614/1M | Not published | $1.5614/1M |
stepfun/step-3.7-flashStepFun: Step 3.7 Flash |
IQ 102#84 | 262,144 | $0.211/1M | $0.0422/1M | $1.21325/1M |
tencent/hy3Tencent: Hy3 |
IQ 100#90 | 262,144 | $0.1477/1M | $0.036925/1M | $0.6119/1M |
xiaomimimo/mimo-v2.5Xiaomi MiMo V2.5 |
IQ 108#67 | 1,048,576 | $0.17724/1M | $0.01/1M | $0.35448/1M |
xiaomimimo/mimo-v2.5-proXiaomi MiMo V2.5 Pro |
IQ 112#55 | 1,048,576 | $0.55071/1M | $0.01/1M | $1.10142/1M |
z-ai/glm-5.2Z.ai: GLM 5.2 |
IQ 120#28 | 1,048,576 | $1.477/1M | $0.2743/1M | $4.642/1M |
z-ai/glm-5.3Z.ai: GLM 5.3 |
IQ 123#18 | 1,048,576 | $1.477/1M | $0.2743/1M | $4.642/1M |
z-ai/glm-5.3-flashZ.ai: GLM 5.3 Flash |
IQ 116#40 | 1,048,576 | $0.15825/1M | $0.03165/1M | $0.5275/1M |
z-ai/glm-5.3-pzai-org/glm-5.3-p |
— | 1,048,576 | $1.477/1M | $0.2743/1M | $4.642/1M |
zai-org/autoglm-phone-9b-multilingualAutoGLM Phone 9B Multilingual |
— | 65,536 | $0.036925/1M | Not published | $0.14559/1M |
zai-org/glm-4.5-airGLM 4.5 Air |
— | 131,072 | $0.13715/1M | $0.026375/1M | $0.89675/1M |
zai-org/glm-4.5vGLM 4.5V |
— | 65,536 | $0.633/1M | $0.11605/1M | $1.899/1M |
zai-org/glm-4.6GLM 4.6 |
— | 204,800 | $0.58025/1M | $0.11605/1M | $2.321/1M |
zai-org/glm-4.6vGLM 4.6V |
— | 131,072 | $0.3165/1M | $0.058025/1M | $0.9495/1M |
zai-org/glm-4.7GLM 4.7 |
IQ 103#81 | 204,800 | $0.633/1M | $0.11605/1M | $2.321/1M |
zai-org/glm-4.7-flashGLM 4.7 Flash |
— | 200,000 | $0.07385/1M | $0.01055/1M | $0.422/1M |
zai-org/glm-5GLM 5 |
IQ 105#75 | 202,800 | $1.055/1M | $0.211/1M | $3.376/1M |
zai-org/glm-5.1GLM 5.1 |
IQ 113#49 | 204,800 | $1.4559/1M | $0.2743/1M | $4.642/1M |
Questions
Does Novita AI have zero data retention?
TrustedRouter does not currently mark Novita AI as provider-level zero data retention. Use trustedrouter/zdr or provider.min_privacy=zdr to select a different eligible route, and review the linked policy source for changes.
Is Novita AI end-to-end encrypted?
TrustedRouter does not currently mark Novita AI as end-to-end encrypted at the provider boundary. The TrustedRouter gateway is still attested, but the selected provider normally receives the request in order to run the model. Use trustedrouter/e2e for the stronger route requirement.
Which Novita AI models are available through TrustedRouter?
This page currently lists 93 public Novita AI models, with live pricing, route count, context length, measured performance when available, and links to each model's provider and benchmark pages.