Complete OpenRouter pricing reference for Nemotron 3 Ultra, GLM 5.2, and GLM 5.3 Flash with provider comparisons, benchmarks, and performance data.
OpenRouter AI Model Pricing - September 2026
Compiled from openrouter.ai on September 9, 2026.
NVIDIA: Nemotron 3 Ultra (free)
| Field |
Value |
| Slug |
nvidia/nemotron-3-ultra-550b-a55b:free |
| Price |
Free |
| Context |
1,000,000 tokens |
| Max Output |
65,536 tokens |
| Released |
June 4, 2026 |
| Architecture |
550B total / 55B active parameters (MoE), hybrid Transformer-Mamba |
| Modalities |
Text in / Text out |
| Tool Calling |
Yes (tools + tool_choice) |
| Structured Output |
No (response_format not supported) |
| Best For |
Long-running agentic workflows, agent orchestration, coding agents, deep research, complex enterprise tasks |
| Performance |
~4 tok/s throughput, ~38.8s latency (P50), 96.44% uptime |
| Top Apps |
Hermes Agent, Kilo Code, Claude Code, Janitor AI, OpenClaw |
Note: Free endpoints are rate limited. Do not upload confidential or personal data. Session data is logged for security and improvement purposes per NVIDIA's API Trial Terms of Service.
Z.ai: GLM 5.2
| Field |
Value |
| Slug |
z-ai/glm-5.2 |
| Input Price |
$0.4875 / 1M tokens (35% off list) |
| Output Price |
$1.56 / 1M tokens |
| Cache Read |
$0.091 / 1M tokens |
| Context |
1,048,576 tokens (~1M) |
| Max Output |
163,840 tokens |
| Released |
June 16, 2026 |
| Modalities |
Text in / Text out |
| Reasoning Effort |
high, xhigh (maps to max reasoning) |
| Tool Calling |
Yes (tools + tool_choice) |
| Structured Output |
Yes (JSON schema via response_format) |
| Best For |
Long-horizon agent workflows, project-level software engineering, complex multi-step automation |
| Performance |
Up to 181 tok/s throughput, 0.41s latency (P50), 100% uptime |
| Providers |
24 providers (DeepInfra, Ambient, StreamLake, NovitaAI, DigitalOcean, CoreWeave, Mistral, Z.ai, etc.) |
| Top Apps |
Claude Code, Hermes Agent, HighLevel, schema-markup-generation, pi |
Provider Pricing Comparison (Standard tier)
| Provider |
Input/M |
Output/M |
Cache/M |
Latency |
Throughput |
Uptime |
| DeepInfra (35% off) |
$0.4875 |
$1.56 |
$0.091 |
11.66s |
48 tps |
98.58% |
| Ambient |
$0.60 |
$2.00 |
$0.15 |
2.32s |
20 tps |
98.98% |
| StreamLake (52% off) |
$0.6762 |
$2.125 |
$0.125 |
61.52s |
58 tps |
99.71% |
| NovitaAI (51% off) |
$0.6832 |
$2.147 |
$0.126 |
61.79s |
55 tps |
99.92% |
| DigitalOcean |
$0.70 |
$2.20 |
$0.105 |
1.06s |
65 tps |
99.75% |
| CoreWeave |
$0.76 |
$2.42 |
$0.14 |
0.74s |
81 tps |
99.91% |
| AtlasCloud (33% off) |
$0.938 |
$2.948 |
$0.174 |
22.14s |
41 tps |
100.00% |
| Alibaba Cloud Int. |
$0.966 |
$3.036 |
$0.193 |
21.61s |
49 tps |
99.91% |
| SiliconFlow (15% off) |
$1.19 |
$3.74 |
$0.22 |
12.29s |
44 tps |
99.98% |
| Inceptron |
$1.25 |
$2.99 |
$0.22 |
0.73s |
40 tps |
99.05% |
| Phala |
$1.26 |
$3.00 |
$0.22 |
2.06s |
31 tps |
99.64% |
| Mistral (ZDR) |
$1.40 |
$4.40 |
$0.14 |
0.99s |
101 tps |
99.97% |
| Mistral |
$1.40 |
$4.40 |
$0.14 |
0.81s |
116 tps |
99.96% |
| Crusoe |
$1.40 |
$4.40 |
$0.26 |
1.20s |
140 tps |
99.98% |
| Together |
$1.40 |
$4.40 |
$0.26 |
0.64s |
51 tps |
96.85% |
| Fireworks |
$1.40 |
$4.40 |
$0.14 |
1.65s |
60 tps |
99.94% |
| Z.ai |
$1.40 |
$4.40 |
$0.26 |
4.32s |
42 tps |
99.93% |
Z.ai: GLM 5.3 Flash
| Field |
Value |
| Slug |
z-ai/glm-5.3-flash |
| Input Price |
$0.075 / 1M tokens (50% off promo through Sep 9, 2026) |
| Output Price |
$0.25 / 1M tokens |
| Cache Read |
$0.015 / 1M tokens |
| List Price |
$0.15 input / $0.50 output / $0.03 cache |
| Context |
1,310,720 tokens (~1.3M) |
| Released |
August 26, 2026 |
| Modalities |
Text, images, video in / Text out (native multimodal) |
| Architecture |
Hybrid sparse and linear attention |
| Tool Calling |
Yes (tools + tool_choice) |
| Structured Output |
Yes (response_format, no JSON-schema enforcement) |
| Best For |
Efficient coding, long-horizon agent tasks |
| Performance |
Up to 178 tok/s throughput, 0.45s latency (P50), 100% uptime |
| Providers |
25 providers (GMICloud, DeepInfra, Z.ai, Relace, Wafer, Baseten, Modal, etc.) |
| Top Apps |
Hermes Agent, Claude Code, Cline, omp, pi |
| Stealth Origin |
Was previously known as "Ox Alpha" before reveal |
Provider Pricing Comparison (Promo 50% off)
| Provider |
Input/M |
Output/M |
Cache/M |
Latency |
Throughput |
Uptime |
| GMICloud (50% off) |
$0.075 |
$0.25 |
$0.015 |
5.11s |
20 tps |
95.57% |
| DeepInfra (50% off) |
$0.075 |
$0.25 |
$0.015 |
1.93s |
29 tps |
98.04% |
| Z.ai (50% off) |
$0.075 |
$0.25 |
$0.015 |
2.52s |
45 tps |
99.06% |
| Relace |
$0.09 |
$0.30 |
$0.018 |
1.01s |
51 tps |
98.63% |
| Wafer |
$0.10 |
$0.35 |
$0.02 |
0.52s |
38 tps |
99.45% |
| Modal (67% off) |
$0.15 |
$0.50 |
$0.03 |
0.45s |
140 tps |
99.87% |
| Baseten |
$0.15 |
$0.50 |
$0.03 |
0.63s |
178 tps |
99.84% |
| Makora |
$0.14 |
$0.47 |
$0.024 |
0.49s |
107 tps |
99.51% |
| Cloudflare |
$0.15 |
$0.50 |
$0.03 |
0.96s |
57 tps |
99.96% |
| CoreWeave |
$0.15 |
$0.50 |
$0.05 |
1.00s |
56 tps |
98.65% |
| Fireworks |
$0.15 |
$0.50 |
$0.03 |
1.22s |
54 tps |
99.74% |
| Friendli |
$0.15 |
$0.50 |
$0.03 |
6.25s |
103 tps |
99.69% |
| Parasail |
$0.15 |
$0.50 |
$0.03 |
2.76s |
100 tps |
95.24% |
Summary Comparison
| Model |
Input/M |
Output/M |
Context |
Max Out |
Price Tier |
Best For |
| Nemotron 3 Ultra (free) |
Free |
Free |
1M |
65K |
Free |
Agentic workflows, orchestration |
| GLM 5.3 Flash |
$0.075 |
$0.25 |
1.3M |
N/A |
Cheapest paid |
Coding, long-horizon agents |
| GLM 5.2 |
$0.4875 |
$1.56 |
1M |
164K |
Mid-range |
Software engineering, automation |
Data sourced from openrouter.ai model pages. Prices reflect promotional discounts available at time of capture. Provider availability and pricing may change.
Comments
No comments yet. Start the discussion.