← Back to Gists

OpenRouter AI Model Pricing - September 2026

📝 Markdown Rendered
retoor
retoor · Level 54802 ·

Complete OpenRouter pricing reference for Nemotron 3 Ultra, GLM 5.2, and GLM 5.3 Flash with provider comparisons, benchmarks, and performance data.

OpenRouter AI Model Pricing - September 2026

Compiled from openrouter.ai on September 9, 2026.


NVIDIA: Nemotron 3 Ultra (free)

Field Value
Slug nvidia/nemotron-3-ultra-550b-a55b:free
Price Free
Context 1,000,000 tokens
Max Output 65,536 tokens
Released June 4, 2026
Architecture 550B total / 55B active parameters (MoE), hybrid Transformer-Mamba
Modalities Text in / Text out
Tool Calling Yes (tools + tool_choice)
Structured Output No (response_format not supported)
Best For Long-running agentic workflows, agent orchestration, coding agents, deep research, complex enterprise tasks
Performance ~4 tok/s throughput, ~38.8s latency (P50), 96.44% uptime
Top Apps Hermes Agent, Kilo Code, Claude Code, Janitor AI, OpenClaw

Note: Free endpoints are rate limited. Do not upload confidential or personal data. Session data is logged for security and improvement purposes per NVIDIA's API Trial Terms of Service.


Z.ai: GLM 5.2

Field Value
Slug z-ai/glm-5.2
Input Price $0.4875 / 1M tokens (35% off list)
Output Price $1.56 / 1M tokens
Cache Read $0.091 / 1M tokens
Context 1,048,576 tokens (~1M)
Max Output 163,840 tokens
Released June 16, 2026
Modalities Text in / Text out
Reasoning Effort high, xhigh (maps to max reasoning)
Tool Calling Yes (tools + tool_choice)
Structured Output Yes (JSON schema via response_format)
Best For Long-horizon agent workflows, project-level software engineering, complex multi-step automation
Performance Up to 181 tok/s throughput, 0.41s latency (P50), 100% uptime
Providers 24 providers (DeepInfra, Ambient, StreamLake, NovitaAI, DigitalOcean, CoreWeave, Mistral, Z.ai, etc.)
Top Apps Claude Code, Hermes Agent, HighLevel, schema-markup-generation, pi

Provider Pricing Comparison (Standard tier)

Provider Input/M Output/M Cache/M Latency Throughput Uptime
DeepInfra (35% off) $0.4875 $1.56 $0.091 11.66s 48 tps 98.58%
Ambient $0.60 $2.00 $0.15 2.32s 20 tps 98.98%
StreamLake (52% off) $0.6762 $2.125 $0.125 61.52s 58 tps 99.71%
NovitaAI (51% off) $0.6832 $2.147 $0.126 61.79s 55 tps 99.92%
DigitalOcean $0.70 $2.20 $0.105 1.06s 65 tps 99.75%
CoreWeave $0.76 $2.42 $0.14 0.74s 81 tps 99.91%
AtlasCloud (33% off) $0.938 $2.948 $0.174 22.14s 41 tps 100.00%
Alibaba Cloud Int. $0.966 $3.036 $0.193 21.61s 49 tps 99.91%
SiliconFlow (15% off) $1.19 $3.74 $0.22 12.29s 44 tps 99.98%
Inceptron $1.25 $2.99 $0.22 0.73s 40 tps 99.05%
Phala $1.26 $3.00 $0.22 2.06s 31 tps 99.64%
Mistral (ZDR) $1.40 $4.40 $0.14 0.99s 101 tps 99.97%
Mistral $1.40 $4.40 $0.14 0.81s 116 tps 99.96%
Crusoe $1.40 $4.40 $0.26 1.20s 140 tps 99.98%
Together $1.40 $4.40 $0.26 0.64s 51 tps 96.85%
Fireworks $1.40 $4.40 $0.14 1.65s 60 tps 99.94%
Z.ai $1.40 $4.40 $0.26 4.32s 42 tps 99.93%

Z.ai: GLM 5.3 Flash

Field Value
Slug z-ai/glm-5.3-flash
Input Price $0.075 / 1M tokens (50% off promo through Sep 9, 2026)
Output Price $0.25 / 1M tokens
Cache Read $0.015 / 1M tokens
List Price $0.15 input / $0.50 output / $0.03 cache
Context 1,310,720 tokens (~1.3M)
Released August 26, 2026
Modalities Text, images, video in / Text out (native multimodal)
Architecture Hybrid sparse and linear attention
Tool Calling Yes (tools + tool_choice)
Structured Output Yes (response_format, no JSON-schema enforcement)
Best For Efficient coding, long-horizon agent tasks
Performance Up to 178 tok/s throughput, 0.45s latency (P50), 100% uptime
Providers 25 providers (GMICloud, DeepInfra, Z.ai, Relace, Wafer, Baseten, Modal, etc.)
Top Apps Hermes Agent, Claude Code, Cline, omp, pi
Stealth Origin Was previously known as "Ox Alpha" before reveal

Provider Pricing Comparison (Promo 50% off)

Provider Input/M Output/M Cache/M Latency Throughput Uptime
GMICloud (50% off) $0.075 $0.25 $0.015 5.11s 20 tps 95.57%
DeepInfra (50% off) $0.075 $0.25 $0.015 1.93s 29 tps 98.04%
Z.ai (50% off) $0.075 $0.25 $0.015 2.52s 45 tps 99.06%
Relace $0.09 $0.30 $0.018 1.01s 51 tps 98.63%
Wafer $0.10 $0.35 $0.02 0.52s 38 tps 99.45%
Modal (67% off) $0.15 $0.50 $0.03 0.45s 140 tps 99.87%
Baseten $0.15 $0.50 $0.03 0.63s 178 tps 99.84%
Makora $0.14 $0.47 $0.024 0.49s 107 tps 99.51%
Cloudflare $0.15 $0.50 $0.03 0.96s 57 tps 99.96%
CoreWeave $0.15 $0.50 $0.05 1.00s 56 tps 98.65%
Fireworks $0.15 $0.50 $0.03 1.22s 54 tps 99.74%
Friendli $0.15 $0.50 $0.03 6.25s 103 tps 99.69%
Parasail $0.15 $0.50 $0.03 2.76s 100 tps 95.24%

Summary Comparison

Model Input/M Output/M Context Max Out Price Tier Best For
Nemotron 3 Ultra (free) Free Free 1M 65K Free Agentic workflows, orchestration
GLM 5.3 Flash $0.075 $0.25 1.3M N/A Cheapest paid Coding, long-horizon agents
GLM 5.2 $0.4875 $1.56 1M 164K Mid-range Software engineering, automation

Data sourced from openrouter.ai model pages. Prices reflect promotional discounts available at time of capture. Provider availability and pricing may change.

Comments

No comments yet. Start the discussion.