GPT-5.6 Luna on Foundry: PTU Sizing, PayGo vs. PTU + Spillover Pricing
A quick note before we start: While this article focuses on GPT-5.6 Luna to make the pricing and PTU calculations concrete, the same methodology applies to other models when their model-specific throughput and pricing values are substituted. Provisioned Throughput provides a dedicated, fixed amount of processing capacity exclusively for your model deployment. Unlike Standard/PayGo, it provides a model-specific latency SLA, and its capacity is not shared across tenants. - PTU is a good fit for predictable, sustained traffic with consistent latency and high-throughput requirements. - PTU quota is model-independent, so the same quota pool can be allocated across supported models. Throughput per PTU remains model- and version-specific. - PTU quota is granted per subscription, region, and deployment type. Quota in East US does not carry over to West Europe, and Global Provisioned quota does not carry over to Data Zone Provisioned. PTU Sizing and Estimation | Input | Description | |---|---| | Model and Version | The model determines which Input TPM per PTU and output-to-input ratio values to use. Each model has a minimum PTU count and specific PTU throughput. | | Deployment type | The provisioned deployment type: Global Provisioned, Data Zone Provisioned, or Regional Provisioned. | | Peak RPM | The expected peak number of calls per minute sent to the model. | | Average prompt size | The average number of input tokens per request. | | Average response size | The average number of output tokens per request. | | Cache rate | The percentage of input tokens served from the prompt cache. Cached tokens don't consume any PTU capacity. | | FORMULAS | |---| Input TPM = Peak RPM ร Average input tokens per request | Output TPM = Peak RPM ร Average output tokens per request | Effective Input TPM = Input TPM ร (1 - Cache rate ) | Normalized TPM = Effective Input TPM + (Output-to-input ratio ร Output TPM ) | Estimated PTUs = Normalized TPM / Input TPM per PTU | Note: Input TPM is the workload-specific calculated volume, whereasInput TPM per PTU is a model-specific sizing constant. For example, some listedInput TPM per PTU values are 30,000 for GPT-5.6 Luna, 3,000 for GPT-5.6 Terra, and 1,200 for GPT-5.6 Sol. Sample Pricing Calculations for GPT-5.6 Luna on Microsoft Foundry Representative sample: Let's suppose your application sends requests at a peak rate of 1,000 RPM, with an average prompt size of 1,200 tokens and an average response size of 200 tokens, using the gpt-5.6-luna model with a Global Provisioned deployment. Based on the Microsoft Foundry PTU sizing table, gpt-5.6-luna has these constants: | GPT-5.6 LUNA SIZING CONSTANT | VALUE | |---|---| | Input TPM per PTU | 30,000 | | Output-to-input ratio | 6 | | Minimum Global Provisioned deployment | 15 PTUs | | Global Provisioned scale increment | 5 PTUs | 1. Without Prompt Caching | CALCULATION | |---| Input TPM = 1,000 ร 1,200 = 1,200,000 | Output TPM = 1,000 ร 200 = 200,000 | Normalized TPM = 1,200,000 + (6 ร 200,000) = 2,400,000 | Estimated PTUs = 2,400,000 / 30,000 = 80 PTUs | PTUs deployed = 80 PTUs | 2. With a 50% Prompt-Cache Hit Rate | CALCULATION | |---| Input TPM = 1,000 ร 1,200 = 1,200,000 | Effective Input TPM = 1,200,000 ร (1 - 0.50) = 600,000 | Output TPM = 1,000 ร 200 = 200,000 | Normalized TPM = 600,000 + (6 ร 200,000) = 1,800,000 | Estimated PTUs = 1,800,000 / 30,000 = 60 PTUs | PTUs deployed = 60 PTUs | | Peak RPM | Prompt size | Response size | Cache rate | Effective Input TPM | Output TPM | Normalized TPM | Estimated PTUs | PTUs deployed | |---|---|---|---|---|---|---|---|---| | 1,000 | 1,200 | 200 | 0% | 1,200,000 | 200,000 | 2,400,000 | 80.00 | 80 | | 1,000 | 1,200 | 200 | 50% | 600,000 | 200,000 | 1,800,000 | 60.00 | 60 | Notes and Remarks - For this simplified example, each representative prompt is assumed to contain 1,200 tokens. This exceeds the 1,024-token minimum for prompt caching. A cache hit also requires at least the first 1,024 tokens to be identical across requests. - Prompt caching is also available for Provisioned Throughput deployments. Cached input tokens don't consume PTU capacity. In this example, caching reduces the calculated requirement from 80 PTUs to 60 PTUs-a reduction of 20 PTUs (25%). Both values are above the 15-PTU minimum and divisible by the 5-PTU scale increment. - No additional rounding is required in this example. Rounding would be required if an estimate fell below the 15-PTU minimum or between supported 5-PTU increments; for example, 62.4 PTUs would be rounded up to 65 PTUs. Handling Spiky Traffic Illustrative 24-Hour RPM Profile and 30-Day Estimate Representative sample: Let's suppose traffic fluctuates between 0 and 2,500 RPM over a typical 24-hour period, with an average input size of 1,200 tokens and an average response size of 200 tokens, using the gpt-5.6-luna model with a Global Standard (pay-as-you-go) deployment. For the 30-day estimate, this daily traffic profile is assumed to repeat every day. We'll also assume that 50% of input tokens are cache reads and the remaining 50% are cache misses. Of all input tokens, 10 percentage points are cache writes, leaving 40 percentage points as regular input that is neither read from nor written to the cache. Prompt caching applies only to input tokens; output tokens are always charged at the regular output-token rate. Illustrative RPM distribution over 24 hours (RPM changes every 4 hours) 2500 โค โโโโโโ 2250 โค โโโโโโ 2000 โค โโโโโโ โโโโโโ 1750 โค โโโโโโ โโโโโโ 1500 โค โโโโโโ โโโโโโ 1250 โค โโโโโโ โโโโโโ 1000 โค โโโโโโ โโโโโโ โโโโโโ โโโโโโ 750 โค โโโโโโ โโโโโโ โโโโโโ โโโโโโ 500 โค โโโโโโ โโโโโโ โโโโโโ โโโโโโ โโโโโโ 250 โค โโโโโโ โโโโโโ โโโโโโ โโโโโโ โโโโโโ 0 โผโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ HOUR โ 00-04 โ 04-08 โ 08-12 โ 12-16 โ 16-20 โ 20-24 RPM โ 0 โ 500 โ 1,000 โ 2,500 โ 2,000 โ 1,000 1. PayGo-Only Cost Calculation For Standard/PayGo rates, as of August 27, 2026, the Azure OpenAI pricing page lists these USD rates for gpt-5.6-luna Global Standard: | Meter | Symbol | Price per 1M tokens | |---|---|---| | Regular input | P_regular | $0.20 | | Cached input | P_cached | $0.02 | | Cache writes | P_write | $0.25 | | Output | P_output | $1.20 | 50% cache reads + 10% cache writes + 40% regular input = 100% | Symbol | Description | |---|---| RPM | Requests per minute | a.i.T , a.o.T | Average input and output tokens per request | Hours | Duration of the batch in hours | r_cached , r_write | Cache-read and cache-write rates | C_batch | Total input-and-output token cost for one batch | | Formula | Calculation | |---|---| | Total input tokens | T_input = RPM ร a.i.T ร 60 ร Hours | | Cached input tokens | T_cached = T_input ร r_cached | | Cache-write input tokens | T_write = T_input ร r_write | | Regular input tokens | T_regular = T_input - T_cached - T_write | | Total output tokens | T_output = RPM ร a.o.T ร 60 ร Hours | | Batch cost | C_batch = (T_regular ยทP_regular + T_cached ยทP_cached + T_write ยทP_write + T_output ยทP_output ) / 1,000,000 | Each column represents one 4-hour batch on a typical day. The daily traffic profile is assumed to repeat itself for 30 days. Token values are shown in billions (B). | Metric | 00-04 | 04-08 | 08-12 | 12-16 | 16-20 | 20-24 | Total per day | |---|---|---|---|---|---|---|---| | Batch duration | 4 hours | 4 hours | 4 hours | 4 hours | 4 hours | 4 hours | 24 hours | | RPM | 0 | 500 | 1,000 | 2,500 | 2,000 | 1,000 | - | T_regular | 0 | 0.05760B | 0.11520B | 0.28800B | 0.23040B | 0.11520B | 0.80640B | T_cached | 0 | 0.07200B | 0.14400B | 0.36000B | 0.28800B | 0.14400B | 1.00800B | T_write | 0 | 0.01440B | 0.02880B | 0.07200B | 0.05760B | 0.02880B | 0.20160B | T_output | 0 | 0.02400B | 0.04800B | 0.12000B | 0.09600B | 0.04800B | 0.33600B | C_4-hour batch | $0.00 | $45.36 | $90.72 | $226.80 | $181.44 | $90.72 | $635.04 | 30-day token totals: - Regular input ( T_regular ): 24.19200B- Cached input ( T_cached ): 30.24000B- Cache writes ( T_write ): 6.04800B- Output ( T_output ): 10.08000BEstimated pay-as-you-go cost per day: $635.04. Estimated cost per 30-day month: $19,051.20. Estimated cost per 365-day year: $231,789.60. 2. PTU + Spillover to PayGo Cost Calculation The following Sweden Central PTU rates were retrieved on August 27, 2026, using the Azure Retail Prices API. Sample queries and commands are included in the Appendix as a reference. | Tier | Retail API rate | Monthly equivalent per PTU | |---|---|---| | Hourly PTU | $1.00/PTU/hour | $720.00 | | Monthly reservation | $260.00/PTU/month | $260.00 | | One-year reservation | $2,652.00/PTU/year | $221.00 | The hourly PTU monthly equivalent assumes 720 hours (30 days). Warning: Hourly, non-reserved PTU is best suited to temporary or uncertain workloads, such as testing, benchmarking, capacity validation, short pilots, or migration exercises. 2.A Provisioned Baseline of 250 RPM Let's take 250 RPM as the provisioned baseline and use Standard/PayGo spillover for bursts above it. | CALCULATION | |---| Input TPM = 250 ร 1,200 = 300,000 | Effective Input TPM = 300,000 ร (1 - 0.50) = 150,000 | Output TPM = 250 ร 200 = 50,000 | Normalized TPM = 150,000 + (6 ร 50,000) = 450,000 | Estimated PTUs = 450,000 / 30,000 = 15 PTUs | Provisioned baseline = 15 PTUs | | Time | Incoming RPM | Normalized TPM Demand | 15-PTU Capacity (Normalized TPM) | Potential Spillover Demand | |---|---|---|---|---| | 00-04 | 0 | 0 | 450,000 | 0 | | 04-08 | 500 | 900,000 | 450,000 | 450,000 | | 08-12 | 1,000 | 1,800,000 | 450,000 | 1,350,000 | | 12-16 | 2,500 | 4,500,000 | 450,000 | 4,050,000 | | 16-20 | 2,000 | 3,600,000 | 450,000 | 3,150,000 | | 20-24 | 1,000 | 1,800,000 | 450,000 | 1,350,000 | Incoming traffic โโโถ 15-PTU deployment (450,000 normalized TPM capacity) X X if throttled (HTTP 429) โ โโโโถ Automated Spillover to Standard/PayGo deployment if configured Important: Spillover is optional and must be configured either for the provisioned deployment or per request. Once configured, Microsoft Foundry automatically route
Comments
No comments yet. Start the discussion.