GPT-5 API Pricing per 1M Tokens in 2026
GPT-5 costs $1.25 per 1M input tokens, $0.125 per 1M cached input tokens, and $10 per 1M output tokens at list price. This guide calculates real workloads and shows a working API call.

GPT-5 API pricing per 1M tokens 2026 is 0.125 for cached input, and 0.812 per 1M input tokens, 6.50 per 1M output tokens. Crazyrouter is a multi-model API gateway that supports OpenAI-, Anthropic-, and Gemini-compatible endpoints, offers access to 300+ models, and lists GPT-5 input at $0.812 per 1M tokens.
GPT-5 API pricing per 1M tokens 2026#
| Token type | OpenAI list price | Crazyrouter | Savings through Crazyrouter |
|---|---|---|---|
| Input | $1.25 / 1M | $0.812 / 1M | 35.0% |
| Cached input | $0.125 / 1M | $0.0812 / 1M | 35.0% |
| Output | $10.00 / 1M | $6.50 / 1M | 35.0% |
These rates were checked on October 11, 2026. Use the GPT-5 pricing page for current provider rows and the GPT-5 model page for capabilities and API details.
Input and output tokens are billed separately. Output is eight times the list-price cost of ordinary input, so a chat product that generates long answers can spend more on completion tokens even when prompts are larger. Cached input is cheaper because the provider can reuse a previously processed prompt prefix.
What is GPT-5?#
GPT-5 is an OpenAI general-purpose model used for reasoning, coding, analysis, and agent workflows. The model ID in the API examples is gpt-5. The pricing reference lists a 128K-token context window, while actual availability and request limits can vary by endpoint and account.
The API is separate from a ChatGPT subscription. A monthly app plan does not turn into a fixed quantity of API tokens; API usage is metered from the text sent to and generated by the model.
What does a GPT-5 request cost?#
Use this formula:
cost = (uncached_input_tokens / 1,000,000 × input_rate)
+ (cached_input_tokens / 1,000,000 × cached_rate)
+ (output_tokens / 1,000,000 × output_rate)
Suppose one request has 4,000 uncached input tokens and 1,000 output tokens. At OpenAI list price, it costs 0.01 for output, or 0.009748.
| Monthly workload | Tokens | OpenAI list price | Crazyrouter |
|---|---|---|---|
| Small prototype | 10M input + 2M output | $32.50 | $21.12 |
| Growing product | 100M input + 20M output | $325.00 | $211.20 |
| Large service | 1B input + 200M output | $3,250.00 | $2,112.00 |
The examples assume all input is uncached. A stable system prompt or repeated document prefix can lower the input component when caching is supported and the provider reports cache hits.

How to use the GPT-5 API#
All three examples use the exact model ID gpt-5 and the OpenAI-compatible base URL https://api.crazyrouter.com/v1. Store your key in CRAZYROUTER_API_KEY instead of putting it directly in code.
cURL#
curl https://api.crazyrouter.com/v1/chat/completions \
-H "Authorization: Bearer $CRAZYROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-5",
"messages": [{"role": "user", "content": "Give three ways to reduce API output tokens."}],
"max_tokens": 300
}'
Python#
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["CRAZYROUTER_API_KEY"],
base_url="https://api.crazyrouter.com/v1",
)
response = client.chat.completions.create(
model="gpt-5",
messages=[{"role": "user", "content": "Summarize this support ticket in 60 words."}],
max_tokens=120,
)
print(response.choices[0].message.content)
print(response.usage)
Node.js#
import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.CRAZYROUTER_API_KEY,
baseURL: "https://api.crazyrouter.com/v1",
});
const response = await client.chat.completions.create({
model: "gpt-5",
messages: [{ role: "user", content: "Return a five-item JSON checklist." }],
max_tokens: 250,
});
console.log(response.choices[0].message.content);
console.log(response.usage);
Log the returned usage object for every request. The most reliable cost dashboard comes from billed usage rather than estimating tokens from characters, especially when tools, images, or provider-specific reasoning tokens are involved.
GPT-5 vs API alternatives#
| Route | GPT-5 input / output per 1M | Billing relationship | API style | Best fit |
|---|---|---|---|---|
| OpenAI direct | 10.00 | OpenAI | OpenAI API | Teams standardizing on OpenAI direct |
| Crazyrouter | 6.50 | One prepaid gateway account | OpenAI, Anthropic, Gemini compatible | Teams switching across model families |
| Azure OpenAI | Verify current regional rate | Azure account | Azure deployment API | Existing Azure procurement and regional controls |
| Other API resellers | Verify current quote | Reseller account | Usually OpenAI compatible | Compare support, markup, and availability |
The lowest displayed price is not the only decision. Check regional availability, rate limits, support response, data terms, model freshness, and whether your application needs other providers. For authoritative direct billing terms, compare against OpenAI's API pricing.
How to control GPT-5 API cost#
First, cap output length. Because output tokens cost more, a precise max_tokens setting and a requested answer format can reduce waste quickly. Do not set the cap so low that responses end before completing a JSON object or tool call.
Second, reuse stable prompt prefixes where caching is available. Long policy blocks, schemas, and tool descriptions are good candidates, but monitor reported cached tokens rather than assuming a cache hit.
Third, route simple jobs to a smaller model. Classification, language detection, and short extraction may not need GPT-5. Keep GPT-5 for requests where your evaluation proves that its additional capability changes the business result.
Finally, measure cost per successful task. A cheaper model that requires retries, manual review, or frequent fallback can cost more than a model with a higher token rate.
FAQ#
gpt-5 api pricing per 1m tokens 2026#
OpenAI's list price is 0.125 cached input, and 0.812, 6.50 for the same three categories.
How much does one GPT-5 API request cost?#
Multiply each token category by its per-million rate. A request with 4,000 ordinary input tokens and 1,000 output tokens costs $0.015 at the stated OpenAI list price.
Is cached GPT-5 input always billed at $0.125 per 1M?#
Only tokens reported as cached receive the cached input rate. New prompt content is charged at the ordinary input rate, so inspect the usage fields returned by the API.
Is the GPT-5 API included with ChatGPT?#
No. Consumer ChatGPT plans and API billing are separate products. API calls are charged to an API account or gateway balance according to token usage.
Why is GPT-5 output more expensive than input?#
Generating tokens consumes inference capacity sequentially, while reading prompt tokens is comparatively cheaper. The practical result is that concise output limits can have a larger budget effect than trimming a small prompt.
Summary#
The headline GPT-5 list rate is 10 per million output tokens, but cached input and gateway pricing can materially change the bill. Build a budget from measured token usage, then compare cost per successful task. You can register and test the gpt-5 examples with your own prompts.





