Back to Blog
EnglishGuide

GPT-5 API Pricing per 1M Tokens in 2026

GPT-5 costs $1.25 per 1M input tokens, $0.125 per 1M cached input tokens, and $10 per 1M output tokens at list price. This guide calculates real workloads and shows a working API call.

C
Crazyrouter Team
October 11, 2026 / 0 views
Share:
GPT-5 API Pricing per 1M Tokens in 2026

GPT-5 API pricing per 1M tokens 2026 is 1.25forinput,1.25 for input, 0.125 for cached input, and 10.00foroutputatOpenAI′slistprice.ThroughCrazyrouter,theOctober2026rateis10.00 for output at OpenAI's list price. Through Crazyrouter, the October 2026 rate is 0.812 per 1M input tokens, 0.0812per1Mcachedinputtokens,and0.0812 per 1M cached input tokens, and 6.50 per 1M output tokens. Crazyrouter is a multi-model API gateway that supports OpenAI-, Anthropic-, and Gemini-compatible endpoints, offers access to 300+ models, and lists GPT-5 input at $0.812 per 1M tokens.

GPT-5 API pricing per 1M tokens 2026#

Token typeOpenAI list priceCrazyrouterSavings through Crazyrouter
Input$1.25 / 1M$0.812 / 1M35.0%
Cached input$0.125 / 1M$0.0812 / 1M35.0%
Output$10.00 / 1M$6.50 / 1M35.0%

These rates were checked on October 11, 2026. Use the GPT-5 pricing page for current provider rows and the GPT-5 model page for capabilities and API details.

Input and output tokens are billed separately. Output is eight times the list-price cost of ordinary input, so a chat product that generates long answers can spend more on completion tokens even when prompts are larger. Cached input is cheaper because the provider can reuse a previously processed prompt prefix.

What is GPT-5?#

GPT-5 is an OpenAI general-purpose model used for reasoning, coding, analysis, and agent workflows. The model ID in the API examples is gpt-5. The pricing reference lists a 128K-token context window, while actual availability and request limits can vary by endpoint and account.

The API is separate from a ChatGPT subscription. A monthly app plan does not turn into a fixed quantity of API tokens; API usage is metered from the text sent to and generated by the model.

What does a GPT-5 request cost?#

Use this formula:

text
cost = (uncached_input_tokens / 1,000,000 × input_rate)
     + (cached_input_tokens / 1,000,000 × cached_rate)
     + (output_tokens / 1,000,000 × output_rate)

Suppose one request has 4,000 uncached input tokens and 1,000 output tokens. At OpenAI list price, it costs 0.005forinputplus0.005 for input plus 0.01 for output, or 0.015total.AttheOctober2026[Crazyrouter](https://docs.crazyrouter.com/en/introduction)rate,thesametokencountscostabout0.015 total. At the October 2026 [Crazyrouter](https://docs.crazyrouter.com/en/introduction) rate, the same token counts cost about 0.009748.

Monthly workloadTokensOpenAI list priceCrazyrouter
Small prototype10M input + 2M output$32.50$21.12
Growing product100M input + 20M output$325.00$211.20
Large service1B input + 200M output$3,250.00$2,112.00

The examples assume all input is uncached. A stable system prompt or repeated document prefix can lower the input component when caching is supported and the provider reports cache hits.

GPT-5 API pricing meter for input, cached input, and output tokens

How to use the GPT-5 API#

All three examples use the exact model ID gpt-5 and the OpenAI-compatible base URL https://api.crazyrouter.com/v1. Store your key in CRAZYROUTER_API_KEY instead of putting it directly in code.

cURL#

bash
curl https://api.crazyrouter.com/v1/chat/completions \
  -H "Authorization: Bearer $CRAZYROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-5",
    "messages": [{"role": "user", "content": "Give three ways to reduce API output tokens."}],
    "max_tokens": 300
  }'

Python#

python
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["CRAZYROUTER_API_KEY"],
    base_url="https://api.crazyrouter.com/v1",
)

response = client.chat.completions.create(
    model="gpt-5",
    messages=[{"role": "user", "content": "Summarize this support ticket in 60 words."}],
    max_tokens=120,
)
print(response.choices[0].message.content)
print(response.usage)

Node.js#

javascript
import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.CRAZYROUTER_API_KEY,
  baseURL: "https://api.crazyrouter.com/v1",
});

const response = await client.chat.completions.create({
  model: "gpt-5",
  messages: [{ role: "user", content: "Return a five-item JSON checklist." }],
  max_tokens: 250,
});
console.log(response.choices[0].message.content);
console.log(response.usage);

Log the returned usage object for every request. The most reliable cost dashboard comes from billed usage rather than estimating tokens from characters, especially when tools, images, or provider-specific reasoning tokens are involved.

GPT-5 vs API alternatives#

RouteGPT-5 input / output per 1MBilling relationshipAPI styleBest fit
OpenAI direct1.25/1.25 / 10.00OpenAIOpenAI APITeams standardizing on OpenAI direct
Crazyrouter0.812/0.812 / 6.50One prepaid gateway accountOpenAI, Anthropic, Gemini compatibleTeams switching across model families
Azure OpenAIVerify current regional rateAzure accountAzure deployment APIExisting Azure procurement and regional controls
Other API resellersVerify current quoteReseller accountUsually OpenAI compatibleCompare support, markup, and availability

The lowest displayed price is not the only decision. Check regional availability, rate limits, support response, data terms, model freshness, and whether your application needs other providers. For authoritative direct billing terms, compare against OpenAI's API pricing.

How to control GPT-5 API cost#

First, cap output length. Because output tokens cost more, a precise max_tokens setting and a requested answer format can reduce waste quickly. Do not set the cap so low that responses end before completing a JSON object or tool call.

Second, reuse stable prompt prefixes where caching is available. Long policy blocks, schemas, and tool descriptions are good candidates, but monitor reported cached tokens rather than assuming a cache hit.

Third, route simple jobs to a smaller model. Classification, language detection, and short extraction may not need GPT-5. Keep GPT-5 for requests where your evaluation proves that its additional capability changes the business result.

Finally, measure cost per successful task. A cheaper model that requires retries, manual review, or frequent fallback can cost more than a model with a higher token rate.

FAQ#

gpt-5 api pricing per 1m tokens 2026#

OpenAI's list price is 1.25input,1.25 input, 0.125 cached input, and 10outputper1MtokensasofOctober2026.Crazyrouterlists10 output per 1M tokens as of October 2026. Crazyrouter lists 0.812, 0.0812,and0.0812, and 6.50 for the same three categories.

How much does one GPT-5 API request cost?#

Multiply each token category by its per-million rate. A request with 4,000 ordinary input tokens and 1,000 output tokens costs $0.015 at the stated OpenAI list price.

Is cached GPT-5 input always billed at $0.125 per 1M?#

Only tokens reported as cached receive the cached input rate. New prompt content is charged at the ordinary input rate, so inspect the usage fields returned by the API.

Is the GPT-5 API included with ChatGPT?#

No. Consumer ChatGPT plans and API billing are separate products. API calls are charged to an API account or gateway balance according to token usage.

Why is GPT-5 output more expensive than input?#

Generating tokens consumes inference capacity sequentially, while reading prompt tokens is comparatively cheaper. The practical result is that concise output limits can have a larger budget effect than trimming a small prompt.

Summary#

The headline GPT-5 list rate is 1.25permillioninputtokensand1.25 per million input tokens and 10 per million output tokens, but cached input and gateway pricing can materially change the bill. Build a budget from measured token usage, then compare cost per successful task. You can register and test the gpt-5 examples with your own prompts.

<!-- crazyrouter-related-links -->

Implementation Guides

Topics

Related Articles

Claude Card Declined? How to Fix API Payment Methods and Billing Issues in 2026Guide

Claude Card Declined? How to Fix API Payment Methods and Billing Issues in 2026

Claude card declined? Learn how Claude API payment methods work, why billing fails, how to check supported billing locations, and what alternatives developers can use when direct Anthropic billing is unavailable.

Jun 20
Claude API Pricing 2026: Every Model's Price per 1M Tokens (Opus 5, Fable 5, Sonnet 5, Haiku 4.5)Guide

Claude API Pricing 2026: Every Model's Price per 1M Tokens (Opus 5, Fable 5, Sonnet 5, Haiku 4.5)

Claude API pricing for every current model in one table: Anthropic list price, cached-input price and the Crazyrouter price per 1M tokens, plus monthly cost examples, how billing works, and how to get started.

Sep 30
Google Veo3 API Guide for Production Video Apps in 2026Guide

Google Veo3 API Guide for Production Video Apps in 2026

A production-focused Google Veo3 API guide with code examples, pricing notes, alternatives, and practical integration advice for developers.

Mar 15
DeepSeek V4 API Pricing and How to Call ItTutorial

DeepSeek V4 API Pricing and How to Call It

DeepSeek V4 Pro costs $1.32 per 1M input tokens and $3.96 per 1M output tokens direct, while Crazyrouter lists $1.34 and $4.00. Use the exact model ID with these cURL, Python, and Node.js examples.

Oct 11
Kimi-K2-Thinking Guide 2026: Evals, Reasoning Workflows, and Cost ControlGuide

Kimi-K2-Thinking Guide 2026: Evals, Reasoning Workflows, and Cost Control

A developer guide to Kimi-K2-Thinking covering what it is, where it performs well, how to build eval pipelines, and how to keep reasoning costs under control.

Mar 24
Google Veo3 API Guide 2026: Build Production Video Pipelines with FallbacksGuide

Google Veo3 API Guide 2026: Build Production Video Pipelines with Fallbacks

A developer-focused June 2026 guide to Google Veo3 API, alternatives, implementation patterns, pricing tradeoffs, and when to use Crazyrouter for unified AI API access.

Jun 4