DeepSeek V4 API Pricing and How to Call It
DeepSeek V4 Pro costs $1.32 per 1M input tokens and $3.96 per 1M output tokens direct, while Crazyrouter lists $1.34 and $4.00. Use the exact model ID with these cURL, Python, and Node.js examples.

DeepSeek V4 API pricing and how to call it: the current Pro model uses ID deepseek-v4-pro, costs 3.96 per 1M output tokens direct from DeepSeek, and supports an OpenAI-compatible chat request. Crazyrouter lists the same model at 0.05 cached input, and 1.34 per 1M tokens.**
What is DeepSeek V4 Pro?#
DeepSeek V4 Pro is the higher-capability model in the current DeepSeek V4 API line. The official quick-start documentation accepts deepseek-v4-pro and describes the service as compatible with both OpenAI and Anthropic request formats. That compatibility lets developers use familiar SDKs while choosing either DeepSeek's endpoint or a compatible gateway.
The live pricing reference lists a 1M-token context window. That capacity is useful for large code repositories, document analysis, and long conversation state, but sending more context still increases cost and can reduce signal quality. Retrieval and prompt selection remain important even when the window is large.
Do not confuse V4 Pro with the old deepseek-v4-flash name. DeepSeek says legacy Flash names are still accepted but are served by the newer Flash model and billed at the Flash rate. This article uses the Pro ID throughout because that is the intent of the query and the model represented by the prices below.
DeepSeek V4 API pricing in October 2026#
| Provider | Input / 1M | Cached input / 1M | Output / 1M | Context |
|---|---|---|---|---|
| DeepSeek direct | $1.32 | $0.044 | $3.96 | 1M |
| Crazyrouter | $1.34 | $0.05 | $4.00 | 1M |
| Azure AI Foundry | $1.74 | $0.145 | $3.48 | 1M |
The direct and gateway rates are close. For 10M ordinary input tokens and 2M output tokens, the listed cost is 21.40 through Crazyrouter. At 100M input and 20M output, those totals become 214.00.
Use the DeepSeek V4 Pro pricing page for the daily snapshot and provider comparison. The DeepSeek API documentation is the authoritative source for its direct endpoint and supported request shape.
DeepSeek V4 Pro vs alternatives#
| Option | Price basis | Endpoint style | Best for |
|---|---|---|---|
| DeepSeek direct | 3.96 out per 1M | OpenAI and Anthropic compatible | Teams that only need DeepSeek and can use a direct account |
| Crazyrouter | 4.00 out per 1M | OpenAI, Anthropic, and Gemini compatible | Products switching among DeepSeek and other model families |
| Azure AI Foundry | 3.48 out per 1M | Azure deployment API | Azure procurement and regional controls |
| Self-hosted open models | Hardware and operations | Deployment-specific | Teams needing infrastructure ownership |
Crazyrouter is not the lowest row by every token type: DeepSeek direct is slightly cheaper for input, while Azure lists lower output at this snapshot. Its value is consolidating model access and billing. Choose from operational requirements first, then calculate cost using your actual input-to-output ratio.

How to call DeepSeek V4 Pro with cURL#
Set CRAZYROUTER_API_KEY in your shell and send a standard chat completion. The base URL is https://api.crazyrouter.com/v1, and the exact model name is deepseek-v4-pro.
curl https://api.crazyrouter.com/v1/chat/completions \
-H "Authorization: Bearer $CRAZYROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek-v4-pro",
"messages": [
{"role": "system", "content": "You are a concise software reviewer."},
{"role": "user", "content": "List three failure modes in a retry loop."}
],
"max_tokens": 400
}'
Check both HTTP status and response content. Production clients should retry rate limits and temporary server errors with exponential backoff, but should not automatically retry authentication errors or malformed requests.
How to call DeepSeek V4 Pro with Python#
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["CRAZYROUTER_API_KEY"],
base_url="https://api.crazyrouter.com/v1",
)
response = client.chat.completions.create(
model="deepseek-v4-pro",
messages=[
{"role": "system", "content": "Return valid JSON with no Markdown fence."},
{"role": "user", "content": "Extract language and risk from: Python script deletes temporary files."},
],
max_tokens=200,
)
print(response.choices[0].message.content)
print(response.usage)
Validate structured output before using it. A prompt that requests JSON is not a schema guarantee; parse the response, reject invalid fields, and log the failure rate.
How to call DeepSeek V4 Pro with Node.js#
import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.CRAZYROUTER_API_KEY,
baseURL: "https://api.crazyrouter.com/v1",
});
const response = await client.chat.completions.create({
model: "deepseek-v4-pro",
messages: [
{ role: "user", content: "Explain optimistic locking in four bullet points." },
],
max_tokens: 350,
});
console.log(response.choices[0].message.content);
console.log(response.usage);
Use an environment variable or secret manager for the API key. Browser code must not call the API with a permanent server key because visitors can extract it.
How to estimate and control cost#
Calculate ordinary input, cached input, and output separately. At the Crazyrouter rates above, 5,000 uncached input tokens and 1,000 output tokens cost about 0.004, so response length is meaningful even though the model's input window is large.
Keep prompts focused, retrieve only relevant passages, set a sensible output cap, and monitor the API usage object. Stable prompt prefixes may benefit from cached-input billing, but use reported cache tokens rather than assuming every repeated string qualifies.
For routing, test a smaller or faster model on simple tasks and send only hard cases to V4 Pro. The useful metric is cost per accepted answer, not price per token alone, because retries and human review can erase a nominal saving.
FAQ#
deepseek v4 api pricing and how to call it#
DeepSeek V4 Pro is 3.96 output per 1M tokens direct as of 2026-10. Call model deepseek-v4-pro with an OpenAI-compatible chat completion; the cURL, Python, and Node.js requests above use Crazyrouter's compatible base URL.
What is the DeepSeek V4 Pro model ID?#
Use deepseek-v4-pro. Do not substitute the retired V4 Flash alias when you intend to call the Pro model.
How much does DeepSeek V4 Pro cost through Crazyrouter?#
The October 2026 snapshot lists 0.05 per 1M cached input tokens, and $4.00 per 1M output tokens. Check the linked pricing page before committing a budget.
Is DeepSeek V4 Pro compatible with the OpenAI SDK?#
Yes. DeepSeek direct and Crazyrouter both expose an OpenAI-compatible request format. Change the base URL, API key, and model ID while keeping the standard chat-completions client shape.
Does DeepSeek V4 Pro support a 1M context window?#
The October 2026 pricing reference lists a 1M-token context. Large requests still need careful retrieval, timeout handling, and cost controls.
Summary#
DeepSeek V4 Pro combines a large context window with an OpenAI-compatible API and direct pricing of 3.96 per million input/output tokens. Crazyrouter's 4.00 rate is close and adds one account for other model families. You can register, run the examples, and compare both routes with a workload that reflects your application.
<!-- crazyrouter-related-links -->




