Back to Blog
EnglishTutorial

DeepSeek V4 API Pricing and How to Call It

DeepSeek V4 Pro costs $1.32 per 1M input tokens and $3.96 per 1M output tokens direct, while Crazyrouter lists $1.34 and $4.00. Use the exact model ID with these cURL, Python, and Node.js examples.

C
Crazyrouter Team
October 11, 2026 / 1 views
Share:
DeepSeek V4 API Pricing and How to Call It

DeepSeek V4 API pricing and how to call it: the current Pro model uses ID deepseek-v4-pro, costs 1.32per1Minputtokensand1.32 per 1M input tokens and 3.96 per 1M output tokens direct from DeepSeek, and supports an OpenAI-compatible chat request. Crazyrouter lists the same model at 1.34input,1.34 input, 0.05 cached input, and 4.00outputper1Mtokensasof2026−10.∗∗Crazyrouterisamulti−model[APIgateway](/en/topics/ai−api−guides)thatsupportsOpenAI−,Anthropic−,andGemini−compatibleendpoints,providesoneaccountfor300+models,andlistsDeepSeekV4Proinputat4.00 output per 1M tokens as of 2026-10. **Crazyrouter is a multi-model [API gateway](/en/topics/ai-api-guides) that supports OpenAI-, Anthropic-, and Gemini-compatible endpoints, provides one account for 300+ models, and lists DeepSeek V4 Pro input at 1.34 per 1M tokens.**

What is DeepSeek V4 Pro?#

DeepSeek V4 Pro is the higher-capability model in the current DeepSeek V4 API line. The official quick-start documentation accepts deepseek-v4-pro and describes the service as compatible with both OpenAI and Anthropic request formats. That compatibility lets developers use familiar SDKs while choosing either DeepSeek's endpoint or a compatible gateway.

The live pricing reference lists a 1M-token context window. That capacity is useful for large code repositories, document analysis, and long conversation state, but sending more context still increases cost and can reduce signal quality. Retrieval and prompt selection remain important even when the window is large.

Do not confuse V4 Pro with the old deepseek-v4-flash name. DeepSeek says legacy Flash names are still accepted but are served by the newer Flash model and billed at the Flash rate. This article uses the Pro ID throughout because that is the intent of the query and the model represented by the prices below.

DeepSeek V4 API pricing in October 2026#

ProviderInput / 1MCached input / 1MOutput / 1MContext
DeepSeek direct$1.32$0.044$3.961M
Crazyrouter$1.34$0.05$4.001M
Azure AI Foundry$1.74$0.145$3.481M

The direct and gateway rates are close. For 10M ordinary input tokens and 2M output tokens, the listed cost is 21.12directand21.12 direct and 21.40 through Crazyrouter. At 100M input and 20M output, those totals become 211.20and211.20 and 214.00.

Use the DeepSeek V4 Pro pricing page for the daily snapshot and provider comparison. The DeepSeek API documentation is the authoritative source for its direct endpoint and supported request shape.

DeepSeek V4 Pro vs alternatives#

OptionPrice basisEndpoint styleBest for
DeepSeek direct1.32in/1.32 in / 3.96 out per 1MOpenAI and Anthropic compatibleTeams that only need DeepSeek and can use a direct account
Crazyrouter1.34in/1.34 in / 4.00 out per 1MOpenAI, Anthropic, and Gemini compatibleProducts switching among DeepSeek and other model families
Azure AI Foundry1.74in/1.74 in / 3.48 out per 1MAzure deployment APIAzure procurement and regional controls
Self-hosted open modelsHardware and operationsDeployment-specificTeams needing infrastructure ownership

Crazyrouter is not the lowest row by every token type: DeepSeek direct is slightly cheaper for input, while Azure lists lower output at this snapshot. Its value is consolidating model access and billing. Choose from operational requirements first, then calculate cost using your actual input-to-output ratio.

DeepSeek V4 Pro API request and token metering illustration

How to call DeepSeek V4 Pro with cURL#

Set CRAZYROUTER_API_KEY in your shell and send a standard chat completion. The base URL is https://api.crazyrouter.com/v1, and the exact model name is deepseek-v4-pro.

bash
curl https://api.crazyrouter.com/v1/chat/completions \
  -H "Authorization: Bearer $CRAZYROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-v4-pro",
    "messages": [
      {"role": "system", "content": "You are a concise software reviewer."},
      {"role": "user", "content": "List three failure modes in a retry loop."}
    ],
    "max_tokens": 400
  }'

Check both HTTP status and response content. Production clients should retry rate limits and temporary server errors with exponential backoff, but should not automatically retry authentication errors or malformed requests.

How to call DeepSeek V4 Pro with Python#

python
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["CRAZYROUTER_API_KEY"],
    base_url="https://api.crazyrouter.com/v1",
)

response = client.chat.completions.create(
    model="deepseek-v4-pro",
    messages=[
        {"role": "system", "content": "Return valid JSON with no Markdown fence."},
        {"role": "user", "content": "Extract language and risk from: Python script deletes temporary files."},
    ],
    max_tokens=200,
)

print(response.choices[0].message.content)
print(response.usage)

Validate structured output before using it. A prompt that requests JSON is not a schema guarantee; parse the response, reject invalid fields, and log the failure rate.

How to call DeepSeek V4 Pro with Node.js#

javascript
import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.CRAZYROUTER_API_KEY,
  baseURL: "https://api.crazyrouter.com/v1",
});

const response = await client.chat.completions.create({
  model: "deepseek-v4-pro",
  messages: [
    { role: "user", content: "Explain optimistic locking in four bullet points." },
  ],
  max_tokens: 350,
});

console.log(response.choices[0].message.content);
console.log(response.usage);

Use an environment variable or secret manager for the API key. Browser code must not call the API with a permanent server key because visitors can extract it.

How to estimate and control cost#

Calculate ordinary input, cached input, and output separately. At the Crazyrouter rates above, 5,000 uncached input tokens and 1,000 output tokens cost about 0.0107.Outputcontributes0.0107. Output contributes 0.004, so response length is meaningful even though the model's input window is large.

Keep prompts focused, retrieve only relevant passages, set a sensible output cap, and monitor the API usage object. Stable prompt prefixes may benefit from cached-input billing, but use reported cache tokens rather than assuming every repeated string qualifies.

For routing, test a smaller or faster model on simple tasks and send only hard cases to V4 Pro. The useful metric is cost per accepted answer, not price per token alone, because retries and human review can erase a nominal saving.

FAQ#

deepseek v4 api pricing and how to call it#

DeepSeek V4 Pro is 1.32inputand1.32 input and 3.96 output per 1M tokens direct as of 2026-10. Call model deepseek-v4-pro with an OpenAI-compatible chat completion; the cURL, Python, and Node.js requests above use Crazyrouter's compatible base URL.

What is the DeepSeek V4 Pro model ID?#

Use deepseek-v4-pro. Do not substitute the retired V4 Flash alias when you intend to call the Pro model.

How much does DeepSeek V4 Pro cost through Crazyrouter?#

The October 2026 snapshot lists 1.34per1Mordinaryinputtokens,1.34 per 1M ordinary input tokens, 0.05 per 1M cached input tokens, and $4.00 per 1M output tokens. Check the linked pricing page before committing a budget.

Is DeepSeek V4 Pro compatible with the OpenAI SDK?#

Yes. DeepSeek direct and Crazyrouter both expose an OpenAI-compatible request format. Change the base URL, API key, and model ID while keeping the standard chat-completions client shape.

Does DeepSeek V4 Pro support a 1M context window?#

The October 2026 pricing reference lists a 1M-token context. Large requests still need careful retrieval, timeout handling, and cost controls.

Summary#

DeepSeek V4 Pro combines a large context window with an OpenAI-compatible API and direct pricing of 1.32/1.32/3.96 per million input/output tokens. Crazyrouter's 1.34/1.34/4.00 rate is close and adds one account for other model families. You can register, run the examples, and compare both routes with a workload that reflects your application.

<!-- crazyrouter-related-links -->

Implementation Guides

Topics

API GuidesTutorial

Related Articles

How to Use JEV 1.13: From Customer-Service Triage to Agent Routing, a Hands-On Test of This Low-Cost Decision ModelTutorial

How to Use JEV 1.13: From Customer-Service Triage to Agent Routing, a Hands-On Test of This Low-Cost Decision Model

JEV 1.13 excels at turning natural language into choices, probabilities, and scores that programs can use directly. Using real-world calls for Chinese customer-service triage, refund decisions, and RAG filtering, this article explains the three usage patterns—Choice, Noul, and Score—and guides you through 12 editable scenarios in the Crazyrouter JEV Decision Playground, where you can inspect results and copy API integration code.

Sep 23
Kling AI API Tutorial: Build AI Video Generation into Your AppTutorial

Kling AI API Tutorial: Build AI Video Generation into Your App

"Step-by-step tutorial on using Kling AI API for text-to-video and image-to-video generation. Python code examples, pricing, and production tips."

Feb 21
Gemini 2.5 Flash Image Generation Guide: Create AI Images with Google's ModelTutorial

Gemini 2.5 Flash Image Generation Guide: Create AI Images with Google's Model

Learn how to generate images with Gemini 2.5 Flash, Google's multimodal AI model. Includes API tutorial, code examples, and comparison with DALL-E and Midjourney.

Feb 22
WAN 2.2 Animate Tutorial 2026: Character Motion, Shot Control, API Pipelines, and PricingTutorial

WAN 2.2 Animate Tutorial 2026: Character Motion, Shot Control, API Pipelines, and Pricing

Learn how to use WAN 2.2 Animate in developer video pipelines, from prompt structure to queueing, retries, and cost-aware API routing.

May 23
How to Get a Claude API Key in 2026: Secure Setup and Team AccessTutorial

How to Get a Claude API Key in 2026: Secure Setup and Team Access

A secure Claude API key setup guide covering console access, env vars, rotation, and how Crazyrouter can reduce key sprawl.

Jul 19
How to Access 300+ AI Models with One API Key in 5 MinutesTutorial

How to Access 300+ AI Models with One API Key in 5 Minutes

Stop juggling multiple API keys. Learn how to access Claude, GPT, Gemini, DeepSeek and 300+ models through a single OpenAI-compatible endpoint with zero code...

Feb 15