Back to Blog
EnglishNews

What Is Claude Haiku 5.5? Anthropic's New Model: Specs, Price, and How to Call It

Claude Haiku 5.5 is Anthropic's fastest current model for high-volume extraction, classification, and routing. This guide covers its verified specs, tiered price, and working API examples.

C
Crazyrouter Team
October 11, 2026 / 1 views
Share:
What Is Claude Haiku 5.5? Anthropic's New Model: Specs, Price, and How to Call It

Claude Haiku 5.5 is Anthropic's fastest current language model, built for high-volume, latency-sensitive work such as classification, extraction, moderation, and request routing. It replaces Haiku 4.5 as the practical default when speed and unit cost matter: the base price is 0.10per1millioninputtokensand0.10 per 1 million input tokens and 0.50 per 1 million output tokens as of 2026-10, while its documented context window reaches 1 million tokens.

What is Claude Haiku 5.5?#

Anthropic positions Claude Haiku 5.5 as the smallest and fastest member of its current Claude family. It accepts text and images, produces text, supports tool use and multilingual tasks, and includes adaptive thinking. The official API ID is claude-haiku-5-5.

This release is not merely a lower-priced Haiku 4.5. Its documented context window rises from 200K to 1M tokens, and its maximum output reaches 128K tokens. That makes the model useful for long document pipelines that still need low latency. Anthropic lists a June 2026 reliable knowledge cutoff and says retirement will be no earlier than October 7, 2027.

The best fit is repetitive work with clear success criteria. Examples include assigning support tickets, extracting fields from invoices, screening content, converting text to JSON, and selecting which larger model should handle a request. For difficult proofs, long-horizon agents, or high-stakes analysis, test a Sonnet or Opus model instead.

Claude Haiku 5.5 specs#

SpecClaude Haiku 5.5
MakerAnthropic
API model IDclaude-haiku-5-5
Context window1M tokens
Maximum output128K tokens
InputsText and images
OutputText
FeaturesTool use, multilingual, vision, adaptive thinking
Reliable knowledge cutoffJune 2026
Best useHigh-volume classification, extraction, and routing

These limits come from Anthropic's model overview. Limits can also depend on provider, account, and endpoint, so production code should still handle length and rate-limit errors.

Claude Haiku 5.5 vs alternatives#

ModelContextCrazyrouter input / output per 1MChoose it for
Claude Haiku 5.51M0.10/0.10 / 0.50 at base tierFast, large-scale structured tasks
Claude Haiku 4.5200K0.65/0.65 / 3.25Existing tested Haiku 4.5 workloads
Claude Sonnet 4.6200KSee live pricingMore demanding coding and reasoning

Haiku 5.5 has a clear paper advantage over 4.5: five times the documented context and a base token rate that is 84.6% lower than the discounted 4.5 rate shown above. Migration should still go through your own evaluation set because model behavior can change even when the endpoint shape stays the same.

Claude Haiku 5.5 high-volume API routing illustration

How to call Claude Haiku 5.5#

The following request was tested against https://api.crazyrouter.com/v1 on October 11, 2026. It returned HTTP 200 and the answer 391. Put your own API key in CRAZYROUTER_API_KEY; do not embed a key in source control.

cURL#

bash
curl https://api.crazyrouter.com/v1/chat/completions \
  -H "Authorization: Bearer $CRAZYROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-haiku-5-5",
    "messages": [
      {"role": "user", "content": "Return only the result: 17 * 23."}
    ],
    "max_tokens": 100
  }'

Python#

python
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["CRAZYROUTER_API_KEY"],
    base_url="https://api.crazyrouter.com/v1",
)

response = client.chat.completions.create(
    model="claude-haiku-5-5",
    messages=[{"role": "user", "content": "Extract the invoice total as JSON."}],
    max_tokens=200,
)
print(response.choices[0].message.content)

Node.js#

javascript
import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.CRAZYROUTER_API_KEY,
  baseURL: "https://api.crazyrouter.com/v1",
});

const response = await client.chat.completions.create({
  model: "claude-haiku-5-5",
  messages: [{ role: "user", content: "Classify this ticket as billing or technical." }],
  max_tokens: 100,
});
console.log(response.choices[0].message.content);

Using an OpenAI-compatible client makes a staged migration simple. Keep 4.5 as a fallback, send a sample of real requests to both models, and compare schema validity, task accuracy, latency, and cost before switching all traffic.

Claude Haiku 5.5 pricing in October 2026#

Both Anthropic's list and Crazyrouter's live base tier show 0.10per1millioninputtokensand0.10 per 1 million input tokens and 0.50 per 1 million output tokens. Cached input is 0.01permillionatthebasetier.Requestsupto100Kinputtokensusethoserates;requestsabove100Kuse0.01 per million at the base tier. Requests up to 100K input tokens use those rates; requests above 100K use 0.50 input and 2.50outputpermillion,with2.50 output per million, with 0.05 cached input.

Input per requestInput / 1MCached input / 1MOutput / 1M
Up to 100K$0.10$0.01$0.50
Above 100K$0.50$0.05$2.50

Use the live Claude Haiku 5.5 price row for the current billing entry and the full pricing catalog when comparing models. There is no subscription requirement for API usage; billing is based on tokens consumed.

FAQ#

Is Claude Haiku 5.5 an Anthropic model?#

Yes. Anthropic makes Claude Haiku 5.5 and lists it in the current Claude family. Its official Claude API alias is claude-haiku-5-5.

What did Claude Haiku 5.5 replace?#

It is the current successor to Claude Haiku 4.5 for fast, high-volume work. Existing applications should migrate through evaluation rather than assuming identical behavior.

How much does Claude Haiku 5.5 cost per 1M tokens?#

As of 2026-10, the base tier costs 0.10permillioninputtokensand0.10 per million input tokens and 0.50 per million output tokens. Prompts above 100K input tokens enter the higher tier shown in the pricing section.

Does Claude Haiku 5.5 have a 1M context window?#

Yes, Anthropic documents a 1M-token context window. The maximum output is 128K tokens, which is a separate limit.

Can I call Claude Haiku 5.5 with the OpenAI SDK?#

Yes. An OpenAI-compatible endpoint can be used with the Python and Node.js OpenAI clients by changing the base URL and model ID. The examples above use https://api.crazyrouter.com/v1 and claude-haiku-5-5.

Summary#

Claude Haiku 5.5 combines a 1M-token context window, 128K maximum output, and a 0.10/0.10/0.50 base token rate. It is a strong candidate for structured, high-throughput tasks, but a real migration decision should come from workload-specific tests. You can create an account and run the tested request against your own examples.

<!-- crazyrouter-related-links -->

Implementation Guides

Related Articles