Anthropic has released Claude Haiku 5.5, describing it as its fastest, cheapest and most capable small model. It is designed for high-volume and cost-sensitive work such as summaries, compaction, database queries, classification, customer support, browser use and narrowly scoped subagent tasks.
The model is available through Anthropic’s platform and across Amazon Web Services, Google Cloud and Microsoft Azure. Developers can call it with the model identifier claude-haiku-5-5, while a migration guide covers changes from earlier versions.
Lower price changes where a model can be used
Anthropic says Haiku 5.5 costs about 75 per cent less to run on average than Haiku 4.5. For prompts up to 100,000 tokens, input is priced at US$0.10 per million tokens and output at US$0.50. Longer prompts use higher rates, making prompt size an important part of workload planning.
Low unit cost can make previously uneconomic tasks practical, especially where an application performs millions of short calls. Teams should still evaluate the total workflow because retries, large context windows, tool calls and validation steps can outweigh the base model price.
Benchmarks show a large generational step
Anthropic reports substantial gains over Haiku 4.5 across knowledge work, computer use, multidisciplinary reasoning, agentic coding and visual reasoning. On the OSWorld 2.1 offline subset, Haiku 5.5 scored 72.4 per cent compared with 15.7 per cent for its predecessor. It also reached 39.2 per cent on Terminal-Bench 4.0.
These are vendor-reported results under specified conditions, not a guarantee for every production workload. The more useful test is a customer’s own representative evaluation, including latency, failure recovery, regional availability and the quality of outputs after the same safeguards are applied.
Adjustable effort arrives in the Haiku tier
Haiku 5.5 is the first model in its class to support an adjustable effort setting. Developers can choose a lower-cost mode for routine classification or extraction and spend more computation when a task needs deeper reasoning. That creates another routing dimension beyond simply selecting a different model.
A good router should use evidence rather than intuition. Teams can classify requests by risk and complexity, sample outcomes, and escalate when the small model is uncertain or fails validation. The cheapest successful path matters more than choosing the cheapest model for every request.
Subagents are a central use case
Anthropic positions Haiku 5.5 as a companion to Sonnet 5.5 and Opus 5.5. A larger lead model can plan difficult work while Haiku handles focused searches, summaries, code inspection or data extraction in parallel. This can reduce cost and latency without asking the small model to own the entire problem.
Multi-model systems need clear task boundaries and provenance. The lead agent should know what evidence a subagent used, what confidence it has and whether a result was independently checked. Fast parallel output is not useful if conflicting findings are merged without review.
Safety settings reflect the model’s role
Anthropic says Haiku 5.5 improved across alignment evaluations and was less willing to cooperate with misuse than Haiku 4.5. Its cybersecurity safeguards allow a wider range of defensive work than Sonnet 5.5, while continuing to block penetration testing and techniques judged more likely to assist attackers.
Biology safeguards match those used for recent Sonnet and Opus models. Organisations with broader legitimate needs can apply to verification programs. Customers should test refusals and unsafe completions in their own domain because model-level policy is only one layer of a production control system.
Sonnet pricing and subscriber credits also change
Alongside Haiku, Anthropic has halved Sonnet 5.5 cache-read pricing to US$0.10 per million tokens. The company estimates this reduces the cost of most agentic Sonnet work by about 20 per cent because cached context represents a large share of consumption.
Claude Max and Team subscribers will also receive monthly API credits. Max 5x users receive US$100, Max 20x users US$200 and Team accounts up to US$500 pooled across users. The credits are intended to help subscribers build tools and agents on the Claude Platform.
Production evaluation remains decisive
Haiku 5.5 expands the space between simple automation and frontier-model reasoning. Its combination of speed, lower price and computer-use capability could make small agents common in support, operations and software workflows.
Teams should adopt it with workload-specific tests, budget controls and fallback paths. The release is compelling not because every task should move to Haiku, but because more tasks can now be routed deliberately across model sizes. That architecture can improve responsiveness while keeping expensive reasoning for the work that genuinely needs it.
Operational teams should also monitor whether adjustable effort and prompt length produce predictable bills. Per-request limits, cached-context policies and quality sampling can keep a fast small model from creating hidden costs through excessive retries or poorly bounded autonomous loops.