Haiku got cheap. The plumbing got cheaper.
Haiku got cheap. The plumbing got cheaper.
- Claude Haiku 5.5. Anthropic released it on Oct 7 and says it costs around 75% less to run than Haiku 4.5: $0.10 per million input tokens, $0.50 output, for prompts up to 100k. Cache reads on Sonnet 5.5 are halved, from $0.20 to $0.10. Anthropic's own chart puts Haiku 5.5 at 39.2% on Terminal-Bench 4.0 and Sonnet 5.5 at 70.6%. It says the bigger models stay better for complex coding.
- GPT-6 in ChatGPT. OpenAI shipped it on Oct 7 with "Intelligent UI": answers you can interact with. OpenAI says ChatGPT has 1.2 billion weekly users. In its own evaluation, web-search answers start 44% sooner.
- Gemini agent for work. Google Cloud announced a "universal agent for work" on Oct 8. The post gives no price and no availability date.
- Anthropic Cyber Mission. A program for critical infrastructure, plus a free open-source scanner (Oct 8). The reports are model-generated and unreviewed. Anthropic expects, but has not measured, a true-positive rate above 90%.
- SynthID Detector. Google DeepMind's tool (Oct 7) checks for Google and partner watermarks. A clean result does not prove a person made the file.
My take. The price war is moving from the headline model to the plumbing around it.
One test for Monday
- Pick your three most expensive agent tasks.
- Re-run each on the cheaper model with the same prompts and checks.
- Record pass rate and cost per finished task, not per token.
- Where the cheaper model fails a check, keep the expensive one and write down why.
Sources
Company announcements only.