Claude Haiku 5.5 and Mistral Large 4: what changes on calls
Two models that matter for AI on the phone launched within two days: Mistral Large 4 on 6 October 2026 and Claude Haiku 5.5 on 7 October. The first is Europe's new flagship model, with open weights and European hosting. The second is Anthropic's fastest model, up to 90% cheaper than its predecessor. Nothing changes overnight for an AI phone assistant, but both point the same way: faster answers, falling costs and more choice over where data stays.
Here is what actually shipped, the other AI voice news of recent weeks and what it means for an AI phone assistant.
Claude Haiku 5.5: fast and cheap
Anthropic launched Claude Haiku 5.5 on 7 October 2026. What matters for a business:
- Speed. Anthropic calls it its fastest model and explicitly recommends it for speed-sensitive tasks such as live customer support. Box, one of the customers quoted in the announcement, reports a score 11 points higher than Haiku 4.5 at about half the latency.
- Cost. For requests up to 100,000 tokens it costs $0.10 per million input tokens and $0.50 per million output tokens, 90% less than Haiku 4.5 ($1 and $5). On average, Anthropic says, it costs about 75% less to run.
- Where to get it. It is available through Anthropic's API and on Amazon Web Services, Google Cloud and Microsoft Azure.
Haiku 5.5 reads and writes text (and images); it is not a voice model. In an AI-handled call it is the "brain": it understands the transcribed request, decides what to do and drafts the answer, which a text-to-speech model then speaks.
Mistral Large 4: Europe's open-weight model
Mistral AI launched Mistral Large 4 on 6 October 2026 as a public preview.
- Open weights. Mistral will publish the weights by the end of October, so companies with the infrastructure can run it on their own servers.
- Europe. It was trained in Mistral's European data centres and has a European deployment that Mistral runs "under European law". For clinics, professional firms and anyone handling sensitive data, that is a practical argument.
- Languages. It was trained on more than 160 languages, including every official EU language.
- Size and cost. It is a large model (1 trillion parameters, 52 billion active per answer) and costs $1.36 per million input tokens and $4.18 per million output tokens.
A model this size is built for complex work: document analysis, agents, coding. On the phone, where every fraction of a second is audible, what matters more is that European models are improving at every size.
Other AI news from recent weeks
- Claude Sonnet 5.5 (28 September): Anthropic's mid-size model, which the announcement says is over 30% faster and up to 30% cheaper than Sonnet 5.
- Eleven v4 (28 September): the new voice model from ElevenLabs. The Turbo version starts speaking in a median 150 milliseconds, supports 90+ languages and clones a voice from 10 seconds of audio.
- Gemini 3.8 Live and 3.5 Transcribe (15 September): Google's real-time voice models, with 97+ languages and the ability to run actions (for example, checking a calendar) while the conversation continues.
- Gemini 4 Argon (30 September): Google's new flagship, so far limited to selected partners. We covered it in this article.
What changes for an AI-handled phone call
An AI call goes through three steps: transcribing what the caller says, the model's reasoning, and the voice that answers. This month's releases improve all three.
- Reasoning keeps getting cheaper. With a fast model like Haiku 5.5 at a tenth of the cost, an assistant can afford more checks during a call (confirming a detail, re-reading calendar availability) without slowing down.
- The biggest model is not the best on the phone. Response time is what callers notice: to book a table or an appointment, a light, fast model almost always beats a huge one.
- Voices sound more natural. Text-to-speech models such as Eleven v4 reply within a few tenths of a second and with more human intonation, and callers hear the difference.
- More choice over where data stays. Mistral Large 4 widens the range of European open-weight models, useful for anyone bound by strict privacy rules.
Who benefits most
- Medical practices: faster models mean bookings handled without waiting, and European models give more options to anyone handling health data.
- Restaurants: at rush hour every second of silence on the line counts. Fast models and natural voices make booking smoother.
- Real estate agencies: cheaper reasoning makes it practical to qualify every lead (area, budget, timing) before passing it to an agent.
What to do now
Nothing urgent: the model is a component your provider updates behind the scenes. yourang.ai, for example, is not tied to one model and can work with models from different providers, chosen for how well they perform on calls. If you are evaluating a phone assistant, ask the provider three questions:
- How long between the end of the caller's sentence and the start of the answer?
- How well does it understand spoken language, including accents?
- Where is call data processed?
The quickest way to find out is to hear it: try the demo and judge from a real call. For the cost of the service, see the pricing page.
Frequently asked questions
What is Claude Haiku 5.5?
It is the lightest model in Anthropic's Claude 5.5 family, released on 7 October 2026. Anthropic calls it its fastest model and recommends it for speed-sensitive work such as live customer support.
How much does Claude Haiku 5.5 cost compared with Haiku 4.5?
For requests up to 100,000 tokens it costs 90% less: $0.10 per million input tokens and $0.50 per million output tokens, against $1 and $5 for Haiku 4.5. Above 100,000 tokens the reduction is 50%.
Can I use Mistral Large 4 today?
Yes, as a public preview since 6 October 2026 through the Mistral Studio API. Mistral says it will release the model weights by the end of October.
Should I switch phone assistant to get these models?
No. The model is a component your provider can update. Judge an assistant on real calls instead: response speed, understanding of spoken language, actions during the call and handover to a person.


