Microsoft ships streaming speech models as AI video hype meets a sceptical crowd
Microsoft AI released MAI-Transcribe-2-Streaming and two text-to-speech models on 2 October, claiming first place for transcription accuracy on Artificial Analysis and a 150 ms latency figure for its MAI-Voice-2.1-Flash variant.

Microsoft AI released MAI-Transcribe-2-Streaming and two text-to-speech models on 2 October, according to The Decoder, which reviewed the launch. Microsoft says the transcription model covers 60 languages and returns partial results in just over 100 milliseconds.
That is not a video generator. But it is the layer that decides whether a synthetic presenter feels like a conversation or a recording, and it lands in the same week that the video side of the industry got a lesson in how hard the last mile is.
The realism claim that is easiest to check
Tavus introduced Griffin on 1 October, calling it the first "Human Interaction Model". The Decoder reported the company's own study: 48 percent of participants believed Griffin was a real person after a one-minute video call, against a previous best of 2 percent. Tavus says an Nvidia test scored Griffin 3.83 points for how human an AI feels in audio-video conversation, with actual humans at 3.92 and the previous best AI model at 2.80.
Only a preview, Griffin-Lite, is available to selected testers, with a fuller version promised once safety concerns are addressed. Tavus was founded in 2020 and has raised about $64 million.
Compare that with Microsoft's voice numbers. MAI-Transcribe-2-Streaming costs $0.54 per hour of audio at an introductory price through the end of the year. MAI-Voice-2.1 speaks 23 languages in the same voice with a native accent in each, according to Microsoft, and the Flash variant costs $15 per million characters instead of $22. Both can clone a voice from a few seconds of reference audio. In one Microsoft test, about half of 4,000 participants thought the voices belonged to a real person.
Video models keep shipping, and so do the questions
The wider release calendar is crowded. The dossier's headline list includes FLUX 3 Image, described as supporting up to 4K generation, subject positioning and partial editing, with an open-source version expected in weeks, plus ByteDance's Seedance 2.5 and Kling 4.0 from Kuaishou's AI spinoff. None of those releases is detailed in the source texts supplied here, so the specifics are thin. What is documented is the adjacent tooling.
Ideogram said on 1 October that Ideogram 4.5 will only touch the part of an image a user selects, at four quality tiers from 0.8 to 22 cents per image and native 2K resolution, with an open-weight release promised. That is a correction to a known failure mode rather than a new capability: repeated edits on models such as GPT-Image 2.5 and Nano Banana still leave artifacts, per The Decoder.
Then there is the money. Dealroom reported on 30 September that Luma scaled 9,000 GPUs in six hours for a Dream Machine launch. Runway moved into robotics with an open-weight model called Praxis-1, according to Robotics & Automation News on 1 October. Both sit outside the ten-source core below because the dossier carries only their headlines.
"Storage systems smooth out prices, but they don't generate electricity," Leonhard Gandhi, project manager for Energy Charts at Fraunhofer ISE, said as his institute published a simulator for German day-ahead power prices.
He was talking about batteries, not GPUs. The analogy is not perfect, but the constraint is the same shape: inference capacity is a physical input, and the dossier shows what happens when it becomes scarce. Denmark's parliament adopted an emergency grid plan on 2 October that replaces first-come, first-served connections with four priority categories, and data centers sit in the least prioritised tier unless tied to a critical societal function.
Meanwhile the model layer keeps fragmenting. Amazon Web Services released Strands Decider 2B, an open-source decision model, in the same week OpenAI announced a similar offering, TechCrunch reported on 1 October. Cloudflare answered with Clef and Clef-flash, open-weight decision models it says beat TypeSafe's Jev on three of four of TypeSafe's own benchmarks, at $0.24 per million tokens against Jev's $0.042. Cloudflare self-reported those scores and they have not been reproduced on the official index, as The Register noted.
None of it settles the question a viewer asks in the first two seconds of a synthetic clip: is this a person. Tavus has one number for that. It is 48 percent, on a one-minute call, in a study the company ran itself.
Sources
15- 01Microsoft AI releases new transcription and text-to-speech models for voice agentsEN
- 02Nearly half of test subjects mistook Tavus' AI video avatar for a real person on a one-minute callEN
- 03Ideogram says its new model can edit part of an image without messing up the restEN
- 04Amazon releases its own Jev clone as decision models flood the webEN
- 05Cloudflare tries to outplay Jev with open-weight Clef modelsEN
- 06TypeSafe AI Releases Jev: A Decision-Only Model That Returns Typed Probabilities Instead of TextEN
- 07Fraunhofer ISE releases electricity price simulatorEN
- 08Denmark approves new grid connection prioritization modelEN
- 09Google releases Gemini 4 Argon, called its most powerful model yetEN
- 10Google rolls out new Gemini AI model but restricts access over safety concernsEN
- 11OpenAI scraps release of new model over safety concerns in internal testingEN
- 12OpenAI says planned GPT-6.1 is too insecure to releaseEN
- 13OpenAI abandons plan to release upcoming model as safety concerns escalateEN
- 14The NHL Is Releasing Its Own 'Hot' Fanfic. Romance Lovers Hate ItEN
- 15The Tesla Model 3 Gets A Range Bump And A New Screen In The U.S.EN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.