Microsoft Unveils New AI Models to Challenge Industry Rivals
View original sourceMicrosoft AI, led by Mustafa Suleyman, has announced the release of three foundational AI models:
- MAI-Transcribe-1: Transcribes speech across 25 languages into text, touted to be 2.5 times faster than Microsoft's previous offering.
- MAI-Voice-1: An audio-generating model that allows the creation of custom voices, producing 60 seconds of audio in one second.
- MAI-Image-2: A video-generating model, previously part of the MAI Playground.
These models are available through Microsoft Foundry and are part of a strategic push to compete in the multimodal AI market, particularly against rivals like Google and OpenAI. Although Microsoft continues its partnership with OpenAI, the renegotiation allows it to delve into superintelligence research independently. This move reflects Microsoft's aim to offer cost-effective solutions, with pricing details outlined in the report.
Mustafa Suleyman emphasized the Humanist AI approach, focused on human-centric, practical AI applications, signalling more releases soon.
The models were developed by the MAI Superintelligence team, formed in November 2025. Microsoft has invested over $13 billion in its AI endeavors and maintains a dual strategy in AI, producing its own models while partnering with industry leaders for resources like chips.