Back to all articles
Technology

Mistral AI Launches New Text-to-Speech Model to Revolutionize Voice AI

View original source

French AI company Mistral unveiled a new open-source text-to-speech model named Voxtral TTS on Thursday, aimed at enhancing voice AI applications such as customer support and AI assistants.

  • Key Features & Benefits:

    • Allows enterprises to build customizable voice agents for sales and customer engagement, competing directly with other major brands like ElevenLabs, Deepgram, and OpenAI.
    • Supports nine languages: English, French, German, Spanish, Dutch, Portuguese, Italian, Hindi, and Arabic.
    • Tailored for efficiency, it is small enough to run on devices like smartwatches and smartphones, optimized for state-of-the-art performance at a lower cost than competitors.
  • Technical Specifications:

    • Capable of capturing nuances such as accents and intonations, with a TTFA of 90ms for 500 characters and an RTF of 6x.
    • Built to maintain voice characteristics during language switches, suitable for applications like dubbing and real-time translation.
  • Company's Vision and Future Plans:

    • Pierre Stock, VP of science operations, stated the aim is to provide an end-to-end platform that integrates audio, text, and image inputs and outputs for comprehensive information retrieval.
    • Mistral's strategic position is bolstered by open-source and customizable models that enhance enterprise adaptability.
  • Background Context:

    • Earlier in the year, Mistral released transcription models for both large batch and real-time processing, indicating a broader strategy to offer a complete voice product suite to enterprises.