•Technology
Cohere Launches New Open-Source Voice Model 'Transcribe' with Record Accuracy
View original sourceCohere, an enterprise AI company, has introduced its first voice model named Transcribe, an open-source automatic speech recognition model. This model targets tasks such as note-taking and speech analysis. Here are the key points from the launch:
- Model Efficiency: Transcribe is designed to operate on consumer-grade GPUs and is lightweight at 2 billion parameters.
- Language Support: It supports 14 languages, including English, French, German, Italian, and more.
- Performance: Transcribe shows superior performance compared to models like Zoom Scribe v1 and IBM Granite 4.0 1B, with an average Word Error Rate (WER) of 5.42, as per the Hugging Face Open ASR leaderboard.
- Accuracy & Limitations: While it holds a 61% average win rate in accuracy over other models, it struggles with Portuguese, German, and Spanish.
- Processing Capabilities: Capable of processing 525 minutes of audio in a minute, showing high efficiency in its class.
- Integration Plans: Cohere plans to integrate Transcribe with its North platform and offers it via its API for free access, enhancing its reach.
- Strategic Moves: There's speculation from Cohere's CEO, Aidan Gomez, about a possible IPO soon, supported by a reported annual recurring revenue of $240 million in 2025.
In a broader context, the release of Transcribe signals growing interest and advancements in speech recognition technologies, aligning with increasing demand for AI-powered note-taking apps.