•Technology
Sarvam Launches Breakthrough Large Language Models to Compete with Global Rivals
View original sourceIndian AI lab Sarvam unveiled a new generation of large language models at the India AI Impact Summit in New Delhi. The launch is part of India's strategy to reduce reliance on foreign AI platforms and develop models suited for local languages and use cases.
- The new lineup includes 30-billion and 105-billion-parameter models, a text-to-speech model, a speech-to-text model, and a vision model. This is a significant enhancement from the earlier 2-billion-parameter Sarvam 1 model.
- These models employ a mixture-of-experts architecture, activating only a subset of parameters, which notably decreases computing expenses.
- The 30B model features a 32,000-token context window for real-time conversational applications, and the 105B model has a 128,000-token window for more intricate reasoning tasks.
- The 30B model was pre-trained on about 16 trillion tokens, while the 105B model covered multiple Indian languages.
- Sarvam's efforts are backed by the IndiaAI Mission with infrastructure support from Yotta and technical support from Nvidia.
- Sarvam plans to open-source these models, though specifics on the training data and code remain undisclosed.
- Further specialization includes developing coding-focused models and enterprise tools under 'Sarvam for Work', as well as a conversational AI agent called Samvaad.
- Founded in 2023, Sarvam has raised over $50 million with investors like Lightspeed Venture Partners, Khosla Ventures, and Peak XV Partners.