Researchers at IISc’s SPIRE Lab, in collaboration with ARTPARK and Google, have released SraVaani, a multilingual speech recognition model designed to address the limited availability of speech-technology resources for several Indian regional and non-scheduled languages.
What is SraVaani?
SraVaani is described as the first multilingual Indian speech recognition model. It extends Automatic Speech Recognition (ASR) capabilities to Indian languages that have traditionally remained underserved by existing speech technologies.
The model covers 20 scheduled languages and 45 regional languages and dialects, thereby attempting to capture a wider range of India's linguistic diversity.
Multilingual and Multiscript Capabilities
One of the important features of SraVaani is its ability to convert spoken words into written text across 10 different scripts.
The model can also automatically identify the language being spoken, meaning that users do not have to manually select the language before using the speech-recognition system.
It supports several languages and dialects that have comparatively limited representation in mainstream speech technology, including Garo, Angika, Chakma, Kokborok, Tulu, Bundeli and Bajjika.
Foundation: Project Vaani
SraVaani is based on Project Vaani, one of the major programmes of the Indian Institute of Science (IISc) aimed at understanding and documenting India's extensive linguistic diversity.
Project Vaani has collected approximately 31,000 hours of speech data from more than 156,000 speakers spread across 165 regions.
This large and geographically diverse speech dataset provides researchers with information on natural conversational speech, which is particularly important for developing speech-recognition systems that can function beyond controlled laboratory conditions.
Technology Behind SraVaani
SraVaani uses a FastConformer-based Automatic Speech Recognition architecture.
The use of such an architecture enables the model to process spoken language and convert it into textual form while supporting a broad range of languages, dialects and scripts.
The combination of a large speech dataset with the FastConformer-based ASR architecture is intended to improve the ability of speech technology to handle India's highly diverse linguistic environment.
Potential Applications
SraVaani could be applied across several sectors where people prefer to interact with digital systems in their regional languages.
In education, it could facilitate voice-based learning and access to educational resources. In digital services and e-governance, it could make government and digital platforms more accessible to people who are more comfortable communicating in regional languages.
The technology could also support banking and healthcare services, where speech-based interfaces can help users interact with digital systems. Similarly, customer-support services could use multilingual speech recognition to communicate with users in a wider range of Indian languages.
Significance for India
India has a highly diverse linguistic landscape, but speech technologies have historically been concentrated around languages with larger datasets and greater commercial demand. SraVaani attempts to address this linguistic technology gap by extending speech recognition to a wider range of scheduled, regional and non-scheduled languages and dialects.
By enabling automatic language identification, multilingual speech-to-text conversion and support for multiple scripts, the model can contribute to more inclusive and accessible digital technology.
We provide offline, online and recorded lectures in the same amount.
Every aspirant is unique and the mentoring is customised according to the strengths and weaknesses of the aspirant.
In every Lecture. Director Sir will provide conceptual understanding with around 800 Mindmaps.
We provide you the best and Comprehensive content which comes directly or indirectly in UPSC Exam.
If you haven’t created your account yet, please Login HERE !
We provide offline, online and recorded lectures in the same amount.
Every aspirant is unique and the mentoring is customised according to the strengths and weaknesses of the aspirant.
In every Lecture. Director Sir will provide conceptual understanding with around 800 Mindmaps.
We provide you the best and Comprehensive content which comes directly or indirectly in UPSC Exam.