September 3, 2026

Innovative Wizards

Innovate, Review, Inspire

Meta Launches New Muse Voice Transcribe With Support for 70+ Languages, Including 5 Indian Languages

Rapid advancements in artificial intelligence are transforming voice transcription technology, making it possible to convert spoken words into text faster and more accurately. As part of this development, Meta introduced a new speech-to-text model called Muse Voice Transcribe on Tuesday, September 1. The system is designed to transcribe conversations in real time while also understanding multiple languages within the same conversation.

Muse Voice Transcribe has been developed by Meta Superintelligence Labs (MSL). According to Meta, it is the first real-time audio perception model released by the lab.

Support for More Than 70 Languages

Meta’s new model has been trained on more than 70 languages. The list includes five major Indian languages: Hindi, Tamil, Telugu, Malayalam, and Kannada.

However, multilingual support is not the model’s only highlight. Muse Voice Transcribe has also been designed to handle code-switching, which refers to the practice of switching between two or more languages during a conversation.

For example, if a person moves between English and Hindi while speaking, the model can recognize and process the language change within the same conversation.

This capability could be particularly valuable for users in India, where mixing English with regional languages is common in everyday communication. Instead of requiring separate models or additional processing whenever the speaker changes languages, Muse Voice Transcribe is designed to manage these transitions within a single model.

Transcription While the Person Is Speaking

One of the key capabilities of Muse Voice Transcribe is its real-time streaming transcription feature.

Users do not have to wait for an entire audio recording to finish before receiving a transcript. The system generates text as the person speaks, allowing conversations to be converted into text almost immediately.

Meta says the model can also distinguish between more than 20 speakers in a single recording. It is additionally capable of processing audio recordings that are longer than one hour.

The company says these capabilities are handled directly by the same model, eliminating the need for a separate post-processing system for speaker separation and other functions.

25 Languages Validated at Launch

Muse Voice Transcribe was trained using data covering more than 70 languages. However, 25 languages had been validated when the model launched.

Meta also claimed that, as of September 1, 2026, Muse Voice Transcribe ranked first on the Artificial Analysis streaming speech-to-text leaderboard.

Designed to Balance Speed and Accuracy

Another notable technical feature of Muse Voice Transcribe is its approach to balancing transcription speed with accuracy.

Many speech-to-text systems listen to audio for a fixed amount of time before generating text. Muse Voice Transcribe takes a different approach by determining how much processing time is needed on a word-by-word basis.

This means the model can quickly process words that are easy to understand while spending additional time on words that are more difficult or unclear. According to Meta, this approach allows the system to deliver faster responses without compromising transcription accuracy.

Available to Developers Through an API

Muse Voice Transcribe is now available through the Meta Model API, giving developers access to the technology for their own applications and services.

Meta said the model is already being used for dictation-related features in Meta AI for Mac and Muse Code.

The API is priced at $3 per 1,000 audio minutes. Meta says this works out to approximately $0.18 per hour of audio.

The pricing could be attractive to developers building applications that require speech-to-text processing at scale, particularly for services involving large amounts of audio content.

Why Could Muse Voice Transcribe Be Important for India?

Using multiple languages within the same conversation is common in India, particularly when speakers move between English and regional languages.

Traditional speech-recognition systems can face additional challenges when the speaker changes languages during a conversation. Muse Voice Transcribe’s ability to handle code-switching through a single model could help address some of these challenges and make real-time transcription more useful for multilingual users.

Support for Hindi, Tamil, Telugu, Malayalam, and Kannada could further expand the potential applications of the technology in India.

Potential use cases include:

  • Real-time meeting transcription
  • Voice assistants
  • Dictation
  • Customer support systems
  • Coding applications
  • Multilingual note-taking
  • Accessibility tools
  • Audio and video content transcription
  • Voice-based AI applications

The Growing Role of Voice-Based AI

AI-powered voice technology is no longer limited to traditional voice assistants. Speech-based AI is increasingly being used for transcription, dictation, coding, productivity tools, accessibility services, and several other applications.

As these systems become more advanced, simply supporting major languages such as English is no longer enough. The ability to understand regional languages, different speech patterns, and natural language mixing is becoming increasingly important.

This is particularly relevant in multilingual markets such as India. If AI systems can seamlessly understand language changes during conversations, voice-based services could become more useful and accessible to a much broader audience.

Alongside the launch, Meta has also published additional technical information about Muse Voice Transcribe through a research blog, providing further details about the model and its capabilities.

With support for more than 70 languages, five major Indian languages, real-time transcription, the ability to distinguish between more than 20 speakers, and code-switching capabilities, Muse Voice Transcribe represents another significant step in Meta’s efforts to advance voice-based AI and make speech recognition more adaptable to real-world multilingual conversations.

Leave a Reply