Meta is making a significant move into the highly competitive real-time speech-to-text market with the introduction of Muse Voice Transcribe, a novel audio perception model. This new offering from Meta Superintelligence Labs uniquely combines streaming transcription, precise endpoint detection, and sophisticated speaker diarization capabilities designed to handle over 20 distinct speakers simultaneously. The model is being rolled out with a public API price point of a mere $0.18 per hour of processed audio, positioning Meta as a formidable contender in this rapidly evolving technological landscape.
Developed with the aim of processing speech as it happens, rather than requiring complete recordings, Muse represents a departure from traditional transcription methods. Meta’s official launch post for Muse Voice Transcribe highlights the model’s advanced features, including its capacity to process lengthy audio exceeding an hour, seamless multilingual code-switching, language and keyword biasing for improved accuracy, and integrated diarization that eliminates the need for a separate post-processing pipeline. The model has been meticulously trained on an extensive dataset encompassing over 70 languages, with a robust 25 of these languages undergoing thorough validation for the initial release. This broad linguistic support is crucial for global applications and diverse user bases.
While the capability to handle more than 20 speakers is substantial and places Muse toward the higher end of current market offerings, it is not an unprecedented figure. A review of existing vendor documentation reveals systems with published ceilings that exceed Meta’s stated capacity. For instance, Speechmatics’ real-time transcription service advertises its ability to identify 50 speakers by default, with an option to increase this limit to 100. Similarly, Amazon Transcribe’s diarization documentation specifies a maximum of 30 unique speakers, even for streaming transcription scenarios. This comparison underscores that while Muse’s speaker count is impressive, the true value proposition likely lies in its integrated approach and pricing strategy.
Despite not setting a new world record for speaker count, Muse is strategically positioned toward the higher end of the market in terms of its feature set. Meta’s overarching proposition appears to be more impactful than the raw maximum speaker number. The integration of high-capacity real-time diarization alongside low-latency transcription, precise endpointing, multilingual code-switching, and a notably aggressive API pricing structure within a single, unified model is a significant development. For enterprise developers engaged in building sophisticated meeting systems, call analytics platforms, live assistant applications, or ambient AI solutions, this comprehensive combination could prove far more beneficial than simply achieving the highest speaker count. The ability to accurately attribute speech to multiple participants in real-time is becoming a critical differentiator in the development of intelligent audio applications.
The increasing importance of speaker diarization is fundamentally changing the architecture of voice processing stacks. Historically, speech recognition systems primarily focused on answering the question: "What was said?" Diarization adds a crucial second layer to this inquiry: "Who said it?" This distinction is of paramount importance as transcribed audio feeds into increasingly complex downstream AI systems. A meeting assistant, for example, might flawlessly transcribe every spoken sentence, but if it incorrectly attributes an approval, a commitment, or a dissenting opinion to the wrong participant, the resulting corporate record becomes unreliable and potentially misleading. The same challenge extends to customer-service analytics, where accurate speaker attribution is vital for performance evaluation and compliance workflows, as well as for AI agents operating in dynamic environments where multiple individuals are speaking concurrently.
Muse fundamentally addresses this by incorporating speaker attribution directly into its autoregressive multimodal architecture. Meta explains that audio data is processed in small, 80-millisecond chunks, translating to approximately 12.5 chunks per second. Each of these chunks is transformed into a "soft token." At each processing step, the model dynamically decides whether to consume additional audio data or to emit transcribed text. Meta refers to this innovative mechanism as "adaptive delay." Instead of adhering to a uniform latency budget for every word, Muse can afford to wait longer when speech is ambiguous, allowing for more context to be gathered before committing to a transcription. Conversely, it can commit earlier when sufficient context is available. Meta states that the model’s behavior is trained through reinforcement learning, which balances rewards for minimizing word error rate and reducing delay. The technical intricacies of Muse’s architecture are further detailed in Meta’s technical explanation.
Within this unified architecture, speaker attribution and endpointing are seamlessly integrated into the same token sequence. A special token, <|start_of_turn|>, signals a potential new speaker turn. Subsequent tokens, such as <|speaker_A|>, are used to identify the specific speaker. Additionally, separate onset and endpoint tokens are employed to precisely delineate the boundaries of speech segments. Meta emphasizes that its approach involves training the Automatic Speech Recognition (ASR), diarization, and endpointing components together, rather than treating speaker clustering as a disconnected, downstream process. This integrated training methodology is key to achieving cohesive and accurate real-time performance.
Meta’s Model API speech-to-text documentation further solidifies this integrated approach by exposing diarization as a first-class operating mode. This is available alongside traditional push-to-talk and endpointing functionalities. The speaker labels, such as ‘A’ and ‘B’, are scoped to individual sessions rather than representing verified identities, and the API provides turn-level timestamps rather than the more granular word-level timestamps. This design choice reflects a focus on practical application in real-time conversational scenarios where identifying the speaker of a complete utterance is often the primary requirement.
When comparing speaker-count capabilities, it is essential to acknowledge the nuanced ways in which different vendors implement diarization and the varying levels of transparency regarding their maximum supported speaker counts. Speechmatics currently holds the strongest explicit real-time capacity claim identified in this review. Their real-time STT documentation indicates that speaker diarization is available live, and their real-time FAQ further specifies that the system supports 50 speakers by default, with the ability to be configured up to 100. Amazon Web Services (AWS) also surpasses Meta’s stated figure, with Amazon Transcribe capable of differentiating a maximum of 30 unique speakers. AWS provides explicit instructions for speaker partitioning within its streaming transcription services.

Other players in the market also offer diarization, though often with different capacities and pricing models. Soniox supports diarization in both real-time and asynchronous processing, but documents a maximum of 15 speakers per session. AssemblyAI’s streaming diarization system allows developers to set the max_speakers parameter between one and ten. Both Soniox and AssemblyAI caution that achieving accurate live speaker attribution is inherently more challenging due to the limited future audio context available to streaming systems compared to offline models. xAI’s current Speech-to-Text API also includes support for speaker diarization in streaming mode. However, its documentation, as reviewed for this story, does not publish a maximum diarized-speaker count, making a direct ceiling comparison with Muse difficult.
Consequently, it would be inaccurate to portray Muse’s "20-plus" speaker capability as establishing a new global record for real-time diarization. The highest explicitly documented real-time number identified in this survey is Speechmatics’ configurable 100-speaker ceiling. Furthermore, Meta’s own launch materials do not showcase demonstrations with over 20 simultaneous participants. Their principal live demonstration features eight speakers, and a long-form recording includes 11 labeled participants. The "20-plus" figure represents a stated model capability rather than the participant count observed in their public demonstrations.
The pricing structure of Muse Voice Transcribe introduces a compelling competitive dynamic into the market. According to Meta’s Muse Voice Transcribe developer page, the service is priced at $3 per 1,000 minutes, which translates to an exceptionally competitive $0.18 per hour. Notably, both streaming and non-streaming transcription are offered at the same price point. Meta also states that zero-data-retention processing is priced identically to standard processing. Billing is applied to the audio that is actually processed and is rounded down to whole seconds, offering predictable cost management.
To provide a clearer perspective on Meta’s pricing strategy, a standardization of publicly posted rates to one hour of streaming audio reveals the following approximate cost comparisons: Soniox’s stt-rt-v5 service is priced around $0.12 per hour and includes diarization for up to 15 speakers. Meta Muse Voice Transcribe, at $0.18 per hour, includes diarization for over 20 speakers. xAI’s Speech to Text API is priced at approximately $0.20 per hour, with diarization supported but a maximum speaker count not publicly stated. Speechmatics’ Real-time Standard service costs around $0.24 per hour and includes diarization for 50 default speakers, configurable up to 100. Alibaba Cloud’s Qwen3 ASR Flash Realtime is approximately $0.324 per hour for international use, with no comparable maximum documented in the reviewed sources for diarization. Deepgram’s Nova-3 Multilingual is around $0.35 per hour as a base rate, with diarization adding an additional $0.12 per hour, bringing the total to approximately $0.47 per hour. ElevenLabs Scribe v2 Realtime is priced at $0.39 per hour on a pay-as-you-go basis, but does not support diarization in real-time. AssemblyAI’s Universal-3.5 Pro Realtime is $0.45 per hour, with a $0.12 per hour add-on for streaming diarization supporting up to 10 speakers. Google’s Gemini 3.5 Transcribe Live has an estimated blended cost of approximately $0.54 per hour, but does not support diarization in live mode. Amazon Transcribe Streaming is approximately $0.60 per hour in AWS’s N. Virginia example, and includes diarization for up to 30 speakers. OpenAI’s GPT Live Transcribe is priced at $1.02 per hour, with diarization not listed as a model capability.
It is important to note that these comparisons are inherently imperfect due to variations in how services are packaged and priced. For instance, Qwen’s pricing varies by deployment geography, with its international real-time rate translating to about $0.324 per hour. Google’s Gemini figure is an estimated blended token cost rather than a flat hourly tariff. AWS pricing also varies by region and usage tier. ElevenLabs lists $0.39 per hour on its API pricing page but advertises lower rates for annual business plans. Deepgram’s pricing model particularly highlights the significance of feature-level comparisons: their Nova-3 Multilingual streaming rate is about $0.35 per hour, but speaker diarization incurs an additional charge of $0.002 per minute, elevating the comparable total to roughly $0.47 per hour. Similarly, AssemblyAI lists $0.45 per hour for its Universal-3.5 Pro Realtime and an additional $0.12 per hour for streaming diarization. Cartesia’s Ink-2 pricing is more complex, packaged through monthly credit plans rather than a straightforward metered hourly rate. Their $5 Pro plan includes approximately nine hours and 16 minutes of Ink-2 transcription, which equates to about $0.54 per transcription hour if all credits are exclusively used for STT. This should not be directly equated to a standalone $0.54 hourly API tariff.
Even with these caveats, Muse’s market positioning is exceptionally clear. While it may not be the absolute cheapest streaming transcription service on the market—Soniox currently publishes a lower equivalent rate—the $0.18 per hour price point, which includes robust diarization capabilities, firmly places Meta toward the lower end of the market. This is particularly true when compared to providers that levy separate charges for speaker attribution. For an enterprise processing 1,000 hours of audio, Meta’s public rate implies transcription charges of approximately $180, representing significant cost savings.
Beyond its aggressive pricing, Meta also leads its launch accuracy benchmarks. Price is a critical factor, but it is rendered less significant if it comes at a substantial accuracy penalty. Meta’s benchmark material strongly argues against this trade-off. On the Artificial Analysis AA-WER Streaming Index, which was supplied with the launch, Muse records a final-transcription word error rate of 3.1%. This performance places it ahead of several key competitors, including Cartesia Ink-2 at 3.4%, ElevenLabs Scribe v2 Realtime at 3.6%, Qwen3 ASR Flash Realtime at 3.7%, GPT Live Transcribe and Grok Speech to Text Streaming at 3.9%, and Gemini 3.5 Transcribe Live and AssemblyAI U3.5 Realtime Pro at 4.0%. Meta proudly points out that Muse secured the number one spot on third-party independent AI benchmarking firm Artificial Analysis’ streaming speech-to-text evaluation as of September 1. Meta published detailed benchmark charts illustrating these findings in its launch post.
The diarization results presented by Meta may be even more relevant to the product’s overall positioning. Meta reports an average diarization error rate of 17.5% across the AMI-IHM, AMI-SDM, and VoxConverse datasets, a figure that is lower than the competing systems depicted in their comparative chart. It is crucial to note that speaker capacity and diarization error rate should not be conflated. A platform capable of representing 100 individuals is not automatically superior in correctly attributing speech than one supporting 20. Moreover, Meta’s benchmark does not test every competitor operating at their advertised maximum speaker counts.
There are also certain deployment trade-offs to consider. Meta’s API currently provides turn-level timestamps but not word-level timestamps. It also does not expose word-level confidence scores, sound-event detection, or emotion detection. The documentation further specifies eight concurrent streams per tenant by default and real-time sessions of up to 60 minutes, after which an application must reconnect.
Nevertheless, Muse’s launch introduces an unusually sharp price-performance proposition to the market. While its 20-plus-speaker diarization capability may not establish a new world record, the record itself might be the less important metric for many enterprise users. For these developers, the more pressing question is whether a service can consistently preserve speaker attribution, deliver accurate text, and provide usable turn boundaries while a complex, real-world conversation is still unfolding. At $0.18 per hour, with integrated 20-plus-speaker diarization within the same real-time model that currently leads Meta’s supplied streaming accuracy benchmarks, Muse Voice Transcribe presents enterprise teams with a compelling new option for enhancing meeting intelligence, enabling live transcription services, and building robust voice-agent infrastructure. This development is expected to put additional pressure on competitors to focus not merely on raw speech recognition capabilities but also on speaker-aware accuracy and overall operating cost.

