Microsoft AI has made a significant splash in the artificial intelligence landscape with the release of MAI-Transcribe-2, a sophisticated speech-recognition model that the company asserts surpasses current offerings from industry titans like OpenAI, Google, and ElevenLabs in terms of speed, accuracy, and cost-effectiveness. Priced at an astonishing 10 cents per hour of audio, this new model represents a dramatic reduction from its predecessor, MAI-Transcribe-1, which was launched just five months prior at a significantly higher rate of $0.36 per hour. This represents a staggering 72% price cut, translating into substantial savings for enterprises. For a large organization processing 100,000 hours of call-center audio annually, the cost plummets from $36,000 to a mere $10,000, effectively transforming transcription from a significant operational expense into a negligible line item.
This strategic release underscores Microsoft’s evolving AI strategy: the meticulous development of its own frontier-class models across various modalities, with the ultimate goal of integrating them into its existing product ecosystem, thereby reducing reliance on external partners like OpenAI, despite its substantial $13 billion investment. Transcription appears to be the modality where this strategy is manifesting most rapidly, with MAI-Transcribe-2 serving as a compelling proof point for Microsoft’s ambition to compete at the forefront of AI innovation without being beholden to its former key collaborator.
MAI-Transcribe-2 boasts an impressive multilingual capability, supporting 60 languages, a significant expansion from the 43 supported by MAI-Transcribe-1.5 released in June and the 25 languages available with the initial MAI-Transcribe-1. This model is readily accessible through Microsoft Foundry, the company’s model marketplace for developers, and within MAI Playground, its dedicated testing environment. Crucially, Microsoft has engineered MAI-Transcribe-2 to excel in real-world business scenarios, adeptly handling the challenges of noisy environments, low-quality recordings, and overlapping speech, rather than being optimized solely for pristine studio conditions.
Beyond its expansive language support, the true value for enterprise buyers lies in the comprehensive feature set integrated into the base product. Speaker diarization, a critical function for multi-participant recordings, accurately identifies and attributes dialogue to specific speakers, transforming a monolithic block of text into a comprehensible and actionable transcript. Word-level timestamps provide precise temporal markers for every word, facilitating efficient search, editing, and synchronization with video content. Keyword biasing empowers developers to improve transcription accuracy for domain-specific jargon, such as pharmaceutical names, product codes, or employee titles, by providing the model with relevant lexicons. Furthermore, automatic language identification eliminates the need for users to manually specify the language of the audio beforehand.
Two standout features further enhance MAI-Transcribe-2’s enterprise appeal. A configurable output style offers a "verbatim" mode that meticulously preserves all speech disfluencies, including "ums," false starts, and stutters, which are invaluable for legal and compliance teams. Conversely, a "clean" mode strips out these fillers to produce more readable captions and notes. The model also exhibits remarkable code-switching capabilities, seamlessly handling conversations that transition between languages within a single sentence. Microsoft specifically highlights Hinglish and Spanglish as examples, acknowledging the linguistic fluidity often encountered in markets like India and the U.S. Hispanic community, where customer service calls can frequently switch languages multiple times. Historically, specialized vendors have commanded premium prices for each of these advanced functionalities; Microsoft, however, bundles them all into MAI-Transcribe-2 for a mere dime per hour.
Microsoft’s performance claims for MAI-Transcribe-2 are substantiated by several benchmark evaluations, each offering a distinct perspective on its capabilities. The first claim positions MAI-Transcribe-2 as the number one model on the FLEURS benchmark across all 60 languages, achieving an average word error rate (WER) of 5.2%. FLEURS, a benchmark developed by Google researchers in 2022, comprises approximately 12 hours of speech per language, with native speakers reading around 2,000 sentences. It serves as a widely accepted standard for multilingual speech recognition, enabling direct comparison of a model’s performance across diverse linguistic contexts. WER, the primary metric, quantifies errors through substitutions, insertions, and deletions against a human-verified reference transcript, with a 5.2% WER indicating roughly one incorrect word for every 20. However, it’s important to note that FLEURS primarily assesses read speech, not spontaneous conversation. The reported average WER for MAI-Transcribe-2 is a slight increase from the 3.7% reported for MAI-Transcribe-1.5 in June. This rise is likely attributable to the broader language coverage, which includes lower-resource languages where all models inherently struggle, rather than a regression in performance. Nevertheless, enterprise buyers are advised to request a per-language breakdown for a more granular understanding of performance.
The second performance claim highlights MAI-Transcribe-2’s second-place ranking on the Artificial Analysis word-error-rate leaderboard, defining its position on the firm’s accuracy-latency Pareto frontier. Artificial Analysis is an independent benchmarking organization that evaluates models through their public APIs, providing insights into real-world performance experienced by customers. Their index incorporates a diverse range of speech data, including simulated agent conversations, European Parliament speeches, and corporate earnings calls, with a significant weighting towards English business speech. In June, MAI-Transcribe-1.5 secured third place with a 2.4% WER, trailing Alibaba’s Fun-Realtime-ASR-preview and ElevenLabs’ Scribe v2, while simultaneously being recognized as the fastest model among the top 10. The climb to second place suggests that Microsoft has successfully surpassed ElevenLabs. The term "Pareto frontier" is particularly significant for practitioners, indicating that no competitor can achieve higher accuracy without sacrificing speed, nor can they achieve greater speed without compromising accuracy.
The third claim focuses on raw speed, asserting that MAI-Transcribe-2 is ten times faster than OpenAI’s GPT-Transcribe, seven times faster than ElevenLabs’ Scribe v2, and five times faster than Google’s Gemini 3.5 Transcribe, according to Artificial Analysis evaluations. In batch transcription scenarios, speed’s importance extends beyond mere waiting times to encompass throughput and cost efficiency. A model operating at 300 times real-time requires a fraction of the GPU hours compared to one operating at 30 times real-time. This enhanced efficiency is precisely what enables Microsoft to offer MAI-Transcribe-2 at a 10-cent price point while maintaining profitability.
The rapid release cadence of Microsoft’s speech models is a story in itself. On April 2nd, MAI-Transcribe-1 was launched with 25 languages at $0.36 per hour. This was followed by MAI-Transcribe-1.5 on June 2nd, which expanded language support to 43, introduced keyword biasing, and achieved a third-place ranking on Artificial Analysis. Now, MAI-Transcribe-2 arrives with 60 languages, speaker diarization, timestamps, code switching capabilities, a second-place ranking, and the remarkably low price of $0.10 per hour. This pace of three releases in five months, each substantially increasing language coverage by approximately 40% while incorporating features that competitors often relegate to premium tiers, suggests a team that has established a stable architecture and is now focusing on data scaling and refinement. This is the phase where speech models typically experience rapid and predictable improvements. It also signals Microsoft’s strategic intent to commoditize transcription services before its competitors can.
The organizational structure behind this rapid development was articulated by Mustafa Suleyman, CEO of Microsoft AI, in an interview with The Verge in April. He attributed the success of the first model to a "small, focused 10-person team" that was "liberated from any of the bureaucracy," supported by a larger group managing vendor relations and data acquisition. Suleyman also highlighted that the initial model operated at "half the GPU cost of the other state-of-the-art models," presenting a "huge cost-saving" for Microsoft. This lean, agile approach mirrors similar flattened structures experimented with by tech giants like Meta, Amazon, Google, and Anthropic. Microsoft’s transcription line serves as a prominent test case for the commercial viability of this organizational model.
Microsoft’s substantial investment of over $13 billion in OpenAI and its hosting of OpenAI’s models across Azure, Office, and Copilot have long raised the question: why develop in-house models? The answer has become increasingly clear over the past year, centered on the pursuit of independence. The recruitment of Mustafa Suleyman from Inflection AI in March 2024, along with a significant portion of Inflection’s staff, was interpreted by Salesforce CEO Marc Benioff as a definitive statement of intent. Benioff predicted that Microsoft would increasingly rely on its own frontier models, moving away from OpenAI, a sentiment driven by his own company’s competitive landscape and investments in rival AI firms like Anthropic.
Subsequent developments have largely validated this perspective. In October 2025, Microsoft and OpenAI renegotiated their partnership, a move that, according to Microsoft’s official announcement, granted Microsoft the unprecedented ability to "independently pursue AGI alone or in partnership with third parties." Suleyman confirmed to The Verge that this renegotiation "unlocked [Microsoft’s] ability to pursue superintelligence," leading to the subsequent announcement of its MAI Superintelligence team. Further amendments to the partnership in April 2026 reportedly ended Microsoft’s exclusive access to OpenAI’s models and eliminated revenue-share payments, according to contemporary reports. Each of these contractual adjustments progressively loosened the ties between the two companies, and each was followed by the introduction of new MAI models.
The second crucial aspect of Microsoft’s in-house AI strategy is profit margin. Every query routed to an OpenAI model incurs a cost for Microsoft. Conversely, queries processed by its own models on its own GPUs incur a significantly lower cost. In July, Bloomberg reported that Microsoft had begun integrating its MAI models into applications like Word and Excel, which were previously advertised as powered by OpenAI and Anthropic. This shift was framed by TechCrunch as part of a broader industry trend of AI cost-cutting, with other major companies like Amazon, Uber, Meta, and Accenture reportedly undertaking similar measures.
Transcription represents a natural starting point for this substitution strategy due to the well-defined nature of the problem and the objective measurability of its output. Microsoft’s ownership of Teams, a platform generating vast amounts of meeting audio, and its acquisition of Nuance, a company whose clinical documentation business relies heavily on speech recognition, provide substantial internal use cases. Furthermore, Azure’s existing speech services cater to thousands of enterprises. Each of these workloads presents an opportunity to transition to MAI-Transcribe-2, thereby eliminating external vendor costs.
Suleyman has openly acknowledged this strategic objective. He stated to The Verge in April that the pursuit of superintelligence is fundamentally about "Are these models capable of delivering product value for the millions of enterprises that depend on us to deliver world-class language models?" While the term "superintelligence" may seem grandiose for a transcription API, the commercial logic is undeniable: develop a capability once, deploy it across a multitude of products, and cease paying a partner that is increasingly becoming a competitor.
Microsoft’s release explicitly names four key rivals: OpenAI’s GPT-Transcribe, Google’s Gemini 3.5 Transcribe, OpenAI’s older Whisper V3-Large, and ElevenLabs’ Scribe v2. Notably absent from this list are established transcription specialists like Deepgram, AssemblyAI, Speechmatics, and Rev, companies that have served enterprise clients for a decade. This strategic framing positions Microsoft against the frontier AI labs rather than the incumbent service providers.
This positioning is partly strategic marketing and partly a reflection of reality. The frontier labs have historically treated speech recognition as a supplementary feature within broader AI platforms, with pricing reflecting this approach. A dedicated model like MAI-Transcribe-2, which surpasses these platforms in speed by a factor of five to ten while matching their accuracy, represents a genuine competitive differentiator. However, the most significant impact of this price disruption will likely be felt by the specialized transcription vendors. At $0.10 per hour, Microsoft’s pricing is at or below the enterprise contract rates offered by many of these specialists, and it includes advanced features like diarization, timestamps, and 60 languages as standard. The remaining competitive advantage for these specialists lies in domain-specific expertise, such as specialized medical or legal vocabularies and industry-specific integrations. Microsoft’s keyword biasing feature directly challenges this moat.
Interestingly, Microsoft does not claim to outperform Alibaba in accuracy, a company whose models have consistently ranked at the top of independent leaderboards throughout 2026. Reports from July indicated that some U.S. companies were exploring Chinese models as more cost-effective alternatives, despite potential security concerns. Microsoft’s implicit pitch to these buyers is clear: comparable accuracy, superior inference speed, a lower price point, and a trusted vendor for compliance.
Despite the detailed performance benchmarks, the MAI-Transcribe-2 release leaves several practical questions for technical decision-makers. Firstly, the $0.10 per hour pricing is presented as a launch offer, with no specified end date or standard rate. Any enterprise developing cost models should secure these details in writing. Secondly, the announcement emphasizes batch throughput and long-form audio transcription, with no mention of real-time transcription capabilities, which are essential for applications like voice agents and live captioning. Artificial Analysis maintains a separate leaderboard for streaming performance, and Microsoft’s silence on this front is notable.
Thirdly, per-language accuracy remains a critical consideration. An average WER of 5.2% across 60 languages could mask significant performance disparities, with major languages potentially achieving 3% WER while lower-resource languages struggle at 12%. Buyers with specific language needs should conduct their own direct testing. Fourthly, the quality of diarization is not quantitatively assessed. While word error rate measures transcription accuracy, it does not evaluate speaker attribution. A transcript could achieve near-perfect WER yet assign dialogue to the wrong speakers. The release provides no diarization error rate or comparable metric.
Fifthly, data handling practices are not fully detailed. Enterprise transcription frequently involves sensitive data, including medical records, legal privilege, and financial disclosures. The release offers no information on data residency, retention policies, or whether audio submitted to Foundry will be used for future model training. Microsoft’s April announcements described training data as a combination of human-curated recordings, contractor-recorded noisy audio, and "vast amounts of data from the open web," a description that warrants careful scrutiny from regulated industries. While these omissions are not unusual for a product launch announcement, they represent precisely the kind of details that differentiate a benchmark achievement from a production-ready deployment.
Stepping back from the specifics of speech recognition, a discernible pattern emerges that extends across Microsoft’s broader AI initiatives. Microsoft AI is now actively developing and releasing models for a diverse range of modalities, including images, voice, transcription, code generation, reasoning, and cybersecurity. At the Build conference in June, the company announced seven new MAI models in a single keynote. The underlying strategy for each of these models is consistent: target a clearly defined modality, aggressively optimize for inference cost, price below the leading frontier labs, distribute through Foundry, and discreetly integrate the models into Microsoft’s own product suite.
This strategy is not an attempt to create a single, all-encompassing model that rivals GPT or Gemini across all tasks. Instead, it focuses on building a portfolio of specialized models designed to enable Microsoft to serve the majority of its enterprise workloads internally, thereby reducing reliance on external vendors. The surplus capacity generated by this approach is then offered to other businesses at prices that specialist providers cannot match. Transcription has proven to be the first modality where this strategy has fully matured, but the release notes for MAI-Transcribe-2 suggest a template for future product announcements.
Mustafa Suleyman has consistently articulated a vision of "humanist superintelligence" and AI assistants that are "accountable to them, on their side." While the language may be aspirational, the execution is grounded in practical business realities. Five months ago, Microsoft was charging 36 cents to convert an hour of speech to text. Today, it is charging a dime, bundling six features that competitors sell separately, and claiming the top spot on the industry’s standard multilingual benchmark. Microsoft, having invested heavily in understanding the cost of frontier AI, has evidently decided that owning the manufacturing process is more advantageous than merely renting the output. Now, it is offering that output at a price lower than the rental cost.
MAI-Transcribe-2 is currently available through Microsoft Foundry and MAI Playground.

