Microsoft AI has unleashed a formidable new contender in the speech recognition arena with the release of MAI-Transcribe-2. This groundbreaking model, announced on Thursday, is positioned by the tech giant as not only faster and more accurate than existing solutions but also significantly cheaper, directly challenging established players like OpenAI, Google, and ElevenLabs. The audacious pricing strategy, set at a mere 10 cents per hour of audio, marks a dramatic shift in the market and underscores Microsoft’s aggressive push to dominate the AI landscape.
This price point warrants careful consideration. It represents a staggering 72% reduction from the company’s previous offering, MAI-Transcribe-1, which was launched just five months prior at a rate of $0.36 per hour. For large enterprises, such as major banks or telecommunications firms processing hundreds of thousands of hours of call-center audio annually, this price cut translates into substantial savings. A hypothetical scenario involving 100,000 hours of audio per year would see costs plummet from $36,000 to a mere $10,000. This drastic reduction effectively removes transcription costs as a significant line item in enterprise budgets, making advanced speech-to-text capabilities accessible and economically viable for a broader range of applications.
The introduction of MAI-Transcribe-2 is a clear manifestation of Microsoft’s evolving AI strategy. Two years ago, the idea of Microsoft independently developing and deploying frontier-class AI models across various modalities would have seemed improbable. Yet, the company is now systematically building its own specialized models, one after another, and integrating them into its existing product ecosystem, gradually replacing reliance on third-party technologies, particularly those from its partner OpenAI. Speech recognition represents the fastest-moving front in this strategic pivot, with MAI-Transcribe-2 serving as its most compelling proof of concept to date. This development also offers a glimpse into how Microsoft intends to maintain its competitive edge in the AI sector without being beholden to the partner for which it has invested a substantial $13 billion.
MAI-Transcribe-2: Unpacking the Enterprise Value Proposition
MAI-Transcribe-2 significantly expands its linguistic capabilities, now supporting 60 languages, a notable increase from the 43 languages offered by MAI-Transcribe-1.5 in June and the initial 25 languages of the April release. This enhanced multilingual support is crucial for global businesses and diverse customer bases. The model is accessible through Microsoft Foundry, the company’s model marketplace for developers, and within MAI Playground, its dedicated testing environment. Microsoft has engineered MAI-Transcribe-2 with real-world business audio in mind, acknowledging the inherent challenges of background noise, low-quality recordings, and overlapping speech, which are far more prevalent than the pristine conditions of studio recordings.
Beyond the sheer number of languages, the integrated features of MAI-Transcribe-2 are particularly impactful for enterprise buyers. Speaker diarization, a critical function for multi-person recordings, accurately attributes speech to individual speakers, transforming a monolithic block of text into a coherent and usable transcript. Word-level timestamps provide precise temporal markers for every word, enabling efficient searching, editing, and synchronization with video content. Keyword biasing allows developers to fine-tune the model by providing lists of domain-specific terms, such as product names, medical jargon, or employee identifiers, thereby mitigating the misinterpretation of specialized vocabulary. Furthermore, automatic language identification eliminates the need for users to manually specify the language of the audio beforehand.
Two features stand out for their particular utility and specificity. The configurable output style offers a "verbatim" mode, which meticulously preserves all conversational nuances, including filler words like "um" and false starts, essential for compliance and legal documentation. Conversely, a "clean" mode removes these fillers to create more readable captions and concise notes. The model also adeptly handles "code switching," enabling seamless transcription of conversations that fluidly transition between languages within the same sentence. Microsoft specifically highlights Hinglish and Spanglish, recognizing the significant markets in India and the U.S. Hispanic community, where customer service calls frequently involve rapid language shifts. Historically, specialized vendors have charged a premium for each of these advanced capabilities; Microsoft, however, includes them all as part of its remarkably low base price.
Decoding Microsoft’s Benchmark Claims: FLEURS and Artificial Analysis
Microsoft’s performance claims for MAI-Transcribe-2 are multifaceted, resting on distinct benchmarks, and technical buyers must understand the nuances of each to fully appreciate the model’s capabilities and limitations.
The first claim asserts that MAI-Transcribe-2 achieves the top ranking on the FLEURS benchmark across 60 languages, with an average word error rate (WER) of 5.2%. FLEURS, a benchmark introduced by Google researchers in 2022, comprises approximately 2,000 sentences read by native speakers in each of 102 languages, totaling around 12 hours of speech per language. It is widely regarded as the standard for evaluating multilingual speech recognition due to its ability to compare a model’s performance on diverse languages using identical content. The WER metric quantifies errors by counting substitutions, insertions, and deletions against a human reference, meaning a 5.2% WER indicates roughly one error for every 20 words. However, it is important to note that FLEURS utilizes read speech, not spontaneous conversation. Microsoft’s reported average WER has actually increased from the 3.7% achieved by MAI-Transcribe-1.5 in June. This rise is likely attributable to the expanded language coverage rather than a performance regression; averaging across 60 languages inherently includes lower-resource languages where all models face greater challenges. Consequently, buyers are advised to request per-language breakdowns for a more granular understanding.
The second performance assertion places MAI-Transcribe-2 second on the Artificial Analysis word-error-rate leaderboard, a position that also defines the model’s accuracy-latency Pareto frontier. Artificial Analysis serves as an independent benchmark, evaluating models through their public APIs to reflect real-world customer experiences. Its benchmark incorporates a blend of simulated agent conversations, European Parliament speeches, and corporate earnings calls, with a significant weighting towards English business speech. In June, MAI-Transcribe-1.5 ranked third on this leaderboard with a WER of 2.4%, trailing behind Alibaba’s Fun-ASR and ElevenLabs’ Scribe v2, while being recognized as the fastest model within the top 10. The climb to second place suggests Microsoft has successfully surpassed ElevenLabs. The term "Pareto frontier" is particularly significant for practitioners, indicating that no competitor can outperform MAI-Transcribe-2 on accuracy without sacrificing speed, nor can any rival match its speed without compromising accuracy.
The third claim focuses on raw speed, reporting that MAI-Transcribe-2 is ten times faster than OpenAI’s GPT-Transcribe, seven times faster than ElevenLabs’ Scribe v2, and five times faster than Google’s Gemini 3.5 Transcribe, according to Artificial Analysis evaluations. In the context of batch transcription, speed’s importance transcends mere waiting time; it directly impacts throughput and, consequently, cost. A model operating at 300 times real-time requires a fraction of the GPU hours compared to one operating at 30 times real-time. This efficiency is the fundamental enabler of Microsoft’s aggressive 10-cent pricing strategy, allowing for profitability while offering a dramatically reduced cost to customers.
A Relentless Cadence: Three Speech Models in Five Months at Microsoft AI
The astonishing pace of development is a story in itself. Microsoft AI has demonstrated an extraordinary release cadence, launching three successive iterations of its transcription model within a span of just five months. On April 2, MAI-Transcribe-1 debuted with 25 languages and a price of $0.36 per hour. This was followed on June 2 by MAI-Transcribe-1.5, which expanded language support to 43, introduced keyword biasing, and secured a third-place ranking on Artificial Analysis. Most recently, on Thursday, MAI-Transcribe-2 was released, boasting 60 languages, speaker diarization, timestamps, code switching capabilities, a second-place benchmark ranking, and the groundbreaking $0.10 per hour price point.
This rapid succession of releases, each enhancing language coverage by approximately 40% while integrating features that competitors typically reserve for premium tiers, points to a team that has solidified its core architecture and is now aggressively scaling data and performance. This is the phase where speech models are known to improve rapidly and predictably. It also signals Microsoft’s intent to commoditize transcription services before its rivals can establish a comparable offering.
The organizational philosophy underpinning this accelerated development was articulated by Mustafa Suleyman, CEO of Microsoft AI, in an April interview with The Verge. He attributed the success of the initial model to "a small, focused 10-person team" that operated with "liberated bureaucracy," supported by a larger group managing vendor relations and data acquisition. Suleyman also noted that the initial model operated at "half the GPU cost of the other state-of-the-art models," highlighting it as a "huge cost-saving" for Microsoft. This lean, agile approach mirrors similar experiments by tech giants like Meta, Amazon, Google, and Anthropic, with Microsoft’s transcription line serving as a prominent test case for the commercial viability of such structures beyond pure research.
Strategic Independence: Microsoft’s In-House AI Model Development Amidst the OpenAI Partnership
Microsoft’s substantial investment of over $13 billion in OpenAI and its integration of OpenAI’s models across Azure, Office, and Copilot have long posed the question: why is Microsoft investing so heavily in developing its own AI models? The answer has become increasingly clear over the past year, and it centers on achieving greater independence and strategic control.
The hiring of Mustafa Suleyman in March 2024 from Inflection AI, along with a significant portion of his team, was widely interpreted as a declaration of intent. Marc Benioff, CEO of Salesforce, a direct competitor and investor in Anthropic, famously stated in January 2025 that Microsoft was "building their own AI and I don’t think Microsoft will use OpenAI in the future. They’ll have their own frontier models. That’s why they hired Mustafa Suleyman." Events have largely corroborated this assessment.
In October 2025, Microsoft and OpenAI restructured their partnership. According to Microsoft’s announcement, this renegotiation granted Microsoft the unprecedented ability to "independently pursue AGI alone or in partnership with third parties." Suleyman confirmed to The Verge that this renegotiation "unlocked [Microsoft’s] ability to pursue superintelligence," a sentiment echoed by Microsoft’s subsequent announcement of its MAI Superintelligence team weeks later. Further amendments to the partnership in April 2026 reportedly ended Microsoft’s exclusive access to OpenAI’s models and eliminated its revenue-share payments, according to reports at the time. Each of these adjustments has progressively loosened the operational ties between the two companies and has been consistently followed by the release of new MAI models.
Margin Maximization: In-House Models Drive Down Costs Across Microsoft’s Product Suite
The second key driver for Microsoft’s in-house AI development is profitability. Every prompt routed to an OpenAI model incurs a direct cost for Microsoft. Conversely, processing prompts using its own models on its own GPUs significantly reduces this expenditure. In July, Bloomberg reported that Microsoft had begun leveraging its MAI models to handle a portion of user prompts within Word and Excel, applications previously advertised as powered by OpenAI and Anthropic. This shift was framed by TechCrunch as part of a broader industry trend of AI cost optimization, with major players like Amazon, Uber, Meta, and Accenture reportedly implementing similar spending reductions.
Transcription is a logical starting point for this cost-saving strategy due to the well-defined nature of the problem and the objective measurability of its outcomes. Microsoft’s ownership of Teams, which generates vast quantities of meeting audio, and its acquisition of Nuance, a leader in clinical documentation relying heavily on speech recognition, position it uniquely. Furthermore, its existing Azure speech services are already utilized by thousands of enterprises. Each of these existing workloads represents a prime candidate for migration to MAI-Transcribe-2, and every hour of audio processed internally translates into direct cost savings for Microsoft, eliminating payments to external providers.
Suleyman has been remarkably transparent about this strategic imperative. He articulated to The Verge in April that the pursuit of "superintelligence" is fundamentally about "Are these models capable of delivering product value for the millions of enterprises that depend on us to deliver world-class language models?" While the term "superintelligence" might seem grand for a transcription API, the underlying commercial logic is undeniable: build a core capability once, deploy it across a multitude of products, and cease making payments to a partner that is increasingly becoming a direct competitor.
MAI-Transcribe-2 vs. the Competition: A Strategic Market Positioning
Microsoft’s announcement explicitly names four key rivals: OpenAI’s GPT-Transcribe, Google’s Gemini 3.5 Transcribe, OpenAI’s older Whisper V3-Large, and ElevenLabs’ Scribe v2. Notably absent from this list are dedicated transcription specialists like Deepgram, AssemblyAI, Speechmatics, and Rev, companies that have been serving the enterprise market for a decade. This strategic omission highlights Microsoft’s intention to position MAI-Transcribe-2 against the frontier AI labs, rather than directly confronting established incumbent transcription providers.
This framing is partly a marketing tactic and partly an accurate reflection of the competitive landscape. The frontier labs have largely treated speech recognition as an ancillary feature within broader AI platforms, with pricing reflecting this status. A dedicated model that demonstrably surpasses these offerings in speed by five to ten times, while matching their accuracy, represents a genuine competitive differentiator. However, the most acute pressure will undoubtedly be felt by the specialized transcription vendors. At $0.10 per hour, Microsoft is pricing at or below the rates many of these specialists charge for high-volume enterprise contracts, and crucially, it is bundling advanced features such as diarization, timestamps, and extensive language support into its base price. The remaining competitive advantage for these specialists lies in their domain-specific expertise—medical vocabularies, legal formatting, and deep industry integrations. Microsoft’s keyword biasing feature, however, directly targets this niche by enabling users to inject their own domain-specific terminology.
One competitor that Microsoft does not claim to surpass in accuracy is Alibaba, whose models have consistently achieved leading positions on independent benchmarks throughout 2026. Reports from July indicated that some U.S. companies were exploring Chinese models as more cost-effective alternatives, despite potential security concerns. Microsoft’s implicit but powerful proposition to these potential buyers is straightforward: comparable accuracy, superior inference speed, a lower price point, and the assurance of a vendor with an established reputation and a compliance framework that enterprises already trust.
Critical Considerations for Technical Decision-Makers Evaluating Transcription Vendors
Despite the detailed benchmark claims, the MAI-Transcribe-2 announcement leaves several practical questions for technical decision-makers to address before committing to a new vendor. The first pertains to pricing duration. Microsoft has designated the $0.10 per hour rate as a "launch offer" without specifying an end date or a subsequent standard rate. Any enterprise developing a cost model should secure these details in writing. The second critical area is streaming capabilities. While the release emphasizes batch throughput and long-form audio processing, it remains silent on real-time transcription, a vital component for applications such as voice agents and live captioning. Artificial Analysis maintains a separate leaderboard for streaming performance, and Microsoft’s omission in this regard is noteworthy.
Thirdly, per-language accuracy remains a concern. An average WER of 5.2% across 60 languages could mask significant disparities, with higher accuracy in major languages and substantially lower accuracy in less common ones. Organizations with specific language requirements should conduct direct testing. Fourth, the quality of diarization requires further clarification. Word error rate does not directly measure the accuracy of speaker attribution. A transcript could achieve near-perfect WER while incorrectly assigning sentences to the wrong speakers. The release provides no specific diarization error rate or comparable metric.
Finally, data handling protocols are paramount. Enterprise transcription often involves sensitive data, including medical records, legally privileged information, and financial disclosures. The announcement offers no details regarding data residency, retention policies, or whether audio submitted to Foundry contributes to future model training. Microsoft’s April announcements described training data as a composite of human-curated recordings, contractor-recorded noisy audio, and "vast amounts of data from the open web," a description that warrants detailed scrutiny from regulated industries. While these omissions are not unusual for a product launch announcement, they represent precisely the critical questions that differentiate a theoretical benchmark win from a practical, production-ready deployment.
A Blueprint for Broader AI Ambitions: Microsoft’s Specialized Model Strategy
Stepping back from the granular details of speech recognition, a clear pattern emerges that extends far beyond transcription. Microsoft’s AI unit is actively developing and deploying specialized models across a wide spectrum of modalities, including images, voice, transcription, code generation, reasoning, and cybersecurity. At its Build conference in June, the company unveiled seven new MAI models in a single keynote. The strategy for each follows a consistent playbook: target a well-defined modality, aggressively optimize for inference cost, price below established frontier AI providers, distribute through Foundry, and systematically integrate these models into Microsoft’s own extensive product portfolio.
This approach is not an attempt to create a single, monolithic model that can outperform GPT or Gemini across all tasks. Instead, it represents a strategic effort to build a diverse portfolio of specialized models. The aggregate effect of this portfolio is to enable Microsoft to serve the majority of its enterprise workloads internally, thereby reducing reliance on external providers. The surplus capacity generated by this strategy is then made available to other customers at prices that specialized vendors find difficult to match. Transcription has emerged as the first modality where this comprehensive strategy has reached full maturity, but the release notes for MAI-Transcribe-2 read less like a product announcement and more like a template for future developments.
Suleyman has consistently spoken about "humanist superintelligence" and the development of AI assistants that are "accountable to them, on their side." While the rhetoric may be lofty, the execution is demonstrably grounded in pragmatic commercial objectives. Five months ago, Microsoft charged 36 cents to convert an hour of speech to text. On Thursday, it announced a price of 10 cents, bundled six advanced features typically offered as premium add-ons by competitors, and claimed the top spot on the industry’s leading multilingual benchmark. The company that invested $13 billion to understand the costs associated with frontier AI has now chosen to own the manufacturing process rather than simply renting the output—and is now selling that output at a price lower than the rental fee.
MAI-Transcribe-2 is currently available for use via Microsoft Foundry and MAI Playground.

