Microsoft AI has dramatically reshaped the landscape of speech recognition with the Thursday release of MAI-Transcribe-2, a groundbreaking model that the tech giant claims surpasses current offerings from industry titans like OpenAI, Google, and ElevenLabs in speed, accuracy, and cost-effectiveness. This ambitious announcement is underscored by a jaw-dropping price point: just 10 cents per hour of audio processed. This figure represents a staggering 72% reduction from the $0.36 per hour charged for the first iteration of the MAI-Transcribe line just five months prior. For an enterprise handling a moderate volume of 100,000 hours of call-center audio annually—a common scenario for large financial institutions or telecommunications companies—this price cut translates into a dramatic savings, slashing annual transcription costs from $36,000 to a mere $10,000. At such a competitive price, transcription expenses are poised to become a negligible concern for businesses.
This strategic release signals Microsoft’s aggressive pivot towards developing its own frontier-class AI models, a strategy that seemed almost unimaginable just two years ago. The company is methodically building and integrating proprietary models across various AI modalities, progressively replacing reliance on OpenAI’s technology within its extensive product suite. Speech transcription has emerged as the domain where this transition has been most rapid and impactful, with MAI-Transcribe-2 serving as its most compelling testament. More significantly, this development offers a clear preview of how the world’s most valuable software company intends to navigate the fiercely competitive AI arena without over-dependence on its partner, a relationship cemented by a substantial $13 billion investment.
MAI-Transcribe-2: Unpacking Features Crucial for Enterprise Adoption
MAI-Transcribe-2’s capabilities extend across an impressive 60 languages, a significant expansion from the 43 supported by MAI-Transcribe-1.5 in June and the 25 languages of the initial April release. The model is accessible through Microsoft Foundry, the company’s model marketplace designed for developers, and within MAI Playground, its dedicated testing environment. Microsoft emphasizes that MAI-Transcribe-2 is engineered to excel with the real-world, often imperfect audio encountered by businesses, including challenges like background noise, low-quality recordings, and overlapping speech, rather than being optimized solely for pristine studio conditions.
Beyond the sheer volume of languages, the integrated features bundled into the base product are of paramount importance to enterprise buyers. Speaker diarization, a critical component, distinguishes between different speakers within a multi-person recording, transforming a monolithic block of text into a structured and comprehensible transcript of meetings or conversations. Word-level timestamps provide precise temporal markers for every word, enabling efficient search functionality, seamless editing, and accurate alignment with video content. Keyword biasing allows developers to provide the model with custom lists of domain-specific terms, such as drug names, product codes, or employee titles, significantly improving the accuracy of jargon transcription and preventing common misinterpretations. Furthermore, automatic language identification eliminates the need for users to pre-select the language, streamlining the transcription process.
Two features, in particular, stand out for their targeted utility. A configurable output style offers a "verbatim" mode, meticulously preserving every "um," false start, and stutter—essential for compliance and legal teams requiring exhaustive records. Conversely, a "clean" mode strips out these disfluencies, producing more readable captions and concise notes. The model also boasts impressive code-switching capabilities, adeptly handling conversations that seamlessly transition between languages within a single sentence. Microsoft explicitly references Hinglish and Spanglish, acknowledging the significant markets in India and the U.S. Hispanic community where customer service interactions may frequently switch between languages. Historically, specialized vendors have levied premium charges for each of these sophisticated functionalities; Microsoft is now bundling them all at an unprecedentedly low price.
Deconstructing Microsoft’s Benchmark Claims: FLEURS and Artificial Analysis
Microsoft’s performance assertions are built upon three distinct claims, each employing a different measurement standard. Technical buyers must critically understand the strengths and limitations of each benchmark.
Firstly, MAI-Transcribe-2 is positioned as the top-ranked model on the FLEURS benchmark across all 60 languages, achieving an average word error rate (WER) of 5.2%. Developed by Google researchers and published in 2022, FLEURS is a widely adopted benchmark for multilingual speech recognition. It comprises approximately 12 hours of recorded speech per language, with native speakers reading around 2,000 sentences. This standardized approach allows for direct comparison of a model’s performance across diverse languages using identical content. WER quantifies errors by counting substitutions, insertions, and deletions against a human-verified reference; a 5.2% WER signifies that, on average, one out of every twenty words is transcribed incorrectly. However, it’s crucial to note that FLEURS primarily evaluates read speech, not spontaneous conversation. Microsoft’s reported average WER has actually increased from the 3.7% achieved by MAI-Transcribe-1.5 in June. This likely reflects the broader language coverage of MAI-Transcribe-2, which includes lower-resource languages where all models typically struggle, rather than a regression in performance on established languages. Consequently, potential buyers should request a per-language breakdown to assess performance in their specific linguistic contexts.
Secondly, the model secures the second position on the Artificial Analysis word-error-rate leaderboard, defining its accuracy-latency Pareto frontier. Artificial Analysis operates as an independent benchmarking service, evaluating models through their public APIs to reflect real-world customer experiences. Its comprehensive index incorporates simulated agent conversations, speeches from the European Parliament, and corporate earnings calls, with a significant weighting towards English business-oriented speech. In June, MAI-Transcribe-1.5 held the third spot on this leaderboard with a 2.4% WER, trailing behind Alibaba’s Fun-ASR and ElevenLabs’ Scribe v2, while simultaneously being recognized as the fastest model among the top ten. The climb to second place suggests Microsoft has effectively surpassed ElevenLabs in this evaluation. The term "Pareto frontier" is particularly relevant for practitioners, indicating that no competitor can outperform MAI-Transcribe-2 on accuracy without sacrificing speed, nor can they beat it on speed without compromising accuracy.
Thirdly, the claim of raw speed is substantial. Artificial Analysis evaluations indicate MAI-Transcribe-2 is ten times faster than OpenAI’s GPT-Transcribe, seven times faster than ElevenLabs’ Scribe v2, and five times faster than Google’s Gemini 3.5 Transcribe. In the context of batch transcription, speed is less about immediate user waiting times and more about maximizing throughput and reducing operational costs. A model operating at 300 times real-time requires a fraction of the GPU hours compared to one operating at 30 times real-time. This inherent efficiency is the bedrock upon which Microsoft can offer its 10-cent-per-hour pricing and still maintain profitability.
The Accelerating Cadence: Three Speech Models in Five Months at Microsoft AI
The rapid release schedule itself is a compelling narrative. On April 2nd, MAI-Transcribe-1 debuted with 25 languages at $0.36 per hour. Just two months later, on June 2nd, MAI-Transcribe-1.5 arrived, expanding language support to 43, introducing keyword biasing, and achieving a third-place ranking on Artificial Analysis. Now, MAI-Transcribe-2 has been launched with 60 languages, integrated diarization and timestamps, code-switching capabilities, a second-place ranking, and the unprecedented price of $0.10 per hour.
This cadence of three significant releases within five months, each boosting language coverage by approximately 40% and incorporating features that competitors often relegate to premium tiers, points to a team that has likely established a stable architectural foundation. They are now aggressively scaling data and optimization efforts—a phase where speech models typically exhibit rapid and predictable improvements. This also signifies Microsoft’s strategic intent to commoditize transcription services before its rivals can effectively respond.
The organizational philosophy underpinning this accelerated pace was articulated by Mustafa Suleyman, CEO of Microsoft AI, in an April interview with The Verge. He attributed the success of the initial model to a "small, focused 10-person team" operating with minimal bureaucracy, supported by a larger group managing vendor relationships and data acquisition. Suleyman also highlighted the cost efficiencies, noting that the model ran at "half the GPU cost of other state-of-the-art models," representing a "huge cost-saving" for Microsoft. This lean, agile approach mirrors experiments seen at Meta, Amazon, Google, and Anthropic, with Microsoft’s transcription line serving as the most prominent real-world test of its commercial viability.
Why Microsoft is Forging Its Own Path in AI Despite the OpenAI Investment
Microsoft’s substantial $13 billion investment in OpenAI and its hosting of OpenAI’s models across Azure, Office, and Copilot have long raised the question: why develop in-house models? The answer has become increasingly clear over the past year, centering on the strategic imperative of independence.
The recruitment of Mustafa Suleyman from Inflection AI in March 2024, along with a significant portion of Inflection’s staff, was interpreted by industry leaders as a definitive statement of Microsoft’s strategic direction. Salesforce CEO Marc Benioff, in a January 2025 interview with CNBC, opined that Microsoft was "building their own AI and I don’t think Microsoft will use OpenAI in the future. They’ll have their own frontier models. That’s why they hired Mustafa Suleyman." While Benioff’s perspective may have been influenced by Salesforce’s own competitive interests and investments in Anthropic, subsequent events have largely validated his assessment.
A pivotal moment occurred in October 2025 when Microsoft and OpenAI restructured their partnership. Microsoft’s own announcement indicated this deal granted them the unprecedented ability to "independently pursue AGI alone or in partnership with third parties." Suleyman elaborated to The Verge that this renegotiation "unlocked [Microsoft’s] ability to pursue superintelligence," a sentiment followed by Microsoft’s announcement of its MAI Superintelligence team weeks later. Further amendments in April 2026 reportedly ended Microsoft’s exclusive access to OpenAI’s models and eliminated its revenue-share payments, according to contemporaneous reports. Each of these contractual shifts progressively loosened Microsoft’s ties to OpenAI, and each was followed by the introduction of new MAI models.
Cost Efficiencies: Integrating In-House Models Across Microsoft’s Product Ecosystem
The second crucial element driving Microsoft’s in-house model development is margin enhancement. Every query directed to an OpenAI model incurs a cost for Microsoft, whereas queries processed by its own models on its own GPUs represent a significantly lower expense. In July, Bloomberg reported that Microsoft had begun integrating MAI models to handle a portion of user prompts within Word and Excel, applications previously marketed as powered by OpenAI and Anthropic. This shift was framed by TechCrunch as part of a broader industry trend of AI spending retrenchment, with major players like Amazon, Uber, Meta, and Accenture reportedly implementing cost-cutting measures.
Transcription represents a natural and logical starting point for this internal substitution strategy. The problem domain is well-defined, and its success can be objectively measured. Microsoft’s ownership of Teams generates a vast amount of meeting audio, and its acquisition of Nuance brings a significant clinical documentation business heavily reliant on speech recognition. Furthermore, Azure’s existing speech services already cater to thousands of enterprises. Each of these existing workloads presents a prime candidate for migration to MAI-Transcribe-2. Every hour of audio transcribed internally translates directly into savings by eliminating payments to external providers.
Suleyman has been remarkably transparent about this strategic objective. He articulated to The Verge in April that the pursuit of "superintelligence" is fundamentally about "Are these models capable of delivering product value for the millions of enterprises that depend on us to deliver world-class language models?" Regardless of how one interprets the application of "superintelligence" to a transcription API, the commercial logic is undeniable: develop a capability once, deploy it across numerous products, and cease making substantial payments to a partner that is increasingly morphing into a competitor.
MAI-Transcribe-2 vs. the Competition: OpenAI, Google, and ElevenLabs
Microsoft’s announcement explicitly names OpenAI’s GPT-Transcribe, Google’s Gemini 3.5 Transcribe, OpenAI’s Whisper V3-Large, and ElevenLabs’ Scribe v2 as its primary rivals. Notably absent from this list are established specialists like Deepgram, AssemblyAI, Speechmatics, and Rev, companies that have served enterprise transcription needs for a decade. Microsoft’s positioning is clearly aimed at the frontier AI labs, rather than the incumbent players in the specialized transcription market.
This strategic framing is both a marketing tactic and a reflection of reality. The frontier labs have largely treated speech recognition as an ancillary feature within broader AI platforms, pricing it accordingly. A dedicated model that demonstrably outperforms these platforms in speed by a factor of five to ten, while matching their accuracy, represents a genuine competitive advantage. However, the price pressure will be most acutely felt by the specialized transcription vendors. At $0.10 per hour, Microsoft is pricing its service at or below the rates many of these specialists charge for high-volume enterprise contracts, and crucially, it is bundling advanced features like diarization, timestamps, and support for 60 languages into this base rate. The remaining competitive moat for these specialists lies in domain-specific expertise—such as medical vocabularies, legal formatting, and industry-specific integrations. Microsoft’s keyword biasing feature directly challenges this advantage by enabling users to tailor the model to their specific jargon.
One notable omission from Microsoft’s claims of superior accuracy is Alibaba, whose models have consistently ranked at the top of independent leaderboards throughout 2026. TechCrunch reported in July that some U.S. companies were beginning to explore Chinese models as more cost-effective alternatives, despite lingering security concerns. Microsoft’s implicit but clear message to these potential buyers is: comparable accuracy, superior speed, a lower price point, and the assurance of a vendor whose compliance standards are already trusted.
Critical Questions for Technical Decision Makers Evaluating Transcription Vendors
Despite the detailed performance claims, the MAI-Transcribe-2 announcement leaves several practical questions for technical decision-makers to consider. Firstly, the pricing: Microsoft describes the $0.10 per hour as a launch offer without specifying an end date or establishing a standard rate. Any organization developing cost models should seek a written commitment for these details. Secondly, real-time transcription: The release heavily emphasizes batch throughput and long-form audio processing, but remains silent on real-time transcription capabilities, which are essential for applications like voice agents and live captioning. Artificial Analysis maintains a separate leaderboard for streaming performance, and Microsoft’s lack of commentary on this front is noteworthy.
Thirdly, per-language accuracy: An average WER of 5.2% across 60 languages could mask significant disparities, potentially indicating 3% accuracy for major languages and as high as 12% for lower-resource ones. Buyers with specific linguistic requirements should conduct direct testing for those languages. Fourthly, diarization quality: While word error rate is a key metric, it does not assess the accuracy of speaker attribution. A transcript can achieve near-perfect WER yet incorrectly assign sentences to the wrong speakers. The announcement provides no diarization error rate or equivalent metric.
Fifthly, data handling: Enterprise transcription often involves sensitive data, including medical records, legal privilege, and financial disclosures. The release offers no details regarding data residency, retention policies, or whether audio submitted through Foundry will be used for future model training. Microsoft’s April announcements described training data as a composite of human-curated recordings, contractor-recorded noisy audio, and "vast amounts of data from the open web," a description that warrants thorough scrutiny from regulated industries. While these omissions are not uncommon for a product launch, they represent the critical factors that differentiate a benchmark success from a seamless production deployment.
Microsoft’s Speech Model Strategy: A Blueprint for Broader AI Ambitions
Stepping back from the technical specifics of speech recognition, a discernible pattern emerges that extends well beyond transcription. Microsoft’s AI unit is now actively developing and releasing models across a wide spectrum of modalities, including images, voice, transcription, code generation, reasoning, and cybersecurity. At its Build conference in June, the company unveiled seven new MAI models in a single keynote. The underlying strategy for each is consistent: target a well-defined modality, aggressively optimize for inference cost, price below the leading frontier labs, distribute through Foundry, and strategically integrate these models into Microsoft’s existing product portfolio.
This approach is not an attempt to create a single, monolithic model that can outperform GPT or Gemini across all tasks. Instead, it’s a strategy to build a comprehensive suite of specialized models. This portfolio aims to enable Microsoft to serve the majority of its enterprise workloads internally, thereby reducing reliance on external providers. The surplus capacity can then be offered to the broader market at prices that smaller, specialized competitors cannot match. Transcription has simply been the first modality where this strategy has fully matured, but the release notes for MAI-Transcribe-2 read less like a product announcement and more like a definitive template for future AI development at Microsoft.
Suleyman has consistently spoken of "humanist superintelligence" and AI assistants that are "accountable to them, on their side." While the rhetoric may be lofty, the execution is grounded in pragmatic, spreadsheet-driven economics. Five months ago, Microsoft charged 36 cents to convert an hour of speech to text. On Thursday, it reduced that price to a dime, included six features that its rivals typically charge extra for, and claimed the top position on the industry’s standard multilingual benchmark. The company that invested $13 billion to understand the cost of frontier AI has now opted to own the manufacturing process rather than simply rent the output—and it is now selling that output for less than the cost of renting.
MAI-Transcribe-2 is immediately accessible via Microsoft Foundry and MAI Playground.

