Microsoft AI on Thursday unveiled MAI-Transcribe-2, a revolutionary speech-recognition model that the technology giant asserts significantly surpasses the capabilities of its competitors, including OpenAI, Google, and ElevenLabs, in terms of speed, accuracy, and cost-effectiveness. The model’s pricing, set at an unprecedented 10 cents per hour of audio, represents a dramatic reduction from previous offerings and signals a strategic shift in Microsoft’s approach to AI development and deployment. This aggressive pricing strategy, coupled with a feature-rich model, positions MAI-Transcribe-2 to disrupt the enterprise transcription market and underscores Microsoft’s ambition to become a dominant force in the AI landscape.
The startling 10-cent per hour price point warrants significant attention. When Microsoft AI introduced the first iteration of this model line just five months prior, the cost stood at $0.36 per hour. The latest release effectively slashes this cost by approximately 72%, a move that could dramatically alter the economics for businesses processing large volumes of audio data. For an enterprise handling 100,000 hours of call-center audio annually – a typical figure for major financial institutions or telecommunications companies – this price reduction translates to savings from $36,000 down to $10,000. At such a competitive price, the cost of transcription is poised to become a negligible concern for many organizations, allowing them to focus on leveraging the insights derived from their audio data.
This release is a clear manifestation of Microsoft’s evolving strategy, one that would have seemed improbable just two years ago. The company is now systematically developing its own frontier-class AI models across various modalities, with the ultimate goal of integrating them into its existing product suite, thereby reducing reliance on external partners like OpenAI, for which it has invested billions. Transcription has emerged as the modality where this strategy is progressing most rapidly, and MAI-Transcribe-2 stands as the most compelling evidence of its success. This development also offers a glimpse into how the world’s most valuable software company intends to navigate the competitive AI arena without being tethered to the partner it has cultivated through a substantial $13 billion investment.
MAI-Transcribe-2’s Capabilities and Enterprise Value Proposition
MAI-Transcribe-2 boasts an impressive array of features designed to meet the demanding needs of enterprise users. The model now supports transcription in 60 languages, a significant expansion from the 43 languages supported by MAI-Transcribe-1.5 in June and the 25 languages of the original April release. It is accessible through Microsoft Foundry, the company’s model marketplace for developers, and within MAI Playground, its dedicated testing environment. Microsoft emphasizes that MAI-Transcribe-2 has been engineered to handle the complexities of real-world business audio, including background noise, low-quality recordings, and overlapping speech, scenarios far removed from controlled studio conditions.
Beyond the expanded language support, the integrated features are particularly valuable for enterprise buyers. Speaker diarization, a crucial capability, distinguishes between different speakers in multi-person recordings, transforming a dense block of text into a coherent and actionable transcript. Word-level timestamps provide precise temporal markers for each word, facilitating efficient search, editing, and synchronization with video content. Keyword biasing allows developers to supply domain-specific terms, such as drug names, product codes, or employee names, thereby improving the model’s accuracy in transcribing industry jargon and specialized vocabulary. Furthermore, automatic language identification eliminates the need for users to manually specify the language of the audio, streamlining the transcription process.
Two features, in particular, highlight the model’s tailored design. A configurable output style offers a "verbatim" mode, which captures every utterance, including filler words like "um" and false starts, crucial for compliance and legal applications. Conversely, a "clean" mode removes these elements to produce more readable captions and notes. Code-switching is another standout feature, adeptly handling conversations that fluidly transition between languages within a single sentence. Microsoft specifically references Hinglish and Spanglish, acknowledging the linguistic diversity in markets like India and the U.S. Hispanic community, where customer service calls might involve numerous language shifts. Historically, specialized vendors have commanded premium prices for each of these advanced capabilities. Microsoft’s inclusion of all these features within the base MAI-Transcribe-2 offering at just ten cents per hour is a disruptive development.
Decoding Microsoft’s Benchmark Claims: FLEURS and Artificial Analysis
Microsoft’s performance claims for MAI-Transcribe-2 are underpinned by three distinct benchmarks, each measuring different aspects of the model’s capabilities. Technical buyers are encouraged to understand the nuances of each metric to fully appreciate the model’s strengths and limitations.
The first claim is that MAI-Transcribe-2 achieves the top ranking on the FLEURS benchmark across 60 languages, with an average word error rate (WER) of 5.2%. FLEURS, a benchmark introduced by Google researchers in 2022, comprises approximately 12 hours of speech per language, with native speakers reading around 2,000 sentences in each of 102 languages. It serves as a standard for multilingual speech recognition by enabling direct comparison of a model’s performance on diverse languages using identical content. WER, the primary metric, quantifies errors by counting substitutions, insertions, and deletions against a human reference. A 5.2% WER indicates that, on average, one in every twenty words is transcribed incorrectly. However, it is important to note that FLEURS relies on read speech rather than natural conversation. The reported average WER of 5.2% for MAI-Transcribe-2 represents an increase from the 3.7% reported for MAI-Transcribe-1.5 in June. This increase is likely attributable to the expanded language coverage, which now includes low-resource languages where all models typically struggle, rather than a regression in performance. Nevertheless, buyers are advised to request a per-language breakdown for a more granular understanding of performance.
The second claim asserts that MAI-Transcribe-2 ranks second on the Artificial Analysis word-error-rate leaderboard, defining its position on the firm’s accuracy-latency Pareto frontier. Artificial Analysis is an independent benchmarking organization that evaluates models via their public APIs, simulating real-world customer experiences. Its comprehensive index incorporates simulated agent conversations, European Parliament speeches, and corporate earnings calls, with a significant emphasis on English business speech. In June, MAI-Transcribe-1.5 held the third position on this leaderboard with a 2.4% WER, trailing only Alibaba’s Fun-ASR and ElevenLabs’ Scribe v2, while being recognized as the fastest model within the top 10. The climb to second place suggests that Microsoft has successfully surpassed ElevenLabs. The term "Pareto frontier" is a critical indicator for practitioners, signifying that no competitor can achieve higher accuracy without sacrificing speed, nor can they achieve greater speed without compromising accuracy.
The third claim focuses on raw inference speed. Artificial Analysis evaluations indicate that MAI-Transcribe-2 is approximately 10 times faster than OpenAI’s GPT-Transcribe, seven times faster than ElevenLabs’ Scribe v2, and five times faster than Google’s Gemini 3.5 Transcribe. In batch transcription, speed is not merely about reducing user wait times but is fundamentally linked to throughput and cost efficiency. A model operating at 300 times real-time requires significantly fewer GPU hours than one running at 30 times real-time. This enhanced efficiency is the bedrock upon which Microsoft can offer its 10-cent per hour pricing while maintaining profitability.
The Unprecedented Cadence of Microsoft AI’s Model Releases
The rapid release cycle of Microsoft’s AI models is as remarkable as the performance of MAI-Transcribe-2 itself. In a span of just five months, the company has launched three distinct iterations of its transcription model. On April 2nd, MAI-Transcribe-1 debuted with 25 languages and a price of $0.36 per hour. This was followed by MAI-Transcribe-1.5 on June 2nd, which expanded language support to 43, introduced keyword biasing, and secured a third-place ranking on Artificial Analysis. The latest release, MAI-Transcribe-2, arrived on Thursday, offering 60 languages, speaker diarization, timestamps, code switching, a second-place ranking, and the groundbreaking $0.10 per hour price point.
This aggressive release cadence, with each iteration enhancing language coverage by approximately 40% while incorporating features that competitors often reserve for premium tiers, suggests a team that has achieved architectural stability and is now focused on scaling data and performance. This is the phase where speech models typically exhibit rapid and predictable improvements. It also reflects a deliberate strategy by Microsoft to commoditize transcription services before its competitors can establish a dominant market position.
The organizational philosophy driving this speed was articulated by Mustafa Suleyman, CEO of Microsoft AI, in an interview with The Verge in April. He attributed the success of the initial model to a "small, focused 10-person team" that operated with minimal bureaucracy, supported by a larger group managing vendor relations and data acquisition. Suleyman also highlighted the cost efficiencies, noting that the model operated at "half the GPU cost of the other state-of-the-art models," a significant advantage for Microsoft. This lean, agile approach mirrors experiments by other major tech players like Meta, Amazon, Google, and Anthropic, and Microsoft’s transcription line serves as a high-profile test case for its commercial viability.
Microsoft’s Strategic Pivot: Building In-House AI Despite OpenAI Investment
Microsoft’s substantial investment of over $13 billion in OpenAI, and its integration of OpenAI models across Azure, Office, and Copilot, has long raised the question: why develop in-house models? The answer has become increasingly clear over the past year, centering on the strategic imperative of independence.
The hiring of Mustafa Suleyman from Inflection AI in March 2024, along with a significant portion of his team, was widely interpreted as a declaration of intent. Marc Benioff, CEO of Salesforce, a competitor and investor in Anthropic, famously stated in January 2025 that "Microsoft is building their own AI and I don’t think Microsoft will use OpenAI in the future. They’ll have their own frontier models. That’s why they hired Mustafa Suleyman." While Benioff had his own strategic motivations, subsequent events have largely validated his assessment.
In October 2025, Microsoft and OpenAI restructured their partnership. According to Microsoft’s announcement, this renegotiation granted Microsoft the ability to "independently pursue AGI alone or in partnership with third parties" for the first time. Suleyman confirmed to The Verge that this restructuring "unlocked [Microsoft’s] ability to pursue superintelligence," leading to the subsequent establishment of Microsoft’s MAI Superintelligence team. Further amendments to the deal in April 2026 reportedly ended Microsoft’s exclusive access to OpenAI’s models and eliminated its revenue-share payments, according to contemporary reports. Each contractual adjustment has incrementally loosened the ties between the two companies, and each has been followed by the introduction of new MAI models.
Cost Efficiencies and Integration Across Microsoft Products
The pursuit of margin is a critical driver behind Microsoft’s in-house AI development. Every query routed to an OpenAI model incurs a cost for Microsoft, whereas queries processed by its own models on its own GPUs carry a significantly lower expense. In July, Bloomberg reported that Microsoft had begun incorporating MAI models to handle a portion of user prompts in Word and Excel, applications previously advertised as powered by OpenAI and Anthropic. This shift was framed by TechCrunch as part of a broader industry trend of AI cost-cutting, with companies like Amazon, Uber, Meta, and Accenture reportedly implementing similar reductions.
Transcription is a logical starting point for this cost-optimization strategy due to the well-defined nature of the problem and the objective measurability of its outcomes. Microsoft possesses significant assets in this domain, including Teams, which generates a vast amount of meeting audio; Nuance, whose clinical documentation services rely heavily on speech recognition; and Azure’s speech services, already utilized by thousands of enterprises. Each of these workloads represents a prime candidate for migration to MAI-Transcribe-2. Every hour of audio transcribed by MAI-Transcribe-2 is an hour for which Microsoft no longer incurs external costs.
Suleyman has been notably candid about this strategic objective. He articulated to The Verge in April that the pursuit of superintelligence is fundamentally about "Are these models capable of delivering product value for the millions of enterprises that depend on us to deliver world-class language models?" Regardless of how one interprets the application of "superintelligence" to a transcription API, the commercial logic is irrefutable: develop the capability internally, deploy it across a wide array of products, and cease incurring expenses to a partner that is increasingly becoming a direct competitor.
MAI-Transcribe-2 vs. the Competition: A New Market Dynamic
Microsoft’s announcement explicitly names four key rivals: OpenAI’s GPT-Transcribe, Google’s Gemini 3.5 Transcribe, OpenAI’s older Whisper V3-Large, and ElevenLabs’ Scribe v2. Notably absent from this list are established specialized transcription providers such as Deepgram, AssemblyAI, Speechmatics, and Rev, companies that have served enterprises for a decade. This strategic framing positions Microsoft not against the incumbents but against the frontier AI labs.
This positioning is partly marketing and partly a reflection of reality. The frontier labs have often treated speech recognition as an ancillary feature within broader AI platforms, with pricing reflecting this. A dedicated model that demonstrably outperforms these offerings in speed by five to ten times, while matching their accuracy, represents a significant competitive advantage. However, the price pressure will be felt most acutely by the specialized vendors. At $0.10 per hour, Microsoft is pricing its service at or below the enterprise contract rates of many of these specialists, and crucially, it includes advanced features like diarization, timestamps, and 60-language support as standard. The remaining competitive moat for these specialists lies in deep domain expertise, such as specialized medical or legal vocabularies and industry-specific integrations. Microsoft’s keyword biasing feature directly addresses this, aiming to mitigate the need for such specialized solutions.
Microsoft conspicuously omits a claim of superior accuracy over Alibaba, whose models have consistently held leading positions on independent leaderboards throughout 2026. Reports from July indicated that some U.S. companies were exploring Chinese models as more cost-effective alternatives, despite potential security concerns. Microsoft’s implicit, yet unmistakable, proposition to these buyers is: comparable accuracy, superior inference speed, a lower price point, and a trusted vendor that already meets stringent compliance requirements.
Critical Questions for Enterprise Decision-Makers Before Vendor Migration
Despite the wealth of detail regarding benchmarks and performance metrics, Microsoft’s announcement leaves several practical questions unanswered for technical decision-makers. The first pertains to pricing duration: Microsoft currently refers to the $0.10 per hour rate as a launch offer without specifying an end date or a standard post-launch rate. Organizations developing cost models should seek written confirmation of these details. The second crucial area is streaming transcription. The release heavily emphasizes batch throughput and long-form audio processing, but offers no information regarding real-time transcription capabilities, which are essential for applications like voice agents and live captioning. Artificial Analysis maintains a separate leaderboard for streaming performance, and Microsoft’s silence on this front is noteworthy.
The third concern relates to per-language accuracy. An average WER of 5.2% across 60 languages could mask significant disparities, with higher error rates in low-resource languages compared to major ones. Buyers with specific language requirements should conduct their own direct testing. Fourth, the quality of diarization is not clearly defined. While word error rate measures transcription accuracy, it does not assess speaker attribution. A transcript can achieve near-perfect WER yet misattribute sentences to the wrong speakers. The release provides no diarization error rate or equivalent metric.
The fifth critical question concerns data handling. Enterprise transcription often involves sensitive data, including medical records, legal privilege information, and financial disclosures. The announcement offers no specifics on data residency, retention policies, or whether audio submitted through Foundry contributes to future model training. Microsoft’s April announcements described training data as a blend of human-curated recordings, contractor-recorded noisy audio, and "vast amounts of data from the open web," a description that warrants careful scrutiny from regulated industries. While these omissions are not unusual for an initial product announcement, they represent key considerations that differentiate a benchmark victory from a successful production deployment.
Microsoft’s Speech Model Strategy: A Blueprint for Broader AI Ambitions
Stepping back from the specifics of speech recognition, a discernible pattern emerges that extends far beyond transcription services. Microsoft’s AI unit is now actively developing and releasing models for a diverse range of modalities, including images, voice, transcription, code generation, reasoning, and cybersecurity. At the Build conference in June, the company announced seven new MAI models in a single keynote address. Each of these models follows a consistent playbook: target a well-defined modality, aggressively optimize for inference cost, price below that of leading AI labs, distribute through the Foundry marketplace, and systematically integrate the model into Microsoft’s existing product ecosystem.
This strategy is not an attempt to create a singular model that can outperform established leaders like GPT or Gemini across all tasks. Instead, it represents an effort to build a comprehensive portfolio of specialized models. This approach aims to enable Microsoft to serve the majority of its enterprise workloads internally, thereby eliminating the need for third-party services, and to offer surplus capacity to external clients at prices that specialized providers cannot match. Transcription has serendipitously become the first modality where this strategy has fully matured, but the release notes for MAI-Transcribe-2 read less like a singular product announcement and more like a comprehensive template for future developments.
Suleyman has consistently spoken of "humanist superintelligence" and AI assistants that are "accountable to them, on their side." While the vocabulary may be aspirational, the execution is demonstrably rooted in practical economics. Five months ago, Microsoft charged 36 cents per hour for speech-to-text conversion. On Thursday, it reduced that price to a dime, bundled six features that its rivals typically charge extra for, and claimed the top position on the industry’s standard multilingual benchmark. The company that invested $13 billion to understand the costs associated with frontier AI has now opted to own the manufacturing process rather than simply rent the output, and is now selling that output at a price lower than the rental cost.
MAI-Transcribe-2 is now accessible through Microsoft Foundry and MAI Playground.

