Meta unveiled its latest AI model, Muse Spark 1.3, yesterday, a release that boasts significant improvements in speed and performance on third-party benchmarks compared to its predecessor. However, the full picture of its capabilities and availability requires a closer examination. Mark Zuckerberg, Meta’s co-founder and CEO, heralded the launch on X, describing Muse Spark 1.3 as a model with “frontier performance almost too cheap to meter” and representing Meta’s “biggest jump yet” in coding and agentic work. This assertion is backed by substantial progress over the previous month’s 1.2 release, particularly in handling long-running agent tasks. The version currently accessible to developers is positioned as one of the most cost-effective options near the apex of independent model rankings.
The most impressive benchmark results for Muse Spark 1.3 stem from its "max reasoning" configuration. Meta has indicated that this version is undergoing final safety testing and is expected to be available "shortly." In the interim, the third-party benchmarking firm Artificial Analysis, which evaluated the "max" configuration during a limited partner preview, currently lists no API provider for this specific setup. The version broadly rolling out this week, accessible through the Muse Code harness and the Meta Model API, utilizes Meta’s previously established reasoning settings, including "xhigh." This distinction is critical for enterprises evaluating the practical deployment of Muse Spark 1.3, shifting the focus from its theoretical frontier capabilities to the actual performance and cost of the deployable version.
The commercially available Muse Spark 1.3 model demonstrates considerable strength, though it doesn’t consistently lead benchmarks in its currently deployed state. Meta openly shares results for both configurations within its evaluation reports, ensuring transparency about the deployable model’s performance. Nevertheless, the company’s launch materials prominently feature the "max" variant, and its associated scores are often the most striking. For instance, Meta reports GDPval-AA v2 scores of 1,754 Elo for "max" versus 1,709 for "xhigh," OSWorld 2.0 scores of 66.9 versus 57.2, and JobBench scores of 64.9 versus 61.2. In some specific tests, the performance difference is negligible or even favors the "xhigh" configuration. DeepSearchQA, for example, shows a tie at 89.4, and "xhigh" scores 89.2 on Terminal-Bench 2.1, narrowly surpassing "max" at 88.8.
Artificial Analysis further contextualizes these results, scoring Muse Spark 1.3 "max" at 62 on its Intelligence Index, with the widely available "xhigh" version achieving a score of 61. This score of 61 places the "xhigh" version in direct competition with other leading models like GPT-5.6 Sol max, Grok 4.6 high, and Claude Opus 5 high. However, Anthropic’s models continue to occupy the top positions on the leaderboard, with Claude Fable 5.1 reaching 66 at "max" and 65 at "xhigh," and Claude Opus 5 achieving 63 at both "max" and "xhigh." Therefore, while Muse Spark 1.3 "xhigh" is firmly within the frontier cluster, it is not currently the undisputed leader.
This represents a significant advancement from Muse Spark 1.2. VentureBeat’s previous coverage highlighted Meta’s emergence as a credible coding challenger with the 1.2 release, but it generally lagged behind Anthropic’s top-tier models. Muse Spark 1.2 scored 82.9% on Terminal-Bench 2.1, compared to Opus 5’s 86.7%, and also finished behind Opus on other key coding comparisons presented by Meta. With the 1.3 iteration, Meta is now actively competing and, in several coding and agentic evaluations, trading wins with industry giants like OpenAI and Anthropic.
Beyond raw performance metrics, Meta emphasizes that Muse Spark 1.3 has been engineered for enhanced usability. The model is now trained to adeptly manage multiple workflows within a single long thread, efficiently gather context using tools, identify and address gaps in its own planning processes, seek user clarification when needed, and obtain confirmation before executing critical actions. Internal comparisons by Meta engineers reveal that Muse Spark 1.3 utilized approximately 20% fewer tool calls and 25% fewer tokens compared to version 1.2 during coding tasks. For enterprises managing thousands or millions of agent interactions, these behavioral improvements in efficiency and workflow management could prove more impactful than marginal gains on benchmark leaderboards.
Zuckerberg’s "almost too cheap to meter" commentary does not translate to a reduction in API pricing for Muse Spark 1.3. Meta has maintained the same pricing structure as for Muse Spark 1.2: $1.25 per million input tokens, $4.25 per million output tokens, and $0.15 per million cached input tokens. This pricing strategy suggests that Meta’s value proposition lies not in lower per-token costs, but in the increased efficiency and capabilities developers can achieve with those tokens.
Artificial Analysis provides data that supports this perspective, while also introducing a layer of complexity. The firm measures Muse Spark 1.3 "xhigh" at 235.2 output tokens per second, estimating a cost of $0.55 per Intelligence Index task. With an Intelligence Index score of 61, this positions Muse Spark 1.3 "xhigh" as having the lowest cost per task among all models currently measured at that intelligence level. For comparison, Muse Spark 1.2 cost only $0.40 per Artificial Analysis task, while scoring 57 on the Intelligence Index. Despite the unchanged per-token pricing, the cost of completing an average task, as measured by Artificial Analysis, has increased from generation to generation. This increase is primarily attributed by Artificial Analysis to a higher consumption of input tokens on agentic evaluations.

This observation does not directly contradict Meta’s claim of 25% lower token usage. Meta’s comparison is focused on its internal coding workflows, whereas Artificial Analysis measures a broader spectrum of reasoning and agentic tasks. However, it highlights the nuanced definition of "cheap" when dealing with advanced AI agents. The actual cost of completing tasks is a confluence of token rates, reasoning effort, conversational turns, tool utilization, and potential retries.
Meta also continues to offer its highly competitive "Contributor" tier. This tier charges a nominal $0.10 per million input tokens and $0.20 per million output tokens, in exchange for permission to use prompts and completions for training purposes. As previously noted with Muse Spark 1.2, this pricing model may be attractive for prototyping and development but presents a significant data governance consideration for enterprises working with proprietary code or sensitive internal information.
Meta’s Chief AI Officer, Alexandr Wang, offered a more assertive assessment of the Muse Spark 1.3 release. Following the publication of Artificial Analysis results, Wang reposted them on X, provocatively asking, "i really hate to say it, but… gemini who? 🙄😒". This pointed remark coincided with Google’s release of Gemini 3.8 Flash on the same day. Gemini 3.8 Flash is positioned for similar workloads, including long-horizon software engineering, autonomous agents, and multi-step professional reasoning. Google touts 3.8 as its most advanced reasoning and coding Flash model to date and marks its third Flash release within a six-week period.
Independent benchmarks offer some justification for Wang’s bold claim, though not a definitive knockout. Artificial Analysis scores Muse Spark 1.3 "xhigh" at 61 on the Intelligence Index with a task cost of $0.55, while Gemini 3.8 Flash at high reasoning scores 59 with a task cost of $0.58. Under these specific settings, Meta demonstrates an advantage over Google in both intelligence and task cost. However, Google decisively leads in throughput. Gemini 3.8 Flash high achieves approximately 305 output tokens per second, significantly outperforming Muse Spark’s 235 tokens per second—a difference of roughly 30%. Furthermore, Gemini currently boasts a lower raw API sticker price: Google is offering an introductory rate of $0.75 per million input tokens and $3.75 per million output tokens, compared to Meta’s $1.25 and $4.25. It’s important to note that this promotional Google pricing is set to expire on December 31, after which it will increase to $1.50 per million input tokens and $7.50 per million output tokens.
This competitive landscape provides a clear snapshot of the evolving economics of frontier AI models. Meta currently holds a slight edge in this independent comparison, leading by two Intelligence Index points and three cents per benchmark task. Google, in contrast, offers substantially higher output throughput and more affordable token prices during its introductory promotion. Wang’s "gemini who?" remark, while spirited executive banter, underscores the intense competition. For enterprise architects, the nuanced answer is that Gemini currently presents a faster option, while Muse Spark, by this independent measure, offers a slightly stronger high-effort agent.
A potentially more significant consideration for many developers lies in Meta’s evolving stance on open-weight models. When Meta initially launched Muse Code and Muse Spark 1.2 in August, it marked a notable departure from the open-weight strategy that had propelled Llama to widespread adoption. Muse Code and Spark 1.2 were proprietary, API-served products, a position that contrasted sharply with Meta’s long-standing advocacy for open AI.
However, Meta reversed course just five days later. On August 10, the company released the 30-billion-parameter Muse Glimmer under an Apache 2.0 license. Zuckerberg also stated Meta’s intention to "open the weights for Muse Spark 1.2" in the coming weeks, a plan corroborated by Reuters. This trajectory now appears to have shifted again with the release of Muse Spark 1.3 as another proprietary model.
While this may not constitute a direct broken promise—"coming weeks" can be interpreted broadly—the current announcement introduces ambiguity into Meta’s roadmap. The latest post no longer specifies Muse Spark 1.2 for open weights; instead, it refers to "the Muse Spark open weights release" without providing details on version, release date, model size, or license. Zuckerberg’s X post also mentions "Muse Spark open weights releases" are forthcoming. This lack of clarity could be a critical factor for teams that have standardized on Llama for its downloadable weights, enabling self-hosting, customization, and control over inference economics. For these developers, the ambiguity surrounding the open-weight Spark model may carry more weight than incremental benchmark gains.
Muse Spark 1.3 undeniably demonstrates Meta’s accelerated iteration pace for proprietary frontier models. The deployed "xhigh" configuration is fast, competitively priced, and has significantly closed the gap with top-tier models in independent rankings. The "max" preview variant further indicates Meta’s capacity to push the family’s capabilities even further when allocated more reasoning compute. The next crucial test for Meta will be its ability to translate this development velocity into a predictable and actionable roadmap for enterprises, including making its most advanced capabilities broadly accessible and delivering on its previous commitment to release open-weight Spark models.

