Meta’s newest AI model, Muse Spark 1.3, unveiled yesterday, is making significant waves in the AI landscape, boasting faster performance and enhanced capabilities on third-party benchmarks compared to its predecessor. This latest iteration, as announced by Meta co-founder and CEO Mark Zuckerberg on X, represents a "biggest jump yet" in coding and agentic work, with Zuckerberg describing its performance as "frontier performance almost too cheap to meter." The claims are substantiated by substantial improvements over the previous month’s 1.2 release, particularly in handling long-running agent tasks. The version now accessible to developers is positioned as one of the most potent price-performance offerings near the apex of independent model rankings.
However, a crucial caveat accompanies these advancements. Meta’s most impressive benchmark results for Muse Spark 1.3 stem from its "max reasoning configuration." This highly capable version is currently undergoing additional safety testing and is slated for release "shortly." Independent benchmarking firm Artificial Analysis evaluated this "max" configuration in a limited partner preview but currently lists no API provider for it. The version broadly rolling out this week, accessible through Meta’s Muse Code harness and the Meta Model API, utilizes Meta’s previously available reasoning settings, including the "xhigh" configuration. This distinction prompts a critical enterprise question: not whether Muse Spark 1.3 can reach frontier territory, but how close the currently deployable model gets and at what real-world cost.
The shipping model, while very good, isn’t the benchmark leader. Meta has been transparent about the performance of both configurations, disclosing results in its underlying evaluation report. However, launch materials prominently highlight the "max" variant, and some of its highest scores are attributed to this configuration. For instance, Meta reports GDPval-AA v2 scores of 1,754 Elo for "max" versus 1,709 for "xhigh," OSWorld 2.0 scores of 66.9 versus 57.2, and JobBench scores of 64.9 versus 61.2. On certain tests, the distinction is negligible or even reversed: DeepSearchQA is tied at 89.4, while "xhigh" scores 89.2 on Terminal-Bench 2.1 compared to "max" at 88.8.
Artificial Analysis scores Muse Spark 1.3 "max" at 62 on its Intelligence Index, with the shipping "xhigh" version at 61. This latter score ties it with GPT-5.6 Sol max, Grok 4.6 high, and Claude Opus 5 high. Nevertheless, Anthropic still holds the top positions on the leaderboard, with Claude Fable 5.1 reaching 66 at "max" and 65 at "xhigh," and Claude Opus 5 achieving 63 at both "max" and "xhigh." Therefore, Muse Spark 1.3 "xhigh" is firmly within the frontier cluster, but it is not currently the model setting the absolute frontier.
This represents a significant leap from Muse Spark 1.2. Previous coverage indicated Meta fielding a credible coding challenger that generally trailed Anthropic’s top-tier models. Muse Spark 1.2 scored 82.9% on Terminal-Bench 2.1, while Opus 5 achieved 86.7%, and Opus also outperformed Spark 1.2 on other key coding comparisons presented by Meta. With the 1.3 release, Meta is no longer merely participating in this competition; it is now trading wins with OpenAI and Anthropic on several coding and agentic evaluations.
Meta emphasizes that the underlying model has also become more user-friendly. Muse Spark 1.3 is engineered to manage multiple workflows within a single, long thread, gather context using tools, identify gaps in its own planning, solicit user clarification when needed, and confirm actions before executing consequential steps. In internal comparisons conducted by Meta engineers, Muse Spark 1.3 used approximately 20% fewer tool calls and 25% fewer tokens than version 1.2 during coding tasks. For enterprises managing thousands or millions of agent loops, these behavioral enhancements could prove more impactful than a marginal increase in leaderboard ranking.
The phrase "almost too cheap to meter" does not signal a price reduction for the Muse Spark 1.3 API. Meta has maintained the exact same pricing structure for Muse Spark 1.3 as it had for Muse Spark 1.2: $1.25 per million input tokens, $4.25 per million output tokens, and $0.15 per million cached input tokens. Zuckerberg’s statement, therefore, is less about lowering token costs and more about Meta’s assessment of what developers can achieve with those tokens.
Artificial Analysis provides data supporting this argument, but also introduces a complication. The firm measures Muse Spark 1.3 "xhigh" at 235.2 output tokens per second, estimating a cost of $0.55 per Intelligence Index task. With an Intelligence Index score of 61, this translates to the lowest cost per task among all currently measured models at that intelligence level. Muse Spark 1.2, scoring 57, cost only $0.40 per Artificial Analysis task. Despite unchanged per-token pricing, the cost of completing an average task on this independent benchmark has therefore increased from generation to generation. Artificial Analysis attributes this rise primarily to increased input-token consumption on agentic evaluations.

This finding doesn’t directly contradict Meta’s claim of 25% lower token use; Meta is referring to comparisons within its own coding workflows, whereas Artificial Analysis measures a broader spectrum of reasoning and agentic tasks. However, it highlights why the notion of "cheap" becomes nuanced once models operate as agents. Actual costs are influenced by a combination of token rates, reasoning effort, the number of turns, tool calls, and retries.
Meta continues to offer its unusually inexpensive "Contributor" tier – $0.10 per million input tokens and $0.20 per million output tokens – in exchange for permission to use prompts and completions for training purposes. As noted previously with Muse Spark 1.2, this tier may be attractive for prototyping but presents a significantly different data governance calculation for enterprises working with proprietary code or sensitive internal information.
Meta’s Chief AI Officer, Alexandr Wang, was more direct in his praise of the release. Following Artificial Analysis’s posting of Muse Spark results, Wang reposted them on X, provocatively asking, "i really hate to say it, but… gemini who? 🙄🤷♀️". This remark was particularly pointed given that Google released Gemini 3.8 Flash on the same day, targeting a very similar class of workloads: long-horizon software engineering, autonomous agents, and multi-step professional reasoning. Google positions 3.8 as its best reasoning and coding Flash model to date and marks its third Flash release in just six weeks.
Independent metrics offer Wang some ammunition, though not a decisive victory. Artificial Analysis assigns Muse Spark 1.3 "xhigh" an Intelligence Index score of 61 at $0.55 per task, compared to Gemini 3.8 Flash at 59 and $0.58 for its "high" reasoning configuration. At these specific settings, Meta thus edges out Google in both intelligence and task cost. Google, however, achieves a decisive win in throughput. Artificial Analysis measures Gemini 3.8 Flash "high" at approximately 305 output tokens per second, compared to 235 for Muse Spark – a roughly 30% speed advantage. Gemini also boasts a lower raw API sticker price for the time being: Google is offering an introductory rate of $0.75 per million input tokens and $3.75 per million output tokens, contrasting with Meta’s $1.25 and $4.25. This promotional Google pricing is set to expire on December 31, after which it will increase to $1.50 per million input tokens and $7.50 per million output tokens.
The outcome provides a valuable snapshot of the intensifying competition and evolving economics at the frontier of model development. Meta currently leads this independent comparison by two Intelligence Index points and three cents per benchmark task; Google counters with substantially higher output throughput and cheaper raw tokens during its launch promotion. Wang’s "gemini who?" jab, while entertaining executive banter, simplifies a more complex reality. For an enterprise architect, the answer is more nuanced: Gemini represents the faster option, while Muse Spark, by this independent measure, is currently the slightly stronger high-effort agent.
A more significant concern for some developers may lie beyond the immediate benchmark race and revolve around Meta’s evolving stance on open weights. When Meta launched Muse Code and Muse Spark 1.2 in August, it marked a notable departure from the open-weight strategy that had propelled Llama to widespread adoption. Muse Code and Spark 1.2 were proprietary, API-served products, a stark contrast to Meta’s long-held advocacy for open AI. However, just five days later, Meta shifted course again. On August 10, the company released the 30-billion-parameter Muse Glimmer under an Apache 2.0 license. Zuckerberg also announced plans to "open the weights for Muse Spark 1.2" in the following weeks, a plan corroborated by separate reports from Reuters.
Now, Meta has launched Muse Spark 1.3 as another proprietary model. While "coming weeks" can encompass a period longer than three weeks, today’s announcement introduces ambiguity into the roadmap. Meta’s latest post no longer specifically mentions Muse Spark 1.2 in relation to open weights. Instead, it refers to its roadmap including "the Muse Spark open weights release," without specifying a version, release date, model size, or license. Zuckerberg echoed this sentiment on X, mentioning "Muse Spark open weights releases" are slated for the near future. For development teams that standardized on Llama due to the availability of downloadable weights for self-hosting, customization, and control over inference economics, this ambiguity may carry more weight than marginal gains on a composite benchmark.
Muse Spark 1.3 demonstrates Meta’s remarkable ability to rapidly iterate on proprietary frontier models. The currently available "xhigh" configuration is fast, competitively priced, and significantly closer to the top of independent rankings than its predecessors. The "max" preview further indicates Meta’s capacity to push the family’s capabilities further when allowed to allocate more reasoning compute. The next critical test for Meta will be its ability to translate this development pace into a clear and actionable roadmap for enterprises, including making its most advanced capabilities broadly deployable and delivering on its promise of open-weight Spark models.

