21 Sep 2026, Mon

Meta’s New AI Model Muse Spark 1.3: Faster, Cheaper, and Sparking Debate

Meta’s latest artificial intelligence model, Muse Spark 1.3, unveiled yesterday, is making significant strides in speed and performance, outperforming its predecessor on third-party benchmarks. This advancement is accompanied by a crucial caveat that redefines the interpretation of its performance and cost-effectiveness. Meta co-founder and CEO Mark Zuckerberg proclaimed on X that "Muse Spark 1.3 is rolling out today with frontier performance almost too cheap to meter," characterizing it as Meta’s "biggest jump" yet in coding and agentic capabilities.

There is substantial evidence supporting both aspects of Zuckerberg’s claim. Muse Spark 1.3 demonstrates considerable improvements over the previous month’s 1.2 release, particularly in handling long-running agent tasks. The current version available to developers positions itself as one of the most potent price-performance offerings near the top of independent model rankings. However, Meta’s most impressive benchmark results for Muse Spark 1.3 stem from its "max reasoning" configuration. This particular version is still undergoing additional safety testing and is slated for release "shortly." Meanwhile, the third-party benchmarking firm Artificial Analysis evaluated the "max" configuration in a limited partner preview and currently lists no API provider for it. The version broadly rolling out this week, accessible through its Muse Code harness and the Meta Model API, utilizes Meta’s previously available reasoning settings, including "xhigh." This distinction shifts the critical enterprise question from whether Muse Spark 1.3 can achieve frontier status to how close the currently deployable model gets and at what real-world cost.

The commercially available version of Muse Spark 1.3 is undeniably powerful, but it doesn’t claim the absolute benchmark leadership. Meta openly discloses results for both configurations within its underlying evaluation report, so this isn’t a case of withholding data on the deployable model. However, its launch materials prominently highlight the "max" variant, and some of its highest scores are attributed to this configuration. For instance, Meta reports GDPval-AA v2 scores of 1,754 Elo for "max" compared to 1,709 for "xhigh," OSWorld 2.0 scores of 66.9 versus 57.2, and JobBench scores of 64.9 versus 61.2. On some tests, the differences are negligible or even reversed: DeepSearchQA is tied at 89.4, while "xhigh" scores 89.2 on Terminal-Bench 2.1, slightly ahead of "max" at 88.8.

Artificial Analysis scores Muse Spark 1.3 "max" at 62 on its Intelligence Index, with the shipping "xhigh" version achieving a score of 61. This latter score ties it with models like GPT-5.6 Sol max, Grok 4.6 high, and Claude Opus 5 high. However, Anthropic still holds the top positions on the leaderboard, with Claude Fable 5.1 reaching 66 at "max" and 65 at "xhigh," and Claude Opus 5 achieving 63 at both "max" and "xhigh." This analysis indicates that Muse Spark 1.3 "xhigh" is firmly within the frontier cluster of AI models, but it is not the current vanguard.

Nevertheless, this represents a significant leap from Muse Spark 1.2. VentureBeat’s previous coverage of the 1.2 launch noted that while Meta had introduced a credible coding challenger, it generally lagged behind Anthropic’s top-tier models. Muse Spark 1.2 scored 82.9% on Terminal-Bench 2.1, falling short of Opus 5’s 86.7%, and also finished behind Opus on other key coding comparisons presented by Meta. With the 1.3 release, Meta is no longer just participating in this competition; it is now trading wins with OpenAI and Anthropic on several coding and agentic evaluations.

Meta further claims that the underlying model has become more user-friendly. Muse Spark 1.3 is engineered to manage multiple workflows within a single long thread, effectively gather context using tools, identify its own planning deficiencies, solicit clarification from users when necessary, and obtain confirmation before executing consequential actions. In internal comparisons conducted by Meta engineers, Muse Spark 1.3 utilized approximately 20% fewer tool calls and 25% fewer tokens than version 1.2 during coding tasks. For enterprises managing thousands or millions of agent loops, these behavioral enhancements could prove more impactful than a marginal increase in leaderboard ranking.

Zuckerberg’s assertion of performance being "almost too cheap to meter" does not translate to a reduction in API prices for Muse Spark 1.3. Meta has maintained the same pricing structure for Muse Spark 1.3 as it had for 1.2: $1.25 per million input tokens, $4.25 per million output tokens, and $0.15 per million cached input tokens. This pricing stability means that Zuckerberg’s remark is less about a decrease in token costs and more about Meta’s belief in the enhanced capabilities and efficiency that developers can achieve with these tokens.

Artificial Analysis provides data supporting this perspective, albeit with a complicating factor. The firm measures Muse Spark 1.3 "xhigh" at 235.2 output tokens per second and estimates a cost of $0.55 per Intelligence Index task. With an Intelligence Index score of 61, this positions Muse Spark 1.3 "xhigh" as having the lowest cost per task among all currently measured models at that intelligence level. In contrast, Muse Spark 1.2, which scored 57 on the Intelligence Index, cost only $0.40 per Artificial Analysis task. This suggests that despite unchanged per-token pricing, the cost of completing an average task, as measured by Artificial Analysis, has increased from generation to generation. This increase is primarily attributed by Artificial Analysis to heavier input-token consumption on agentic evaluations.

Meta says Muse Spark 1.3 has frontier performance — but its best results come from a model developers can’t broadly use yet

This observation does not directly contradict Meta’s claim of 25% lower token use, as Meta is referring to its internal coding workflows, while Artificial Analysis measures a broader spectrum of reasoning and agentic tasks. However, it highlights why the concept of "cheap" becomes nuanced once models operate as agents. The actual cost of completing work is a complex interplay of token rates, reasoning effort, the number of interaction turns, tool calls, and any necessary retries.

Meta also continues to offer its unusually affordable "Contributor" tier, priced at $0.10 per million input tokens and $0.20 per million output tokens. This pricing is in exchange for permission to use prompts and completions for training purposes. As previously noted with Muse Spark 1.2, this tier may be attractive for prototyping but introduces significant data governance considerations for enterprises handling proprietary code or sensitive internal information.

Meta’s chief AI officer, Alexandr Wang, made a more direct and less qualified endorsement of the new release. Following Artificial Analysis’s posting of the Muse Spark results, Wang reposted them on X, provocatively asking, "i really hate to say it, but… gemini who? 🙄🤷‍♀️." This jab was particularly pointed as Google released its Gemini 3.8 Flash model on the same day, targeting a similar class of workloads: long-horizon software engineering, autonomous agents, and multi-step professional reasoning. Google has described 3.8 as its best reasoning and coding Flash model to date and marks its third Flash release in just six weeks.

Independent benchmarks offer Wang some ammunition, though not a decisive victory. Artificial Analysis scores Muse Spark 1.3 "xhigh" at 61 on the Intelligence Index with a cost of $0.55 per task, while Gemini 3.8 Flash at high reasoning scores 59 with a cost of $0.58. This indicates that Meta’s model slightly edges out Google’s on both intelligence and task cost at these specific settings. However, Google holds a clear advantage in throughput. Artificial Analysis measures Gemini 3.8 Flash high at approximately 305 output tokens per second, compared to 235 for Muse Spark, representing roughly a 30% speed advantage. Gemini also boasts a lower initial API sticker price during its promotional period: Google is charging $0.75 per million input tokens and $3.75 per million output tokens, contrasting with Meta’s $1.25 and $4.25. This promotional pricing from Google is set to expire on December 31, after which it will increase to $1.50 per million input tokens and $7.50 per million output tokens.

The current situation presents a snapshot of the intensifying competition and rapidly evolving economics in the frontier model space. Meta currently leads this specific independent comparison by two Intelligence Index points and three cents per benchmark task. Google, in contrast, offers substantially higher output throughput and more affordable raw tokens during its introductory promotion. Wang’s "gemini who?" remark, while entertaining executive trash talk, simplifies a more complex reality for enterprise architects. Gemini currently offers a faster option, while Muse, by this independent measure, currently provides a slightly more capable high-effort agent.

A more significant consideration for some developers may lie beyond the immediate benchmark race and revolve around Meta’s evolving stance on open-weight models. When Meta launched Muse Code and Muse Spark 1.2 in August, it marked a notable departure from the open-weight strategy that had propelled Llama to widespread adoption. Muse Code and Spark 1.2 were proprietary, API-served products, a significant shift for a company that had long advocated for open AI as the path forward. However, Meta pivoted again just five days later. On August 10, it released the 30-billion-parameter Muse Glimmer under an Apache 2.0 license. Zuckerberg also announced plans to "open the weights for Muse Spark 1.2" in the coming weeks, a move independently reported by Reuters.

Now, Meta has released Muse Spark 1.3 as another proprietary model. While "coming weeks" can encompass a broad timeframe, this latest announcement introduces ambiguity into the roadmap rather than clarifying it. Meta’s new post no longer specifically mentions Muse Spark 1.2’s open weights. Instead, its roadmap now includes "the Muse Spark open weights release," without specifying a version, release date, model size, or license. Zuckerberg echoed this on X, referring to "Muse Spark open weights releases" as forthcoming. For development teams that standardized on Llama due to the availability of downloadable weights, enabling self-hosting, customization, and control over inference economics, this ambiguity could be more impactful than a fractional gain on a composite benchmark.

Muse Spark 1.3 undeniably demonstrates Meta’s capacity to iterate on proprietary frontier models with remarkable speed. The commercially available "xhigh" configuration is fast, competitively priced, and significantly closer to the top of independent rankings than its predecessors. The "max" preview further indicates Meta’s ability to push the model’s capabilities when greater reasoning compute is allocated. The next critical challenge for Meta will be to translate this rapid iteration into a clear and actionable roadmap that enterprises can rely on for strategic planning. This includes making its most advanced capabilities broadly deployable and, crucially, delivering the open-weight Spark model that has already been alluded to.

By admin

Leave a Reply

Your email address will not be published. Required fields are marked *