1 Oct 2026, Thu

Meta’s New AI Model Muse Spark 1.3: Faster, Cheaper, and Ready to Challenge the Frontier, But Open Weights Remain a Question Mark

Meta has unveiled its latest artificial intelligence model, Muse Spark 1.3, a significant leap forward in performance and efficiency that positions it as a formidable contender in the rapidly evolving AI landscape. Launched yesterday, the new model demonstrably outperforms its predecessor across third-party benchmarks, boasting enhanced speed and improved capabilities, particularly in complex agentic tasks. This advancement has been lauded by Meta’s leadership, with CEO Mark Zuckerberg proclaiming it Meta’s "biggest jump yet" in coding and agentic work, emphasizing its "frontier performance almost too cheap to meter."

The claims of significant progress are well-substantiated. Muse Spark 1.3 builds upon the foundation laid by last month’s 1.2 release, showcasing marked improvements, especially in handling long-running agent tasks. The version now accessible to developers represents one of the most compelling price-performance offerings among the top-tier independent model rankings. Meta’s most impressive benchmark results for Muse Spark 1.3 emerge from its "max reasoning" configuration. However, this top-tier version is currently undergoing additional safety testing and is slated for release "shortly." Independent benchmarking firm Artificial Analysis has confirmed evaluating this "max" configuration in a limited partner preview but currently lists no API provider for it. The version broadly rolling out this week, accessible through the Muse Code harness and the Meta Model API, utilizes Meta’s existing reasoning settings, including the "xhigh" configuration. This distinction is crucial for enterprises evaluating the practical deployment of Muse Spark 1.3, shifting the focus from its theoretical frontier capabilities to its currently deployable performance and associated costs.

The shipping model, while exceptionally capable, does not hold the absolute benchmark leadership. Meta transparently discloses results for both configurations in its underlying evaluation reports, ensuring that the performance of the deployable model is not obscured. However, launch materials tend to highlight the "max" variant, which often garners the highest scores. For instance, Meta reports GDPval-AA v2 scores of 1,754 Elo for "max" versus 1,709 for "xhigh," OSWorld 2.0 scores of 66.9 versus 57.2, and JobBench scores of 64.9 versus 61.2. In some specific tests, the performance difference is negligible or even favors the "xhigh" version. For example, DeepSearchQA shows a tie at 89.4, and "xhigh" scores 89.2 on Terminal-Bench 2.1, slightly surpassing the "max" configuration’s 88.8.

Artificial Analysis scores Muse Spark 1.3 "max" at 62 on its Intelligence Index, with the currently available "xhigh" version scoring 61. This 61 score places Muse Spark 1.3 "xhigh" in parity with leading models like GPT-5.6 Sol max, Grok 4.6 high, and Claude Opus 5 high. Nevertheless, Anthropic models continue to occupy the apex of the leaderboard, with Claude Fable 5.1 reaching 66 at its maximum configuration and 65 at "xhigh," while Claude Opus 5 scores 63 at both "max" and "xhigh." Therefore, while Muse Spark 1.3 "xhigh" is firmly within the frontier cluster, it does not currently set the absolute frontier.

This represents a substantial upgrade from Muse Spark 1.2. Previous coverage of the 1.2 release indicated that Meta was fielding a competitive coding model, but it generally lagged behind Anthropic’s top offerings. Muse Spark 1.2 scored 82.9% on Terminal-Bench 2.1, compared to Opus 5’s 86.7%, and also fell behind Opus in other key coding comparisons presented by Meta. With the 1.3 iteration, Meta is no longer just a participant in this competition; on several coding and agentic evaluations, it is now trading blows with industry giants like OpenAI and Anthropic.

Meta emphasizes that the underlying model has also become more user-friendly. Muse Spark 1.3 is engineered to efficiently manage multiple workflows within extended conversational threads, leverage tools for context gathering, identify its own planning deficiencies, proactively seek user clarification when needed, and obtain confirmation before executing critical actions. Internal comparisons by Meta engineers revealed that Muse Spark 1.3 utilized approximately 20% fewer tool calls and 25% fewer tokens than its predecessor during coding tasks. For enterprises that rely on thousands or millions of agent interactions, these behavioral enhancements could prove more impactful than marginal gains on benchmark leaderboards.

Zuckerberg’s declaration of Muse Spark 1.3 being "almost too cheap to meter" does not translate to a price reduction for API access. Meta has maintained the same Standard pricing for Muse Spark 1.3 as for Muse Spark 1.2: $1.25 per million input tokens, $4.25 per million output tokens, and $0.15 per million cached input tokens. This pricing structure makes Zuckerberg’s statement less about reduced per-token costs and more about the perceived value and capabilities that developers can unlock with those tokens.

Artificial Analysis provides compelling evidence to support this value proposition, but also introduces a nuance. Their metrics show Muse Spark 1.3 "xhigh" achieving 235.2 output tokens per second, with an estimated cost of $0.55 per Intelligence Index task. At a score of 61 on the Intelligence Index, this positions Muse Spark 1.3 "xhigh" as having the lowest cost per task among all models currently evaluated at that intelligence level. This is a notable improvement from Muse Spark 1.2, which cost $0.40 per task while scoring 57. Despite unchanged per-token pricing, the cost of completing an average task, as measured by Artificial Analysis, has effectively increased from generation to generation. This rise is primarily attributed by Artificial Analysis to increased input-token consumption on agentic evaluations.

Meta says Muse Spark 1.3 has frontier performance — but its best results come from a model developers can’t broadly use yet

This observation does not directly contradict Meta’s reported 25% reduction in token usage. Meta’s figure likely refers to specific internal coding workflows, whereas Artificial Analysis measures a broader spectrum of reasoning and agentic tasks. However, it highlights the complexity of "cheapness" when models operate as agents. The actual cost of completing a task is a confluence of token rates, reasoning effort, the number of conversational turns, tool usage, and potential retries.

Meta also continues to offer its highly competitive "Contributor" tier. This tier provides exceptionally low rates of $0.10 per million input tokens and $0.20 per million output tokens, in exchange for permission to use prompts and completions for training purposes. As previously noted with Muse Spark 1.2, this tier may be attractive for prototyping and development but introduces significant data governance considerations for enterprises handling proprietary code or sensitive internal information.

Meta’s Chief AI Officer, Alexandr Wang, offered a more direct and less qualified endorsement of the new model’s capabilities. Following the release of Muse Spark 1.3’s benchmark results, Wang reposted them on X with the provocative comment: "i really hate to say it, but… gemini who? 🙄🤦‍♀️". This jab was particularly pointed, as Google concurrently released its Gemini 3.8 Flash model on the same day. Google positioned Gemini 3.8 Flash for remarkably similar workloads: long-horizon software engineering, autonomous agents, and multi-step professional reasoning. Google touts 3.8 as its most advanced reasoning and coding Flash model to date and marks its third Flash release in just six weeks.

Independent benchmarks offer Wang’s assertion some backing, though not a definitive knockout blow. Artificial Analysis scores Muse Spark 1.3 "xhigh" at 61 on the Intelligence Index with a cost of $0.55 per task. In comparison, Gemini 3.8 Flash at high reasoning scores 59 with a cost of $0.58 per task. At these specific settings, Meta demonstrates a slight edge over Google in both intelligence and task cost. However, Google holds a decisive advantage in throughput. Artificial Analysis measures Gemini 3.8 Flash high at approximately 305 output tokens per second, significantly faster than Muse Spark’s 235 tokens per second – a difference of roughly 30%. Furthermore, Gemini currently boasts a lower raw API sticker price. Google is offering an introductory rate of $0.75 per million input tokens and $3.75 per million output tokens, compared to Meta’s $1.25 and $4.25. It is important to note that this promotional Google pricing is set to expire on December 31, after which it will increase to $1.50 per million input tokens and $7.50 per million output tokens.

The current comparison presents a compelling snapshot of the intensely competitive economics at the frontier of model development. Meta currently leads this independent comparison by two Intelligence Index points and three cents per benchmark task. Google, on the other hand, offers substantially higher output throughput and more affordable raw tokens during its launch promotion. Wang’s "gemini who?" remark is entertaining executive banter. For an enterprise architect, the situation is more nuanced: Gemini currently offers a faster option, while Muse, by this independent measure, represents a slightly more capable high-effort agent.

A more significant consideration for many developers may lie beyond the immediate benchmark race and pricing strategies. When Meta launched Muse Code and Muse Spark 1.2 in August, it marked a notable departure from its established open-weight strategy, which had been instrumental in the widespread adoption of its Llama models. Muse Code and Spark 1.2 were proprietary, API-served products, a stark contrast to Meta’s previous advocacy for open AI.

However, Meta shifted its approach again just five days later. On August 10, it released the 30-billion-parameter Muse Glimmer under an Apache 2.0 license, signaling a return to open-source principles. Zuckerberg also announced plans to "open the weights for Muse Spark 1.2" in the coming weeks, a plan corroborated by separate reports from Reuters. The current release of Muse Spark 1.3 as another proprietary model introduces ambiguity into this evolving roadmap.

Meta’s latest announcement no longer specifically mentions Muse Spark 1.2 for open weights. Instead, its roadmap now includes "the Muse Spark open weights release," without specifying a version, release date, model size, or license. Zuckerberg echoed this sentiment on X, stating that "Muse Spark open weights releases" are imminent. For development teams that previously standardized on Llama due to the ability to download weights for self-hosting, customization, and fine-grained control over inference economics, this lack of clarity may be a more critical factor than incremental benchmark improvements.

Muse Spark 1.3 undeniably demonstrates Meta’s accelerated pace in iterating proprietary frontier models. The currently available "xhigh" configuration is fast, competitively priced, and significantly closer to the top of independent rankings than its predecessors. The "max" preview version further illustrates Meta’s capacity to push the model’s capabilities when afforded more reasoning compute. The crucial next step for Meta is to translate this development velocity into a clear and predictable roadmap that enterprises can rely on for planning. This includes making its most advanced capabilities broadly deployable and, crucially, delivering the promised open-weight Spark model.

By admin

Leave a Reply

Your email address will not be published. Required fields are marked *