Meta’s latest artificial intelligence model, Muse Spark 1.3, unveiled yesterday, represents a significant leap forward in speed and performance, particularly on third-party benchmarks, but with a crucial caveat regarding its most potent configuration. Mark Zuckerberg, Meta’s co-founder and CEO, heralded the release on X, calling it Meta’s "biggest jump" yet in coding and agentic work, and optimistically noting that Muse Spark 1.3 is "rolling out today with frontier performance almost too cheap to meter." This pronouncement is substantiated by substantial improvements over the previous month’s 1.2 release, especially in handling long-running agent tasks. The version now accessible to developers stands as one of the most compelling price-performance offerings near the apex of independent model rankings.
The most striking benchmark results for Muse Spark 1.3 are achieved with its "max reasoning" configuration. However, Meta has indicated that this particular version is still undergoing additional safety testing and is slated for release "shortly." The independent benchmarking firm Artificial Analysis confirmed it evaluated the "max" configuration in a limited partner preview but currently lists no API provider for it, underscoring its restricted availability. The version being broadly rolled out this week, accessible through Meta’s Muse Code harness and the Meta Model API, utilizes the company’s previously available reasoning settings, including "xhigh." This distinction shifts the enterprise-focused question from whether Muse Spark 1.3 can reach frontier territory to how close the currently deployable model gets and at what actual cost.
The broadly available Muse Spark 1.3 model is indeed impressive, though not the definitive benchmark leader. Meta has transparently disclosed results for both configurations within its evaluation report, dispelling any notion of data concealment. Nevertheless, the launch materials prominently feature the "max" variant, and several of its highest scores are attributed to this configuration. For instance, Meta reports GDPval-AA v2 scores of 1,754 Elo for "max" compared to 1,709 for "xhigh," OSWorld 2.0 scores of 66.9 versus 57.2, and JobBench scores of 64.9 versus 61.2. On certain tests, the performance difference is negligible or even reversed: DeepSearchQA is tied at 89.4, while "xhigh" scores 89.2 on Terminal-Bench 2.1, slightly edging out "max" at 88.8.
Artificial Analysis places Muse Spark 1.3 "max" at 62 on its Intelligence Index, with the shipping "xhigh" version at 61. This latter score matches GPT-5.6 Sol max, Grok 4.6 high, and Claude Opus 5 high. However, Anthropic’s models still hold the top positions on the leaderboard, with Claude Fable 5.1 reaching 66 at "max" and 65 at "xhigh," and Claude Opus 5 achieving 63 at both "max" and "xhigh." This positions Muse Spark 1.3 "xhigh" firmly within the frontier cluster but not as the vanguard.
This represents a significant advancement from Muse Spark 1.2. VentureBeat’s previous coverage of that release highlighted Meta’s credible entry into the coding AI arena, yet it generally trailed Anthropic’s leading models. Muse Spark 1.2 achieved an 82.9% score on Terminal-Bench 2.1, falling short of Opus 5’s 86.7%, and also lagged behind Opus in other key coding comparisons presented by Meta. With the 1.3 iteration, Meta is no longer merely participating; it is actively competing and trading wins with OpenAI and Anthropic on several coding and agentic evaluations.
Beyond raw performance, Meta emphasizes that the underlying model has become more user-friendly. Muse Spark 1.3 is engineered to manage multiple workflows within a single long thread, effectively gather context using tools, identify gaps in its own planning, solicit user clarification when needed, and confirm actions before execution. Internal comparisons by Meta engineers indicate that it utilized approximately 20% fewer tool calls and 25% fewer tokens than 1.2 during coding tasks. For enterprises managing thousands or millions of agentic loops, these behavioral enhancements could hold more practical value than incremental gains on leaderboards.
Zuckerberg’s assertion of being "almost too cheap to meter" does not translate to a reduction in API pricing for Muse Spark 1.3. Meta has maintained the standard pricing established for Muse Spark 1.2: $1.25 per million input tokens, $4.25 per million output tokens, and $0.15 per million cached input tokens. This pricing structure means Zuckerberg’s comment is less about lowering token costs and more about the perceived value developers can extract from those tokens.
Artificial Analysis provides data that supports this perspective, albeit with a complication. The firm measures Muse Spark 1.3 "xhigh" at 235.2 output tokens per second and estimates a cost of $0.55 per Intelligence Index task. With an Intelligence Index score of 61, this translates to the lowest cost per task among all currently measured models at that intelligence level. In contrast, Muse Spark 1.2 cost only $0.40 per Artificial Analysis task while scoring 57. Despite the unchanged per-token pricing, the cost of completing an average task, as measured by the independent benchmark, has consequently increased from one generation to the next.

Artificial Analysis attributes this increase primarily to a higher consumption of input tokens on agentic evaluations. This finding does not directly contradict Meta’s claim of 25% lower token usage; Meta’s figure pertains to its internal coding workflows, whereas Artificial Analysis measures a broader spectrum of reasoning and agentic tasks. However, it illustrates how the definition of "cheap" becomes more complex once models function as agents. Token rates, reasoning complexity, the number of conversational turns, tool calls, and potential retries all contribute to the actual cost of task completion.
Meta also continues to offer its unusually inexpensive "Contributor" tier, priced at $0.10 per million input tokens and $0.20 per million output tokens. This tier is contingent on granting Meta permission to use prompts and completions for training purposes. As VentureBeat observed with Muse Spark 1.2, this tier may be attractive for prototyping but presents a significantly different data governance calculation for enterprises handling proprietary code or sensitive internal information.
Meta’s Chief AI Officer, Alexandr Wang, expressed a more unreserved enthusiasm for the release. Following Artificial Analysis’s publication of Muse Spark results, Wang reposted them on X, provocatively adding, "i really hate to say it, but… gemini who? 🙄🤷‍♀️". This jab was particularly pointed given that Google released Gemini 3.8 Flash on the same day, positioning it for similar workloads such as long-horizon software engineering, autonomous agents, and multi-step professional reasoning. Google describes 3.8 as its best reasoning and coding Flash model to date and its third Flash release in just six weeks.
Independent metrics offer Wang some ammunition, though not a decisive victory. Artificial Analysis scores Muse Spark 1.3 "xhigh" at 61 on the Intelligence Index with a cost of $0.55 per task. This compares to Gemini 3.8 Flash’s score of 59 and a cost of $0.58 at high reasoning. Under these specific settings, Meta edges out Google in both intelligence and task cost. However, Google demonstrates a decisive advantage in throughput. Artificial Analysis measures Gemini 3.8 Flash high at approximately 305 output tokens per second, substantially exceeding Muse Spark’s 235 tokens per second—a roughly 30% faster rate. Furthermore, Gemini boasts a lower raw API sticker price during its introductory period: Google is charging $0.75 per million input tokens and $3.75 per million output tokens, compared to Meta’s $1.25 and $4.25. This promotional pricing from Google is set to expire on December 31, after which it will increase to $1.50 per million input tokens and $7.50 per million output tokens.
The current situation provides a valuable snapshot of the increasingly competitive economics of frontier models. Meta currently leads this independent comparison by two Intelligence Index points and three cents per benchmark task. Google, on the other hand, offers significantly higher output throughput and more affordable token rates during its launch promotion. Wang’s "gemini who?" remark, while entertaining executive banter, simplifies a more nuanced reality for enterprise architects: Gemini presents a faster option, while Muse, by this independent measure, currently operates as a slightly more robust high-effort agent.
Perhaps more consequential for some developers than the immediate benchmark race is Meta’s evolving stance on open weights. When Meta initially launched Muse Code and Muse Spark 1.2 in August, VentureBeat noted a significant departure from the open-weight strategy that had propelled Llama to ubiquity. Muse Code and Spark 1.2 were proprietary, API-served products, a striking contrast to Meta’s prior advocacy for open AI.
However, Meta shifted course again just five days later. On August 10, the company released the 30-billion-parameter Muse Glimmer under an Apache 2.0 license. Zuckerberg also stated Meta’s intention to "open the weights for Muse Spark 1.2" in the ensuing weeks, a plan also reported by Reuters. Now, with the release of Muse Spark 1.3 as another proprietary model, Meta’s roadmap appears less clear.
The latest announcement no longer specifies Muse Spark 1.2 for open weights. Instead, the roadmap mentions "the Muse Spark open weights release" without specifying a version, release date, model size, or license. Zuckerberg echoed this ambiguity on X, referring to "Muse Spark open weights releases" coming soon. For development teams that have standardized on Llama due to the ability to self-host, customize, and control inference economics with downloadable weights, this ambiguity may carry more weight than a fractional gain on a composite benchmark.
Muse Spark 1.3 demonstrates Meta’s accelerated iteration cycle for proprietary frontier models. The currently available "xhigh" configuration is fast, competitively priced, and significantly closer to the top of independent rankings than its predecessors. The "max" preview version indicates Meta’s capability to push the model’s performance further when allowed to allocate more reasoning compute. The critical next challenge for Meta is to translate this rapid development pace into a predictable roadmap for enterprises, ensuring its most advanced capabilities are broadly deployable and delivering the promised open-weight Spark model.

