10 Sep 2026, Thu

Meta’s Newest AI Model, Muse Spark 1.3, Offers Frontier Performance at an Unprecedented Price Point, But With a Crucial Caveat

Meta’s latest artificial intelligence model, Muse Spark 1.3, unveiled yesterday, is making waves in the AI development community, boasting significant improvements in speed and performance on third-party benchmarks compared to its predecessor. This advancement, however, comes with a notable caveat that enterprises and developers must carefully consider. Mark Zuckerberg, Meta’s co-founder and CEO, enthusiastically announced the release on X, describing it as "frontier performance almost too cheap to meter" and heralding it as Meta’s "biggest jump" yet in coding and agentic capabilities.

The substance behind Zuckerberg’s bold claims is evident. Muse Spark 1.3 demonstrates substantial gains over the previous iteration, Muse Spark 1.2, which was released just last month. These improvements are particularly pronounced in the realm of long-running agent tasks, a critical area for developing sophisticated AI applications. The current version accessible to developers positions itself as one of the most compelling price-performance offerings among leading independent model rankings.

However, Meta’s most impressive benchmark results for Muse Spark 1.3 stem from its "max reasoning configuration." This top-tier version is currently undergoing additional safety testing and is slated for release "shortly." While the third-party benchmarking firm Artificial Analysis has evaluated this "max" configuration in a limited partner preview, it currently lists no API provider for this specific setup, indicating its restricted availability. The version that is broadly rolling out this week, accessible through Meta’s Muse Code harness and the Meta Model API, utilizes the company’s previously available reasoning settings, including the "xhigh" configuration. This distinction raises the pertinent enterprise question: not if Muse Spark 1.3 can reach frontier territory, but how close the currently deployable version gets, and at what actual cost.

The Shipping Model: Impressive, Yet Not the Benchmark Leader

Meta has been transparent about the performance of both configurations, disclosing results for each in its underlying evaluation report. This is not a scenario where the company is withholding data on its deployable model. However, its launch materials prominently feature the "max" variant, and some of its most significant benchmark scores are attributed to this configuration. For instance, Meta reports GDPval-AA v2 scores of 1,754 Elo for "max" versus 1,709 for "xhigh," OSWorld 2.0 scores of 66.9 versus 57.2, and JobBench scores of 64.9 versus 61.2.

In some instances, the performance difference between the two configurations is negligible or even reversed. DeepSearchQA scores are tied at 89.4, and the "xhigh" version scores 89.2 on Terminal-Bench 2.1, slightly edging out the "max" configuration at 88.8.

Artificial Analysis corroborates this, scoring Muse Spark 1.3 "max" at 62 on its Intelligence Index, while the widely available "xhigh" version scores 61. This latter score ties it with notable models like GPT-5.6 Sol max, Grok 4.6 high, and Claude Opus 5 high. However, Anthropic’s models continue to hold the top positions on the leaderboard, with Claude Fable 5.1 reaching 66 at its "max" setting and 65 at "xhigh," and Claude Opus 5 achieving 63 at both its "max" and "xhigh" configurations.

Therefore, while Muse Spark 1.3 "xhigh" is undoubtedly operating within the frontier cluster of AI models, it is not currently setting the pace. This represents a significant leap forward from Muse Spark 1.2. VentureBeat’s coverage of the previous month’s launch highlighted Meta’s emergence as a credible coding challenger, yet it generally lagged behind Anthropic’s top-tier models. Muse Spark 1.2 achieved an 82.9% score on Terminal-Bench 2.1, compared to Opus 5’s 86.7%, and also fell behind Opus on other key coding comparisons presented by Meta.

With the release of Muse Spark 1.3, Meta is no longer merely participating in this competitive landscape; it is actively trading wins with industry giants like OpenAI and Anthropic across several coding and agentic evaluations.

Meta also emphasizes that the underlying model has become more user-friendly and capable. Muse Spark 1.3 is engineered to manage multiple workflows within a single, extended conversation thread, efficiently gather context using tools, identify deficiencies in its own planning processes, solicit user clarification when necessary, and confirm actions before executing them. In internal comparisons conducted by Meta engineers, Muse Spark 1.3 utilized approximately 20% fewer tool calls and 25% fewer tokens than Muse Spark 1.2 during coding tasks. For enterprises operating at scale, processing thousands or even millions of agent loops, these behavioral enhancements could prove more impactful than marginal gains on benchmark leaderboards.

"Almost Too Cheap to Meter": A Matter of Value, Not Price Cuts

Despite Zuckerberg’s evocative phrase, the release of Muse Spark 1.3 did not come with a reduction in API pricing. Meta has maintained the exact same pricing structure as for Muse Spark 1.2: $1.25 per million input tokens, $4.25 per million output tokens, and $0.15 per million cached input tokens. This pricing strategy suggests that Zuckerberg’s "almost too cheap to meter" comment is less about lowering token costs and more about the perceived value and capabilities developers can unlock with those tokens.

The table below illustrates the pricing landscape for various AI models, highlighting Muse Spark 1.3’s position:

Model Input ($/1M) Output ($/1M) Total ($/1M) Source
Muse Spark 1.2 / 1.3 Contributor $0.10 $0.20 $0.30 Meta
MiMo-V2.5 Flash $0.10 $0.30 $0.40 Xiaomi
DeepSeek-V4-Flash – off-peak $0.22 $0.66 $0.88 DeepSeek
GPT-5.6 Luna $0.20 $1.20 $1.40 OpenAI
MiniMax-M3 $0.30 $1.20 $1.50 MiniMax
LongCat-2.0 – limited-time promo $0.30 $1.20 $1.50 LongCat
DeepSeek-V4-Flash – peak hours $0.44 $1.32 $1.76 DeepSeek
MiMo-V2.5 $0.40 $2.00 $2.40 Xiaomi
DeepSeek-V4-Pro – off-peak $0.66 $1.98 $2.64 DeepSeek
LongCat-2.0 – standard $0.75 $2.95 $3.70 LongCat
MiMo-V2.5 Pro (≤256K) $1.00 $3.00 $4.00 Xiaomi
Gemini 3.7 Flash – through Dec. 31, 2026 $0.75 $3.75 $4.50 Google
Gemini 3.8 Flash – through Dec. 31, 2026 $0.75 $3.75 $4.50 Google
DeepSeek-V4-Pro – peak hours $1.32 $3.96 $5.28 DeepSeek
Muse Spark 1.1 / 1.2 / 1.3 $1.25 $4.25 $5.50 Meta
GLM-5.3 $1.40 $4.40 $5.80 Z.AI
Grok 4.6 – <200K prompt tokens $2.00 $6.00 $8.00 xAI
MiMo-V2.5 Pro (>256K) $2.00 $6.00 $8.00 Xiaomi
Qwen3.8-Max $2.00 $6.00 $8.00 QwenCloud
Gemini 3.7 Flash – starting Jan. 1, 2027 $1.50 $7.50 $9.00 Google
Gemini 3.8 Flash – starting Jan. 1, 2027 $1.50 $7.50 $9.00 Google
GPT-5.6 Terra $2.00 $12.00 $14.00 OpenAI
Grok 4.6 – ≥200K prompt tokens $4.00 $12.00 $16.00 xAI
GPT-5.4 $2.50 $15.00 $17.50 OpenAI
Kimi K3 $3.00 $15.00 $18.00 Moonshot AI
Claude Opus 5 $5.00 $25.00 $30.00 Anthropic
Sakana Fugu Ultra (≤272K) $5.00 $30.00 $35.00 Sakana AI
GPT-5.6 Sol – Standard mode $5.00 $30.00 $35.00 OpenAI
Claude Fable 5 / Claude Mythos 5 $10.00 $50.00 $60.00 Anthropic
Claude Fable 5.1 / Claude Mythos 5.1 $10.00 $50.00 $60.00 Anthropic
GPT-5.6 Sol – Fast mode $10.00 $60.00 $70.00 OpenAI

Artificial Analysis provides data that supports Meta’s value proposition. It measures Muse Spark 1.3 "xhigh" at an impressive 235.2 output tokens per second and estimates a cost of $0.55 per Intelligence Index task. Given its Intelligence Index score of 61, this results in the lowest cost per task among all currently measured models at that intelligence level.

This is a significant improvement from Muse Spark 1.2, which cost $0.40 per Artificial Analysis task while scoring 57 on the Intelligence Index. Despite the unchanged per-token pricing, the cost of completing an average task, as measured by Artificial Analysis, has actually increased from generation to generation. This increase is primarily attributed to heavier input-token consumption observed in agentic evaluations.

Meta says Muse Spark 1.3 has frontier performance — but its best results come from a model developers can’t broadly use yet

This observation does not directly contradict Meta’s reported 25% reduction in token use. Meta’s figure pertains to its internal coding workflows, whereas Artificial Analysis measures a broader spectrum of reasoning and agentic tasks. However, it vividly illustrates the complexity of "cheap" when models begin operating as agents. The actual cost of completing a task is a confluence of token rates, reasoning effort, conversational turns, tool invocations, and potential retries.

Meta also continues to offer its exceptionally low-cost "Contributor" tier, priced at $0.10 per million input tokens and $0.20 per million output tokens. This tier requires users to grant Meta permission to use their prompts and completions for model training. As noted with Muse Spark 1.2, this tier may be attractive for prototyping but presents a considerably different data governance calculation for enterprises handling proprietary code or sensitive internal information.

Wang’s "Gemini Who?" Remark Lands Near a Close Competitor

Meta’s Chief AI Officer, Alexandr Wang, was less reserved in his praise for the new model. Following the release of Muse Spark’s benchmark results, Wang reposted them on X with the provocative caption: "i really hate to say it, but… gemini who? 🙄🤷‍♂️".

This pointed jab comes on the heels of Google’s simultaneous release of Gemini 3.8 Flash, a model positioned to compete in precisely the same arena: long-horizon software engineering, autonomous agents, and multi-step professional reasoning. Google touts Gemini 3.8 as its best reasoning and coding Flash model to date and marks its third Flash release in just six weeks.

Independent benchmarks provide some ammunition for Wang’s assertion, though not a decisive victory. Artificial Analysis scores Muse Spark 1.3 "xhigh" at 61 on the Intelligence Index with a cost of $0.55 per task. In comparison, Gemini 3.8 Flash, at its "high" reasoning setting, scores 59 with a cost of $0.58 per task. Under these specific parameters, Meta’s model edges out Google’s in both intelligence and task cost.

However, Google demonstrably wins on throughput. Artificial Analysis measures Gemini 3.8 Flash "high" at approximately 305 output tokens per second, significantly outpacing Muse Spark’s 235 tokens per second – a difference of roughly 30%. Furthermore, Gemini currently boasts a lower raw API sticker price. Google is offering an introductory rate of $0.75 per million input tokens and $3.75 per million output tokens, which is more competitive than Meta’s $1.25 and $4.25 respectively. It’s important to note that this promotional Google pricing is set to expire on December 31, after which it will increase to $1.50 per million input tokens and $7.50 per million output tokens.

This comparison offers a valuable snapshot of the intensely competitive economics at the frontier of AI model development. Currently, Meta holds a slight advantage in independent benchmarks, leading Google by two Intelligence Index points and three cents per benchmark task. Conversely, Google offers substantially higher output throughput and a more attractive token price during its launch promotion.

Wang’s "Gemini who?" remark, while entertaining executive banter, frames the enterprise decision-maker’s dilemma. Gemini presents itself as the faster option, while Muse Spark, by this independent measure, currently offers a slightly more capable high-effort agent.

Meta’s Evolving Stance on Open Weights

For a segment of developers, the more pressing concern may lie beyond the immediate benchmark race and more in Meta’s evolving strategy regarding open-weight models. When Meta initially launched Muse Code and Muse Spark 1.2 in August, it marked a significant departure from the open-weight philosophy that propelled Llama to widespread adoption. Muse Code and Spark 1.2 were exclusively proprietary, API-served products, a stark contrast to Meta’s previous advocacy for open AI.

However, Meta shifted its approach again just five days later. On August 10, it released the 30-billion-parameter Muse Glimmer under an Apache 2.0 license, signaling a return to its open-source roots. At the time, Zuckerberg also stated Meta’s intention to "open the weights for Muse Spark 1.2" in the coming weeks, a plan that was also reported by Reuters.

Now, with the release of Muse Spark 1.3 as another proprietary model, the roadmap has become less clear. Meta’s latest announcement no longer specifies Muse Spark 1.2 for its open-weights release. Instead, it mentions "the Muse Spark open weights release" without providing details on version, release date, model size, or license. Zuckerberg’s recent comments on X similarly refer to "Muse Spark open weights releases" coming soon.

This ambiguity could be a significant factor for teams that have standardized on Llama due to the benefits of downloadable weights, enabling self-hosting, customization, and granular control over inference economics. The uncertainty surrounding the release of open-weight Spark models may hold more weight for these developers than a marginal improvement on a composite benchmark.

Muse Spark 1.3 undeniably demonstrates Meta’s capacity for rapid iteration on proprietary frontier models. The shipping "xhigh" configuration is fast, competitively priced, and has significantly narrowed the gap with top-ranked models. The "max" preview suggests Meta can achieve even greater performance by allocating more reasoning compute. The next crucial test for Meta will be its ability to translate this development pace into a predictable roadmap that enterprises can rely on for planning. This includes making its most advanced capabilities broadly accessible and fulfilling its commitment to release the open-weight Spark model that has already been announced.

By admin

Leave a Reply

Your email address will not be published. Required fields are marked *