6 Aug 2026, Thu

Qwen 3.8-Max Arrives with a Bold Claim: It Outperforms GPT-5, 6-Sol-Max, and Fable-5 on Agentic Computer Use.

Alibaba’s recent unveiling of Qwen 3.8-Max has ignited a fervent debate within the AI community, particularly concerning its performance claims on complex agentic computer use tasks. While Alibaba initially marketed the preview version as a near-peer to Claude Fable 5, with their launch-day table suggesting leadership on one of twelve coding-agent rows, an independent benchmark run by VulcanBench has cast a shadow of doubt. This independent assessment, reportedly using the same preview version, placed Qwen 3.8-Max’s best-effort setting in the mid-pack and its default setting at the very bottom. This stark divergence in results, however, is not necessarily indicative of a flawed model but rather a testament to the profound impact of unstated or overlooked parameters like token and time budgets in AI performance evaluation.

The discrepancy between Alibaba’s published results and the independent benchmark is largely attributable to differences in computational resource allocation, specifically token and time budgets. Alibaba’s own footnotes reveal generous time allowances for its coding benchmarks: a five-hour timeout for general coding tasks and an extensive up to twelve hours per run on the PaperBench dataset. In contrast, the independent VulcanBench harness operated with significantly tighter constraints, allowing only between 45 and 60 minutes of wall clock time per task. This disparity of a five to sixteen times larger time budget on Alibaba’s side offers a compelling explanation for the substantial variation in reported performance. When evaluating AI models, particularly those designed for agentic tasks requiring iterative reasoning and problem-solving, these budgetary constraints are not mere technicalities; they are fundamental determinants of success.

The implications of this revelation are far-reaching, necessitating a fundamental shift in how we approach AI model selection and evaluation. Two critical adjustments are paramount. Firstly, the primary metric for assessing AI performance must evolve from raw accuracy or speed to a more holistic measure: cost per successful task. This metric inherently accounts for all resources expended, including the cumulative cost of failed attempts, divided by the number of tasks that have successfully met predefined acceptance criteria. This approach provides a truer reflection of an AI’s economic viability and practical utility. Secondly, and equally crucial, time and token budgets must be elevated from hidden implementation details to explicit components of acceptance criteria. They should no longer be an afterthought but a foundational element in any evaluation framework, ensuring that performance is assessed within realistic and clearly defined operational parameters.

The traditional approach of evaluating AI models primarily through price comparisons has become increasingly unreliable, especially for sophisticated reasoning models like Qwen. In the initial week following Qwen 3.8-Max’s release, the public discourse was dominated by pricing discussions, as it was the most readily available data point. Qwen 3.8-Max, with its listed prices of $2 for input tokens and $6 for output tokens, is demonstrably more expensive than competitors like DeepSeek-V4-Flash-0731, which is priced at 14 cents per million input tokens and 28 cents per million output tokens. Kimi K3 further highlights this price disparity, with a $3 input and $15 output token cost. However, these figures, while informative regarding raw API costs, fail to capture the true economic picture for reasoning-intensive tasks.

The reason for this disconnect lies in the inherent nature of advanced AI models. Achieving a coherent and accurate result often requires the model to expend a significant portion of its token allowance on internal reasoning processes. This "thinking time" can lead to the model reaching its token limit before it can formulate and output the final answer. The outcome is an empty or incomplete result, functionally indistinguishable from a complete failure, yet incurring the full cost of a completed run. This phenomenon underscores the inadequacy of simple price-per-token metrics.

Artificial Analysis provides a compelling illustration of this issue with its Intelligence Index benchmarks. Their assessment of DeepSeek-V4-Flash on maximum effort revealed an astonishing 210 million output tokens consumed, far exceeding the class median of 100 million. While the absolute cost remained relatively low due to DeepSeek’s inexpensive tokens, the significant verbosity highlights a critical trade-off: excessive token consumption directly translates into increased time expenditure. Depending on the specific application, this time cost can be a far more detrimental factor than the monetary expense.

Therefore, the need for a more nuanced evaluation metric is clear. Cost per successful task emerges as the indispensable metric, encompassing all expenditures, including those on failed attempts, and normalizing them against the tasks that have demonstrably met predefined quality and completeness standards within specified time and token constraints. This metric provides a genuine measure of an AI’s effectiveness and efficiency in achieving desired outcomes.

The concept of "failure rate" in AI performance is not a monolithic entity; it is, to a significant degree, a configurable setting dictated by the operational parameters of the benchmark or application. A run that yields an incorrect answer and a run that exhausts its allocated budget are fundamentally different events, requiring distinct diagnostic approaches and remedial actions. Astonishingly, the vast majority of AI evaluation harnesses fail to differentiate between these scenarios, and consequently, most leaderboards report a generalized failure rate without dissecting the underlying causes. This oversight was a significant challenge encountered during the development of an agent benchmark by the author, where initial failure logs lacked crucial detail. The necessity to manually introduce this distinction became apparent, revealing that budget exhaustion frequently emerged as the dominant cause of failure.

The Long-Horizon-Terminal-Bench, published in July, offers empirical evidence supporting this assertion. This study evaluated seventeen frontier AI models across forty-six distinct tasks, each with a single ninety-minute attempt. The findings were striking: timeouts accounted for an overwhelming 79% of unresolved runs. Agents that self-terminated represented a mere 19%, with harness errors comprising a negligible 3%. While the authors prudently caution against the assumption that additional time would have guaranteed success – noting that timed-out runs exhibited low mean reward (between 0.10 and 0.35) – the underlying message is unambiguous. Benchmarks, by their very nature, are implicitly measuring time efficiency, regardless of whether this aspect is explicitly emphasized.

A particularly lucid illustration of this mechanism can be found in the work of VulcanBench, the same open-source harness responsible for the Qwen performance chart. A report dated July 26th highlighted a fascinating trade-off with Claude Opus 5. At its lowest effort setting, the model achieved a superior performance, successfully completing 20 out of 23 tasks. In contrast, its high-effort setting, while returning fewer incorrect answers (one versus three), struggled with time constraints, leading to timeouts on tasks that the low-effort setting successfully navigated. While the high-effort setting theoretically offered more sophisticated reasoning, its benefits were nullified by the clock. Given unlimited time, the high-effort setting merely matched its cheaper counterpart in task completion, but at a staggering 3.1 times the cost.

This dynamic has direct and significant implications for the design of AI routing ladders, a common architectural pattern where simpler, cheaper models handle initial requests, escalating to more complex and expensive models when the initial attempt fails. The prevailing assumption is that the escalated model offers superior performance and merely incurs a higher cost. However, for a substantial proportion of model and task combinations, this assumption proves erroneous. Instead of achieving improved results, users end up paying the premium for a more sophisticated model only to encounter a timeout or reach a computational cap, effectively nullifying any potential benefit.

The growing recognition of these shortcomings has spurred several research groups and organizations to independently develop and adopt the cost-per-successful-task metric. This convergence of independent efforts strongly signals its imminent emergence as a standard evaluation paradigm. VulcanBench, for instance, has consistently reported dollars per solved task as a prominent column in its benchmarks since its earliest publications, demonstrating a long-standing commitment to this more practical metric. The Long-Horizon-Terminal-Bench also publishes per-task cost alongside accuracy figures, providing insightful comparisons. A particularly striking example is the comparison of GPT-5.4, which costs approximately $26 per task but achieves a lower pass rate than Grok 4.5 at about $11 per task. TestEvo-Bench further emphasizes this by running agents under a strict cost cap, revealing a significant drop in Claude Code’s test-generation score from 71% to 44% when subjected to tighter budget constraints.

Vendors are not lagging behind in recognizing the value of this outcome-oriented pricing model. HubSpot, for example, transitioned its Breeze Customer Agent in April to a pricing structure of 50 cents per resolved conversation, a significant improvement from its previous $1 per handled conversation. Zendesk has also adopted a per-automated-resolution billing model, aligning its pricing with tangible outcomes. Fin charges 99 cents per outcome, with billing contingent only upon end-to-end resolution, further solidifying the industry’s pivot towards value-based pricing.

In light of these evolving insights, several actionable changes are recommended for immediate implementation. Firstly, when evaluating AI models, prioritize cost-per-successful-task as the primary metric, integrating all expenses, including failed attempts, and dividing by demonstrably completed tasks. Secondly, explicitly define time and token budgets as integral components of acceptance criteria, moving them from implicit technical details to explicit performance benchmarks. This ensures that evaluations are conducted within realistic operational constraints and that the true cost-effectiveness of AI models can be accurately assessed. By embracing these shifts, organizations can move beyond superficial performance claims and make informed decisions that align with practical utility and economic efficiency in the rapidly advancing landscape of artificial intelligence.

By admin

Leave a Reply

Your email address will not be published. Required fields are marked *