Alibaba Unveils Qwen3.8-Max: A 2.4 Trillion Parameter Giant for AI Agents
Alibaba’s Qwen team has announced Qwen3.8-Max, a massive 2.4 trillion parameter flagship model designed to move beyond simple chat prompts toward long-horizon autonomous reasoning. By focusing on multi-day workflows and complex task execution, this model marks a significant shift from reactive LLMs to proactive AI agents.
Scaling Parameters for Autonomous Intelligence
Qwen3.8-Max is a technical powerhouse, utilizing a total of 2.4 trillion parameters with 95 billion active parameters per query. Building upon the Qwen3.5 architecture, the model is specifically optimized for "long-horizon" tasks—complex objectives that require the AI to operate independently over days or even weeks.
In a significant move for the open-source community, Alibaba has confirmed that the weights for the Qwen-Max class will be made publicly available next week, challenging the closed-door dominance of frontier models like GPT and Claude.
Proving Autonomy Through Coding and Research
To demonstrate its agentic capabilities, Alibaba presented several high-stakes case studies where the model operated without human intervention:
- Software Engineering: Over a 16-day period, Qwen3.8-Max autonomously developed the command-line tool
oh-my-cli. It managed its own GitHub issues, wrote code, ran tests, and executed 265 commits and 127 pull requests. - Scientific Reproduction: Tasked with reproducing the results of the "Unified Data Selection for LLM Reasoning" paper, the model spent 125 hours of compute time writing 7,600 lines of code. It not only replicated the original research but improved the AIME24 math benchmark score by 2.7 points.
- Competitive Data Science: In the WWW2025 Multimodal Dialogue Intent Recognition Challenge, the model outperformed 458 out of 526 human teams, scaling its accuracy from 0.60 to 0.853 within 24 hours.
Beyond Software: Chip Design and Fiscal Simulation
Qwen3.8-Max's reasoning extends into hardware and economics. In a cryptographic circuit design task, the model iteratively reduced a "bloated" design from 8,298 logic gates to just 678 gates. Using the OpenROAD tool, this resulted in an 81% reduction in physical chip area.
In a complex retail simulation called E-Commerce-Bench, the model managed multiple online stores through a simulated fiscal year. Starting with 100,000 yuan, it navigated supplier negotiations, identified 152 scammers, and managed supply chain disruptions to end with 416,252 yuan—quadrupling its initial capital and significantly outperforming the runner-up, GLM 5.2.
Multimodal Frontiers and Benchmark Dominance
The model also pushes the boundaries of multimodal processing, capable of handling 200-page documents and videos exceeding 100 hours. Alibaba is also introducing RecreationBench, a benchmark that tests an agent's ability to reconstruct running applications (on Windows, macOS, Android, etc.) solely through visual and keyboard interaction, without access to source code.
According to internal benchmarks, Qwen3.8-Max competes directly with top-tier models, landing near or above Claude Opus 4.8 and GPT-5.6 Sol in several categories, including a leading score of 93 on PaperBench.
Key Takeaways
- Massive Scale: Qwen3.8-Max features 2.4 trillion total parameters and 95 billion active parameters, optimized for multi-day autonomous workflows.
- Agentic Mastery: The model has demonstrated the ability to handle complex software engineering, chip design optimization, and full-scale fiscal management autonomously.
- Open-Weight Impact: Unlike many frontier models, Alibaba plans to release the weights for the Qwen-Max class, providing developers with unprecedented access to high-level agentic intelligence.
