HomeAlibaba Qwen3.8-Max claims 16-day autonomous coding runUncategorizedAlibaba Qwen3.8-Max claims 16-day autonomous coding run

Alibaba Qwen3.8-Max claims 16-day autonomous coding run

Alibaba’s Qwen team has released Qwen3.8-Max, a 2.4 trillion parameter model with 95 billion active parameters, pitched at coding, office work, research, and long-horizon tasks run with minimal supervision.

The model is available now through QwenCloud. Open weights follow next week, marking the first time Alibaba has released a Max-class Qwen model outside a closed API.

Qwen3.8-Max sits on the architecture introduced with Qwen 3.5, and the company’s own account of what it can do leans heavily on multi-day autonomous runs rather than single-turn answers. That’s the pitch, at least. What the company documents is a set of internal test cases: a 16-day coding project, a research paper reproduction, an online competition entry, a chip design exercise, and a year-long simulated e-commerce operation.

Ten days building software that upgrades itself

The most detailed case involves a project called oh-my-cli, built from an empty repository over what Qwen describes as a 10+ day autonomous run.

The model combined an issue state machine, a dispatcher, a monitor, and a watchdog into one loop: a new requirement lands in GitHub Issues, an agent claims it, moves it through ready, leased, and active states, then triggers end-to-end tests and CI checks before merging the pull request.

Self-testing ran on every update – build, unit test, end-to-end, and desktop lifecycle checks – with failures routed back to the originating issue for another pass. By 30 July 2026, after roughly 16 days of continuous operation, Qwen reports the repository had accumulated 265 commits, 127 pull requests, and 151 issues. 

The trace is public on GitHub under qwen-code-dev-bot/oh-my-cli, which at least gives outside parties something to inspect rather than take on faith.

Reproducing a paper, then beating it

A second case handed Qwen3.8-Max a published paper, ’Unified Data Selection for LLM Reasoning’, with instructions to reproduce its results and then improve on them. No starter code. No pipeline. Just the paper and GPU access.

Working alone for roughly 125 hours, the model wrote about 7,600 lines of code across 1,100 actions and 33 rounds of training. The first 37 hours went into rebuilding the paper’s pipeline and confirming its six main findings, including a +7.7 percent gain over random data selection on the AIME24 maths benchmark. 

Then it kept going. Over the following 88 hours it ran four rounds of hypothesis-test-analyse cycles, trying 18 of its own improvement ideas. The winning idea – counting “hard decision points” more precisely – pushed the AIME24 score from a 49.58 percent baseline to 52.29 percent, a 2.71-point gain Qwen attributes directly to the model’s own iteration rather than the original paper’s method.

A real contest, a real leaderboard

The clearest test against outside competition came from the WWW2025 Multimodal Dialogue Intent Recognition Challenge on Alibaba Cloud’s Tianchi platform, where 526 human teams entered. 

Qwen3.8-Max built a solution combining fine-tuned Chinese language models (BERT, MacBERT, and RoBERTa) for text, paired with a fine-tuned Qwen2.5-VL-7B and a Chinese-CLIP fallback for screenshots.

Across 45 submissions in 24 hours, accuracy climbed from 0.60 to 0.853. That placed it ahead of 458 of the 526 teams, or 87 percent of the field, according to Qwen. Of course, a leaderboard position among a fixed competitor pool at a single point in time isn’t the same claim as general superiority over human analysts on open-ended tasks.

Professions tested, productivity claimed

Qwen also ran the model against several hundred of what it calls high-value professions, and the write-up leans on comparison figures throughout: a compliance review completed in under an hour against a week for a paralegal team, a banking app prototype delivered in one pass against three to five revision rounds, a 26-dish restaurant menu built from a hundred supply briefs.

Every one of these comparisons is Qwen’s own framing of “traditional workflow” time, not a documented client engagement or a controlled study. Treat the multiples as vendor-supplied context, not measured productivity gains.

The quant research example follows the same pattern but with harder numbers attached. Qwen3.8-Max reportedly built an ETF-rotation strategy end to end, then parallelised factor mining across roughly 330 sub-agents running around 6,000 backtests from six seed descriptions.

The resulting factors carried excess Sharpe ratios between 0.64 and 1.48, with information coefficients between 0.010 and 0.014. Backtested performance. Not live trading.

Chip design and a 365-day simulation

Two cases push further into synthetic territory. The first tasked the model with designing a GCD/RSA cryptographic accelerator from a stub RTL template, running entirely inside a sandbox wired to Iverilog, Yosys, and OpenROAD.

Over roughly 500 turns and 71 evaluations, the model cut its first working design from 8,298 gates to 678, then carried that through physical layout—shrinking the die from 106×106 µm² to 46×46 µm² and closing timing at 500MHz. The single largest jump came at turn 22, when the model swapped a hardware modulo divider for an iterative shift-subtract design, cutting 6,288 gates in one move.

The second case, E-Commerce Bench, is a full year of simulated retail operation built on desensitised Taobao and Tmall transaction data, complete with 152 planted fraudulent suppliers and randomised supply shocks. Qwen3.8-Max finished that year with a simulated balance of ¥416,252 from ¥100,000 starting capital, ahead of second-place GLM 5.2 by 38 percent.

Availability and what’s still unverified

Qwen3.8-Max ships with a reasoning_effort parameter (xhigh, medium, or low) for trading off cost against depth, and preserve_thinking switched on by default. QwenCloud supports OpenAI-compatible chat completions and responses APIs alongside an Anthropic-compatible interface, and the model plugs into Claude Code, Codex, Qoder CLI, Qwen Code, and OpenClaw.

None of the headline figures in this release come from a third-party benchmark authority or an audited enterprise deployment. Every number traces back to Qwen’s own test harnesses.

Open weights land next week, which is the point at which outside labs can actually start checking whether the gate counts, the accuracy curves, and the multi-day autonomous runs hold up outside Alibaba’s own sandboxes.

See also: VulnCheck data questions AI vulnerability discovery risk

Banner for AI & Big Data Expo by TechEx events.

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information.

Developer is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

Home
Services
Careers
Call Us
Contact