Alibaba says 33 automated cycles lifted Qwen3.8-Max
Alibaba made the claim in its 22 September 2026 press release. The release says Alibaba "has revealed progress in RSI (Recursive Self-Improvement) driven by empirical feedback." It then gives the central claim:
"Over a month of fully automated runs - spanning pipeline design, data validation, iterative experimentation, and error diagnosis - Qwen3.8-Max completed 33 iterative cycles. Through autonomous training optimisation and post-training techniques, the updated Qwen3.8-Max boosted its Artificial Analysis score from 40 to 45."
The claim has two parts. The first part is a score change on a public leaderboard. The second part says the model's own automated work produced that change. We checked each part separately.
A keynote release at the Apsara Conference
Alibaba published the release from Hangzhou during its annual Apsara Conference. The same release made three other vendor claims, which we label here as Alibaba's statements only.
- Alibaba says a chip design experiment ran for over 60 hours and made more than 10,000 EDA tool calls. It says the result "reduced chip area by 42% with zero compromise in performance."
- Alibaba says Qwen 4 is in training. It says Qwen 4.5 and Qwen 5 are "projected to scale up to 5 to 10 trillion parameters."
- Alibaba's chip unit T-Head says the Zhenwu V900 delivers three times the performance of the Zhenwu M890. T-Head lists 216 GB of GPU memory, 1,200 GB per second of inter-chip bandwidth, and FP8 and FP4 support. It plans mass production and commercial release in Q1 2027.
Alibaba published no paper, benchmark report or design files for any of these claims. We found no independent test of the V900.
The leaderboard confirms the jump from 40 to 45
The score change checks out. The Artificial Analysis leaderboard lists two Qwen3.8 Max builds as of 28 September.
| Artificial Analysis field | Qwen3.8 Max (3 August build) | Qwen3.8 Max (0902 build) |
|---|---|---|
| Intelligence Index | 40.2 | 45.4 |
| Terminal-Bench 2.1 | 81.3% | 88.8% |
| Terminal-Bench 4.0 | 18.7% | 38.9% |
| GDPval (normalised) | 0.548 | 0.584 |
| Omniscience | 3.4 | 12.0 |
| Humanity's Last Exam | 43.0% | 43.1% |
| GPQA Diamond | 92.7% | 92.8% |
| SciCode | 53.2% | 52.1% |
| CritPt | 20.0% | 17.7% |
| Cost to run the index, per task | $2.67 | $5.41 |
Artificial Analysis marks neither index score as estimated. It lists the older build as deprecated. It lists the 0902 build as closed weights, priced at $2 per million input tokens and $6 per million output tokens.
The gain is uneven. Terminal-Bench 4.0 more than doubled, and Terminal-Bench 2.1 rose by 7.5 points. Science and reasoning scores barely moved or fell slightly. The newer build also cost about twice as much to run through the index. Its per-token prices stayed the same, so it used more tokens.
The index changed recently. Artificial Analysis says version 4.3 added AutomationBench-AA, removed the tau-cubed Banking test and moved to Terminal-Bench 4.0. The largest single gain sits on a benchmark the index now counts.
The dates also line up. A Model Studio notice says the qwen3.8-max endpoint switched to the qwen3.8-max-0902 snapshot on 5 September 2026, Beijing time. It describes gains in coding, agentic collaboration and visual understanding. The notice does not mention self-improvement.
Five facts Alibaba has not published
For the second part of the claim to hold, several things would have to be true. Alibaba has published evidence for none of them.
- The 0902 snapshot would have to come from the automated loop. Alibaba does not say which training steps the loop controlled.
- People would have to stay out of the loop. The release does not say who chose data, set goals, picked checkpoints or approved the final release.
- A "cycle" would need a definition. The release does not say what one cycle contains or how much compute it used.
- The loop would have to avoid tuning toward the leaderboard itself. The release frames the goal as the Artificial Analysis score. If the loop used similar benchmarks as feedback, the score would measure the target and overstate general gains.
- The chip result would need a baseline. The release does not name the original design or the performance tests behind "zero compromise."
The month-long window fits the release dates. Artificial Analysis dates the first build to 3 August and the second to 2 September. Timing alone shows nothing about who or what made the changes.
Verdict: the score is supported and the method is too early to judge
Verdict on the score: supported. Artificial Analysis lists Qwen3.8-Max rising from 40.2 to 45.4, and neither score is estimated.
Verdict on the self-improvement attribution: too early. Alibaba has released no paper, logs, cycle definitions or human-oversight details. Outsiders cannot yet separate the loop's contribution from normal post-training by Alibaba staff.
Verdict on the chip design and Zhenwu V900 claims: unsupported so far. Both rest on Alibaba's statements alone.
Attribute the loop to Alibaba and pin the 0902 snapshot
We will write: "Artificial Analysis scores the September build of Qwen3.8-Max at 45, up from 40. Alibaba says an automated training loop produced the update. It has published no method."
We will avoid "Qwen improved itself" and "recursive self-improvement" as plain statements of fact. We will attribute both phrases to Alibaba.
Teams that call qwen3.8-max without a pinned snapshot already run the 0902 build. They can pin a snapshot name and rerun their own evaluations. They should also budget for more output tokens, since the index cost per task doubled.
