Japan AISI publishes two ExploitBench notes
Japan's AI Safety Institute (J-AISI), part of the Information-technology Promotion Agency (IPA), published two evaluation notes on 2 October 2026. Both use ExploitBench, a capability ladder benchmark that J-AISI says Carnegie Mellon University researchers developed.
- Cyber Capability Evaluation of Claude Opus 4.8 Using ExploitBench compares Anthropic's Opus 4.8 with Opus 4.7. J-AISI ran it in June 2026.
- The Proliferation of Advanced AI Cyber Capabilities to Open Models compares Z.ai's open-weight GLM-5.2 with both Opus models. J-AISI ran it in July 2026.
Each note appears in Japanese and English on the same page.
IPA's safety institute ran the tests
J-AISI says it tested 41 known vulnerabilities in V8, the JavaScript engine in Google Chrome. Each model got the V8 source, a build environment, a debugger and other tools, then worked through tool calls on its own.
ExploitBench scores 16 capability items across five tiers. T5 is identifying the vulnerable code. T1, the most severe, covers program counter control or arbitrary operations outside the system.
In the main comparisons, J-AISI ran each task once, with seed 1, up to 300 turns and up to five hours. The scores come from single runs, so they are not averages.
For the Opus 4.8 note, J-AISI says it disabled the real-time cyber safeguards Anthropic applies in normal deployment. It says the results do not show behavior under generally available conditions.
GLM-5.2 trails both Opus models
On all 41 tasks, J-AISI reports these average scores out of 16:
| Model | Average score | Highest | Lowest |
|---|---|---|---|
| Claude Opus 4.8 | 5.22 | 10 | 2 |
| Claude Opus 4.7 | 3.63 | 8 | 0 |
| GLM-5.2 | 2.80 | 8 | 0 |
The open-models note says GLM-5.2 topped out at T3, while Opus 4.8 reached T2 and T1 on some tasks. Opus 4.7 hit its maximum of 8 on five tasks, against two for GLM-5.2.
J-AISI found GLM-5.2 cheaper per result. It compared the 36 tasks where all three models reached T5. Token cost to that point was $7.34 for Opus 4.8, $4.60 for Opus 4.7 and $3.71 for GLM-5.2.
J-AISI also tested GLM-5.2 with nudges on 10 tasks, five seeds each. A nudge is advice the system sends to keep an agent from stalling or giving up. Its average stayed at 3.0 with or without them.
J-AISI says GLM-5.2 made no safety refusals in its environment. The Opus models did refuse at times, which it says suggests different default safeguards.
J-AISI could compare Opus 4.8 with Opus 4.7 on 40 tasks. Opus 4.8 reached a higher tier on 15, the same tier on 23 and a lower tier on two.
One Opus 4.8 result did not reproduce
On one task, crbug-1509576, Opus 4.8 gained program counter control, one of the two items needed for T1. It did not gain the other, and J-AISI says the model concluded it could not escape the sandbox.
A retry on the same task did not reproduce the result.
The two notes come from separate runs in June and July. The June note highlights one task where Opus 4.8 stopped short of T1. The July note says Opus 4.8 reached T1 on some tasks. J-AISI does not reconcile the two. J-AISI says scores and tiers alone should not be used to judge capability, and that notable runs need log review and repeat trials.
IPA's disclaimer says the notes are not a government view, certification or endorsement of any model.
The GLM-5.2 note covers an older model than the GLM-5.3 release Anthropic assessed on 29 September. See Anthropic says Z.ai GLM-5.3 brings Mythos-class cyber skills with weak open-weight safeguards.
Security teams weighing open-weight risk
J-AISI's summary says the data do not support a definitive conclusion that advanced cyber capabilities have broadly spread to open-weight models. It adds that open-weight cyber skills could improve sharply over short periods and are hard to restrict after release.
- Security leads can use the notes as a second government data point next to the UK's recent work. See UK AISI finds GPT-6 Astra completed unsanctioned supply-chain attacks in 29.2% of simulated cyber runs.
- Teams running Opus 4.8 can compare J-AISI's findings with Anthropic's own claims in our Opus 4.8 launch coverage.
- Evaluators can take J-AISI's method point directly: rerun standout results before reporting them.
