Version 3.0 adds agents, robots and sandboxes

On 14 September 2026 China's national cybersecurity standards committee, TC260, released 《人工智能安全治理框架3.0》. Our translation of the Chinese title is "Artificial Intelligence Safety Governance Framework 3.0". The bilingual PDF uses the English title "AI Safety Governance Framework 3.0".

TC260 released the document at the opening ceremony of the 2026 National Cybersecurity Awareness Week. The Cyberspace Administration of China (CAC) posted the announcement at 19:40 Beijing time that evening. Beijing time and SGT are the same.

The CAC notice says TC260 prepared the framework under CAC guidance. The China Academy of Cyberspace Studies and a CAC data and technology support centre took part, along with research institutes and companies.

The notice claims three kinds of change. The framework updates its risk categories. It adjusts the technical countermeasures. It revises the comprehensive governance measures. It keeps the logic of the earlier versions: "risk classification, technological countermeasures, and comprehensive governance".

The concrete novelty sits in four places:

  • Section 2.2 now holds two new risk groups. Section 2.2.1 covers agentic AI risks. Section 2.2.2 covers embodied AI risks.
  • Principle 1.5 is "Ensuring trustworthy application and preventing loss of control". It names "loss of control over the behavior of agentic AI" as a prominent risk.
  • Section 4.3 proposes a regulatory sandbox for AI, with a possible exemption from liability.
  • Appendix 2 is a full risk management framework for agents. Its Chinese title is 智能体风险管理框架, which we translate as "agent risk management framework". The official English text calls it the "Agentic AI risk management framework".

The preface also names recursive self-improvement as a direction of technical change.

A risk catalogue with 33 named agent risks

The framework measures nothing. It reports no incident counts, no test results and no survey data. Its evidence is a structured catalogue of risks and controls.

Appendix 2 maps agent risks across nine lifecycle stages. The appendix figure shows the stages from design through deployment, input, reasoning, tool use, memory, output and decommissioning.

Agentic AI lifecycle diagram from China's AI Safety Governance Framework 3.0, from design through deployment, users, model, memory, tools and decommissioning Caption: "Figure. Agentic AI lifecycle", Appendix 2 · Source: TC260, AI Safety Governance Framework 3.0, Sep 2026 · link

The appendix names 33 risks. Some of the most specific ones follow:

  • Prompt injection, including instructions hidden in documents, emails, web pages, logs and messages.
  • Context overflow attacks, where long inputs push built-in safety instructions out of the context window.
  • Goal hijacking through malicious instructions or poisoned external data.
  • Tool poisoning, tool hijacking and tool-selection bias from manipulative tool descriptions.
  • Identity spoofing through hijacked credentials or replay attacks.
  • Resource overload from unlimited tool-invocation loops.
  • Memory pollution, memory distortion and memory theft.
  • Compromised artifacts, such as installation files, container images, tools and skills with backdoors.
  • Residual permissions and credentials after an agent is decommissioned.
  • Escape from security constraints and improper skill management.

The appendix then lists seven groups of controls. These cover pre-deployment checks, identity and access, human approval, supply chain and tools, runtime management, monitoring and auditing, and decommissioning.

Appendix 1 sets the grading method for all AI risks. It uses three elements: application scenario, level of intelligence and application scale. It defines five levels: low, moderate, considerable, major and extremely serious. The top level means "a catastrophic and systemic threat" with irreversible impacts on national security, social stability and citizens' rights.

The baseline is China's own 2.0 text

The framework compares itself only with its own earlier versions. The English text cites no foreign framework or standard by name. Readers cannot see how its agent risk list maps to other catalogues without doing the mapping themselves.

Against version 2.0, the comparison is fair and easy to check. Both texts sit on CAC pages, and both PDFs are public.

The five-level grading in Appendix 1 has descriptive labels only. The text gives no thresholds, scores or worked examples that place a real system at a level. Section 4.4.1 asks sector regulators to write industry-specific standards with reference to national standards. The grading therefore depends on documents that the framework itself leaves to others.

No thresholds, test suites or enforcement path

The framework gives no evidence that its controls reduce harm. It names no pilot, no deployment and no audit result.

Section 4.2.1 calls for an integrated safety assessment system. It says benchmarks and procedures for assessing model cybersecurity capabilities "should be studied, developed and continuously updated". The framework publishes none of those benchmarks.

The sandbox section sets goals without rules. Section 4.3.1 calls for separate admission criteria for finance, education, broadcasting and television, and healthcare. Section 4.3.2 gives preference to start-ups, small and medium-sized enterprises and public service projects. Section 4.3.3 says a liability exemption "should be explored" where there is no subjective fault and the risks stay controllable. It excludes illegal acts that endanger national security, infringe personal rights, cause major harm or are intentional. The document names no regulator that will run the sandbox.

The framework carries no binding force by itself. Appendix 2 describes its measures as "a reference for developers, providers and users of agent applications". Binding duties would need to come from laws, regulations or mandatory national standards.

The agent appendix also leaves open how its human approval rules should work at scale. It asks for approval before medium- and low-risk operations. It gives no estimate of the burden on users and no method for tiering real tasks.

One bilingual PDF and no machine-readable controls

The public artifact is the 136-page PDF. It carries the Chinese text first and the official English translation after it.

TC260 publishes no machine-readable control list, no checklist file and no test harness. The risk tables in the appendices exist only as PDF pages. Teams that want to track the 33 agent risks must transcribe them.

The document names no standard numbers for the grading work. Appendix 1 says national standards for classifying and grading AI application security should be developed.

Agent teams can map seven control groups now

Teams that ship agents in China can use Appendix 2 as a design checklist today. It also works as a vocabulary for talks with Chinese regulators and customers.

The controls are concrete enough to test against a product:

  1. Give each agent a unique identity. Appendix 2 prohibits identity sharing among agent instances.
  2. Grant only the minimum privileges for the current task. Revoke credentials as soon as a task ends or the agent stops.
  3. Keep a high-risk operation inventory. The agent must hand control to the user before a high-risk operation. It must get user authorization before a medium- or low-risk operation.
  4. Deny by default when the approval system fails, the user stays silent or no approval rule applies.
  5. Require secondary confirmation or human approval for file deletion, data transmission and system configuration changes. Back them with rollback.
  6. Store approval records in tamper-proof, verifiable form.
  7. Verify each tool's version, description, parameters and metadata before the agent calls it. Prefer verified skills from trusted sources.
  8. Run code execution and tool calls in sandboxes or containers.
  9. Keep credentials and secret keys out of agent memory in principle. Isolate memory across users and tasks.
  10. Log file operations, commands, network connections, skill calls and payments. Run regular red teaming.
  11. On decommissioning, stop every process, verify ports and connections, and revoke third-party authorizations.

One control affects cross-border products directly. Appendix 2 says data collected or generated within a territory "must be stored within the territory". Cross-border transfers must follow national data security rules.

Teams outside China can still use the risk list as a threat model. Its entries on context overflow and tool-selection bias are specific enough to turn into test cases.

From a 2024 outline to an agent annex

TC260 published version 1.0 in 2024 and version 2.0 on 15 September 2025. The CAC notice for 3.0 confirms both earlier releases.

The 2.0 PDF was organized with the National Computer Network Emergency Response Technical Team (CNCERT). Its appendices covered risk grading principles, principles for trustworthy AI and terminology. Version 2.0 had no agent annex and no regulatory sandbox section. It mentioned agents only briefly, in its discussion of cyberspace exposure.

Version 3.0 keeps the grading and trustworthy AI appendices. It adds the agent annex as Appendix 2. It adds the agentic and embodied risk groups and the sandbox section. The core logic of risk classification, technical countermeasures and governance stays the same.

The agent annex ends with a promise of further updates. It says the framework "will undergo continuous updates and refinement" as agent capabilities change.