Provision record
Anthropic · Anthropic Responsible Scaling Policy · View original document ↗

ASL Capability Thresholds and Mandatory Safeguard Upgrades

Medium severity High confidence Explicit document language Unique · 0 of 352 platforms
Stay ahead of the changes
Track Anthropic and get the diff the day its terms change.
Share 𝕏 Share in Share 🔒 PDF
Document Record

What it is

The RSP establishes defined capability thresholds (ASL levels) that, when crossed by a model, require Anthropic to implement specified upgraded security and deployment safeguards and to document and publish affirmative risk mitigation cases.

ⓘ

This analysis describes what Anthropic's agreement states, permits, or reserves. It does not constitute a legal determination about enforceability. Regulatory applicability and practical outcomes may vary by jurisdiction, enforcement context, and individual circumstances. Read our methodology

ConductAtlas Analysis

Why it matters (compliance & governance perspective)

This provision establishes the core operational trigger mechanism of the RSP: internally defined capability benchmarks that create mandatory internal obligations for Anthropic to upgrade safeguards and publish risk assessments before continuing deployment of models crossing those thresholds.

Consumer impact (what this means for users)

Under this provision, the deployment conditions for Anthropic's frontier models are governed by internally assessed capability thresholds; when those thresholds are crossed, the terms require Anthropic to implement upgraded safeguards and publish corresponding risk documentation before continued deployment.

Cross-platform context

See how other platforms handle ASL Capability Thresholds and Mandatory Safeguard Upgrades and similar clauses.

Compare across platforms →
▸ View Original Clause Language DOCUMENT RECORD
"
Our RSP requires that once models cross the AI R&D-4 capability threshold, we develop an affirmative case identifying the most immediate and relevant misalignment risks from models pursuing misaligned goals and explaining how we have mitigated them.

Excerpt from Anthropic's Responsible Scaling Policy

ConductAtlas Analysis

Institutional analysis (regulatory & governance intelligence)

1) REGULATORY LANDSCAPE: This provision engages with the EU AI Act, which requires providers of general-purpose AI models with systemic risk to conduct model evaluations and report serious incidents.

Insight

Unlock the full institutional analysis

Enforcement risk, jurisdiction flags, contract triggers, and due diligence action items.

Applicable agencies

  • Federal Trade Commission (ftc)
    Oversees unfair or deceptive business practices and can investigate companies that mislead consumers about data collection, sharing, or use.
    Who can file: Anyone affected by the company's practices (US or international)
    What you need: Your account details, a timeline of relevant events, and a description of the specific issue
    What to expect: Complaints inform FTC enforcement priorities and investigations but do not result in individual resolution or compensation
    File a complaint →

Provision details

Document information
Document
Anthropic Responsible Scaling Policy
Entity
Anthropic
Document last updated
May 12, 2026
Tracking information
First tracked
July 12, 2026
Last verified
July 12, 2026
Record ID
CA-P-074224
Document ID
CA-D-00823
Evidence Provenance
Source URL
Wayback Machine
Content hash (SHA-256)
9ad6c66902eb2b771ccc40a9d0e1c2d4664c197372e6833be51b902e16754d55
Analysis generated
July 12, 2026 14:38 UTC
Methodology
Evidence
✓ Snapshot stored   ✓ Hash verified
Citation Record
Entity: Anthropic
Document: Anthropic Responsible Scaling Policy
Record ID: CA-P-074224
Captured: 2026-07-12 14:38:45 UTC
SHA-256: 9ad6c66902eb2b77…
URL: https://conductatlas.com/platform/anthropic/anthropic-responsible-scaling-policy/provision/CA-P-074224/asl-capability-thresholds-and-mandatory-safeguard-upgrades/
Accessed: Sept. 26, 2026
Permanent archival reference. Stable identifier suitable for legal filings, compliance documentation, and research citation.
Classification
Severity
Medium
Categories

Other risks in this policy

Get the research letter

Companies change their terms quietly. We read every version and catch what actually changed. One email a week on the changes that matter and what they mean.

Frequently Asked Questions

What does Anthropic's ASL Capability Thresholds and Mandatory Safeguard Upgrades clause do?

This provision establishes the core operational trigger mechanism of the RSP: internally defined capability benchmarks that create mandatory internal obligations for Anthropic to upgrade safeguards and publish risk assessments before continuing deployment of models crossing those thresholds.

How does this clause affect you?

Under this provision, the deployment conditions for Anthropic's frontier models are governed by internally assessed capability thresholds; when those thresholds are crossed, the terms require Anthropic to implement upgraded safeguards and publish corresponding risk documentation before continued deployment.

Is ConductAtlas affiliated with Anthropic?

No. ConductAtlas is an independent monitoring service. We are not affiliated with, endorsed by, or sponsored by Anthropic.