Provision record
Anthropic · Claude Sonnet 5 System Card · View original document ↗

Multi-Agent Trust and Prompt Injection Risk

High severity High confidence Explicit document language Unique · 0 of 352 platforms
Stay ahead of the changes
Track Anthropic and get the diff the day its terms change.
Share 𝕏 Share in Share 🔒 PDF
Document Record

What it is

The document discloses that in multi-agent deployments where Claude acts as a subagent receiving instructions from an orchestrating system, it cannot verify the identity or integrity of the orchestrator and must apply its safety standards regardless, and that Claude should be vigilant about prompt injection attacks from external content.

ⓘ

This analysis describes what Anthropic's agreement states, permits, or reserves. It does not constitute a legal determination about enforceability. Regulatory applicability and practical outcomes may vary by jurisdiction, enforcement context, and individual circumstances. Read our methodology

ConductAtlas Analysis

Why it matters (compliance & governance perspective)

This provision discloses a structural trust limitation in multi-agent architectures that is operationally significant for enterprises building complex AI pipelines, as it means Claude will not automatically trust instructions received through automated orchestration channels and may refuse or pause tasks based on its own safety assessment.

Consumer impact (what this means for users)

Under this provision, Claude Sonnet 4.5 deployed as a subagent in multi-agent systems is designed to apply independent safety judgment to orchestrator instructions, which may affect the reliability and predictability of automated pipeline behavior in enterprise deployments.

Cross-platform context

See how other platforms handle Multi-Agent Trust and Prompt Injection Risk and similar clauses.

Compare across platforms →
▸ View Original Clause Language DOCUMENT RECORD
"
When Claude operates as an agent being orchestrated by an orchestrator, it should behave safely and ethically regardless of the instruction source, since it has no way to verify that it is talking with Claude or that the Claude model it's talking with has not been compromised.

Excerpt from Anthropic's Claude Sonnet 5 System Card

ConductAtlas Analysis

Institutional analysis (regulatory & governance intelligence)

1.

Insight

Unlock the full institutional analysis

Enforcement risk, jurisdiction flags, contract triggers, and due diligence action items.

Applicable agencies

  • Federal Trade Commission (ftc)
    Oversees unfair or deceptive business practices and can investigate companies that mislead consumers about data collection, sharing, or use.
    Who can file: Anyone affected by the company's practices (US or international)
    What you need: Your account details, a timeline of relevant events, and a description of the specific issue
    What to expect: Complaints inform FTC enforcement priorities and investigations but do not result in individual resolution or compensation
    File a complaint →

Provision details

Document information
Document
Claude Sonnet 5 System Card
Entity
Anthropic
Document last updated
July 6, 2026
Tracking information
First tracked
July 6, 2026
Last verified
July 6, 2026
Record ID
CA-P-013377
Document ID
CA-D-00921
Evidence Provenance
Source URL
Wayback Machine
Content hash (SHA-256)
cdf033accb28f24cdba717a62a97f374bbf1637875644a05d1e43ef6c3f7fa1a
Analysis generated
July 6, 2026 21:57 UTC
Methodology
Evidence
✓ Snapshot stored   ✓ Hash verified
Citation Record
Entity: Anthropic
Document: Claude Sonnet 5 System Card
Record ID: CA-P-013377
Captured: 2026-07-06 21:57:58 UTC
SHA-256: cdf033accb28f24c…
URL: https://conductatlas.com/platform/anthropic/claude-sonnet-5-system-card/provision/CA-P-013377/multi-agent-trust-and-prompt-injection-risk/
Accessed: Oct. 2, 2026
Permanent archival reference. Stable identifier suitable for legal filings, compliance documentation, and research citation.
Classification
Severity
High
Categories

Other risks in this policy

Get the research letter

Companies change their terms quietly. We read every version and catch what actually changed. One email a week on the changes that matter and what they mean.

Frequently Asked Questions

What does Anthropic's Multi-Agent Trust and Prompt Injection Risk clause do?

This provision discloses a structural trust limitation in multi-agent architectures that is operationally significant for enterprises building complex AI pipelines, as it means Claude will not automatically trust instructions received through automated orchestration channels and may refuse or pause tasks based on its own safety assessment.

How does this clause affect you?

Under this provision, Claude Sonnet 4.5 deployed as a subagent in multi-agent systems is designed to apply independent safety judgment to orchestrator instructions, which may affect the reliability and predictability of automated pipeline behavior in enterprise deployments.

Is ConductAtlas affiliated with Anthropic?

No. ConductAtlas is an independent monitoring service. We are not affiliated with, endorsed by, or sponsored by Anthropic.