Get the weekly research letter
Companies change their terms quietly. We read every version and catch what actually changed. One email a week on the changes that matter and what they mean. No account.
The document discloses that Command R and Command R+ models may generate toxic text including obscenities, sexually explicit content, and content that stereotypes groups, despite safeguards, particularly in extended multi-turn conversations.
This analysis describes what Cohere's agreement states, permits, or reserves. It does not constitute a legal determination about enforceability. Regulatory applicability and practical outcomes may vary by jurisdiction, enforcement context, and individual circumstances. Read our methodology
This disclosure establishes that Cohere's safeguards do not eliminate the risk of toxic output generation, which has direct implications for customers deploying Command R in consumer-facing applications, particularly those accessible to minors or in regulated content environments.
The agreement discloses that Command R models may generate toxic text, sexually explicit content, and content that stereotypes groups of people, even with safeguards in place, with heightened risk in long multi-turn conversations. Customers building consumer-facing applications are responsible for implementing additional content filtering appropriate to their deployment context.
Cross-platform context
See how other platforms handle Toxic Degeneration Disclosure and similar clauses.
Compare across platforms →Monitoring
Cohere has changed this document before.
Receive same-day alerts, structured change summaries, and monitoring for up to 20 platforms.
"Models have been trained on a wide variety of text from many sources that contain toxic content (see Luccioni and Viviano, 2021). As a result, models may generate toxic text. This may include obscenities, sexually explicit content, and messages which mischaracterize or stereotype groups of people based on problematic historical biases perpetuated by internet communities (see Gehman et al., 2020 for more about toxic language model degeneration). We have put safeguards in place to avoid generating harmful text, and while they are effective (see the 'Safety Benchmarks' section above), it is still possible to encounter toxicity, especially over long conversations with multiple turns.Excerpt from Cohere's Responsible Use Policy
1. REGULATORY LANDSCAPE: For platforms accessible to minors, this disclosure engages COPPA and the FTC's enforcement authority over child-directed platforms. In the EU, the Digital Services Act requires platforms to assess and mitigate systemic risks including harmful content generation. The FTC Act's prohibition on unfair practices may apply if toxic content generation causes consumer harm in the absence of adequate downstream safeguards by the deploying customer. 2. GOVERNANCE EXPOSURE: Medium to High depending on deployment context. For consumer-facing applications, particularly those accessible to minors or in health, education, or financial services contexts, the disclosed residual risk of toxic output generation creates product liability and regulatory exposure that customers must address through their own content moderation layers. 3. JURISDICTION FLAGS: EU Digital Services Act obligations apply to platform providers distributing AI-generated content at scale. California's Age-Appropriate Design Code may impose additional obligations for platforms accessible to minors. The UK Online Safety Act imposes content safety duties on user-to-user and search services, which may engage where AI-generated content is distributed to users. 4. CONTRACT AND VENDOR IMPLICATIONS: Customers should assess whether their downstream terms of service and content policies account for the possibility of toxic AI-generated outputs and whether their content moderation architecture is sufficient to meet applicable platform obligations. Vendor risk assessments should document the residual risk disclosed in this model card. 5. COMPLIANCE CONSIDERATIONS: Compliance teams building consumer-facing applications on Command R should implement output filtering, content moderation review, and incident response procedures that address the residual toxic generation risk disclosed in this document. For applications accessible to minors, a COPPA compliance review and age-appropriate design assessment are warranted.
This disclosure establishes that Cohere's safeguards do not eliminate the risk of toxic output generation, which has direct implications for customers deploying Command R in consumer-facing applications, particularly those accessible to minors or in regulated content environments.
The agreement discloses that Command R models may generate toxic text, sexually explicit content, and content that stereotypes groups of people, even with safeguards in place, with heightened risk in long multi-turn conversations. Customers building consumer-facing applications are responsible for implementing additional content filtering appropriate to their deployment context.
No. ConductAtlas is an independent monitoring service. We are not affiliated with, endorsed by, or sponsored by Cohere.