Provision record
Hugging Face · Hugging Face Model Card Guidelines · View original document ↗

Training Data Attribution Fields

Medium severity Medium confidence Explicit document language Common · 273 of 352 platforms
Stay ahead of the changes
Track Hugging Face and get the diff the day its terms change.
Share 𝕏 Share in Share 🔒 PDF
Document Record

What it is

The model card YAML metadata includes a datasets field for attributing training datasets used to develop the model, with Hub-hosted datasets linked to their dataset pages. This field is parsed by the Hub to create dataset-model linkages in the platform's discovery infrastructure.

This analysis describes what Hugging Face's agreement states, permits, or reserves. It does not constitute a legal determination about enforceability. Regulatory applicability and practical outcomes may vary by jurisdiction, enforcement context, and individual circumstances. Read our methodology

ConductAtlas Analysis

Why it matters (compliance & governance perspective)

This provision establishes the mechanism through which training data provenance is disclosed on the Hub, which downstream users, auditors, and regulators may rely upon to assess data sourcing practices, potential bias origins, copyright implications, and compliance with data governance requirements.

Interpretive note: The document does not specify whether training dataset attributions are subject to any verification or accuracy obligation on the part of model publishers, and does not address disclosure obligations where training data sourcing is proprietary or partially undisclosed.

Clause Stability Stable

0
Changes
3
Months Monitored
May 21, 2026
First Seen
May 22, 2026
Last Seen
This clause type exists across 1424 other provisions on other platforms.

Change history

removed Jul 21, 2026

Removal of explicit datasets field guidance represents de-emphasis on granular training data attribution in favor of higher-level ethical considerations.

View full change record →
modified May 21, 2026

Provision was renamed, converted from general guidance to specific YAML field syntax, and expanded to include requirements for listing datasets separately and linking to Hub datasets.

View full change record →

Consumer impact (what this means for users)

Under this framework, the datasets field in model card metadata is the primary mechanism through which training data provenance is disclosed to Hub users. Users assessing models for use in regulated or sensitive applications can reference this field to evaluate data sourcing, though the document does not state that Hugging Face independently verifies training dataset attributions.

How other platforms handle this

Glassdoor Medium

In certain situations, Glassdoor may be required to disclose personal data in response to lawful requests by public authorities, including to meet national security or law enforcement requirements.

Shopify Medium

If you would like to submit a legally binding request to demand someone else's Personal Data (for example, if you have a subpoena or court order), please review our Guidelines for Legal Requests.

Amazon Associates Medium

You will include a date/time stamp adjacent to your display of pricing or availability information on your application if you obtain Product Advertising Content from Data Feeds, or if you call Creators API, PA API or refresh the Product Advertising Content...less frequently than hourly.

See all platforms with this clause type →
▸ View Original Clause Language DOCUMENT RECORD
"
datasets: This field is used to indicate the datasets used to train the model. Each dataset should be listed as a separate item. If the dataset is available on the Hub, it should be linked to the dataset page.

Excerpt from Hugging Face's Model Card Guidelines

ConductAtlas Analysis

Institutional analysis (regulatory & governance intelligence)

(1) REGULATORY LANDSCAPE: Training data attribution disclosures engage GDPR and other data protection regulations where training datasets contain personal data, as well as emerging AI-specific transparency requirements under the EU AI Act regarding training data …

Insight

Unlock the full institutional analysis

Enforcement risk, jurisdiction flags, contract triggers, and due diligence action items.

Applicable agencies

  • Federal Trade Commission (ftc)
    Oversees unfair or deceptive business practices and can investigate companies that mislead consumers about data collection, sharing, or use.
    Who can file: Anyone affected by the company's practices (US or international)
    What you need: Your account details, a timeline of relevant events, and a description of the specific issue
    What to expect: Complaints inform FTC enforcement priorities and investigations but do not result in individual resolution or compensation
    File a complaint →

Provision details

Document information
Document
Hugging Face Model Card Guidelines
Entity
Hugging Face
Document last updated
May 12, 2026
Tracking information
First tracked
May 21, 2026
Last verified
May 21, 2026
Record ID
CA-P-013102
Document ID
CA-D-00842
Evidence Provenance
Source URL
Wayback Machine
Content hash (SHA-256)
66b6b488c95d3920fe9e1acec75ede720f6f4f4162de5fd0577053fc630bdcb3
Analysis generated
May 21, 2026 05:03 UTC
Methodology
Evidence
✓ Snapshot stored   ✓ Hash verified
Citation Record
Entity: Hugging Face
Document: Hugging Face Model Card Guidelines
Record ID: CA-P-013102
Captured: 2026-05-21 05:03:11 UTC
SHA-256: 66b6b488c95d3920…
URL: https://conductatlas.com/platform/hugging-face/hugging-face-model-card-guidelines/provision/CA-P-013102/training-data-attribution-fields/
Accessed: Sept. 8, 2026
Permanent archival reference. Stable identifier suitable for legal filings, compliance documentation, and research citation.
Classification
Severity
Medium
Categories

Other risks in this policy

Get the research letter

Companies change their terms quietly. We read every version and catch what actually changed. One email a week on the changes that matter and what they mean.

Frequently Asked Questions

What does Hugging Face's Training Data Attribution Fields clause do?

This provision establishes the mechanism through which training data provenance is disclosed on the Hub, which downstream users, auditors, and regulators may rely upon to assess data sourcing practices, potential bias origins, copyright implications, and compliance with data governance requirements.

How does this clause affect you?

Under this framework, the datasets field in model card metadata is the primary mechanism through which training data provenance is disclosed to Hub users. Users assessing models for use in regulated or sensitive applications can reference this field to evaluate data sourcing, though the document does not state that Hugging Face independently verifies training dataset attributions.

How many platforms have this type of clause?

ConductAtlas has identified this type of provision across 273 platforms. See the full comparison.

Is ConductAtlas affiliated with Hugging Face?

No. ConductAtlas is an independent monitoring service. We are not affiliated with, endorsed by, or sponsored by Hugging Face.