Model publishers are encouraged to disclose what datasets were used to train their model, which helps users assess potential biases, data provenance issues, and licensing implications of the training data.
This analysis describes what Hugging Face's agreement states, permits, or reserves. It does not constitute a legal determination about enforceability. Regulatory applicability and practical outcomes may vary by jurisdiction, enforcement context, and individual circumstances. Read our methodology
Training data disclosure is directly relevant to intellectual property compliance, data provenance assessments, and bias risk evaluation, particularly as regulatory frameworks increasingly require transparency about AI training data sources.
Interpretive note: Training data disclosure is described as a recommendation rather than a mandatory field, so completeness and accuracy depend on individual model publisher behavior and cannot be assumed.
The training data section of a model card, when completed by the publisher, provides users with information about where the model's capabilities come from, including whether training data may have included copyrighted material, personal data, or datasets with known demographic biases.
How other platforms handle this
to request that your data be transferred to a third party (data portability)
Your organization may allow you to access and export your data in order to back it up or transfer it to a service outside of Google.
Further, you may take legal actions in relation to any potential breach of your rights regarding the processing of your Personal Information, as well as to lodge complaints before the competent data prot...
"Model cards should include information about the datasets used to train the model. This information helps users understand the potential biases and limitations of the model.Excerpt from Hugging Face's Model Card Guidelines
(1) REGULATORY LANDSCAPE: Training data disclosure engages the EU AI Act's requirements for data governance and transparency for AI systems.
Enforcement risk, jurisdiction flags, contract triggers, and due diligence action items.
Ad personalization controls removed. Contact scanning added. Advertiser data partnerships quietly dropped. A timeline of every change.
Get the research letter
Companies change their terms quietly. We read every version and catch what actually changed. One email a week on the changes that matter and what they mean.
Training data disclosure is directly relevant to intellectual property compliance, data provenance assessments, and bias risk evaluation, particularly as regulatory frameworks increasingly require transparency about AI training data sources.
The training data section of a model card, when completed by the publisher, provides users with information about where the model's capabilities come from, including whether training data may have included copyrighted material, personal data, or datasets with known demographic biases.
ConductAtlas has identified this type of provision across 290 platforms. See the full comparison.
No. ConductAtlas is an independent monitoring service. We are not affiliated with, endorsed by, or sponsored by Hugging Face.