Provision record
Hugging Face · Hugging Face Model Card Guidelines · View original document ↗

Evaluation Results Structured Reporting

Low severity Medium confidence Explicit document language Common · 273 of 352 platforms
Stay ahead of the changes
Track Hugging Face and get the diff the day its terms change.
Share 𝕏 Share in Share 🔒 PDF
Document Record

What it is

The model card metadata schema includes a structured evaluation results section that allows model publishers to report benchmark performance metrics linked to specific tasks, datasets, and configuration parameters. These structured results are parsed by the Hub and used to populate model comparison and leaderboard features.

This analysis describes what Hugging Face's agreement states, permits, or reserves. It does not constitute a legal determination about enforceability. Regulatory applicability and practical outcomes may vary by jurisdiction, enforcement context, and individual circumstances. Read our methodology

ConductAtlas Analysis

Why it matters (compliance & governance perspective)

This provision establishes the structured format through which model performance claims are disclosed and indexed on the Hub, making the accuracy and completeness of evaluation result fields relevant to how users and automated systems assess and compare model capabilities.

Interpretive note: The document describes the evaluation results schema but does not specify verification standards or accuracy obligations for reported metric values.

Clause Stability Stable

0
Changes
3
Months Monitored
May 12, 2026
First Seen
May 22, 2026
Last Seen
This clause type exists across 1424 other provisions on other platforms.

Change history

modified May 21, 2026

Severity downgraded from medium to low, and guidance shifted from general description to specific YAML field structure (model-index) with detailed subfield requirements.

View full change record →

Consumer impact (what this means for users)

Under this framework, structured evaluation results in model card metadata are surfaced in Hub search and comparison features, meaning users may rely on these fields when selecting models for specific tasks. The document does not state that Hugging Face independently verifies or audits the accuracy of reported evaluation metrics.

How other platforms handle this

NVIDIA NIM Medium

Customer must report upon NVIDIA's email request, no more than monthly, the Software in use by Customer Personnel and Customer End Users Customer enabled, quantity, start and end dates, and any other reasonably required information...

Weights & Biases Medium

If Customer disables the usage tracker within the Software or Service, Customer will, no later than the end of each calendar quarter...provide W&B with information reasonably requested...to verify compliance...

Venmo Medium

we are required to report your business name and the name of your beneficial owners and/or principals to: a) the MATCH listing maintained by Mastercard...and b) the VMAS database upheld by Visa

See all platforms with this clause type →
▸ View Original Clause Language DOCUMENT RECORD
"
model-index contains results which is a list of evaluation results. Each result includes: task, dataset, and metrics fields. The metrics field contains a list of metric results. Each metric result includes: type, value, name, and config fields.

Excerpt from Hugging Face's Model Card Guidelines

ConductAtlas Analysis

Institutional analysis (regulatory & governance intelligence)

(1) REGULATORY LANDSCAPE: Accuracy of evaluation result claims in model card metadata may engage FTC guidance on truthful representation of AI system performance, particularly where metric values are used in commercial contexts to represent model …

Insight

Unlock the full institutional analysis

Enforcement risk, jurisdiction flags, contract triggers, and due diligence action items.

Applicable agencies

  • Federal Trade Commission (ftc)
    Oversees unfair or deceptive business practices and can investigate companies that mislead consumers about data collection, sharing, or use.
    Who can file: Anyone affected by the company's practices (US or international)
    What you need: Your account details, a timeline of relevant events, and a description of the specific issue
    What to expect: Complaints inform FTC enforcement priorities and investigations but do not result in individual resolution or compensation
    File a complaint →

Provision details

Document information
Document
Hugging Face Model Card Guidelines
Entity
Hugging Face
Document last updated
May 12, 2026
Tracking information
First tracked
May 21, 2026
Last verified
May 21, 2026
Record ID
CA-P-012040
Document ID
CA-D-00842
Evidence Provenance
Source URL
Wayback Machine
Content hash (SHA-256)
66b6b488c95d3920fe9e1acec75ede720f6f4f4162de5fd0577053fc630bdcb3
Analysis generated
May 21, 2026 05:03 UTC
Methodology
Evidence
✓ Snapshot stored   ✓ Hash verified
Citation Record
Entity: Hugging Face
Document: Hugging Face Model Card Guidelines
Record ID: CA-P-012040
Captured: 2026-05-21 05:03:11 UTC
SHA-256: 66b6b488c95d3920…
URL: https://conductatlas.com/platform/hugging-face/hugging-face-model-card-guidelines/provision/CA-P-012040/evaluation-results-structured-reporting/
Accessed: Sept. 8, 2026
Permanent archival reference. Stable identifier suitable for legal filings, compliance documentation, and research citation.
Classification
Severity
Low
Categories

Other risks in this policy

Get the research letter

Companies change their terms quietly. We read every version and catch what actually changed. One email a week on the changes that matter and what they mean.

Frequently Asked Questions

What does Hugging Face's Evaluation Results Structured Reporting clause do?

This provision establishes the structured format through which model performance claims are disclosed and indexed on the Hub, making the accuracy and completeness of evaluation result fields relevant to how users and automated systems assess and compare model capabilities.

How does this clause affect you?

Under this framework, structured evaluation results in model card metadata are surfaced in Hub search and comparison features, meaning users may rely on these fields when selecting models for specific tasks. The document does not state that Hugging Face independently verifies or audits the accuracy of reported evaluation metrics.

How many platforms have this type of clause?

ConductAtlas has identified this type of provision across 273 platforms. See the full comparison.

Is ConductAtlas affiliated with Hugging Face?

No. ConductAtlas is an independent monitoring service. We are not affiliated with, endorsed by, or sponsored by Hugging Face.