The document specifies a structured YAML format for embedding evaluation results in model card metadata, which the Hub uses to display benchmark scores and, where applicable, to index results into Papers with Code leaderboards.
This analysis describes what Hugging Face's agreement states, permits, or reserves. It does not constitute a legal determination about enforceability. Regulatory applicability and practical outcomes may vary by jurisdiction, enforcement context, and individual circumstances. Read our methodology
This provision establishes that evaluation results declared in model card metadata may be propagated to external leaderboards, meaning metadata accuracy directly affects how models are represented in third-party benchmark comparisons.
Interpretive note: The extent to which evaluation result metadata accuracy obligations are enforceable depends on commercial context and how results are represented to downstream users.
Changed from technical specification of result field structure to explanation of historical basis and leaderboard integration capability, removing granular field documentation.
View full change record →This provision establishes that structured evaluation results in model card metadata are displayed on the model page and may be indexed into Papers with Code leaderboards, creating a dependency between contributor-provided metadata accuracy and the benchmark representations visible to downstream users.
Cross-platform context
See how other platforms handle Structured Evaluation Results Metadata and Leaderboard Indexing and similar clauses.
Compare across platforms →"The initial metadata spec was based on Papers with code's model-index specification. This allowed us to directly index the results into Papers with code's leaderboards when appropriate. You could also link the source from where the eval results has been computed.Excerpt from Hugging Face's Model Card Guidelines
1) REGULATORY LANDSCAPE: This provision does not directly engage specific regulatory frameworks.
Enforcement risk, jurisdiction flags, contract triggers, and due diligence action items.
Get the research letter
Companies change their terms quietly. We read every version and catch what actually changed. One email a week on the changes that matter and what they mean.
This provision establishes that evaluation results declared in model card metadata may be propagated to external leaderboards, meaning metadata accuracy directly affects how models are represented in third-party benchmark comparisons.
This provision establishes that structured evaluation results in model card metadata are displayed on the model page and may be indexed into Papers with Code leaderboards, creating a dependency between contributor-provided metadata accuracy and the benchmark representations visible to downstream users.
No. ConductAtlas is an independent monitoring service. We are not affiliated with, endorsed by, or sponsored by Hugging Face.