The document discloses that Gemini 3.1 Pro demonstrated measurable gains in machine learning R&D capabilities relative to Gemini 3 Pro on RE-Bench, including a specific benchmark result showing the model reduced a fine-tuning script runtime to 47 seconds compared to a human reference solution of 94 seconds, while stating average performance remains below alert thresholds.
This analysis describes what Google Gemini's agreement states, permits, or reserves. It does not constitute a legal determination about enforceability. Regulatory applicability and practical outcomes may vary by jurisdiction, enforcement context, and individual circumstances. Read our methodology
This provision discloses specific quantified machine learning R&D capability metrics that are operationally relevant for AI governance assessments, academic research communities, and regulatory evaluations of automation-level AI capabilities. The specific RE-Bench result demonstrating performance exceeding the human reference solution on one challenge is a material disclosure for AI capability assessment purposes.
The document discloses that Gemini 3.1 Pro demonstrated gains in ML R&D task performance, including one benchmark where it outperformed the human reference solution by approximately 50%, while stating overall performance remains below frontier safety alert thresholds. This disclosure is primarily relevant to research organizations, AI developers, and governance teams assessing autonomous AI capabilities.
Cross-platform context
See how other platforms handle Machine Learning R&D Capability Disclosure and similar clauses.
Compare across platforms →"The model shows gains on RE-Bench compared to Gemini 3 Pro, with a human-normalised average score of 1.27 compared to Gemini 3 Pro's score of 1.04. On one particular challenge, Optimise LLM Foundry, it scores double the human-normalised baseline score (reducing the runtime of a fine-tuning script from 300 seconds to 47 seconds, compared to the human reference solution of 94 seconds). However, the model's average performance across all challenges remains beneath the alert threshold for the CCLs.Excerpt from Google Gemini's Gemini 3.1 Pro Model Card
(1) REGULATORY LANDSCAPE: The ML R&D capability disclosure engages EU AI Act provisions on general-purpose AI models with systemic risk, particularly those addressing automation of scientific research and AI self-improvement capabilities.
Enforcement risk, jurisdiction flags, contract triggers, and due diligence action items.
Get the research letter
Companies change their terms quietly. We read every version and catch what actually changed. One email a week on the changes that matter and what they mean.
This provision discloses specific quantified machine learning R&D capability metrics that are operationally relevant for AI governance assessments, academic research communities, and regulatory evaluations of automation-level AI capabilities. The specific RE-Bench result demonstrating performance exceeding the human reference solution on one challenge is a material disclosure for AI capability assessment purposes.
The document discloses that Gemini 3.1 Pro demonstrated gains in ML R&D task performance, including one benchmark where it outperformed the human reference solution by approximately 50%, while stating overall performance remains below frontier safety alert thresholds. This disclosure is primarily relevant to research organizations, AI developers, and governance teams assessing autonomous AI capabilities.
No. ConductAtlas is an independent monitoring service. We are not affiliated with, endorsed by, or sponsored by Google Gemini.