Get the weekly research letter
Companies change their terms quietly. We read every version and catch what actually changed. One email a week on the changes that matter and what they mean. No account.
The document states that GPT-5.5 Pro safety evaluations are generally based on GPT-5.5 evaluation results used as proxies, with separate evaluation conducted only where OpenAI judges that the parallel test-time compute setting could materially affect risk or safeguard posture.
This analysis describes what OpenAI's agreement states, permits, or reserves. It does not constitute a legal determination about enforceability. Regulatory applicability and practical outcomes may vary by jurisdiction, enforcement context, and individual circumstances. Read our methodology
This provision establishes that GPT-5.5 Pro, a distinct product configuration, relies primarily on safety evaluations conducted for a different operational setting. Compliance teams assessing AI model governance should evaluate whether this proxy methodology meets the independent evaluation expectations of applicable frameworks, particularly for higher-compute configurations that may exhibit different capability profiles.
Interpretive note: The document does not specify which capability categories or risk domains triggered separate GPT-5.5 Pro evaluation, limiting assessment of the proxy methodology's adequacy for specific deployment contexts.
Under these terms, users of GPT-5.5 Pro operate under a safety posture that is primarily informed by evaluations conducted on GPT-5.5 rather than on GPT-5.5 Pro directly, except in cases where OpenAI determines the compute setting materially affects risk. The document does not specify which capability categories triggered separate GPT-5.5 Pro evaluation.
Cross-platform context
See how other platforms handle GPT-5.5 Pro Proxy Safety Evaluation Methodology and similar clauses.
Compare across platforms →Monitoring
OpenAI has changed this document before.
Receive same-day alerts, structured change summaries, and monitoring for up to 25 platforms.
"We generally treat GPT‑5.5's safety results as strong proxies for GPT‑5.5 Pro, which is the same underlying model using a setting that makes use of parallel test time compute. As noted below, we separately evaluate GPT‑5.5 Pro in certain cases because we judge that the setting could materially impact the relevant risks or appropriate safeguards posture.Excerpt from OpenAI's GPT-5.5 System Card
(1) REGULATORY LANDSCAPE: The proxy evaluation methodology engages with EU AI Act conformity assessment requirements, which may require that safety evaluations reflect the actual operational configuration of the deployed system. The UK AI Safety Institute's model evaluation protocols similarly focus on the specific deployed configuration. Where GPT-5.5 Pro exhibits materially different capability profiles due to parallel test-time compute, proxy evaluations may not satisfy these requirements. (2) GOVERNANCE EXPOSURE: Medium to High. The document acknowledges that the parallel test-time compute setting could materially impact relevant risks, yet the default evaluation posture relies on proxy results. This creates a documented gap between the evaluated configuration and the deployed configuration that compliance and governance teams should formally assess. (3) JURISDICTION FLAGS: EU and UK jurisdictions present the highest exposure given mandatory conformity assessment and incident reporting obligations for high-risk AI systems. The proxy methodology may face scrutiny under any regulatory framework requiring that safety assessments reflect the specific system being deployed. (4) CONTRACT AND VENDOR IMPLICATIONS: Enterprise customers deploying GPT-5.5 Pro via API should assess whether vendor representations about model safety are adequately grounded in direct evaluation of the GPT-5.5 Pro configuration. Procurement contracts that reference OpenAI safety assurances may warrant amendment to specify which evaluation basis those assurances reflect. (5) COMPLIANCE CONSIDERATIONS: Legal teams should request clarification from OpenAI on which capability categories triggered separate GPT-5.5 Pro evaluation, and document that inquiry as part of AI vendor due diligence. Organizations subject to AI risk management frameworks such as NIST AI RMF should assess whether the proxy methodology is consistent with their internal model governance requirements.
This provision establishes that GPT-5.5 Pro, a distinct product configuration, relies primarily on safety evaluations conducted for a different operational setting. Compliance teams assessing AI model governance should evaluate whether this proxy methodology meets the independent evaluation expectations of applicable frameworks, particularly for higher-compute configurations that may exhibit different capability profiles.
Under these terms, users of GPT-5.5 Pro operate under a safety posture that is primarily informed by evaluations conducted on GPT-5.5 rather than on GPT-5.5 Pro directly, except in cases where OpenAI determines the compute setting materially affects risk. The document does not specify which capability categories triggered separate GPT-5.5 Pro evaluation.
No. ConductAtlas is an independent monitoring service. We are not affiliated with, endorsed by, or sponsored by OpenAI.