GPT-4o's audio output is restricted to a preset list of approved voices created with voice actors, enforced by a real-time streaming classifier that blocks any output deviating from the approved voice list, with the document stating the system currently catches 100% of meaningful deviations based on internal evaluations.
This analysis describes what OpenAI's agreement states, permits, or reserves. It does not constitute a legal determination about enforceability. Regulatory applicability and practical outcomes may vary by jurisdiction, enforcement context, and individual circumstances. Read our methodology
This provision establishes a technical enforcement mechanism restricting voice output to a defined set of pre-approved voices, directly addressing fraud, impersonation, and voice cloning risks identified during red teaming and relevant to the broader landscape of synthetic media regulation.
Under this provision, GPT-4o will not generate audio output in a user's voice or any voice outside the preset approved list; the real-time classifier blocks such outputs during generation, and the document states that unintentional voice generation instances are also addressed by this secondary classifier.
Cross-platform context
See how other platforms handle Unauthorized Voice Generation Restriction and Classifier and similar clauses.
Compare across platforms →"We addressed voice generation related-risks by allowing only the preset voices we created in collaboration with voice actors to be used. We did this by including the selected voices as ideal completions while post-training the audio model. Additionally, we built a standalone output classifier to detect if the GPT-4o output is using a voice that's different from our approved list. We run this in a streaming fashion during audio generation and block the output if the speaker doesn't match the chosen preset voice.Excerpt from OpenAI's GPT-4o System Card (PDF)
(1) REGULATORY LANDSCAPE: This provision is relevant to emerging synthetic media and voice cloning regulations at the state level in the US, including laws in states such as California and Tennessee addressing unauthorized use of …
Enforcement risk, jurisdiction flags, contract triggers, and due diligence action items.
Get the research letter
Companies change their terms quietly. We read every version and catch what actually changed. One email a week on the changes that matter and what they mean.
This provision establishes a technical enforcement mechanism restricting voice output to a defined set of pre-approved voices, directly addressing fraud, impersonation, and voice cloning risks identified during red teaming and relevant to the broader landscape of synthetic media regulation.
Under this provision, GPT-4o will not generate audio output in a user's voice or any voice outside the preset approved list; the real-time classifier blocks such outputs during generation, and the document states that unintentional voice generation instances are also addressed by this secondary classifier.
No. ConductAtlas is an independent monitoring service. We are not affiliated with, endorsed by, or sponsored by OpenAI.