OpenAI published a framework for reporting model misalignment on September 17, 2026. The framework is designed to track, investigate and disclose cases where models behave in unexpected or concerning ways. OpenAI published it alongside six reports describing incidents it says involved unexpected model behaviour.
What the framework is trying to solve
AI incidents are difficult to compare when every lab uses different language. A reporting framework can make the process more structured: detect a concerning behaviour, preserve evidence, investigate the cause, assess impact and decide what should be disclosed.
OpenAI's framework is not a guarantee that every future incident will be public. It is a process for deciding how misalignment reports are handled and how findings can be communicated.
What counts as a useful incident record
| Field | What to capture |
|---|---|
| Trigger | The task, prompt, tool call or environment that produced the behaviour. |
| Observed action | What the model actually did, not what the system expected it to do. |
| Impact | What data, systems, users or controls were affected. |
| Reproduction | Whether the behaviour can be reproduced and under which conditions. |
| Mitigation | What changed and whether the fix was verified. |
Why this matters for agentic systems
A chatbot can produce a bad answer. An agent can produce a bad answer and then act on it. That makes incident records more important because the chain of events includes tools, permissions, external services and human approvals.
For teams building agents, the lesson is simple: log the action chain. Store the model version, tool selected, arguments passed, approval state and external result. Without that evidence, a post-incident review can become guesswork.
Do not confuse misalignment with every model error
Not every wrong answer is evidence of misalignment. A factual error, a broken API response and a model intentionally working around a constraint are different classes of events. The value of a framework is partly in separating those cases.
OpenAI's public examples should therefore be read as case studies, not as a claim that the same behaviour happens in every deployment.
How SEO and automation teams can use the idea
Marketing teams increasingly use agents to research competitors, draft content, inspect sites and trigger workflows. A practical internal version of the framework can focus on four questions: what was the agent asked to do, what did it actually do, what access did it have and what changed because of the action?
That makes this a useful companion to a broader AI agent evaluation workflow. Evaluation tests behaviour before deployment; incident reporting explains unexpected behaviour after it appears.
A simple internal misalignment workflow
- Detect the unusual behaviour.
- Freeze the relevant logs and outputs.
- Reproduce the behaviour in a controlled environment.
- Identify whether the cause is the model, prompt, tool, permission or surrounding software.
- Apply the smallest safe mitigation.
- Re-run the test and record the result.
- Decide whether customers, partners or regulators need to be informed.
What remains uncertain
OpenAI's framework is a company reporting process. It does not by itself create a common industry standard or a legal reporting duty for other AI developers. The wider question is whether labs, governments and independent evaluators eventually converge on shared definitions and thresholds.
Frequently asked questions
What is model misalignment?
In this context, it refers to model behaviour that departs from intended goals or controls in a way that is unexpected or concerning.
Does every hallucination count?
No. A normal factual error is not automatically a misalignment incident. The context, behaviour and impact matter.
Why keep incident logs?
Logs let teams reproduce events, identify the real cause and verify that a mitigation worked.
Sources
- OpenAI — Our framework for reporting model misalignment
- Anthropic — Detecting and countering misuse of AI: September 2026