OpenAI is releasing a framework for detecting, reviewing, and disclosing unexpected or concerning behaviors in its models, along with six initial misalignment reports. The goal is to document these incidents using a method followed over time, rather than through general statements about AI safety alone.

In its publication “Our framework for reporting model misalignment”, the lab describes misalignment as a gap between a model’s expected behavior and that observed during internal evaluations or investigations. The process covers the detection of problematic behavior, its investigation, and then its disclosure in a report.

Six cases to test transparency

The six reports are the first concrete cases published under this framework. They are intended to distinguish situations linked to an evaluation weakness, a control failure, a particular interaction with a user, or a more general property of the model.

The value of the initiative will nevertheless depend on the content of the reports: the severity of the cases selected, test context, classification criteria, corrective measures, and unpublished incidents. Publishing a framework alone does not make it possible to determine what share of observed behaviors is made public.

These documents may inform the debate on the safety evidence that AI providers should provide, particularly in France and Europe. The decisive factor will be the regularity of disclosures and their ability to shed light on risks.

Back to all news

Comments· 2 comments

  1. Olivia Davis· 17 septembre 2026

    The framework sounds useful, but I’d want to see the underlying methodology before judging the six cases. What thresholds define “concerning” behavior, how reproducible are the evaluations, and were the models tested by independent reviewers?

    1. Anna Taylor· 17 septembre 2026

      Those are the key questions to look for in the reports: evaluation prompts and scoring criteria, sample sizes and uncertainty, whether results were replicated across runs or model versions, and what safeguards were used to prevent cherry-picking. Independent external review would also make the disclosures more persuasive.

Leave a comment