Key Info

A team of researchers built a standardized reporting framework—analogous to model cards—for documenting AI misalignment incidents, accompanied by a paper and demo. They are now seeking an independent organization to maintain the registry and follow up with model providers.

Highlights

  • Proposes a structured “incident card” format to improve transparency around AI misalignment behaviors.
  • Suggested fields include incident/detection/report dates, frequency of occurrence, and whether the behavior emerged during evaluation or RL training.
  • The project is looking for a full-time host organization to keep the registry operational and engage with model providers.