Kudos to OpenAI for sharing their recent experiences with a misaligned internal model, where they encountered problems sufficiently severe they were forced to take the model offline to work on new mitigations and defense-to-depth.
Firehose
Filtered to People, tagged “internal model misalignment” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News