Skip to content
AI Lehel Briefing
← Back to latest
Models & research OpenAI

Toward understanding and preventing misalignment generalization

We study how training on incorrect responses can cause broader misalignment in language models and identify an internal feature driving this behavior—one that can be reversed with minimal fine-tuning.

Excerpt supplied by the publisher’s feed

Read the original article at OpenAI Opens in a new tab