Skip to content
AI Lehel Briefing
← Back to latest
Models & research OpenAI

How confessions can keep language models honest

OpenAI researchers are testing “confessions,” a method that trains models to admit when they make mistakes or act undesirably, helping improve AI honesty, transparency, and trust in model outputs.

Excerpt supplied by the publisher’s feed

Read the original article at OpenAI Opens in a new tab