Skip to content
AI Lehel Briefing
← Back to latest
Models & research OpenAI

Improving mathematical reasoning with process supervision

We’ve trained a model to achieve a new state-of-the-art in mathematical problem solving by rewarding each correct step of reasoning (“process supervision”) instead of simply rewarding the correct final answer (“outcome supervision”). In addition to boosting performance relative to outcome supervision, process supervision also has an important alignment benefit: it directly trains the model to produce a chain-of-thought that is endorsed by humans.

Excerpt supplied by the publisher’s feed

Read the original article at OpenAI Opens in a new tab