Skip to content
AI Lehel Briefing
← Back to latest
Models & research OpenAI

Reinforcement learning with prediction-based rewards

We’ve developed Random Network Distillation (RND), a prediction-based method for encouraging reinforcement learning agents to explore their environments through curiosity, which for the first time exceeds average human performance on Montezuma’s Revenge.

Excerpt supplied by the publisher’s feed

Read the original article at OpenAI Opens in a new tab