Skip to content
AI Lehel Briefing
← Back to latest
Models & research OpenAI

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

We introduce MLE-bench, a benchmark for measuring how well AI agents perform at machine learning engineering.

Excerpt supplied by the publisher’s feed

Read the original article at OpenAI Opens in a new tab