Why we no longer evaluate SWE-bench Verified
SWE-bench Verified is increasingly contaminated and mismeasures frontier coding progress. Our analysis shows flawed tests and training leakage. We recommend SWE-bench Pro.
Independent signal, primary sources
Models, research, coding tools, open source, infrastructure, and major product releases—ordered by publication time.
Chronological feed
SWE-bench Verified is increasingly contaminated and mismeasures frontier coding progress. Our analysis shows flawed tests and training leakage. We recommend SWE-bench Pro.
We share our AI model’s proof attempts for the First Proof math challenge, testing research-grade reasoning on expert-level problems.
3.1 Pro is designed for tasks where a simple answer isn’t enough.
OpenAI commits $7.5M to The Alignment Project to fund independent AI alignment research, strengthening global efforts to address AGI safety and security risks.
The Gemini app now features our most advanced music generation model Lyria 3, empowering anyone to make 30-second tracks using text or images.
Machine Perception
Google DeepMind brings National Partnerships for AI initiative to India, scaling AI for science and education
A new preprint shows GPT-5.2 proposing a new formula for a gluon amplitude, later formally proved and verified by OpenAI and academic collaborators.
How OpenAI built a real-time access system combining rate limits, usage tracking, and credits to power continuous access to Sora and Codex.
GABRIEL is a new open-source toolkit from OpenAI that uses GPT to turn qualitative text and images into quantitative data, helping social scientists analyze research at scale.
Our most specialized reasoning mode is now updated to solve modern science, research and engineering challenges.
Algorithms & Theory