Strategic Report  ·  2026-07-31

Can AI agents conduct open-ended AI research? Early evidence from two case studies

Strategic ReportMedium impactGlobal
Researchers from Princeton's CITP (Sayash Kapoor, Arvind Narayanan and collaborators, including UK AISI-affiliated co-authors) introduce 'shadow evaluations' — giving frontier AI agents the central open-ended research question from an unpublished, high-quality paper, six days, and thousands of dollars of compute, then having the original human authors blind-grade the AI's output against their own. Across two unpublished NeurIPS 2026 submissions, agents completed all engineering work (literature review, debugging, hundreds of experiments, camera-ready LaTeX) without human help but were 'unambiguously rejected' by the original authors on research substance. The team identifies five recurring failure modes: poor judgment about the publication bar, uncreative responses to design flaws, ineffective backtracking, poor resource awareness, and instruction drift; a robustness check with a second model/scaffold reproduced the pattern. This is a preprint, not peer-reviewed.
Directly informs the recursive-self-improvement/AGI-timeline debate that boards and policy teams increasingly ask about: agents show fluent engineering execution but a genuine research-judgment gap, tempering near-term AI R&D automation risk assumptions used in scaling and safety-policy discussions.
Share with technical AI safety/research leadership evaluating claims of autonomous AI R&D capability; incorporate the failure-mode taxonomy into internal agent-capability risk assessments.
Can AI agents conduct open-ended AI research? Early evidence from two case studies (arXiv:2607.27191)
See this in the live feed Explore related AI security and governance findings — updated every morning.
Open the feed →