Tuesday, August 11, 2026

This should not come as a surprise…

https://www.bespacific.com/ai-agents-cant-yet-do-open-ended-ai-research/

AI agents can’t yet do open-ended AI research

Academics Sayash Kapoor and Arvind Narayanan in their latest AI As Normal Technology newsletter introduce their recent research evaluating whether AI agents can conduct open-ended research. After analysing hundreds of hours of agent logs, they identified recurring failures: agents lacked judgment for open-ended research, weren’t aware of available resources (spending less than 50% of budgets), failed to respond creatively to feedback, didn’t effectively backtrack from failed approaches, and ignored concrete instructions. The goal of leading  AI labs is recursive self-improvement (RSI): the automation of AI research using AI agents. RSI also underpins forecasts of explosive  AI  progress. How can we assess if we are close to this milestone? One way is to use benchmarks that test if agents can conduct AI research. Given the AI community’s focus on benchmarks, they have been the dominant way to evaluate progress towards RSI. Over the last year, many such evaluations have found that agents are now able to make progress on tasks where success is easily verifiable, prompting speculation that we are on the verge of RSI. But while these evaluations are helpful, they are limited to narrow, verifiable tasks. AI research can be much more open-ended. Success is often not immediately clear or verifiable, and to make progress, researchers need to test promising hypotheses, backtrack, or consider new or unconventional approaches. How can we evaluate agents’ ability to conduct open-ended AI research?  We take our first step towards answering this question in a new paper.  We partnered with the authors of two unpublished AI papers and asked them to draft their papers’ main research questions. We then tasked frontier AI agents with conducting research to answer these questions, and gave them thousands of dollars of API credits and compute, and six days of wall-clock time. The original authors reviewed the agents’ papers.

The authors unambiguously rejected both agent papers.  To better understand these results, our team spent over a hundred hours analyzing the agents’ logs. Our main takeaways:

  • The agents lacked the judgment for conducting open-ended research.  While the agents proposed directions the expert reviewers found impressive, they quickly rejected their proposed directions based on low-quality or synthetic data.

  • The agents lacked awareness about the resources available to them.  Both runs ended with less than 50% of the API budget spent and with hours left before the deadline, even though the agents could monitor their usage and were encouraged to spend down their budgets.

  • The agents did not creatively respond to feedback.  Despite the agents’ own AI self-reviews surfacing many of the issues that the expert reviewers later raised, the agents did not creatively address these concerns. When faced with negative feedback they responded by adding caveats to existing findings, and doubled down on unpromising research directions.

  • The agents did not effectively backtrack.  They retired their most ambitious research targets within the first day of the experiment, and neither agent fundamentally shifted its approach after that point.

  • The agents did not follow concrete instructions.  They ignored explicit rules about how much time to spend on exploration, how often to get reviews from AI self-review tools, and strict limits on paper length…





Fun watching this one.

https://thenextweb.com/news/social-media-addiction-lawsuits-ninth-circuit-section-230

A US court just cleared thousands of social media harm lawsuits to proceed

… At the heart of the case is Section 230 of the Communications Decency Act, the provision that has long protected online platforms from liability for content their users post.

Writing for the panel, Judge Jacqueline Nguyen concluded in a 24-page opinion that the statute offers “a defense against liability, not blanket immunity from being sued.”



No comments: