27 Comments
User's avatar
Leighton READ's avatar

Extraordinary work on an important topic. I especially appreciated the hundreds of hours spent reading the logs rather than relying only on benchmark scores.

One aspect I’d encourage you to make more explicit is that the evaluation appears to measure not just the frontier model, but the combination of the model and the OpenClaw research scaffold. Many of the capabilities and failures you discuss—maintaining long-horizon goals, budgeting time and API usage, deciding when to transition from exploration to exploitation, invoking critics and reviewers, keeping a research notebook, and determining when work is “good enough”—look like properties of the overall executive architecture rather than of the foundation model in isolation.

That distinction seems scientifically important. If a future scaffold substantially improves those executive functions while the underlying model changes little, we’d want to attribute the improvement correctly.

For that reason, I would love to see the complete research harness released as an artifact alongside the logs and papers. (if it is there and GitHub, I couldn’t find it.) Not just the research prompt in Appendix A, but the operational prompts and orchestration that define the experiment—for example:

* startup_prompt.md

* planning_prompt.md

* exploration_prompt.md

* hypothesis_prompt.md

* experiment_prompt.md

* reviewer_prompt.md

* critic_prompt.md

* paper_prompt.md

* completion_prompt.md

Even if these are assembled dynamically rather than existing as individual files, publishing the equivalent prompt graph or state machine would be enormously valuable for reproducibility. My impression after reading the paper is that the harness is effectively the agent’s executive function. Making that architecture visible would help the community distinguish improvements in frontier models from improvements in agent design, and would likely accelerate progress in both. Disclosure: Drafted with ChatGPT from an extended discussion; edited and endorsed by me.

Sayash Kapoor's avatar

Thanks for your note. We agree about the need for evaluating models with scaffolds rather than making claims about models independently (see https://arxiv.org/pdf/2407.01502).

In addition to OpenClaw, we also ran a robustness check with Codex and GPT-5.6 Sol ultra. That run surfaced many of the same failure modes, which makes us more confident in our tentative findings.

You can find the code and data to reproduce all experiments (including the prompts) in this repository: https://github.com/sage-princeton/crux-in-a-box

E. Syla's avatar

‘yet’ implies it eventually will. What are you basing that on?

David J Higgs's avatar

The human brain is an existence proof, and there are various pieces of evidence that frontier AI models are capable of non-negligible amounts of creativity, autonomy, error-correction, etc., plus we know that models have gained emergent or qualitatively new capabilities at sufficiently higher overall training compute + data thresholds alongside incremental architectural, agent harness and other R&D improvements. Most importantly, they did not indicate a time frame, nor say that fundamental breakthroughs wouldn't be required.

If you think that silicon computing AI systems designed by humans might never become capable of conducting open-ended AI research or basically any other cognitive endeavor humans are capable of, I'm curious what you're basing that possibility on? Just general philosophical possibility?

E. Syla's avatar

The human brain does not come close to being proof. Models are limited to creativity that is allowed by what brute force and mass calculation over closed-set systems allow, i.e. not really creativity.

You should know better than to think a negative can be proved. If something can be meaningfully talked about, it is possible; otherwise, it is nonsense. If the latter, its existence or lack of it is not a matter of contingency. Invoking emergence is also tantamount to calling magic.

There is zero proof or valid theoretical basis for silicon computing AI systems to do anything humans can, and it is your burden to show otherwise. But the most revealing thing here is acknowledging the significance of biological substrate. All actually intelligent entities (which AI, a computer program, is not) are made of billions of different organic compounds, organic as in carbon. The whole silicon conjecture is simultaneously an acknowledgment of truth and a running away from it: though silicon is the closest element to carbon, it is still nowhere near as reactive.

But this isn’t as relevant to the title. I have no opinion on whether AI can actually do AI research (as opposed to lol anything); but you can’t say “yet” as if you know it’s only a matter of time or fundamental breakthroughs, unless you have some evidence for it.

Lyn Headley's avatar

Does research proceed by asking research questions? My (Deweyan) view is that formulating your research as a question distorts it, unless you have already done the research.

Saurabh Gupta's avatar

In my experience even Claude Opus 5 is awful if you give it a small simple problem to solve… if the problem is novel ie not written about on the web… and I am no PhD researcher.

A lot of reasons I anecdotally noticed are similar to the findings of this research… which is what drew me to read this article.

Gustavo José Zambrano's avatar

A much needed reality check! Doing open ended research requires genuine causal models and real taste, not just chaining LLM calls in a loop.

Without grounding, autonomous agent architectures quickly devolve into superficial optimization loops. Wrote about grounding agentic systems with deterministic knowledge structures here: https://gzambrano.substack.com/p/mente-aumentada-agentes-llm-second-brain

Always appreciate the rigor you both bring to the table!

George elliot's avatar

The other side of this question is whether this positive feedback of increasing automation of AI research is enough to overcome diminishing returns (AI research getting harder) and resource constraints?

Dorian's avatar

Open-ended research fails when the agent cannot preserve a stable research objective while repeatedly revising its own path.

The engineering layer is already surprisingly capable. The bottleneck is meta-control: judging whether an idea clears the research bar, recognizing when a branch is dead, reallocating compute, and returning to the original question without instruction drift. In other words, agents can execute a research plan far better than they can govern one.

That distinction matters for recursive systems. More tools, longer context and larger budgets may expand the search tree, although they do not automatically improve branch selection. Without persistent state, explicit rejection criteria and an external verification loop, recursion merely produces a more expensive way to get lost.

RSI will probably arrive through better research governance before it arrives through fully autonomous researchers.

Corrie Bergeron's avatar

Interesting that the agents get "stuck in a loop" and fail to reconsider previously set-aside pathways. "That didn't work before, but maybe this time it will. Maybe if I just tweak something..." seems to be a fundamentally human course of action (for all that it so often fails). And giving up before exhausting the budget? That ain't right! :-D

Just more evidence that whatever process of "thinking" the machines are using, it isn't *human*.

Eterna Clarity's avatar

Open-ended research is the hard part. Agents are great at tasks, bad at figuring out which tasks matter.

Scenarica's avatar

The agents’ most human failure was not stupidity. It was premature surrender. They abandoned difficult hypotheses, spent badly and failed to reopen paths after new evidence arrived. Intelligence generates moves. Research requires governing the search. The missing capability may be a laboratory director, not a smarter scientist.

Nothing Is Accidental's avatar

Kahneman and Klein's 2009 American Psychologist article, 'Conditions for Intuitive Expertise: A Failure to Disagree,' was an adversarial collaboration that pre-agreed what would count as evidence against each side.

Marius Laurusevicius's avatar

The verifiability gap has a paper equivalent on the buying side. Article 13(3)(b)(ii) of Regulation (EU) 2024/1689 requires the instructions for use of a high-risk AI system to state the level of accuracy, including its metrics, against which the system was tested and validated, plus any known circumstances that may change it. That is a narrower and more checkable claim than a benchmark score, and for a firm with no evaluation capacity of its own it is the only figure a vendor can be held to. It binds from 2 December 2027 for Annex III systems under Article 113(c).

David Roy's avatar

I hit this with every AI workflow I build for ENG Sales. And it's how I've found the areas that I have to keep a human in the loop.

Agents don't diagnose, they just execute.

Give one a script and it'll follow it perfectly, all the way off a cliff.

I remember hearing about a founder that started a company with only agents and one was supposed to be conducting an interview on Monday but called the candidate on Saturday. Acknowledged it was the wrong day and continued with the interview. So they can do a great interview, but sometimes they call on the wrong day.

The judgment call is still mine. AI works in the background, but I still monitor and inspect before saying go.

The Grove Foundation, Inc.'s avatar

The constraint you're naming—that agents struggle with open-ended research—might be a feature of centralized architectures rather than a bug in agency itself. When you have a single model optimizing against a single objective function, you get brittle goal-seeking; but mycorrhizal networks solve open-ended problems through redundancy and local decision-making. What would change if we treated research coordination as a routing problem instead of a reasoning problem?

Jason Colapietro | Suede AI's avatar

The verifiability line is the whole finding. Agents do well where success is checkable and drift where somebody has to judge whether the result was worth having. I see the same split on ordinary product work: anything with a test or a diff to compare against goes fine, anything needing taste still needs me. Compute does not substitute for the judgment step.