Discussion about this post

User's avatar
deusexmachina's avatar

I have never clicked on a Substack article this fast

rasmi's avatar

Hi Sayash and Arvind, excellent essay as always!

A few thoughts:

"Notably, despite its usefulness, auto mode was introduced only in March 2026 — over a year after the release of Claude Code, and well after coding agents became mainstream, and despite the fact that it didn’t require new technical breakthroughs."

I would argue that it did require technical breakthroughs in the form of fast, cheap, and accurate classifiers: https://www.anthropic.com/engineering/claude-code-auto-mode

And similarly, resilience against prompt injection/jailbreaks has required significant research investment: https://x.com/bcherny/status/2086520950259118464

I think the AI safety crowd would consider this an accomplishment of alignment research (in the sense of improving model robustness against attacks, and improving classification/detection capabilities), though it would also fall under AI control in your framework (in that these improvements have now been deployed as guardrails in products).

One more:

"In contrast, turning to the other risks, we stand by our prediction that superhuman persuasion ability is largely a myth."

This may be a semantic distinction as you point out, but we've observed broader proliferation of deepfake scams against organizations as of 2023: https://arxiv.org/abs/2406.13843 -- It isn't too difficult to imagine self-directed agents pulling similar attacks (or using money/extortion to persuade people to take action on their behalf) at larger scales.

I think the fundamental tension here (which you point out) is that if you earnestly believe that superintelligence is around the corner (as many AI company employees and leadership do), then the controls ARE effectively a stopgap. In that world, what do you do when highly capable (but "misaligned") open source models/agents are broadly deployed, either maliciously or not? One could imagine a cascade of such incidents that are difficult to defend against or recover from. I don't mean to be defeatist, but the policy/liability frameworks haven't caught up, and I think the tendency to view alignment as essential relies in part on an assumption that our institutions won't catch up by the time the risks have proliferated.

41 more comments...

No posts

Ready for more?