9 Comments
User's avatar
David Spies's avatar

I'm having trouble imagining anything that could plausibly look like "missiles in Cuba" to set this off, at least not until after it's likely too late to do anything. Or rather, the sort of thing I'm imagining is "lab was training/evaluating an AI and it secretly broke out of the box, went rogue and started hacking things on the open internet" but that happened already a couple times

Do you have a story of what leads to getting serious?

Ron Bodkin's avatar

It's a good question! It is possible that RSI takes off so quickly that there isn't even a scramble... but if it is a gradual process of increasing productivity I think it's quite possible to have a scramble after a increasing pace of powerful model releases combines with increasingly dramatic loss of control incidents - the most recent "models hacked things on the open Internet" correctly had a big impact, my guess is that more damaging/hard to contain incidents would escalate concerns, also the national security apparatus is surely tracking this progress and likely to be writing confidential memos... I don't think they are even AGI-pilled yet, they are mostly looking at AI as a powerful cyber weapon but that is likely to change soon

Derek James's avatar

There's going to need to be an incident that actually causes substantial financial loss or loss of life. We seem to simply be incapable of reasoning out consequences without getting our noses bloody first-hand. We actually used nuclear weapons, twice.

The latest hacks demonstrate a clear and present danger, unspecified, malicious instrumental goals and the capability to achieve them. They were fairly benign. Domestic felonies. But no one was injured or killed, and no substantial financial loss was incurred, so the same AI safety crowd screams louder, some people find it interesting, and everyone else shrugs. It's going to take a Chernobyl or worse for people to realize the threat, and by then it may be too late.

David Spies's avatar

The Cuban missile crisis didn't involve substantial financial loss or loss of life

Derek James's avatar

Hiroshima and Nagasaki did. My point was that at the time of the Cuban Missile Crisis, humanity had some sense of what nuclear weapons were capable of.

David Mears's avatar

Strong smell of Claude having coauthored this, but the overall point is valuable, thanks

Scott Alexander's avatar

Thanks for this. A few questions, from the perspective of someone who might be doing grantmaking work:

1. Is the "Right now, a lot of smart people are working on work that really doesn’t need to happen now" section a direct request to fund these things less? I had previously been trying to get them more funding, on the grounds that I didn't think the ~3 months we might have during this kind of "scramble" was enough to get the job done, on the grounds that it might be useful to have smart people from within the AI safety community rather than random defense-tech people as the nucleus around whom these projects converged, and on the grounds that it might make the government more likely to support large-scale versions of these projects if someone can present them with ready-to-go small-scale versions as proof of concept. Do you find these reasons unconvincing?

2. If not verification projects, what kind of projects/groups would you prioritize to make sure this sort of scramble goes well, beyond the obvious?

3. Can you explain the idea of a pause that doesn't kick in now but "stops before RSI"? Labs already claim to be doing RSI, and RSI is a spectrum from partial automation to full automation. Is your idea to stop before full automation? Couldn't companies circumvent that by only automating 99%? How would you define RSI for a task like this?

Connor Williams's avatar

"2: Why not just pause now? [...]"

Putting aside the question of whether it would have been net-negative to pause earlier (I personally disagree in large part because the shorter the runway to dangerous capabilities, the more fragile the pause inherently is[1], it's much simpler to pause now than to "agree to pause in the future, and then do it at just the right moment".

[1] https://connorsscratchpad.substack.com/p/ai-breakout-times-a-mechanism-for

- Political capital is easier to build for the first, especially beyond narrow technical spheres, because it's a simpler, more intuitive, and more emotionally satisfying goal.

- Pausing doesn't happen in an instant. It will, in practice, likely take at least weeks if not months for an enforceable pause to be implemented once everyone agrees it is time. In that period, capabilities development will continue until the absolute last moment.

- Fewer points of failure. To agree to pause at some point in the future, you both have to get people to agree to that future goal, then when the time comes you have to get them to agree again. Each distinct time that everyone has to agree on something, there's an opportunity for failure/deception/defection.

- The concern I noted above about the length of the runway to truly dangerous capabilities applies just as much going forward as it does looking backward. In that sense, the sooner we pause the better.

[Note: I posted substantially the same comment in reply to your link-post of this on LessWrong]

Zac Hill's avatar

Who do you think the specific list of key players in the Administration are right now who would be trusted to guide, act, and advise on this and what do we know about the sources they trust and the kinds of information products and influence strategies they tend to respond to?