Interesting read, really appreciate these interviews. I am a bit puzzled about interviewees pointing to the METR graph. Was that mainly just a rhetorical device or do they really think that the benchmark is a proxy for when RSI is reached? For one, the benchmark is dominated by tasks that are self-contained, well-specified, low in org context and easily-automatically gradable. In contrast, actual AI research includes asking relevant questions, noticing that an experiment leads to ambiguous evidence, assessing quality, considering what the most valuable avenue is - and so forth. Arvind Narayanan goes into greater detail here: https://www.cs.princeton.edu/~arvindn/talks/icml-2026-annotated-slides/ I am sure you are aware of this talk of course.
Good article, but not a fan of the recommendations:
1. Convene public hearings that put AI companies under oath on recursive AI improvement. --> Good :)
2. Track the threshold: Direct the Center for AI Security and Innovation (CAISI) to maintain a government-run task-horizon benchmark and publish recurring capability forecasts. --> We are already past this point, too late: it is increasingly hard for METR to conduct those time-horizon evals; this would only reproduce what METR painfully created, and wouldn't change a thing
Lots of great stuff here - the load-bearing ness of empirical problem definition to RSI; the lack of commercialization incentive to deploy internal S-tier tools publicly (in stark contrast to eg Amazon). But imo I’m most compelled by the imperative towards third-party verification institutions, which themselves I think will have to be RSI enabled. How IYO do you think we can crack the nut on the internal talent needed to execute against that brief sustainably? Do you think the will towards contribution along that axis is sufficient to overcome potential conflicts of interest (eg ‘donated time’ from labs?)
Thanks! I do think this is a concern. IIUC — the talent concern is that the AI companies pay vastly more than government? (ie. 500k+ vs. <100k)
At this point I think the leading AI companies have pretty good will, even up to Sam Altman / Dario Amodei.. but their relationships may become increasingly adversarial with the government as they become larger ..
One key point might be ensuring government access to (1) models, including internal deployments, (2) logs of what those models have done (e.g. chat history, agent rollouts, etc. things that the companies already keep) — this might be a bigger bottleneck than oversight capacity itself!
Right, yeah, exactly - like it's precisely the equilibrial dynamic changing that is the thing to hedge against. But I do think that a lot of what can help is the kind of stuff you are talking about here.
The widening gap between what researchers are producing in AI company settings and what the public understands is a significant vulnerability.
Interesting read, really appreciate these interviews. I am a bit puzzled about interviewees pointing to the METR graph. Was that mainly just a rhetorical device or do they really think that the benchmark is a proxy for when RSI is reached? For one, the benchmark is dominated by tasks that are self-contained, well-specified, low in org context and easily-automatically gradable. In contrast, actual AI research includes asking relevant questions, noticing that an experiment leads to ambiguous evidence, assessing quality, considering what the most valuable avenue is - and so forth. Arvind Narayanan goes into greater detail here: https://www.cs.princeton.edu/~arvindn/talks/icml-2026-annotated-slides/ I am sure you are aware of this talk of course.
Good article, but not a fan of the recommendations:
1. Convene public hearings that put AI companies under oath on recursive AI improvement. --> Good :)
2. Track the threshold: Direct the Center for AI Security and Innovation (CAISI) to maintain a government-run task-horizon benchmark and publish recurring capability forecasts. --> We are already past this point, too late: it is increasingly hard for METR to conduct those time-horizon evals; this would only reproduce what METR painfully created, and wouldn't change a thing
3. "This field is underfunded and undeveloped, but without it, any international limits on AI development are practically unenforceable." --> I think that it is harmful at this point to amplify this narrative of impossibility when you have not even started doing the obvious damn thing: https://ai-frontiers.org/articles/an-international-ai-slowdown-is-ready-whenever-politicians-are + https://x.com/CRSegerie/status/2003122396799398031
Lots of great stuff here - the load-bearing ness of empirical problem definition to RSI; the lack of commercialization incentive to deploy internal S-tier tools publicly (in stark contrast to eg Amazon). But imo I’m most compelled by the imperative towards third-party verification institutions, which themselves I think will have to be RSI enabled. How IYO do you think we can crack the nut on the internal talent needed to execute against that brief sustainably? Do you think the will towards contribution along that axis is sufficient to overcome potential conflicts of interest (eg ‘donated time’ from labs?)
Thanks! I do think this is a concern. IIUC — the talent concern is that the AI companies pay vastly more than government? (ie. 500k+ vs. <100k)
At this point I think the leading AI companies have pretty good will, even up to Sam Altman / Dario Amodei.. but their relationships may become increasingly adversarial with the government as they become larger ..
One key point might be ensuring government access to (1) models, including internal deployments, (2) logs of what those models have done (e.g. chat history, agent rollouts, etc. things that the companies already keep) — this might be a bigger bottleneck than oversight capacity itself!
Right, yeah, exactly - like it's precisely the equilibrial dynamic changing that is the thing to hedge against. But I do think that a lot of what can help is the kind of stuff you are talking about here.