Discussion about this post

User's avatar
Kevin's avatar

The widening gap between what researchers are producing in AI company settings and what the public understands is a significant vulnerability.

Carsten Bergenholtz's avatar

Interesting read, really appreciate these interviews. I am a bit puzzled about interviewees pointing to the METR graph. Was that mainly just a rhetorical device or do they really think that the benchmark is a proxy for when RSI is reached? For one, the benchmark is dominated by tasks that are self-contained, well-specified, low in org context and easily-automatically gradable. In contrast, actual AI research includes asking relevant questions, noticing that an experiment leads to ambiguous evidence, assessing quality, considering what the most valuable avenue is - and so forth. Arvind Narayanan goes into greater detail here: https://www.cs.princeton.edu/~arvindn/talks/icml-2026-annotated-slides/ I am sure you are aware of this talk of course.

4 more comments...

No posts

Ready for more?