At which point is the human using the model as a tool, and at which point is the model using human as a tool?
I can already see the border shift even for mundane tasks I have Claude working on. Increasingly, I'm just setting a high-level goal, and then checking progress and occasionally answering questions or doing something like configuring a system Claude can't easily reach itself (e.g. recording a bunch of traces through my normal use of a system that Claude deemed too fragile to risk operating on its own). Of course, I get detailed instructions to help me - "go there, do this and that, then press this to capture recording, run through this script here to process, attach result to next message". In those cases, Claude is effectively using me as a tool to call.
"improvement" is the word Id put the most emphasis on.
weve seen some improvement from the LLMs unattended, maybe, but will it actually keep improving vs needing a human to bring it back on track?
the recursive part is that it keeps improving on itself, but we really have no example of that. if it does it 30 times with improvements, then maybe, but even then, to actually be relevant it has to do better than paying scientists to do the work for the same cost, consistently.
RSI still means nothing if it costs 1000x the cost to get the same improvements as a human researcher