logoalt Hacker News

dannyw • today at 12:10 PM • 0 replies • view on HN

I think that's debatable. Anthropic's interpretability research suggest models do have self-introspection ability, at least in the "J-Space": https://transformer-circuits.pub/2026/workspace/ ; and this private working/'introspection' space is distinct and distinguishable from the tokens they output (CoT tokens are output too).