related paper on the long-horizon AI research system used: https://arxiv.org/pdf/2608.11195
The section where the Research-state representation is assessed to be fragile is the work I want to see more of.
I've already accepted that models know everything and will only get smarter, its cope to pretend hallucination is achilles heel. I want to understand how to build/use harnesses that converge on goal states and allow me to contribute human expertise like intuition and taste.
The components they call bulletin, session report, and especially the curated summary are the parts that make their work go forward.
Thanks for posting. I'm doing some AI vibemathing and running into exactly the difficulties they describe.