logoalt Hacker News

fnordpiglettoday at 12:47 PM0 repliesview on HN

This research is not current, and is based on a flawed premise - generated agents.md are used. There isn’t a hand curated agents.md that has some knowledge passed to it that is withheld then measured against.

I think there is overall something here for current Claude which is there appears to be a hierarchy of conformance that breaks progressive disclosure and the utility of skills. It seems to honor the system prompt, user instructions, tool call results, and dead last skills. It applies a large amount of discretion as to whether to honor what skills say in the imperative and progressive disclosure seems to have at best a 20-30% recall. Other models like codex gpt 5.6 seem to be the exact opposite and slavishly adhere to the Agent/skills/plugins, to the point of being wasteful and dangerous. It feels clear there’s a tension being RL’ed around between conformance and skeptical behavior that neither has quite found the balance for yet, and is almost certainly an over constrained problem. I just find it funny Anthropic is the one you can’t trust with your wallet while OpenAI does precisely what you and your harness tell it to.

But this “science” and its editorializing are based on flawed techniques, don’t lead to the conclusion let alone the editorialized extrapolation, and are l