From recent ChatGPT (GPT5.6) conversations where I've seen occasional reasoning leaks into the UI, it's clear that something like this is already implemented, and I'd speculate that this is the majority of recent claims of less token usage. Not sure if they are literally prompting for cablese of course.
"Need check output vs prev. Ran script, results fine, need prep next step. Ready? Go."
Would be pretty funny if OpenAI's models are so token efficient now due to hidden caveman prompt.
I can confirm I’ve seen this with Deepseek V4.1, I imagine it’s with other models too.
I suspect you might be right. My conclusions from the digging are that this kind of compression works best with settled instructions/data for machine to machine talk. For something like OpenClaw (which I use a lot), that might mean the AGENTS.md, TOOLS.md, etc. Compression there would free up the context for the agent.