It's very hard in practice. Even without optimization against it, it's about as hard as understanding an activation layer today and will probably get harder in the future.
There's also nothing stopping the secondary LLM from confabulating bullshit and there are less checks on it than CoT (harder for either humans or other models to externally verify).