I think the frequency view is true, but a bit over-simplified. We can force certain output features to be true on tasks where the output is directly measurable, e.g. code compiles, lean proof is valid, benchmarking CUDA.
I think we can at least partially put constraints on other tasks that look fuzzy to us now, but haven't figured out how yet.