continuing with the "we need open training data" thread
how much of the training data had thinking traces that dont make sense to people as being actually a description of why the output should be that way?