Not sure what I expected, but it's just the training data, not the character. It'd be so cool if such systems had natural curiosity at this checkpoint. Eg:
> Me: "What's semiotic crystallography? > Response: "I don't know, what is it?"
Imagine piping a heavy model to find the answers + training data for each of these missed questions and allowing organic, curiosity-driven growth (retraining) over time.
It would lose knowledge about existing subjects unless it’s continually retrained on those too. It could help inform the next training dataset though.