Such a massive leap at averaging possible use cases from previous data collected.
My guess is : collect all the prompt and their satisfaction score. group them by similarity . For each group pretrain the next model on that . Get these results ready.
Next model generation feed them back those answers.