No amount of prompt engineering will give you the true probabilities for the model producing a certain response; this is something you can only get by inspecting the internal state at inference time.
For most use cases is that actually needed though? Just having it choose between predefined responses seems like enough but I'm curious about specific use cases because I do feel like I'm missing something
Does this give you the true probabilities for a certain response either? How does it work exactly? The probability of an overall positive answer isn't just the probability that the next token is "Yes"