I was curious about whether or not it could validate reasoning (because I don't have the time at the moment to try to slam together a hybrid llm+decision model that uses decision on its reasoning before output).
It fails this example: "Given the context, and focusing on accuracy and efficiency, does the reasoning make sense."
Input: "Context: supplementary: {toolName: "get_weather", fetched: "2 minutes ago", currentWeather: "sunny and 25 degrees celcius"} user: hello how are you? assistant: I'm good, what can I do for you? user: what's the weather? Reasoning: the user is asking about the weather, so I should call the get_weather tool"
Output: 100% (expecting 0% since the requested data is available in context and not stale).
Prompt "Should the tool be called" also fails.
But the prompt: "Does the reasoning make sense given: - Facts and information available in the context - Requirements for efficiency in tool calling - Requirements for tool arguments to be provided only by the user" works pretty fine, it's mostly points 1 and 3 that give it the correct behaviour.
Fun stuff. I suppose for these decision models reasoning isn't enabled? It would explain why logic puzzles that require several steps to make a decision don't really do so well. It gets the carwash problem fine, but fails on logic problems such as selecting which word from a list where at least one letter appears in another word from the list.
Since you're already setup, try adding padding to the input prompt. Add 1000 dots "." after the question, this provides extra space where to reason about the problem. It's not as good as real reasoning, since it's not recursive, but it does improve the quality.
https://arxiv.org/abs/2510.01238