Don’t even bother. These people almost always have some goofy ass, non standard, fluid definition of “thinking” or “reasoning” that cannot ever be met.
My opinion of their reasoning capability is based in part on (proprietary, non-published, only internally peer reviewed) experiments on GPT-series models’ ability to perform a suite of formal and informal inference and deduction tasks.
Perhaps you could argue that “appropriately applies syllogism to arrive at correct conclusions” is too high a bar to set, but I don’t think it would be fair to call it a “goofy-ass”, “non-standard” or “fluid” element of a reasoning capacity assessment.
My opinion of their reasoning capability is based in part on (proprietary, non-published, only internally peer reviewed) experiments on GPT-series models’ ability to perform a suite of formal and informal inference and deduction tasks.
Perhaps you could argue that “appropriately applies syllogism to arrive at correct conclusions” is too high a bar to set, but I don’t think it would be fair to call it a “goofy-ass”, “non-standard” or “fluid” element of a reasoning capacity assessment.