People also misunderstand the bar that was set by the test. It was more subtle than “Can the computer convincingly carry one side of a dialogue?”
His “imitation game” had three participants: a human participant, a computer participant, and an interrogator. The interrogator’s job was to talk to the participants and try to determine which participant is human and which is a computer.
He wasn’t interested in computers being able to fool the interrogator on occasion. The point where he thought the question of whether machines can think becomes moot is when the interrogator is unable to do much better than chance over many trials.
That’s a pretty high bar, and I don’t actually believe that LLMs have closed the gap with it by all that much. They still have so many obvious tells. And those tells are something Turing anticipated and accounted for. He explicitly considered deliberate deception as an essential part of the test, right there on the second page of a 30-odd page paper.
>That’s a pretty high bar, and I don’t actually believe that LLMs have closed the gap with it by all that much. They still have so many obvious tells.
Frontier Labs are not interested in having LLMs being able to pass as humans. If anything, they explicitly train them not to. In many ways, this ability has regressed severely since the original GPT-3 with no instruct tuning or RL. How many 'tells' would there be really if a frontier model trained with frontier techniques is optimized to pass this test? I think this was something Turing did not quite forsee. That such machines might be created but not really care about this specific shape of the test. Regardless, i think his broader point about functional equivalence is spot on.