That is not a good summary. In all their versions of the experiment, they presented the tool as AI, and used actual LLM answers. In the first example, participants were directly interacting with a real, though small, LLM, but they had technical issues because of that - 10% of the time the LLM setup failed to present an answer at all. So they redid the experiment with pre-generated answers - the participants saw the same UI, but when they asked the question of the AI, they instead got one of 3 pre-generated answers from that AI, to avoid the technical issues.
im not sure which part of my summary you take issue with then? the reason for study 1b is not relevant to answer the question asked in the comment i replied to