> Participants were randomly assigned to receive assistance from an LLM (GPT-4o, Llama 3, Command R+)
These are pretty old. I'd be curious how performance compares with the latest frontier models.