We have been running a lot of agentic benchmarks with the various loops and tool calls on longer threads - we routinely see 90%+
Just checking now: recent runs tau3[1] was at 96% and toolathlon[2] was at 90%
[1] https://www.induction.ai/docs/benchmarks/tau3 [2] https://www.induction.ai/docs/benchmarks/toolathlon