the benchmark I trust most is whether the model can explain its own pricing page without getting confused
Not even humans can do that, you're literally asking for something beyond AGI
Not even humans can do that, you're literally asking for something beyond AGI