We've seen models perform worse at higher efforts in our vuln detection evals. For example IIRC gpt 5.5 and 5.6 both scored better or high as compared to xhigh.