In the third vision bench result, Sol is 100% correct but the expected has 1 error. Seems like an oversight.
In the next bench, Sol looks like it’s correct again but the bboxes are rotated 90 degrees for some reason.
Seems to be due to the detection area being not fully accurate. Green vs red shows the difference between actual and detected
Hi! I’m the author of this blog and benchmark. You’re right. I’ll fix it in the ground-truth dataset. Thanks for pointing it out.