The way I run the pelican benchmark prevents it from checking its own work.
It gets one API call to return an SVG.
If I ran the benchmark in Claude Code or a similar harness it could render the SVG as an image, look at what it created, then make tweaks to it.
Ah, thanks for clarifying!