so then what I’m really interested in are benchmarks of OSS models vs closed source frontier models, using Claude Code as a harness