I was waiting for this.
Planning on doing flash analyses of PRs that impact evals in some way, and then post comments on GitHub whenever there’s flaws in them
( https://evalship.com )