It is hard to be fair I agree. We tried to be open about what we do here: github.com/keenableai/needle
The queries from what I can tell are not trivial. The actual github repo of the benchmark has a judgement/query browser where you can inspect the different query streams: https://keenableai.github.io/needle/