I think your reply has a somewhat familiar structure -- "it doesn't work for you because you used an [old / suboptimal / non-frontier] model. If you use X you'll see that it works". You might be completely correct! But these sorts of claims push the onus back onto the other person, without accepting any work for yourself. It gets tiresome to retest with the newest model every other week. Is there any data you can provide to support your claim, or any result you can contribute here?
The frontier is advancing really rapidly. The models are getting better faster, especially on RSI related tasks. The best way would be to try astra or fable on some hard problems.
Other than that I'd look at some of the more unique benchmarks for astra, like playing factorio or using blender. It's an entirely different beast.