There is always going to be a problem when we must judge value. You mention gains, short and long term.
Knowing whether something is valuable, a gain, requires a judge. I the case of these tests: the judging is inadequate.
In economics, each of us plays the judge by choosing whether or not to pay for a service. The decision was yours: if you gave money, you must have deemed the service valuable.
There's no such judgement with these model tests. The only judgement is the final score.
It feels a lot like externalities. Like planned obsolescence increases profit at the expense of the environment. Is our judgement lacking because our scoring is failing to account for these externalities? I could be completely off base here and am out of my depth but I find this whole thread fascinating.