Continuing on CRI’s limitations, since 2015 the new and improved metric is TM-30’s “fidelity index” (Rf) which is calculated based on 99 samples, among many other improvements.
The larger number of samples makes it harder for an LED manufacturer to “cheat” by optimizing for the specific reflectance spectrums used to score CRI, getting a higher score while not actually rendering colors of most real world materials that well.
https://www.energystar.gov/sites/default/files/asset/documen...
Unfortnately it’s still much less popular than CRI.
I like the goal of making it harder to overfit, but there are a couple significant weaknesses.
A big one is TM-30 Rf is still an average, and with a larger number of samples, the average minimizes the effect of strong deviations that only affect one or two samples. Actually looking at the color vector graphic reveals it, but some formats, like putting a bunch of different light sources in a table make that information harder to convey.
Probably worse when it comes to LED sources is that it doesn't include a primary red or primary blue. R9 (strong red) is a classic weak point for LEDs, with many "high-CRI" LEDs measuring in the 50s, and some other white LEDs having negative scores. R12 (strong blue) is a different kind of weak point; negative scores aren't common, but neither are values over 80.
So I'd like to a standard that uses more values, but it needs to include the red and blue primaries, and it needs to clearly report significant deviations - "Rworst" or some such.