Because the industry has plenty of experience with referece counting as the very first GC algorithm, already in the early 1960's, in early Lisp implementations, BASIC, Cedar, and several other languages.
The predictable runtime performance is also a myth, because they never take into account the use of NUMA memory, lock contention, possible stack overflow and stop the world in the case of cascaded deletions in naive implementations.
Lock contention yes, it's fundamentally a RC issue, but makes me think per-thread objects make more sense
I'm not sure most GC implementations worry about the rest neither (as by the several complaints we see going around)
Can you elaborate on how cascaded deletions "stop the world" with reference counting? I understand how GC could lead to arbitrarily large latencies for whatever task triggers that kind of cascaded deletions, but as I understand it, "stop the world" usually means that no threads are allowed to execute application code.