No, the Chinese tech companies routinely super optimize for a given 'famous' or 'established' or 'mainstream' benchmarks. Any other concern is secondary.
In other words, they look good on surface, but suck whenever anyone put them to any serious use. That's why they are cheap, they have to be cheap because they suck.
how can you say that with a username of "ngl999", that doesn't make sense
(IIUC, -ngl [NUM_LAYERS] specify number of layers to offload to GPU in llama.cpp, 999 on most use cases might as well be -1)