I'm very convinced this is what's going on. Qualia are super hard to convey with language! If I had to guess these self-reported instruments are also measuring, indirectly, conscientiousness (do I _really_ see it? it'd be wrong to lie on a survey!) as well as neuroticism (oh God what does that mean? do I _really_ see it? do other people _really_ see it? am I broken?).
It’s very hard to prove, yes. But you need to deal with the association with SDAM, and the differences in outcomes on the https://aphantasia.com/study/vviq.
It is Qualia only until they can measure it. Which they can now. Blaming everything you do not understand on neuroticism? Might as well go back to dunking witches.
https://www.sciencedaily.com/releases/2022/04/220420092150.h...
I think a couple of the comments in this very thread above, e.g. https://news.ycombinator.com/item?id=49467695 and https://news.ycombinator.com/user?id=slibhb are quite relevant to your case. Thinking this is just semantics is a strong indicator you are likely highly aphantasic.
It is highly unlikely to just be semantics, because when it comes to visualizing complex objects as a non-aphantasic, the specific details spring forth unbidden as an image first, and then you attach the language, and it isn't the reverse order. I suspect also tests of visual memory of complex pictures should reveal similar things here, because e.g. language and other things don't really seem to be able to represent such details, and neurobiologically, there isn't much plausible other way your brain could be memorizing a complex photo without retaining / having activation patterns similar to the actual viewing (and this is what we see in fMRI studies).
If you are careful at noticing your thoughts (and not aphantasic), you can also notice if you are a highly visual thinker, you might sometimes confuse certain memories of things (spellings, place names, other things) that are similar only visually, but not logically or in any other sense modality representation. This would suggest the representations are genuinely visual in these cases.
But yes, no way to be 100% certain, but the "it is just semantics" case is very weak and can't explain much.