I agree with that historically auto-vertorization does not seem to work reliably. I'm not sure about your broad claim.
Thoughts on an abstraction over ARM and x86, at 128, 256, and 512-bit widths which, either in a manual or automatic way (The latter more challenging) makes your floating point computations 4-16x faster with minimal restructuring? I think that's doable, and a nice goal of SIMD.