It's doubly funny knowing that basically nobody bothers doing these kinds of workloads on CPU if they can help it. And everything you have to do to get GPU floating point performance also makes GPUs really, really bad for normal CPU code. Hell, at one point AMD actually was shipping VLIW for shader code...
I think that’s really illustrating how the problem isn’t what they expected: nobody was doing GPU computing in the 90s and when it became huge that was specialized for certain classes of work and involved custom toolchains. The fact that even AMD’s VLIW ended up diverging on later models makes me think Itanium was doomed even if it had shipped on time and budget.
> It's doubly funny knowing that basically nobody bothers doing these kinds of workloads on CPU if they can help it.
When the Itanium was developed and introduced (2001), nobody was thinking about general-purpose computations. DirectX 8.0, which introduced Shader Model 1.1 (which was far away from being suitable for GPGPU; Shader Model 1.1 was rather about strongly (also size-)limited programs for the vertex and pixel processing stage), was only introduced in 2000, the first release of CUDA was in 2007, and the first release of OpenCL was in 2009.