The M6 doubled the neural engine from 16 to 32 cores. I would expect that the M7 doubles that again to 64 from 32? That would make sense.
I believe that the CPUs are actually limited by ram bandwidth more than the neural engine right when it comes to LLM processing?
Maybe the M7 introduces something new to get around the current ram bandwidth problems on the non-Ultra chips.
Apple's biggest bottleneck for real-world inference is prefill processing. They need a better GPGPU architecture, which is what I'm expecting M7 to reveal.