The Cranelift backend is extremely fast. Most of the slowness is all the crazy optimizations that LLVM does + other things (generating debug info etc).
Cranelift's biggest blocker for me is that it completely breaks debuggers:
https://github.com/rust-lang/rustc_codegen_cranelift/issues/...
For symbol heavy projects, linking is a surprising bottleneck.
Some ideas to speed up compilation by not evaluating items that are not used might pan out significantly for big crates in your dep tree (that behavior might never be stable because that would allow items with compile errors in a crate that would still let your application compile, which is against the Rust approach). The same work to do that would also allow overlapping of crate evaluation between different rustc instances called by cargo, as it would require partial evaluation of crates (to do name res only and gather the symbols needed from its deps).
Another thing is that stable rust doesn't treat macros as idempotent (because that wasn't a requirement from the start, there are crates that do dynamic IO to generate types), but if they are then incr comp can be faster by not evaluating them unnecessarily.
I know people are working on a bunch of different strategies to improve both first and incremental compile times, and I'm looking forward to the fruit of their labor.
Compile times are rarely the bottleneck for me, but that doesn't mean I won't welcome any improvements on that front.