Yup sorry if I mislead people. I usually use 10 hexdigits (40 bits) but no matter how many there are, it's the x bits of the beginning of the checksum that are verified against the hash (in the examples I gave 40 bits).
I wasn't very clear.
So at 3 million different files you have a 98.3% chance of a hash collision. Wouldnt that cause problems in real datasets?
So at 3 million different files you have a 98.3% chance of a hash collision. Wouldnt that cause problems in real datasets?