So at 3 million different files you have a 98.3% chance of a hash collision. Wouldnt that cause problems in real datasets?