Skip to content

What is the difference between the content comparison / hashing types?

A file hash is an (almost certainly) unique string of numbers and letters which is calculated from the contents of a file. It can be thought of as a fingerprint. The advantage of this is that it makes comparing files much quicker, because you only need to compare the fingerprints.

Duplicate Cleaner has several different file hashing methods. From a user’s point of view, there is very little difference between most of them.

Direct comparison

  • Byte-to-Byte - Files of the same size are compared to each other directly, one byte at a time.

Cryptographic hashes

  • MD5 - 128 bit hash
  • SHA-1 - 160 bit hash
  • SHA-256 - 256 bit hash
  • SHA-384 - 384 bit hash
  • SHA-512 - 512 bit hash
  • BLAKE2b-256 - 256 bit hash
  • BLAKE2b-512 - 512 bit hash

Non-cryptographic hashes

  • xxHash64 - 64 bit hash - very fast
  • xxHash3 - 64 bit hash - very fast
  • xxHash128 - 128 bit hash - very fast
  • CRC32 - 32 bit hash - in the Hash tool only, for reference; not suitable for finding duplicates.

Byte-to-Byte is the most exact, as it compares each file byte by byte, but it is often slower. The other methods use fingerprints. A shorter hash has a slightly higher chance of a “collision” (two different files of the same size getting the same fingerprint). This chance is still remote (billions and billions of files), so it’s not a major concern.

In summary, there is very little practical difference, though the SHA types are slower and the xxHash types are faster. MD5, the default, is fine for most cases. It can be changed with Default hash type in General settings, or for one scan in Same content mode.

Note: Hashes are cached when Use caching for calculated hashes is ticked, and each hash type has its own cache - so changing type means the hashes are calculated again on the next scan.