Skip to content
Skip to main content
Logs, Records & Provider Evidence Technical Explainer

What is a file hash value?

A file hash value is a fixed-length value calculated from the bytes of a file using a named hash algorithm. It provides a reproducible way to compare the file's digital content, provided the same input scope and algorithm are used.

Hash values are often displayed as a long sequence of hexadecimal characters.

The short version

The same algorithm over the same bytes should produce the same value. A match strongly supports byte-for-byte equality of the hashed inputs, not authorship, origin or meaning.

How hashing works

A hash function takes digital data as its input and produces an output of a defined length.

For example, SHA-256 always produces a 256-bit hash value, whether the input is a short text file or a large video. The displayed hexadecimal version contains 64 characters because each character represents four bits.

Hashing is designed to be one-way: the original file cannot practically be reconstructed from the hash value alone.

Comparing like-for-like inputs

If two files produce the same value using the same suitable hash algorithm, that is strong evidence that the bytes supplied to the algorithm were the same.

If the values differ, the inputs differed somewhere. Even a small change - such as one byte, a changed embedded field or a different file wrapper - will normally produce a very different hash.

The comparison only works when:

  • the same algorithm was used;
  • the complete values were recorded accurately; and
  • the same scope of data was hashed.

A hash of an exported file cannot be directly compared with a hash of a whole device image and expected to match. They are different inputs.

Content bytes and metadata

Whether metadata affects a file hash depends on where that metadata is stored.

Metadata embedded within the file forms part of its bytes and therefore affects the hash. File-system metadata stored outside the file - such as some permissions or directory timestamps - will not normally be included when only the file itself is hashed.

Two files can therefore look identical to a person but have different hashes because their hidden or structural bytes differ. Conversely, two identical copies can have the same file hash while existing under different filenames or in different folders.

Algorithms matter

Common names include MD5, SHA-1, SHA-256 and SHA-512. They do not produce interchangeable values.

MD5 and SHA-1 have known collision weaknesses and should not be relied on for modern cryptographic security. They may still appear in older systems, tools and datasets. Recording the algorithm alongside the value is essential.

For straightforward file comparison, a modern algorithm such as SHA-256 is commonly used. Local procedures or specialist systems may require particular algorithms or more than one value.

A hash is not guaranteed to be unique

A hash output has a fixed number of possible values while potential inputs are unlimited. In theory, different inputs must sometimes produce the same value. This is called a collision.

Secure hash functions are designed to make finding a useful collision computationally impractical. It is still more accurate to describe a hash as a strong content-comparison value than as a guaranteed unique identity.

What a file hash can and cannot prove

A matching hash can support the conclusion that the hashed bytes match. It does not by itself prove:

  • who created or possessed the file;
  • where the file came from;
  • that the filename or extension is correct;
  • that the content is safe or lawful;
  • when the file was created; or
  • that the hashing process used the intended input.

Those questions require provenance, handling records and other evidence.

Other uses of the word “hash”

The word appears in several different contexts.

  • A password hash is derived from password-related input and used within authentication systems. Pass-the-hash is a particular attack technique; it is not ordinary file hashing.
  • A transaction hash identifies a blockchain transaction within a particular network. It does not mean that the transaction itself is a file.
  • A threat-intelligence service may use file hashes as lookup identifiers for known or previously analysed files.

The surrounding system determines what was hashed and what the value represents.

The point to remember

A file hash is a reproducible value derived from file bytes. It is powerful for comparing content, but it is not proof of authorship, origin or meaning.

Explore related guidance

Go deeper

Investigator First - back to the investigation

Reference: LOG-156Logs, Records & Provider Evidence