aboutsummaryrefslogtreecommitdiff

duplicate-finder

A command-line tool for finding duplicate files using metadata grouping and content hashes.

Files are discovered recursively and first grouped by size and, by default, file extension. Paths in each candidate set are also grouped by filesystem identity so hard links can be handled efficiently. Only remaining candidates are hashed to identify files with matching contents.

Installation

pip install duplicate-finder

Usage

Usage: duplicate-finder [OPTIONS] [PATHS]...

  Scan paths and list duplicate files.

  Files are first grouped into candidate sets by size and, by default,
  case-insensitive file extension. --exclude-file-extension groups candidates
  by size only. Candidate files are then hashed to identify matching content.

Options:
  --read-from FILE                Read files and directories to scan from FILE,
                                  one path per line. Use '-' for standard
                                  input. When set, positional paths are
                                  ignored.
  --follow-symlinks / --no-follow-symlinks
                                  Follow symlinks when scanning directories.
  -H, --include-hardlinks / --exclude-hardlinks
                                  Include hard links in duplicate results, so
                                  multiple paths to the same underlying file
                                  may be listed as duplicates.
  -G, --minsize SIZE              Consider only files >= SIZE bytes. Supports
                                  suffixes such as 10M, 1G, and 500K.
  -L, --maxsize SIZE              Consider only files <= SIZE bytes. Supports
                                  suffixes such as 10M, 1G, and 500K.
  --include-empty / --exclude-empty
                                  Include files that are 0 bytes in size.
  --include-file-extension / --exclude-file-extension
                                  Include file extensions when grouping
                                  duplicate candidates. Enabled by default;
                                  excluding extensions groups candidates by
                                  size only.
  --hash HASH                     Select which hash function to use.
                                  [default: xxh128]
  -j, --jobs INTEGER              Maximum number of worker threads
                                  (0 = use default).
  -o, --output-format [human|plain|table|json]
                                  Select the output format for duplicate
                                  groups. 'human' shows a readable text list,
                                  'plain' prints each duplicate group as a
                                  space-separated line, 'table' prints a
                                  column-aligned overview, and 'json' outputs
                                  machine-readable structured data.
                                  [default: human]
  -S, --size                      Show size of duplicate files
                                  (human/plain/table output only).
  --human-readable / --no-human-readable
                                  Display file sizes in human-readable form.
  -q, --quiet                     Suppress all non-error output.
  -v, --verbose                   Enable verbose output.
  --color [auto|always|never]     Control colorized output.
  --no-color                      Disable colorized output.
  --version                       Show the version and exit.
  -h, --help, -?                  Show this message and exit.

If no path is specified, the current directory is scanned.

Duplicate detection

Duplicate detection is split into inexpensive candidate selection followed by content hashing.

Files are initially grouped by:

  • file size
  • case-insensitive file extension

Groups that cannot contain duplicates are discarded before their contents are hashed.

File extensions can be excluded from candidate grouping with --exclude-file-extension. In that case, files are grouped by size alone. This allows identical files with different or incorrect extensions to be detected, at the cost of potentially hashing more files.

Hard links are recognized by filesystem identity. By default, multiple paths referring to the same file are treated as one file and are not reported as duplicates of each other. Use --include-hardlinks to include those paths in duplicate groups.

Input

One or more files or directories can be supplied as positional arguments:

duplicate-finder ~/Documents ~/Downloads

Directories are scanned recursively.

If no paths are supplied, duplicate-finder scans the current directory:

duplicate-finder

Paths can also be read from a file using --read-from. The file must contain one path per line:

duplicate-finder --read-from paths.txt

Use - to read paths from standard input:

find /data -type d -print | duplicate-finder --read-from -

When --read-from is used, positional paths are ignored.

File selection

Empty files are excluded by default. Include them with:

duplicate-finder --include-empty

Limit duplicate detection to files at least a certain size:

duplicate-finder --minsize 10M /data

Or set an upper limit:

duplicate-finder --maxsize 1G /data

The two options can be combined:

duplicate-finder --minsize 10M --maxsize 1G /data

Size suffixes such as K, M, and G are supported.

Symlinks are not followed by default. Enable symlink traversal with:

duplicate-finder --follow-symlinks /data

Additional filtering

duplicate-finder intentionally provides only a small set of file-selection options directly. For other criteria, use an appropriate external tool to select the paths and pass them to duplicate-finder with --read-from.

For example, use find to consider only files modified within the last 30 days:

find /data -type f -mtime -30 -print | duplicate-finder --read-from -

The same approach can be used for filtering by filename, modification time, ownership, permissions, or other filesystem metadata supported by the tool producing the path list.

--read-from accepts one path per line, either from a file or from standard input.

Hash functions

The hash function used for content comparison can be selected with --hash.

The default is xxh128:

duplicate-finder --hash xxh128 /data

xxh32 and xxh64 are also supported, together with the hash algorithms available through Python's hashlib implementation.

For example:

duplicate-finder --hash sha256 /data

Parallel hashing

Candidate files can be hashed using multiple worker threads.

By default, --jobs=0 lets the application use its default worker count:

duplicate-finder /data

Set an explicit maximum number of workers with:

duplicate-finder --jobs 4 /data

Output formats

Four output formats are available.

Human

human is the default format and is intended for interactive use:

duplicate-finder /data

Plain

plain prints each duplicate group as a space-separated line and is intended for further command-line processing:

duplicate-finder --output-format plain /data

The short form is:

duplicate-finder -o plain /data

Table

table displays duplicate groups in a column-aligned overview:

duplicate-finder --output-format table /data

JSON

json provides machine-readable structured output:

duplicate-finder --output-format json /data

For example, it can be passed directly to jq:

duplicate-finder -o json /data | jq

File sizes

Use --size to include the size of duplicate files in human, plain, and table output:

duplicate-finder --size /data

The short form is:

duplicate-finder -S /data

Combine it with --human-readable to display sizes in a human-readable form:

duplicate-finder --size --human-readable /data

Examples

Find duplicates in the current directory:

duplicate-finder

Scan several locations:

duplicate-finder ~/Documents ~/Downloads

Detect identical files even when their extensions differ:

duplicate-finder --exclude-file-extension /data

Include hard links in duplicate groups:

duplicate-finder --include-hardlinks /data

Include empty files:

duplicate-finder --include-empty /data

Search only files between 1 MiB and 1 GiB:

duplicate-finder --minsize 1M --maxsize 1G /data

Use SHA-256 instead of the default xxh128 hash:

duplicate-finder --hash sha256 /data

Hash candidates using at most eight worker threads:

duplicate-finder --jobs 8 /data

Produce plain output including human-readable file sizes:

duplicate-finder --output-format plain --size --human-readable /data

Read the paths to scan from standard input:

printf '%s\n' ~/Documents ~/Downloads | duplicate-finder --read-from -

Exit status

The command exits successfully after duplicate results have been rendered.

If no duplicates are found, non-JSON output formats print:

No duplicates found!

and exit with status 2.

JSON output instead renders the empty result normally.

Environment

duplicate-finder respects the NO_COLOR convention.

Set DEBUG to 1, true, or yes to enable diagnostic output describing the effective configuration, candidate grouping, duplicate detection, and output rendering.

License

BSD-3-Clause. See LICENSE for details.