duplicate-finder
A command-line tool for finding duplicate files using metadata grouping and content hashes.
Files are discovered recursively and first grouped by size and, by default, file extension. Paths in each candidate set are also grouped by filesystem identity so hard links can be handled efficiently. Only remaining candidates are hashed to identify files with matching contents.
Installation
pip install duplicate-finder
Usage
Usage: duplicate-finder [OPTIONS] [PATHS]...
Scan paths and list duplicate files.
Files are first grouped into candidate sets by size and, by default,
case-insensitive file extension. --exclude-file-extension groups candidates
by size only. Candidate files are then hashed to identify matching content.
Options:
--read-from FILE Read files and directories to scan from FILE,
one path per line. Use '-' for standard
input. When set, positional paths are
ignored.
--follow-symlinks / --no-follow-symlinks
Follow symlinks when scanning directories.
-H, --include-hardlinks / --exclude-hardlinks
Include hard links in duplicate results, so
multiple paths to the same underlying file
may be listed as duplicates.
-G, --minsize SIZE Consider only files >= SIZE bytes. Supports
suffixes such as 10M, 1G, and 500K.
-L, --maxsize SIZE Consider only files <= SIZE bytes. Supports
suffixes such as 10M, 1G, and 500K.
--include-empty / --exclude-empty
Include files that are 0 bytes in size.
--include-file-extension / --exclude-file-extension
Include file extensions when grouping
duplicate candidates. Enabled by default;
excluding extensions groups candidates by
size only.
--hash HASH Select which hash function to use.
[default: xxh128]
-j, --jobs INTEGER Maximum number of worker threads
(0 = use default).
-o, --output-format [human|plain|table|json]
Select the output format for duplicate
groups. 'human' shows a readable text list,
'plain' prints each duplicate group as a
space-separated line, 'table' prints a
column-aligned overview, and 'json' outputs
machine-readable structured data.
[default: human]
-S, --size Show size of duplicate files
(human/plain/table output only).
--human-readable / --no-human-readable
Display file sizes in human-readable form.
-q, --quiet Suppress all non-error output.
-v, --verbose Enable verbose output.
--color [auto|always|never] Control colorized output.
--no-color Disable colorized output.
--version Show the version and exit.
-h, --help, -? Show this message and exit.
If no path is specified, the current directory is scanned.
Duplicate detection
Duplicate detection is split into inexpensive candidate selection followed by content hashing.
Files are initially grouped by:
- file size
- case-insensitive file extension
Groups that cannot contain duplicates are discarded before their contents are hashed.
File extensions can be excluded from candidate grouping with
--exclude-file-extension. In that case, files are grouped by size alone. This
allows identical files with different or incorrect extensions to be detected,
at the cost of potentially hashing more files.
Hard links are recognized by filesystem identity. By default, multiple paths
referring to the same file are treated as one file and are not reported as
duplicates of each other. Use --include-hardlinks to include those paths in
duplicate groups.
Input
One or more files or directories can be supplied as positional arguments:
duplicate-finder ~/Documents ~/Downloads
Directories are scanned recursively.
If no paths are supplied, duplicate-finder scans the current directory:
duplicate-finder
Paths can also be read from a file using --read-from. The file must contain
one path per line:
duplicate-finder --read-from paths.txt
Use - to read paths from standard input:
find /data -type d -print | duplicate-finder --read-from -
When --read-from is used, positional paths are ignored.
File selection
Empty files are excluded by default. Include them with:
duplicate-finder --include-empty
Limit duplicate detection to files at least a certain size:
duplicate-finder --minsize 10M /data
Or set an upper limit:
duplicate-finder --maxsize 1G /data
The two options can be combined:
duplicate-finder --minsize 10M --maxsize 1G /data
Size suffixes such as K, M, and G are supported.
Symlinks are not followed by default. Enable symlink traversal with:
duplicate-finder --follow-symlinks /data
Additional filtering
duplicate-finder intentionally provides only a small set of file-selection
options directly. For other criteria, use an appropriate external tool to
select the paths and pass them to duplicate-finder with --read-from.
For example, use find to consider only files modified within the last 30
days:
find /data -type f -mtime -30 -print | duplicate-finder --read-from -
The same approach can be used for filtering by filename, modification time, ownership, permissions, or other filesystem metadata supported by the tool producing the path list.
--read-from accepts one path per line, either from a file or from standard
input.
Hash functions
The hash function used for content comparison can be selected with --hash.
The default is xxh128:
duplicate-finder --hash xxh128 /data
xxh32 and xxh64 are also supported, together with the hash algorithms
available through Python's hashlib implementation.
For example:
duplicate-finder --hash sha256 /data
Parallel hashing
Candidate files can be hashed using multiple worker threads.
By default, --jobs=0 lets the application use its default worker count:
duplicate-finder /data
Set an explicit maximum number of workers with:
duplicate-finder --jobs 4 /data
Output formats
Four output formats are available.
Human
human is the default format and is intended for interactive use:
duplicate-finder /data
Plain
plain prints each duplicate group as a space-separated line and is intended
for further command-line processing:
duplicate-finder --output-format plain /data
The short form is:
duplicate-finder -o plain /data
Table
table displays duplicate groups in a column-aligned overview:
duplicate-finder --output-format table /data
JSON
json provides machine-readable structured output:
duplicate-finder --output-format json /data
For example, it can be passed directly to jq:
duplicate-finder -o json /data | jq
File sizes
Use --size to include the size of duplicate files in human, plain, and table
output:
duplicate-finder --size /data
The short form is:
duplicate-finder -S /data
Combine it with --human-readable to display sizes in a human-readable form:
duplicate-finder --size --human-readable /data
Examples
Find duplicates in the current directory:
duplicate-finder
Scan several locations:
duplicate-finder ~/Documents ~/Downloads
Detect identical files even when their extensions differ:
duplicate-finder --exclude-file-extension /data
Include hard links in duplicate groups:
duplicate-finder --include-hardlinks /data
Include empty files:
duplicate-finder --include-empty /data
Search only files between 1 MiB and 1 GiB:
duplicate-finder --minsize 1M --maxsize 1G /data
Use SHA-256 instead of the default xxh128 hash:
duplicate-finder --hash sha256 /data
Hash candidates using at most eight worker threads:
duplicate-finder --jobs 8 /data
Produce plain output including human-readable file sizes:
duplicate-finder --output-format plain --size --human-readable /data
Read the paths to scan from standard input:
printf '%s\n' ~/Documents ~/Downloads | duplicate-finder --read-from -
Exit status
The command exits successfully after duplicate results have been rendered.
If no duplicates are found, non-JSON output formats print:
No duplicates found!
and exit with status 2.
JSON output instead renders the empty result normally.
Environment
duplicate-finder respects the
NO_COLOR convention.
Set DEBUG to 1, true, or yes to enable diagnostic output describing the
effective configuration, candidate grouping, duplicate detection, and output
rendering.
License
BSD-3-Clause. See LICENSE for details.
