# duplicate-finder A command-line tool for finding duplicate files using metadata grouping and content hashes. Files are discovered recursively and first grouped by size and, by default, file extension. Paths in each candidate set are also grouped by filesystem identity so hard links can be handled efficiently. Only remaining candidates are hashed to identify files with matching contents. ## Installation ```bash pip install duplicate-finder ``` ## Usage ```text Usage: duplicate-finder [OPTIONS] [PATHS]... Scan paths and list duplicate files. Files are first grouped into candidate sets by size and, by default, case-insensitive file extension. --exclude-file-extension groups candidates by size only. Candidate files are then hashed to identify matching content. Options: --read-from FILE Read files and directories to scan from FILE, one path per line. Use '-' for standard input. When set, positional paths are ignored. --follow-symlinks / --no-follow-symlinks Follow symlinks when scanning directories. -H, --include-hardlinks / --exclude-hardlinks Include hard links in duplicate results, so multiple paths to the same underlying file may be listed as duplicates. -G, --minsize SIZE Consider only files >= SIZE bytes. Supports suffixes such as 10M, 1G, and 500K. -L, --maxsize SIZE Consider only files <= SIZE bytes. Supports suffixes such as 10M, 1G, and 500K. --include-empty / --exclude-empty Include files that are 0 bytes in size. --include-file-extension / --exclude-file-extension Include file extensions when grouping duplicate candidates. Enabled by default; excluding extensions groups candidates by size only. --hash HASH Select which hash function to use. [default: xxh128] -j, --jobs INTEGER Maximum number of worker threads (0 = use default). -o, --output-format [human|plain|table|json] Select the output format for duplicate groups. 'human' shows a readable text list, 'plain' prints each duplicate group as a space-separated line, 'table' prints a column-aligned overview, and 'json' outputs machine-readable structured data. [default: human] -S, --size Show size of duplicate files (human/plain/table output only). --human-readable / --no-human-readable Display file sizes in human-readable form. -q, --quiet Suppress all non-error output. -v, --verbose Enable verbose output. --color [auto|always|never] Control colorized output. --no-color Disable colorized output. --version Show the version and exit. -h, --help, -? Show this message and exit. ``` If no path is specified, the current directory is scanned. ## Duplicate detection Duplicate detection is split into inexpensive candidate selection followed by content hashing. Files are initially grouped by: - file size - case-insensitive file extension Groups that cannot contain duplicates are discarded before their contents are hashed. File extensions can be excluded from candidate grouping with `--exclude-file-extension`. In that case, files are grouped by size alone. This allows identical files with different or incorrect extensions to be detected, at the cost of potentially hashing more files. Hard links are recognized by filesystem identity. By default, multiple paths referring to the same file are treated as one file and are not reported as duplicates of each other. Use `--include-hardlinks` to include those paths in duplicate groups. ## Input One or more files or directories can be supplied as positional arguments: ```bash duplicate-finder ~/Documents ~/Downloads ``` Directories are scanned recursively. If no paths are supplied, `duplicate-finder` scans the current directory: ```bash duplicate-finder ``` Paths can also be read from a file using `--read-from`. The file must contain one path per line: ```bash duplicate-finder --read-from paths.txt ``` Use `-` to read paths from standard input: ```bash find /data -type d -print | duplicate-finder --read-from - ``` When `--read-from` is used, positional paths are ignored. ## File selection Empty files are excluded by default. Include them with: ```bash duplicate-finder --include-empty ``` Limit duplicate detection to files at least a certain size: ```bash duplicate-finder --minsize 10M /data ``` Or set an upper limit: ```bash duplicate-finder --maxsize 1G /data ``` The two options can be combined: ```bash duplicate-finder --minsize 10M --maxsize 1G /data ``` Size suffixes such as `K`, `M`, and `G` are supported. Symlinks are not followed by default. Enable symlink traversal with: ```bash duplicate-finder --follow-symlinks /data ``` ### Additional filtering `duplicate-finder` intentionally provides only a small set of file-selection options directly. For other criteria, use an appropriate external tool to select the paths and pass them to `duplicate-finder` with `--read-from`. For example, use `find` to consider only files modified within the last 30 days: ```bash find /data -type f -mtime -30 -print | duplicate-finder --read-from - ``` The same approach can be used for filtering by filename, modification time, ownership, permissions, or other filesystem metadata supported by the tool producing the path list. `--read-from` accepts one path per line, either from a file or from standard input. ## Hash functions The hash function used for content comparison can be selected with `--hash`. The default is `xxh128`: ```bash duplicate-finder --hash xxh128 /data ``` `xxh32` and `xxh64` are also supported, together with the hash algorithms available through Python's `hashlib` implementation. For example: ```bash duplicate-finder --hash sha256 /data ``` ## Parallel hashing Candidate files can be hashed using multiple worker threads. By default, `--jobs=0` lets the application use its default worker count: ```bash duplicate-finder /data ``` Set an explicit maximum number of workers with: ```bash duplicate-finder --jobs 4 /data ``` ## Output formats Four output formats are available. ### Human `human` is the default format and is intended for interactive use: ```bash duplicate-finder /data ``` ### Plain `plain` prints each duplicate group as a space-separated line and is intended for further command-line processing: ```bash duplicate-finder --output-format plain /data ``` The short form is: ```bash duplicate-finder -o plain /data ``` ### Table `table` displays duplicate groups in a column-aligned overview: ```bash duplicate-finder --output-format table /data ``` ### JSON `json` provides machine-readable structured output: ```bash duplicate-finder --output-format json /data ``` For example, it can be passed directly to `jq`: ```bash duplicate-finder -o json /data | jq ``` ## File sizes Use `--size` to include the size of duplicate files in human, plain, and table output: ```bash duplicate-finder --size /data ``` The short form is: ```bash duplicate-finder -S /data ``` Combine it with `--human-readable` to display sizes in a human-readable form: ```bash duplicate-finder --size --human-readable /data ``` ## Examples Find duplicates in the current directory: ```bash duplicate-finder ``` Scan several locations: ```bash duplicate-finder ~/Documents ~/Downloads ``` Detect identical files even when their extensions differ: ```bash duplicate-finder --exclude-file-extension /data ``` Include hard links in duplicate groups: ```bash duplicate-finder --include-hardlinks /data ``` Include empty files: ```bash duplicate-finder --include-empty /data ``` Search only files between 1 MiB and 1 GiB: ```bash duplicate-finder --minsize 1M --maxsize 1G /data ``` Use SHA-256 instead of the default xxh128 hash: ```bash duplicate-finder --hash sha256 /data ``` Hash candidates using at most eight worker threads: ```bash duplicate-finder --jobs 8 /data ``` Produce plain output including human-readable file sizes: ```bash duplicate-finder --output-format plain --size --human-readable /data ``` Read the paths to scan from standard input: ```bash printf '%s\n' ~/Documents ~/Downloads | duplicate-finder --read-from - ``` ## Exit status The command exits successfully after duplicate results have been rendered. If no duplicates are found, non-JSON output formats print: ```text No duplicates found! ``` and exit with status `2`. JSON output instead renders the empty result normally. ## Environment `duplicate-finder` respects the [NO_COLOR](https://no-color.org/) convention. Set `DEBUG` to `1`, `true`, or `yes` to enable diagnostic output describing the effective configuration, candidate grouping, duplicate detection, and output rendering. ## License BSD-3-Clause. See [LICENSE](LICENSES/BSD-3-Clause.txt) for details.