diff options
Diffstat (limited to '')
| -rw-r--r-- | README.md | 382 |
1 files changed, 382 insertions, 0 deletions
diff --git a/README.md b/README.md new file mode 100644 index 0000000..c93b843 --- /dev/null +++ b/README.md @@ -0,0 +1,382 @@ +<!-- +SPDX-FileCopyrightText: 2026 Dennis Fink <me+coding@dennisfink.me> + +SPDX-License-Identifier: BSD-3-Clause +--> + +# duplicate-finder + +A command-line tool for finding duplicate files using metadata grouping and +content hashes. + +Files are discovered recursively and first grouped by size and, by default, +file extension. Paths in each candidate set are also grouped by filesystem +identity so hard links can be handled efficiently. Only remaining candidates +are hashed to identify files with matching contents. + +## Installation + +```bash +pip install duplicate-finder +``` + +## Usage + +```text +Usage: duplicate-finder [OPTIONS] [PATHS]... + + Scan paths and list duplicate files. + + Files are first grouped into candidate sets by size and, by default, + case-insensitive file extension. --exclude-file-extension groups candidates + by size only. Candidate files are then hashed to identify matching content. + +Options: + --read-from FILE Read files and directories to scan from FILE, + one path per line. Use '-' for standard + input. When set, positional paths are + ignored. + --follow-symlinks / --no-follow-symlinks + Follow symlinks when scanning directories. + -H, --include-hardlinks / --exclude-hardlinks + Include hard links in duplicate results, so + multiple paths to the same underlying file + may be listed as duplicates. + -G, --minsize SIZE Consider only files >= SIZE bytes. Supports + suffixes such as 10M, 1G, and 500K. + -L, --maxsize SIZE Consider only files <= SIZE bytes. Supports + suffixes such as 10M, 1G, and 500K. + --include-empty / --exclude-empty + Include files that are 0 bytes in size. + --include-file-extension / --exclude-file-extension + Include file extensions when grouping + duplicate candidates. Enabled by default; + excluding extensions groups candidates by + size only. + --hash HASH Select which hash function to use. + [default: xxh128] + -j, --jobs INTEGER Maximum number of worker threads + (0 = use default). + -o, --output-format [human|plain|table|json] + Select the output format for duplicate + groups. 'human' shows a readable text list, + 'plain' prints each duplicate group as a + space-separated line, 'table' prints a + column-aligned overview, and 'json' outputs + machine-readable structured data. + [default: human] + -S, --size Show size of duplicate files + (human/plain/table output only). + --human-readable / --no-human-readable + Display file sizes in human-readable form. + -q, --quiet Suppress all non-error output. + -v, --verbose Enable verbose output. + --color [auto|always|never] Control colorized output. + --no-color Disable colorized output. + --version Show the version and exit. + -h, --help, -? Show this message and exit. +``` + +If no path is specified, the current directory is scanned. + +## Duplicate detection + +Duplicate detection is split into inexpensive candidate selection followed by +content hashing. + +Files are initially grouped by: + +- file size +- case-insensitive file extension + +Groups that cannot contain duplicates are discarded before their contents are +hashed. + +File extensions can be excluded from candidate grouping with +`--exclude-file-extension`. In that case, files are grouped by size alone. This +allows identical files with different or incorrect extensions to be detected, +at the cost of potentially hashing more files. + +Hard links are recognized by filesystem identity. By default, multiple paths +referring to the same file are treated as one file and are not reported as +duplicates of each other. Use `--include-hardlinks` to include those paths in +duplicate groups. + +## Input + +One or more files or directories can be supplied as positional arguments: + +```bash +duplicate-finder ~/Documents ~/Downloads +``` + +Directories are scanned recursively. + +If no paths are supplied, `duplicate-finder` scans the current directory: + +```bash +duplicate-finder +``` + +Paths can also be read from a file using `--read-from`. The file must contain +one path per line: + +```bash +duplicate-finder --read-from paths.txt +``` + +Use `-` to read paths from standard input: + +```bash +find /data -type d -print | duplicate-finder --read-from - +``` + +When `--read-from` is used, positional paths are ignored. + +## File selection + +Empty files are excluded by default. Include them with: + +```bash +duplicate-finder --include-empty +``` + +Limit duplicate detection to files at least a certain size: + +```bash +duplicate-finder --minsize 10M /data +``` + +Or set an upper limit: + +```bash +duplicate-finder --maxsize 1G /data +``` + +The two options can be combined: + +```bash +duplicate-finder --minsize 10M --maxsize 1G /data +``` + +Size suffixes such as `K`, `M`, and `G` are supported. + +Symlinks are not followed by default. Enable symlink traversal with: + +```bash +duplicate-finder --follow-symlinks /data +``` + +### Additional filtering + +`duplicate-finder` intentionally provides only a small set of file-selection +options directly. For other criteria, use an appropriate external tool to +select the paths and pass them to `duplicate-finder` with `--read-from`. + +For example, use `find` to consider only files modified within the last 30 +days: + +```bash +find /data -type f -mtime -30 -print | duplicate-finder --read-from - +``` + +The same approach can be used for filtering by filename, modification time, +ownership, permissions, or other filesystem metadata supported by the tool +producing the path list. + +`--read-from` accepts one path per line, either from a file or from standard +input. + +## Hash functions + +The hash function used for content comparison can be selected with `--hash`. + +The default is `xxh128`: + +```bash +duplicate-finder --hash xxh128 /data +``` + +`xxh32` and `xxh64` are also supported, together with the hash algorithms +available through Python's `hashlib` implementation. + +For example: + +```bash +duplicate-finder --hash sha256 /data +``` + +## Parallel hashing + +Candidate files can be hashed using multiple worker threads. + +By default, `--jobs=0` lets the application use its default worker count: + +```bash +duplicate-finder /data +``` + +Set an explicit maximum number of workers with: + +```bash +duplicate-finder --jobs 4 /data +``` + +## Output formats + +Four output formats are available. + +### Human + +`human` is the default format and is intended for interactive use: + +```bash +duplicate-finder /data +``` + +### Plain + +`plain` prints each duplicate group as a space-separated line and is intended +for further command-line processing: + +```bash +duplicate-finder --output-format plain /data +``` + +The short form is: + +```bash +duplicate-finder -o plain /data +``` + +### Table + +`table` displays duplicate groups in a column-aligned overview: + +```bash +duplicate-finder --output-format table /data +``` + +### JSON + +`json` provides machine-readable structured output: + +```bash +duplicate-finder --output-format json /data +``` + +For example, it can be passed directly to `jq`: + +```bash +duplicate-finder -o json /data | jq +``` + +## File sizes + +Use `--size` to include the size of duplicate files in human, plain, and table +output: + +```bash +duplicate-finder --size /data +``` + +The short form is: + +```bash +duplicate-finder -S /data +``` + +Combine it with `--human-readable` to display sizes in a human-readable form: + +```bash +duplicate-finder --size --human-readable /data +``` + +## Examples + +Find duplicates in the current directory: + +```bash +duplicate-finder +``` + +Scan several locations: + +```bash +duplicate-finder ~/Documents ~/Downloads +``` + +Detect identical files even when their extensions differ: + +```bash +duplicate-finder --exclude-file-extension /data +``` + +Include hard links in duplicate groups: + +```bash +duplicate-finder --include-hardlinks /data +``` + +Include empty files: + +```bash +duplicate-finder --include-empty /data +``` + +Search only files between 1 MiB and 1 GiB: + +```bash +duplicate-finder --minsize 1M --maxsize 1G /data +``` + +Use SHA-256 instead of the default xxh128 hash: + +```bash +duplicate-finder --hash sha256 /data +``` + +Hash candidates using at most eight worker threads: + +```bash +duplicate-finder --jobs 8 /data +``` + +Produce plain output including human-readable file sizes: + +```bash +duplicate-finder --output-format plain --size --human-readable /data +``` + +Read the paths to scan from standard input: + +```bash +printf '%s\n' ~/Documents ~/Downloads | duplicate-finder --read-from - +``` + +## Exit status + +The command exits successfully after duplicate results have been rendered. + +If no duplicates are found, non-JSON output formats print: + +```text +No duplicates found! +``` + +and exit with status `2`. + +JSON output instead renders the empty result normally. + +## Environment + +`duplicate-finder` respects the +[NO_COLOR](https://no-color.org/) convention. + +Set `DEBUG` to `1`, `true`, or `yes` to enable diagnostic output describing the +effective configuration, candidate grouping, duplicate detection, and output +rendering. + +## License + +BSD-3-Clause. See [LICENSE](LICENSES/BSD-3-Clause.txt) for details. |
