aboutsummaryrefslogtreecommitdiff
path: root/README.md
diff options
context:
space:
mode:
Diffstat (limited to '')
-rw-r--r--README.md382
1 files changed, 382 insertions, 0 deletions
diff --git a/README.md b/README.md
new file mode 100644
index 0000000..c93b843
--- /dev/null
+++ b/README.md
@@ -0,0 +1,382 @@
+<!--
+SPDX-FileCopyrightText: 2026 Dennis Fink <me+coding@dennisfink.me>
+
+SPDX-License-Identifier: BSD-3-Clause
+-->
+
+# duplicate-finder
+
+A command-line tool for finding duplicate files using metadata grouping and
+content hashes.
+
+Files are discovered recursively and first grouped by size and, by default,
+file extension. Paths in each candidate set are also grouped by filesystem
+identity so hard links can be handled efficiently. Only remaining candidates
+are hashed to identify files with matching contents.
+
+## Installation
+
+```bash
+pip install duplicate-finder
+```
+
+## Usage
+
+```text
+Usage: duplicate-finder [OPTIONS] [PATHS]...
+
+ Scan paths and list duplicate files.
+
+ Files are first grouped into candidate sets by size and, by default,
+ case-insensitive file extension. --exclude-file-extension groups candidates
+ by size only. Candidate files are then hashed to identify matching content.
+
+Options:
+ --read-from FILE Read files and directories to scan from FILE,
+ one path per line. Use '-' for standard
+ input. When set, positional paths are
+ ignored.
+ --follow-symlinks / --no-follow-symlinks
+ Follow symlinks when scanning directories.
+ -H, --include-hardlinks / --exclude-hardlinks
+ Include hard links in duplicate results, so
+ multiple paths to the same underlying file
+ may be listed as duplicates.
+ -G, --minsize SIZE Consider only files >= SIZE bytes. Supports
+ suffixes such as 10M, 1G, and 500K.
+ -L, --maxsize SIZE Consider only files <= SIZE bytes. Supports
+ suffixes such as 10M, 1G, and 500K.
+ --include-empty / --exclude-empty
+ Include files that are 0 bytes in size.
+ --include-file-extension / --exclude-file-extension
+ Include file extensions when grouping
+ duplicate candidates. Enabled by default;
+ excluding extensions groups candidates by
+ size only.
+ --hash HASH Select which hash function to use.
+ [default: xxh128]
+ -j, --jobs INTEGER Maximum number of worker threads
+ (0 = use default).
+ -o, --output-format [human|plain|table|json]
+ Select the output format for duplicate
+ groups. 'human' shows a readable text list,
+ 'plain' prints each duplicate group as a
+ space-separated line, 'table' prints a
+ column-aligned overview, and 'json' outputs
+ machine-readable structured data.
+ [default: human]
+ -S, --size Show size of duplicate files
+ (human/plain/table output only).
+ --human-readable / --no-human-readable
+ Display file sizes in human-readable form.
+ -q, --quiet Suppress all non-error output.
+ -v, --verbose Enable verbose output.
+ --color [auto|always|never] Control colorized output.
+ --no-color Disable colorized output.
+ --version Show the version and exit.
+ -h, --help, -? Show this message and exit.
+```
+
+If no path is specified, the current directory is scanned.
+
+## Duplicate detection
+
+Duplicate detection is split into inexpensive candidate selection followed by
+content hashing.
+
+Files are initially grouped by:
+
+- file size
+- case-insensitive file extension
+
+Groups that cannot contain duplicates are discarded before their contents are
+hashed.
+
+File extensions can be excluded from candidate grouping with
+`--exclude-file-extension`. In that case, files are grouped by size alone. This
+allows identical files with different or incorrect extensions to be detected,
+at the cost of potentially hashing more files.
+
+Hard links are recognized by filesystem identity. By default, multiple paths
+referring to the same file are treated as one file and are not reported as
+duplicates of each other. Use `--include-hardlinks` to include those paths in
+duplicate groups.
+
+## Input
+
+One or more files or directories can be supplied as positional arguments:
+
+```bash
+duplicate-finder ~/Documents ~/Downloads
+```
+
+Directories are scanned recursively.
+
+If no paths are supplied, `duplicate-finder` scans the current directory:
+
+```bash
+duplicate-finder
+```
+
+Paths can also be read from a file using `--read-from`. The file must contain
+one path per line:
+
+```bash
+duplicate-finder --read-from paths.txt
+```
+
+Use `-` to read paths from standard input:
+
+```bash
+find /data -type d -print | duplicate-finder --read-from -
+```
+
+When `--read-from` is used, positional paths are ignored.
+
+## File selection
+
+Empty files are excluded by default. Include them with:
+
+```bash
+duplicate-finder --include-empty
+```
+
+Limit duplicate detection to files at least a certain size:
+
+```bash
+duplicate-finder --minsize 10M /data
+```
+
+Or set an upper limit:
+
+```bash
+duplicate-finder --maxsize 1G /data
+```
+
+The two options can be combined:
+
+```bash
+duplicate-finder --minsize 10M --maxsize 1G /data
+```
+
+Size suffixes such as `K`, `M`, and `G` are supported.
+
+Symlinks are not followed by default. Enable symlink traversal with:
+
+```bash
+duplicate-finder --follow-symlinks /data
+```
+
+### Additional filtering
+
+`duplicate-finder` intentionally provides only a small set of file-selection
+options directly. For other criteria, use an appropriate external tool to
+select the paths and pass them to `duplicate-finder` with `--read-from`.
+
+For example, use `find` to consider only files modified within the last 30
+days:
+
+```bash
+find /data -type f -mtime -30 -print | duplicate-finder --read-from -
+```
+
+The same approach can be used for filtering by filename, modification time,
+ownership, permissions, or other filesystem metadata supported by the tool
+producing the path list.
+
+`--read-from` accepts one path per line, either from a file or from standard
+input.
+
+## Hash functions
+
+The hash function used for content comparison can be selected with `--hash`.
+
+The default is `xxh128`:
+
+```bash
+duplicate-finder --hash xxh128 /data
+```
+
+`xxh32` and `xxh64` are also supported, together with the hash algorithms
+available through Python's `hashlib` implementation.
+
+For example:
+
+```bash
+duplicate-finder --hash sha256 /data
+```
+
+## Parallel hashing
+
+Candidate files can be hashed using multiple worker threads.
+
+By default, `--jobs=0` lets the application use its default worker count:
+
+```bash
+duplicate-finder /data
+```
+
+Set an explicit maximum number of workers with:
+
+```bash
+duplicate-finder --jobs 4 /data
+```
+
+## Output formats
+
+Four output formats are available.
+
+### Human
+
+`human` is the default format and is intended for interactive use:
+
+```bash
+duplicate-finder /data
+```
+
+### Plain
+
+`plain` prints each duplicate group as a space-separated line and is intended
+for further command-line processing:
+
+```bash
+duplicate-finder --output-format plain /data
+```
+
+The short form is:
+
+```bash
+duplicate-finder -o plain /data
+```
+
+### Table
+
+`table` displays duplicate groups in a column-aligned overview:
+
+```bash
+duplicate-finder --output-format table /data
+```
+
+### JSON
+
+`json` provides machine-readable structured output:
+
+```bash
+duplicate-finder --output-format json /data
+```
+
+For example, it can be passed directly to `jq`:
+
+```bash
+duplicate-finder -o json /data | jq
+```
+
+## File sizes
+
+Use `--size` to include the size of duplicate files in human, plain, and table
+output:
+
+```bash
+duplicate-finder --size /data
+```
+
+The short form is:
+
+```bash
+duplicate-finder -S /data
+```
+
+Combine it with `--human-readable` to display sizes in a human-readable form:
+
+```bash
+duplicate-finder --size --human-readable /data
+```
+
+## Examples
+
+Find duplicates in the current directory:
+
+```bash
+duplicate-finder
+```
+
+Scan several locations:
+
+```bash
+duplicate-finder ~/Documents ~/Downloads
+```
+
+Detect identical files even when their extensions differ:
+
+```bash
+duplicate-finder --exclude-file-extension /data
+```
+
+Include hard links in duplicate groups:
+
+```bash
+duplicate-finder --include-hardlinks /data
+```
+
+Include empty files:
+
+```bash
+duplicate-finder --include-empty /data
+```
+
+Search only files between 1 MiB and 1 GiB:
+
+```bash
+duplicate-finder --minsize 1M --maxsize 1G /data
+```
+
+Use SHA-256 instead of the default xxh128 hash:
+
+```bash
+duplicate-finder --hash sha256 /data
+```
+
+Hash candidates using at most eight worker threads:
+
+```bash
+duplicate-finder --jobs 8 /data
+```
+
+Produce plain output including human-readable file sizes:
+
+```bash
+duplicate-finder --output-format plain --size --human-readable /data
+```
+
+Read the paths to scan from standard input:
+
+```bash
+printf '%s\n' ~/Documents ~/Downloads | duplicate-finder --read-from -
+```
+
+## Exit status
+
+The command exits successfully after duplicate results have been rendered.
+
+If no duplicates are found, non-JSON output formats print:
+
+```text
+No duplicates found!
+```
+
+and exit with status `2`.
+
+JSON output instead renders the empty result normally.
+
+## Environment
+
+`duplicate-finder` respects the
+[NO_COLOR](https://no-color.org/) convention.
+
+Set `DEBUG` to `1`, `true`, or `yes` to enable diagnostic output describing the
+effective configuration, candidate grouping, duplicate detection, and output
+rendering.
+
+## License
+
+BSD-3-Clause. See [LICENSE](LICENSES/BSD-3-Clause.txt) for details.