aboutsummaryrefslogtreecommitdiff
path: root/README.md
diff options
context:
space:
mode:
authorDennis Fink2026-09-20 20:04:13 +0200
committerDennis Fink2026-09-20 20:04:13 +0200
commitb751174b4ed7c23786ae54a1bbaf332f653faa81 (patch)
treec1951961b7bdd171aa4903d9e2d5ad4f2a063b51 /README.md
downloadduplicate-finder-b751174b4ed7c23786ae54a1bbaf332f653faa81.tar.gz
duplicate-finder-b751174b4ed7c23786ae54a1bbaf332f653faa81.zip
feat: release duplicate-finder 1.0.0HEADv1.0.0main
Introduce the first public release of duplicate-finder, a command-line tool for detecting duplicate files through metadata grouping and content hashing. Support recursive path scanning, optional symlink traversal, hard-link handling, size and extension-based candidate grouping, and parallel content hashing with xxHash or hashlib algorithms. Provide human, plain, table, and JSON output formats together with size filtering, path input from files or stdin, progress reporting, diagnostic output, and configurable terminal colors. Add Python packaging for Python 3.14, project documentation, dependency locking, development tooling, and BSD-3-Clause licensing.
Diffstat (limited to 'README.md')
-rw-r--r--README.md382
1 files changed, 382 insertions, 0 deletions
diff --git a/README.md b/README.md
new file mode 100644
index 0000000..c93b843
--- /dev/null
+++ b/README.md
@@ -0,0 +1,382 @@
+<!--
+SPDX-FileCopyrightText: 2026 Dennis Fink <me+coding@dennisfink.me>
+
+SPDX-License-Identifier: BSD-3-Clause
+-->
+
+# duplicate-finder
+
+A command-line tool for finding duplicate files using metadata grouping and
+content hashes.
+
+Files are discovered recursively and first grouped by size and, by default,
+file extension. Paths in each candidate set are also grouped by filesystem
+identity so hard links can be handled efficiently. Only remaining candidates
+are hashed to identify files with matching contents.
+
+## Installation
+
+```bash
+pip install duplicate-finder
+```
+
+## Usage
+
+```text
+Usage: duplicate-finder [OPTIONS] [PATHS]...
+
+ Scan paths and list duplicate files.
+
+ Files are first grouped into candidate sets by size and, by default,
+ case-insensitive file extension. --exclude-file-extension groups candidates
+ by size only. Candidate files are then hashed to identify matching content.
+
+Options:
+ --read-from FILE Read files and directories to scan from FILE,
+ one path per line. Use '-' for standard
+ input. When set, positional paths are
+ ignored.
+ --follow-symlinks / --no-follow-symlinks
+ Follow symlinks when scanning directories.
+ -H, --include-hardlinks / --exclude-hardlinks
+ Include hard links in duplicate results, so
+ multiple paths to the same underlying file
+ may be listed as duplicates.
+ -G, --minsize SIZE Consider only files >= SIZE bytes. Supports
+ suffixes such as 10M, 1G, and 500K.
+ -L, --maxsize SIZE Consider only files <= SIZE bytes. Supports
+ suffixes such as 10M, 1G, and 500K.
+ --include-empty / --exclude-empty
+ Include files that are 0 bytes in size.
+ --include-file-extension / --exclude-file-extension
+ Include file extensions when grouping
+ duplicate candidates. Enabled by default;
+ excluding extensions groups candidates by
+ size only.
+ --hash HASH Select which hash function to use.
+ [default: xxh128]
+ -j, --jobs INTEGER Maximum number of worker threads
+ (0 = use default).
+ -o, --output-format [human|plain|table|json]
+ Select the output format for duplicate
+ groups. 'human' shows a readable text list,
+ 'plain' prints each duplicate group as a
+ space-separated line, 'table' prints a
+ column-aligned overview, and 'json' outputs
+ machine-readable structured data.
+ [default: human]
+ -S, --size Show size of duplicate files
+ (human/plain/table output only).
+ --human-readable / --no-human-readable
+ Display file sizes in human-readable form.
+ -q, --quiet Suppress all non-error output.
+ -v, --verbose Enable verbose output.
+ --color [auto|always|never] Control colorized output.
+ --no-color Disable colorized output.
+ --version Show the version and exit.
+ -h, --help, -? Show this message and exit.
+```
+
+If no path is specified, the current directory is scanned.
+
+## Duplicate detection
+
+Duplicate detection is split into inexpensive candidate selection followed by
+content hashing.
+
+Files are initially grouped by:
+
+- file size
+- case-insensitive file extension
+
+Groups that cannot contain duplicates are discarded before their contents are
+hashed.
+
+File extensions can be excluded from candidate grouping with
+`--exclude-file-extension`. In that case, files are grouped by size alone. This
+allows identical files with different or incorrect extensions to be detected,
+at the cost of potentially hashing more files.
+
+Hard links are recognized by filesystem identity. By default, multiple paths
+referring to the same file are treated as one file and are not reported as
+duplicates of each other. Use `--include-hardlinks` to include those paths in
+duplicate groups.
+
+## Input
+
+One or more files or directories can be supplied as positional arguments:
+
+```bash
+duplicate-finder ~/Documents ~/Downloads
+```
+
+Directories are scanned recursively.
+
+If no paths are supplied, `duplicate-finder` scans the current directory:
+
+```bash
+duplicate-finder
+```
+
+Paths can also be read from a file using `--read-from`. The file must contain
+one path per line:
+
+```bash
+duplicate-finder --read-from paths.txt
+```
+
+Use `-` to read paths from standard input:
+
+```bash
+find /data -type d -print | duplicate-finder --read-from -
+```
+
+When `--read-from` is used, positional paths are ignored.
+
+## File selection
+
+Empty files are excluded by default. Include them with:
+
+```bash
+duplicate-finder --include-empty
+```
+
+Limit duplicate detection to files at least a certain size:
+
+```bash
+duplicate-finder --minsize 10M /data
+```
+
+Or set an upper limit:
+
+```bash
+duplicate-finder --maxsize 1G /data
+```
+
+The two options can be combined:
+
+```bash
+duplicate-finder --minsize 10M --maxsize 1G /data
+```
+
+Size suffixes such as `K`, `M`, and `G` are supported.
+
+Symlinks are not followed by default. Enable symlink traversal with:
+
+```bash
+duplicate-finder --follow-symlinks /data
+```
+
+### Additional filtering
+
+`duplicate-finder` intentionally provides only a small set of file-selection
+options directly. For other criteria, use an appropriate external tool to
+select the paths and pass them to `duplicate-finder` with `--read-from`.
+
+For example, use `find` to consider only files modified within the last 30
+days:
+
+```bash
+find /data -type f -mtime -30 -print | duplicate-finder --read-from -
+```
+
+The same approach can be used for filtering by filename, modification time,
+ownership, permissions, or other filesystem metadata supported by the tool
+producing the path list.
+
+`--read-from` accepts one path per line, either from a file or from standard
+input.
+
+## Hash functions
+
+The hash function used for content comparison can be selected with `--hash`.
+
+The default is `xxh128`:
+
+```bash
+duplicate-finder --hash xxh128 /data
+```
+
+`xxh32` and `xxh64` are also supported, together with the hash algorithms
+available through Python's `hashlib` implementation.
+
+For example:
+
+```bash
+duplicate-finder --hash sha256 /data
+```
+
+## Parallel hashing
+
+Candidate files can be hashed using multiple worker threads.
+
+By default, `--jobs=0` lets the application use its default worker count:
+
+```bash
+duplicate-finder /data
+```
+
+Set an explicit maximum number of workers with:
+
+```bash
+duplicate-finder --jobs 4 /data
+```
+
+## Output formats
+
+Four output formats are available.
+
+### Human
+
+`human` is the default format and is intended for interactive use:
+
+```bash
+duplicate-finder /data
+```
+
+### Plain
+
+`plain` prints each duplicate group as a space-separated line and is intended
+for further command-line processing:
+
+```bash
+duplicate-finder --output-format plain /data
+```
+
+The short form is:
+
+```bash
+duplicate-finder -o plain /data
+```
+
+### Table
+
+`table` displays duplicate groups in a column-aligned overview:
+
+```bash
+duplicate-finder --output-format table /data
+```
+
+### JSON
+
+`json` provides machine-readable structured output:
+
+```bash
+duplicate-finder --output-format json /data
+```
+
+For example, it can be passed directly to `jq`:
+
+```bash
+duplicate-finder -o json /data | jq
+```
+
+## File sizes
+
+Use `--size` to include the size of duplicate files in human, plain, and table
+output:
+
+```bash
+duplicate-finder --size /data
+```
+
+The short form is:
+
+```bash
+duplicate-finder -S /data
+```
+
+Combine it with `--human-readable` to display sizes in a human-readable form:
+
+```bash
+duplicate-finder --size --human-readable /data
+```
+
+## Examples
+
+Find duplicates in the current directory:
+
+```bash
+duplicate-finder
+```
+
+Scan several locations:
+
+```bash
+duplicate-finder ~/Documents ~/Downloads
+```
+
+Detect identical files even when their extensions differ:
+
+```bash
+duplicate-finder --exclude-file-extension /data
+```
+
+Include hard links in duplicate groups:
+
+```bash
+duplicate-finder --include-hardlinks /data
+```
+
+Include empty files:
+
+```bash
+duplicate-finder --include-empty /data
+```
+
+Search only files between 1 MiB and 1 GiB:
+
+```bash
+duplicate-finder --minsize 1M --maxsize 1G /data
+```
+
+Use SHA-256 instead of the default xxh128 hash:
+
+```bash
+duplicate-finder --hash sha256 /data
+```
+
+Hash candidates using at most eight worker threads:
+
+```bash
+duplicate-finder --jobs 8 /data
+```
+
+Produce plain output including human-readable file sizes:
+
+```bash
+duplicate-finder --output-format plain --size --human-readable /data
+```
+
+Read the paths to scan from standard input:
+
+```bash
+printf '%s\n' ~/Documents ~/Downloads | duplicate-finder --read-from -
+```
+
+## Exit status
+
+The command exits successfully after duplicate results have been rendered.
+
+If no duplicates are found, non-JSON output formats print:
+
+```text
+No duplicates found!
+```
+
+and exit with status `2`.
+
+JSON output instead renders the empty result normally.
+
+## Environment
+
+`duplicate-finder` respects the
+[NO_COLOR](https://no-color.org/) convention.
+
+Set `DEBUG` to `1`, `true`, or `yes` to enable diagnostic output describing the
+effective configuration, candidate grouping, duplicate detection, and output
+rendering.
+
+## License
+
+BSD-3-Clause. See [LICENSE](LICENSES/BSD-3-Clause.txt) for details.