aboutsummaryrefslogtreecommitdiff
path: root/README.md
blob: c93b843180d2d63fd57fb5c9719a15d437dab62a (plain) (blame)
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
<!--
SPDX-FileCopyrightText: 2026 Dennis Fink <me+coding@dennisfink.me>

SPDX-License-Identifier: BSD-3-Clause
-->

# duplicate-finder

A command-line tool for finding duplicate files using metadata grouping and
content hashes.

Files are discovered recursively and first grouped by size and, by default,
file extension. Paths in each candidate set are also grouped by filesystem
identity so hard links can be handled efficiently. Only remaining candidates
are hashed to identify files with matching contents.

## Installation

```bash
pip install duplicate-finder
```

## Usage

```text
Usage: duplicate-finder [OPTIONS] [PATHS]...

  Scan paths and list duplicate files.

  Files are first grouped into candidate sets by size and, by default,
  case-insensitive file extension. --exclude-file-extension groups candidates
  by size only. Candidate files are then hashed to identify matching content.

Options:
  --read-from FILE                Read files and directories to scan from FILE,
                                  one path per line. Use '-' for standard
                                  input. When set, positional paths are
                                  ignored.
  --follow-symlinks / --no-follow-symlinks
                                  Follow symlinks when scanning directories.
  -H, --include-hardlinks / --exclude-hardlinks
                                  Include hard links in duplicate results, so
                                  multiple paths to the same underlying file
                                  may be listed as duplicates.
  -G, --minsize SIZE              Consider only files >= SIZE bytes. Supports
                                  suffixes such as 10M, 1G, and 500K.
  -L, --maxsize SIZE              Consider only files <= SIZE bytes. Supports
                                  suffixes such as 10M, 1G, and 500K.
  --include-empty / --exclude-empty
                                  Include files that are 0 bytes in size.
  --include-file-extension / --exclude-file-extension
                                  Include file extensions when grouping
                                  duplicate candidates. Enabled by default;
                                  excluding extensions groups candidates by
                                  size only.
  --hash HASH                     Select which hash function to use.
                                  [default: xxh128]
  -j, --jobs INTEGER              Maximum number of worker threads
                                  (0 = use default).
  -o, --output-format [human|plain|table|json]
                                  Select the output format for duplicate
                                  groups. 'human' shows a readable text list,
                                  'plain' prints each duplicate group as a
                                  space-separated line, 'table' prints a
                                  column-aligned overview, and 'json' outputs
                                  machine-readable structured data.
                                  [default: human]
  -S, --size                      Show size of duplicate files
                                  (human/plain/table output only).
  --human-readable / --no-human-readable
                                  Display file sizes in human-readable form.
  -q, --quiet                     Suppress all non-error output.
  -v, --verbose                   Enable verbose output.
  --color [auto|always|never]     Control colorized output.
  --no-color                      Disable colorized output.
  --version                       Show the version and exit.
  -h, --help, -?                  Show this message and exit.
```

If no path is specified, the current directory is scanned.

## Duplicate detection

Duplicate detection is split into inexpensive candidate selection followed by
content hashing.

Files are initially grouped by:

- file size
- case-insensitive file extension

Groups that cannot contain duplicates are discarded before their contents are
hashed.

File extensions can be excluded from candidate grouping with
`--exclude-file-extension`. In that case, files are grouped by size alone. This
allows identical files with different or incorrect extensions to be detected,
at the cost of potentially hashing more files.

Hard links are recognized by filesystem identity. By default, multiple paths
referring to the same file are treated as one file and are not reported as
duplicates of each other. Use `--include-hardlinks` to include those paths in
duplicate groups.

## Input

One or more files or directories can be supplied as positional arguments:

```bash
duplicate-finder ~/Documents ~/Downloads
```

Directories are scanned recursively.

If no paths are supplied, `duplicate-finder` scans the current directory:

```bash
duplicate-finder
```

Paths can also be read from a file using `--read-from`. The file must contain
one path per line:

```bash
duplicate-finder --read-from paths.txt
```

Use `-` to read paths from standard input:

```bash
find /data -type d -print | duplicate-finder --read-from -
```

When `--read-from` is used, positional paths are ignored.

## File selection

Empty files are excluded by default. Include them with:

```bash
duplicate-finder --include-empty
```

Limit duplicate detection to files at least a certain size:

```bash
duplicate-finder --minsize 10M /data
```

Or set an upper limit:

```bash
duplicate-finder --maxsize 1G /data
```

The two options can be combined:

```bash
duplicate-finder --minsize 10M --maxsize 1G /data
```

Size suffixes such as `K`, `M`, and `G` are supported.

Symlinks are not followed by default. Enable symlink traversal with:

```bash
duplicate-finder --follow-symlinks /data
```

### Additional filtering

`duplicate-finder` intentionally provides only a small set of file-selection
options directly. For other criteria, use an appropriate external tool to
select the paths and pass them to `duplicate-finder` with `--read-from`.

For example, use `find` to consider only files modified within the last 30
days:

```bash
find /data -type f -mtime -30 -print | duplicate-finder --read-from -
```

The same approach can be used for filtering by filename, modification time,
ownership, permissions, or other filesystem metadata supported by the tool
producing the path list.

`--read-from` accepts one path per line, either from a file or from standard
input.

## Hash functions

The hash function used for content comparison can be selected with `--hash`.

The default is `xxh128`:

```bash
duplicate-finder --hash xxh128 /data
```

`xxh32` and `xxh64` are also supported, together with the hash algorithms
available through Python's `hashlib` implementation.

For example:

```bash
duplicate-finder --hash sha256 /data
```

## Parallel hashing

Candidate files can be hashed using multiple worker threads.

By default, `--jobs=0` lets the application use its default worker count:

```bash
duplicate-finder /data
```

Set an explicit maximum number of workers with:

```bash
duplicate-finder --jobs 4 /data
```

## Output formats

Four output formats are available.

### Human

`human` is the default format and is intended for interactive use:

```bash
duplicate-finder /data
```

### Plain

`plain` prints each duplicate group as a space-separated line and is intended
for further command-line processing:

```bash
duplicate-finder --output-format plain /data
```

The short form is:

```bash
duplicate-finder -o plain /data
```

### Table

`table` displays duplicate groups in a column-aligned overview:

```bash
duplicate-finder --output-format table /data
```

### JSON

`json` provides machine-readable structured output:

```bash
duplicate-finder --output-format json /data
```

For example, it can be passed directly to `jq`:

```bash
duplicate-finder -o json /data | jq
```

## File sizes

Use `--size` to include the size of duplicate files in human, plain, and table
output:

```bash
duplicate-finder --size /data
```

The short form is:

```bash
duplicate-finder -S /data
```

Combine it with `--human-readable` to display sizes in a human-readable form:

```bash
duplicate-finder --size --human-readable /data
```

## Examples

Find duplicates in the current directory:

```bash
duplicate-finder
```

Scan several locations:

```bash
duplicate-finder ~/Documents ~/Downloads
```

Detect identical files even when their extensions differ:

```bash
duplicate-finder --exclude-file-extension /data
```

Include hard links in duplicate groups:

```bash
duplicate-finder --include-hardlinks /data
```

Include empty files:

```bash
duplicate-finder --include-empty /data
```

Search only files between 1 MiB and 1 GiB:

```bash
duplicate-finder --minsize 1M --maxsize 1G /data
```

Use SHA-256 instead of the default xxh128 hash:

```bash
duplicate-finder --hash sha256 /data
```

Hash candidates using at most eight worker threads:

```bash
duplicate-finder --jobs 8 /data
```

Produce plain output including human-readable file sizes:

```bash
duplicate-finder --output-format plain --size --human-readable /data
```

Read the paths to scan from standard input:

```bash
printf '%s\n' ~/Documents ~/Downloads | duplicate-finder --read-from -
```

## Exit status

The command exits successfully after duplicate results have been rendered.

If no duplicates are found, non-JSON output formats print:

```text
No duplicates found!
```

and exit with status `2`.

JSON output instead renders the empty result normally.

## Environment

`duplicate-finder` respects the
[NO_COLOR](https://no-color.org/) convention.

Set `DEBUG` to `1`, `true`, or `yes` to enable diagnostic output describing the
effective configuration, candidate grouping, duplicate detection, and output
rendering.

## License

BSD-3-Clause. See [LICENSE](LICENSES/BSD-3-Clause.txt) for details.