dupeGuru is a cross-platform (Linux, OS X, Windows) tool to find duplicate files in a system. It is written mostly in Python 3 and uses qt for the UI.
This is a personal fork of arsenetar/dupeguru, maintained at haggyroth/dupeguru. Upstream is no longer actively maintained; this fork exists to carry fixes and features we want for our own use.
All issues, pull requests, and discussion belong on this fork. Nothing here is intended to be contributed upstream. If you are looking for the original project, follow the upstream link instead.
Changes in this fork over upstream include large-scan performance work (parallel hashing, BK-tree photo matching, WAL-mode caches), safety guards around deletion, rule-based marking, and a headless command-line interface.
This folder contains the source for dupeGuru. Its documentation is in help, but is also
available online in its built form. Here's how this source tree is organized:
- core: Contains the core logic code for dupeGuru. It's Python code.
- qt: UI code for the Qt toolkit. It's written in Python and uses PyQt.
- images: Images used by the different UI codebases.
- pkg: Skeleton files required to create different packages
- help: Help document, written for Sphinx.
- locale: .po files for localization.
- hscommon: A collection of helpers used across HS applications.
For windows instructions see the Windows Instructions.
For macos instructions (qt version) see the macOS Instructions.
- Python 3.10+ — parts of the codebase use PEP 604 (
X | None) annotations that are evaluated at runtime, so earlier versions will not import. - PyQt6 (installed by
requirements.txt; see Qt bindings for the PyQt5 fallback)
The GUI runs on PyQt6 by default. PyQt5 is supported as a fallback — nothing in the tree imports a binding directly, everything goes through qtpy, which selects whichever binding is installed. CI runs a PyQt5 leg so the fallback cannot rot unnoticed.
To use PyQt5 instead:
$ pip install -r requirements.txt -r requirements-pyqt5.txt
$ QT_API=pyqt5 python run.py
QT_API is only needed when both bindings are installed; qtpy picks whichever it finds
otherwise. The CLI needs no Qt binding for standard or music scans — the import is deferred —
but --mode picture decodes images through Qt, so a binding is required for that and the CLI
says so rather than failing obscurely.
When running in a linux based environment the following system packages or equivalents are needed to build:
- python3-venv (only if using a virtual environment)
- python3-dev
- build-essential
Qt wheels bundle Qt itself but not the system libraries it links against. On a bare Linux
image — a CI runner or a container — the Qt GUI, including the offscreen platform used by the
test suite, additionally needs libegl1, libgl1, libxkbcommon-x11-0 and libdbus-1-3 or
their equivalents. A normal desktop install will already have them.
Note: pyqt5-dev-tools used to be required here, because the Qt resources were compiled by
pyrcc5 and that tool is packaged separately on some distributions. Images are now embedded
in qt/resources_data.py, which is committed, so no resource compiler is needed to build or
run dupeGuru. Regenerating it after changing an image is python build.py --resources.
To create packages the following are also needed:
- python3-setuptools
- debhelper
dupeGuru comes with a makefile that can be used to build and run:
$ make && make run
$ cd <dupeGuru directory>
$ python3 -m venv --system-site-packages ./env
$ source ./env/bin/activate
$ pip install -r requirements.txt
$ python build.py
$ python run.py
To generate packages the extra requirements in requirements-extra.txt must be installed, the steps are as follows:
$ cd <dupeGuru directory>
$ python3 -m venv --system-site-packages ./env
$ source ./env/bin/activate
$ pip install -r requirements.txt -r requirements-extra.txt
$ python build.py --clean
$ python package.py
This can be made a one-liner (once in the directory) as:
$ bash -c "python3 -m venv --system-site-packages env && source env/bin/activate && pip install -r requirements.txt -r requirements-extra.txt && python build.py --clean && python package.py"
This fork ships a headless CLI for scripted and automated scans. It is installed as the
dupeguru-scan console script, and can also be run directly from a source checkout:
$ dupeguru-scan <folder> [<folder> ...] [options]
$ python cli.py <folder> [<folder> ...] [options]
There is no scan subcommand — folders are positional arguments.
It also ships with the packaged application, as of 4.15.0. On macOS it lives inside the bundle, so dragging dupeGuru to Applications brings it along; on Windows it is installed alongside the application:
/Applications/dupeGuru.app/Contents/Resources/cli/dupeguru-scan/dupeguru-scan
C:\Program Files\dupeGuru\cli\dupeguru-scan\dupeguru-scan.exe
The packaged build excludes Qt, which takes it from 194 MB to 24 MB. The cost is --mode picture, which needs an image decoder: it reports that and exits rather than failing obscurely,
and a standard scan asked to also match pictures carries on without them. Both work normally from
a source checkout with a Qt binding installed.
--version prints the version and exits, without needing a folder argument:
$ dupeguru-scan --version
dupeguru-scan 4.15.0
python -m dupeguru also works, but only from the directory containing the checkout, since
that is where dupeguru is importable as a package:
$ cd <parent of dupeGuru directory>
$ python -m dupeguru <folder> [<folder> ...] [options]
--mode selects what is being compared. It defaults to standard.
$ dupeguru-scan ~/Photos --mode picture # visually similar images
$ dupeguru-scan ~/Music --mode music # tracks, by tag
$ dupeguru-scan ~/data # standard: file contents
--scan-type picks the algorithm within a mode; the default suits the mode
(standard → contents, music → tag, picture → picture-contents). folders compares
whole directories rather than files.
$ dupeguru-scan ~/Backups --scan-type folders
$ dupeguru-scan ~/Music --scan-type fields-noorder
Picture mode compares images of the same dimensions unless told otherwise, so a resized copy
is not reported at any --min-match value. --match-scaled lifts that, at the cost of a larger
comparison space:
$ dupeguru-scan ~/Photos --mode picture --match-scaled --min-match 90
Files in a --ref folder are scanned but never offered for deletion. This is how you protect an
original when scanning it alongside a backup:
$ dupeguru-scan /backup --ref /originals
Re-reading folder listings is the slow part of scanning an external or network drive, and in
picture mode the comparisons are slower still. --file-list-cache reuses both when nothing has
changed. Pass a path, or the flag alone to use the default location alongside the other caches:
$ dupeguru-scan /media/photos --mode picture --file-list-cache
Folders whose contents have not changed are not read again, and in picture mode the previous comparison results are reused. Files added, removed or renamed are still noticed; a file edited in place without its folder changing may be missed until the next full scan. Nothing is deleted on the basis of stale information — every file is re-checked immediately before removal.
The GUI equivalent is Preferences → "Remember scan results between scans". Both are off by default, because they trade a possible missed match for the speed.
| Option | Effect |
|---|---|
--min-match PERCENT |
minimum match percentage (default 80) |
--word-weighting |
weight word matches by frequency (filename/fields modes) |
--match-similar |
match similar words, not just identical ones |
--mix-file-kind |
allow matches between different file extensions |
--filter-hardlinks / --no-filter-hardlinks |
whether hardlinks to the same file count as duplicates |
--trust-cache-ignore-mtime |
reuse cached hashes even when mtime changed |
--verbose |
progress to stderr |
Scan a folder and write JSON results:
$ dupeguru-scan ~/Photos --output results.json
Stream newline-delimited JSON for large result sets:
$ dupeguru-scan /data --ndjson | jq 'select(.type == "group")'
Machine-readable progress on stderr, results on stdout:
$ dupeguru-scan /data --ndjson --progress-json > results.ndjson
Re-use a previous scan instead of rescanning:
$ dupeguru-scan --from-results results.json
Without exclusions the scan walks everything under the given folders, node_modules and
.venv included:
$ dupeguru-scan ~/code --exclude '^node_modules$' --exclude '^\.venv$'
$ dupeguru-scan ~/code --exclude-from excludes.txt
Patterns are regexes matched against the file or folder name, or against the full path when
the pattern contains a path separator. --exclude-from reads one per line, ignoring blank
lines and # comments.
Careful: adding any exclusion replaces the built-in "skip folders whose name starts with a dot" fallback, so
--exclude '^node_modules$'on its own will start descending into.git. Pass--exclude-defaultsalongside it to keep that behaviour — it applies the same set as the GUI's Restore Defaults button (OS metadata, trash folders, dot-prefixed names).
An ignore list saved by the GUI can be loaded too; pairs recorded in it are never reported as matching each other:
$ dupeguru-scan ~/Photos --ignore-list ~/.local/share/dupeguru/ignore_list.xml
| Code | Meaning |
|---|---|
| 0 | Completed, no duplicates found (or nothing deleted) |
| 1 | Completed, duplicates found (or files deleted) |
| 2 | Bad arguments or startup error |
| 3 | Scan failed, or errors were encountered during deletion |
Deletion from the CLI requires an explicit --yes; --delete alone will refuse to run.
$ dupeguru-scan /data --delete --yes
--delete sends files to the system trash. --direct-delete permanently removes them instead.
Files are re-validated against their recorded size and modification time immediately before
removal, and anything that changed since the scan is skipped and reported.
Add --dry-run to see what would be removed without removing it. It takes precedence over
--delete, does not require --yes, and still emits the normal results on stdout:
$ dupeguru-scan /data --delete --yes --dry-run
DRY RUN: no files have been deleted.
would send to trash 412 file(s) in 198 group(s), reclaiming 3.71 GB
412 matched on full content
re-run without --dry-run to execute.
--plan answers a different question from the results: not "what matched" but "what would
actually be removed, and what would not". It implies no mutation, needs no --delete, and
writes a per-file plan as JSON to stdout (or --output) in place of the normal results:
$ dupeguru-scan /data --plan
DELETION PLAN: no files have been deleted.
would send to trash 3,881 file(s) in 1,284 group(s), reclaiming 41.2 GB
3,860 matched on full content
21 matched on a partial (sampled) hash only and would be refused without --allow-partial-matches
1,190 group(s) corroborated: identical contents, and something independent agrees
73 group(s) content only: identical contents, and nothing else corroborates
21 group(s) unconfirmed: the contents were not compared in full
4 would be skipped: file changed since last scan
512.00 MB would not be reclaimed because of those skips
2 are on a different volume from their reference (hardlink replacement would fail)
nothing has been deleted. Re-run with --delete --yes to execute.
Every candidate carries a verdict:
{"path": "/data/b.txt", "size": 9, "mtime": 1785821371.97,
"would_delete": false, "match_confidence": "full",
"blocked_reason": "file changed since last scan"}match_confidence is "full" or "partial" — see partial matches.
blocked_reason appears only when would_delete is false.
Each group also carries a confidence tier and the reason for it, matching the Confidence
column in the application, so a scripted cleanup can act on exactly the set you reviewed in the
window:
{"confidence": "corroborated", "confidence_reason": "every copy has the same filename", ...}corroborated means the contents were compared in full and something independent agrees — a
copy in a --ref folder, or one filename across the whole group. content means the contents
matched and nothing else corroborates. unconfirmed means the contents were never compared in
full, which covers sampled-hash matches and visually similar images. The names describe the
evidence, not a level of safety: none of them means "safe to delete".
A large scan returns groups in the order they were found, which spreads your attention evenly over
groups that free wildly different amounts. --sort-by reclaimable puts the biggest wins first:
$ dupeguru-scan /data --sort-by reclaimable
Reclaimable space is what the duplicates free — the reference stays — so it is not the same as
ranking by file size. Every group carries reclaimable_bytes, with the not-fully-confirmed
portion split out as reclaimable_partial_bytes, and the stats carry a cumulative curve
whichever order you asked for:
first 10 groups -> 292895 bytes (76.8%)
first 20 groups -> 359768 bytes (94.3%)
first 25 groups -> 381587 bytes (100.0%)
Ten of those twenty-five groups give three quarters of the space; in scanner order the same ten gave twenty per cent.
The plan is not an estimate. Each file is re-validated with the same predicate the deletion
itself uses, so a file that changed since the scan is reported as skipped rather than counted
as reclaimable — and the set of paths the plan says will go is exactly the set the subsequent
--delete removes. --plan works against a saved results file too, which is where it earns
its keep: results from last week may describe a directory that has moved on.
$ dupeguru-scan --from-results results.json --plan
If any marked file was matched on a partial (sampled) hash rather than full content — only
possible when --partial-hash-threshold is in use — --delete refuses and exits 2. Those are
probable duplicates, not confirmed ones. Pass --allow-partial-matches to delete them anyway,
or drop --partial-hash-threshold to compare full contents. The same refusal applies when
deleting from a saved results file with --from-results.
--partial-hash-threshold speeds up scans by comparing three sampled chunks of a large file
instead of its full contents. That is a real trade: two different files can agree on every
sampled chunk. Because such a pair still scores 100%, the match percentage alone cannot tell
you which results are certain, so every duplicate entry carries a partial_match flag and
each stats record carries a partial_matches count:
$ dupeguru-scan /data --partial-hash-threshold 64 | jq '.stats.partial_matches'
21
To resolve them rather than just see them, --full-verify re-reads only the files involved in
partial matches, compares them in full, and discards any pair that turns out not to match:
$ dupeguru-scan /data --partial-hash-threshold 64 --full-verify
Full verification: 20 partial match(es) confirmed on full content, 1 discarded as false positive(s).
Verified matches are no longer partial, so --delete accepts them without
--allow-partial-matches. The cost is one extra read of just those files, which is far less
than scanning without a threshold at all.
Results saved before this fork recorded partial_match cannot be checked. Deleting from such
a file warns and proceeds rather than reporting a false zero; re-scan if you need the flag.
Run dupeguru-scan --help for the full option list, including the scanner knobs
(--min-match, --min-size, --max-size, --partial-hash-threshold, and others).
The complete test suite is run with Tox 1.7+. If you have it installed system-wide, you
don't even need to set up a virtualenv. Just cd into the root project folder and run tox.
If you don't have Tox system-wide, install it in your virtualenv with pip install tox and then
run tox.
You can also run automated tests without Tox. Extra requirements for running tests are in
requirements-extra.txt. So, you can do pip install -r requirements-extra.txt inside your
virtualenv and then py.test core hscommon
$ pytest core hscommon --cov=core --cov=hscommon --cov=cli --cov-report=term-missing
CI runs the same command and uploads coverage.xml as a build artifact.
black and flake8 are enforced in CI through pre-commit. Install the hooks locally so they
run before each commit:
$ pip install pre-commit && pre-commit install