One billion files. One index. Answers in seconds.
A filesystem indexer and query suite for storage at HPC scale. Pure Rust. No daemon, no database server, no runtime.
version 0.5.1 · MIT licence · Linux & macOS · x86-64 & arm64
du(1) was written for filesystems you could walk while the
coffee brewed. A modern research storage system holds hundreds of millions to billions of
files, and every du or find re-walks the whole tree to answer one
question — hours of metadata traffic, discarded the moment the prompt comes back.
xdu walks the tree once and writes down what it learned: a persistent, Hive-partitioned Parquet index of every file's path, size, owner, group, permission bits and three timestamps. Everything after that is a query. Who is consuming the quota; what has not been touched in two years; which permission bits are wrong; where the whale directories are hiding — answered in seconds, thousands of times, against the same index.
Crawl /scratch overnight. Audit it all week.
Four commands. One writes the index; three read it.
| Command | What it does |
|---|---|
xdu |
The crawler. Walks a tree with a shared work-stealing thread pool and writes the eight-column Parquet index, one directory of chunks per partition. |
xdu-find |
The query CLI. Glob or regex on paths, size and age windows, owner, group and
octal permission filters; --count, --top, and
path / size / atime / csv / json
output. Pipes into anything. |
xdu-view |
The explorer. An ncdu-style terminal UI over the index, with a list
view and a Miller-columns tree view, file-type detection, a text preview pane, and
filters and sorts you can change without leaving it. Strictly read-only. |
xdu-rm |
The enforcer. Parallel bulk deletion of exactly the set an index query selected,
with --dry-run, a confirmation prompt, and a --safe mode
that re-stats each file and refuses to delete one a user touched since
the index was built. |
Every reader takes its index from -i or from
$XDU_INDEX; every parallel command takes its thread count from -j
or $XDU_JOBS. Man pages and shell completions for all four ship in the release
tarball.
$ xdu /scratch -o /index/scratch -j 32 Finished __root__ (312 files, 1.44 GiB) Finished x-lentner (8.4M files, 214.66 TiB) Finished bioinfo (41.2M files, 1044.31 TiB, 118 vanished) ⠙ 160.1M files, 3882.38 TiB | 412.8k files/s (peak: 501.3k files/s) ⠴ genomics: 92.4M files, 2192.38 TiB | 284.1k files/s (peak: 331.5k files/s) [T3] ⠋ climate: 18.1M files, 431.02 TiB | 96.4k files/s (peak: 140.2k files/s) [T7] ⠒ cms-data: scanning...
Completed 1.0B files (8614.40 TiB) in 3218.44s, 2143 vanished┌─ genomics/rnaseq [older:90d] [min:1.00 MiB] ─────────────────────────────────┐ │ ▸ alignments 412.83 TiB 18.4M files 4 months ago │ │ ▸ fastq 184.22 TiB 2.1M files 7 months ago │ │ ▸ counts 1.42 TiB 884.2K files 3 months ago │ │ ▸ qc 884.10 GiB 12.9K files 8 days ago │ │ ▸ .snakemake 41.06 GiB 308.4K files 2 years ago │ │ run_manifest.tsv 4.50 KiB 1 file today │ └──────────────────────────────────────────────────────────────────────────────┘ 1.4K entries in 0.18s (filtered) │ sort:size-desc mode:list │ q:quit jk↑↓:nav
xdu-view -i /index/scratch -u genomics --older-than 90 --min-size 1M -s size-desc
— the list view, filtered to stale data over a megabyte and sorted by weight. Active
filters are named in the title bar; /, o, > and
< change them in place. Nothing here can write to your filesystem.┌─scratch────────────┐┌─genomics─────────┐┌─rnaseq───────────┐┌─env.yml────────┐ │ ▸ bioinfo 1044 TiB ││ ▸ rnaseq 412 TiB ││ ▸ 2024 88 TiB ││ name: rnaseq │ │ ▸ climate 431 TiB ││ ▸ wgs 988 TiB ││ ▸ 2025 204 TiB ││ channels: │ │ ▸ cms-data 2145 TiB││ ▸ atac 284 TiB ││ README.md 4 KiB││ - bioconda │ │ ▸ genomics 2192 TiB││ ▸ ref 12 TiB ││ env.yml 2 KiB││ - conda-forge │ │ ▸ x-lentner 214 TiB││ ││ ││ dependencies: │ │ ││ ││ ││ - star=2.7.11 │ │ ││ ││ ││ - salmon=1.10 │ └────────────────────┘└──────────────────┘└──────────────────┘└────────────────┘ genomics/rnaseq/env.yml │ 2.05 KiB │ 1 file │ 6 days ago │ mode:tree
# which partitions hold the most files, biggest first $ xdu-find -i /index/scratch --top 5 genomics cms-data bioinfo climate x-lentner # a gigabyte or more, untouched for two years, owned by a departed user # (-f size emits bytes and a tab, ready for cut, awk or sort -n) $ xdu-find -i /index/scratch --owner jdoe --min-size 1G --older-than 730 -f size 4629938784665 /scratch/genomics/wgs/NA12878.cram 2023949056512 /scratch/genomics/wgs/NA12891.cram … # how much of the shared project is world-writable? $ xdu-find -i /index/scratch -u bioinfo --mode /002 --count 1884203 # hand the whole slice to your own tooling $ xdu-find -i /index/scratch -u climate --mtime-older-than 365 -f csv > stale.csv # enforce the retention policy: preview, then commit $ xdu-rm -i /index/scratch --older-than 60 --dry-run | tail -1 847231 file(s) would be deleted. $ xdu-rm -i /index/scratch --older-than 60 --safe -j 16 --force -v SKIP (accessed since index): /scratch/bioinfo/refs/hg38.fa DELETE: /scratch/climate/cmip6/tmp/run0417.nc DELETE: /scratch/bioinfo/tmp/sorted.bam.tmp … Deleted: 847203 Skipped (safe mode): 28
Eight columns, one row per file, Snappy-compressed Parquet. Nothing clever — the point is that it is boring, small, and readable by every tool in the modern data stack.
| Column | Type | Meaning |
|---|---|---|
path | UTF‑8 | Absolute file path |
size | INT64 | Disk usage from st_blocks, or apparent length under --apparent-size |
uid | INT64 | Owning user id |
gid | INT64 | Owning group id |
mode | INT64 | Permission bits (st_mode & 07777) |
atime | INT64 | Last access, Unix epoch seconds |
mtime | INT64 | Last modification |
ctime | INT64 | Last inode change |
One directory per top-level subdirectory of the indexed root, holding numbered chunks.
Loose files directly under the root land in a reserved __root__ partition.
/index/scratch/
├── .xdu-complete # run attestation: xdu, files, bytes, vanished, errors, format
├── __root__/
│ └── 000000.parquet
├── bioinfo/
│ ├── 000000.parquet
│ ├── 000001.parquet
│ └── 000002.parquet
├── climate/
│ └── 000000.parquet
└── genomics/
└── 000000.parquet
That layout is doing real work. It gives DuckDB partition pruning, so
-u alice reads one directory instead of scanning a petabyte-scale index; it lets
partitions be written in parallel and re-indexed incrementally
(xdu /scratch -o /index/scratch --partition alice, which also retires that
partition's stale chunks); and because it is just Hive-partitioned Parquet on a filesystem, it
maps unchanged onto object-store key prefixes.
Every chunk is written as NNNNNN.parquet.partial and renamed into place, so a
reader never sees half a file. But per-file atomicity cannot say whether the run
finished — a crawl that died after three partitions leaves three perfectly valid
directories. So xdu clears .xdu-complete before it writes and
restores it only on the success path, recording the file and byte counts, the files that
vanished mid-walk, any tolerated errors, and the index format version.
All three readers check that marker first. An index whose format they do not recognise is
refused — no rows, and in xdu-rm's case no deletions. An
index whose crawl skipped an unreadable region still answers, but says so on stderr. By
default an unreadable path fails the crawl outright; --allow-errors is how you
opt into indexing what is reachable, and the marker remembers that you did.
stat. Metadata calls happen inside the pool, which
is what matters on a network or parallel filesystem where each one is a round trip.-B, default 100,000 rows), keeping write syscalls off the hot path.| Shape | Files | Threads | Wall | Files/s | Peak RSS |
|---|---|---|---|---|---|
| mixed fan-out | 819,216 | 1 | 5.87 s | 139,560 | 32 MiB |
| mixed fan-out | 819,216 | 2 | 3.59 s | 228,194 | 38 MiB |
| mixed fan-out | 819,216 | 4 | 2.49 s | 329,002 | 61 MiB |
| mixed fan-out | 819,216 | 8 | 2.61 s | 313,876 | 102 MiB |
| 1,000 partitions | 400,000 | 4 | 1.28 s | 312,500 | 12 MiB |
| one flat directory | 400,000 | 4 | 3.17 s | 126,183 | 141 MiB |
Read that table for its shape, not its absolute numbers. Local
APFS on a laptop saturates around four threads; a parallel filesystem with real metadata
latency wants far more, which is what -j 32 is for. The flat-directory row is the
honest worst case: a single directory is the unit of parallelism, so one enormous flat
directory cannot be split. And the harness's own noise floor is 3–28% depending on
shape, so no single figure here is a promise. Reproduce it with
sh bench/run.sh baseline.
$ curl -sSfL https://xdu-project.org/install.sh | sh
Fetches the release binaries for your platform and installs them, with man pages and shell
completions, under ~/.local/bin. Set XDU_INSTALL=/usr/local/bin to
put them somewhere else, or XDU_VERSION=v0.5.1 to pin a version.
From source, with the Rust toolchain the repository pins
(rust-toolchain.toml, currently 1.97.1 — rustup honours it
automatically):
$ cargo install --git https://github.com/xdu-project/xdu.git
Or run it from a container on main. Docker and
Apptainer/Singularity specs are generated from a single
HPCCM recipe in hpccm/:
$ make -C hpccm sif # xdu.sif, via apptainer $ make -C hpccm image # the docker image
They install the published release binaries rather than building from source, so they finish in seconds, and they follow release tags rather than the working tree.
Native .deb and .rpm packages are also
published † v0.6.3, built from
the same release layout.
| Platform | Triple |
|---|---|
| Linux, x86-64 | x86_64-unknown-linux-gnu |
| Linux, arm64 | aarch64-unknown-linux-gnu |
| macOS, Intel | x86_64-apple-darwin |
| macOS, Apple silicon | aarch64-apple-darwin |
The Linux tarballs are built on Ubuntu 24.04 and currently want glibc 2.39, so they will not exec on RHEL8, RHEL9 or bookworm — use the container specs on those hosts until the portable baseline lands † v0.6.
Unix only: xdu reads ownership, mode and timestamps through
std::os::unix::fs::MetadataExt.
xdu-find and xdu-view cover the common questions. For anything
else, the index is just Parquet — point DuckDB, Polars, pandas or Spark at it and write
whatever you want.
-- total usage per user, by partition, no full scan SELECT regexp_extract(path, '/scratch/([^/]+)/', 1) AS project, sum(size) / 1e12 AS tb, count(*) AS files FROM read_parquet('/index/scratch/*/*.parquet') GROUP BY project ORDER BY tb DESC; -- the top ten cold whales SELECT path, size, to_timestamp(atime) AS last_read FROM read_parquet('/index/scratch/*/*.parquet') WHERE atime < epoch(now()) - 86400 * 180 ORDER BY size DESC LIMIT 10; -- exposure audit: world-writable files, by owner SELECT uid, count(*) AS n, sum(size) / 1e9 AS gb FROM read_parquet('/index/scratch/*/*.parquet') WHERE mode & 2 != 0 GROUP BY uid ORDER BY n DESC; -- one partition only: DuckDB never opens the others SELECT sum(size) / 1e12 AS tb FROM read_parquet('/index/scratch/genomics/*.parquet');
Crawl on a schedule:
0 2 * * * /usr/local/bin/xdu /scratch -o /index/scratch -j 32 >> /var/log/xdu.log 2>&1
On a multi-petabyte ZFS filesystem, where a full crawl runs for hours and users keep working the whole time, index a snapshot instead. The snapshot is instantaneous and read-only, so the index describes an exact moment rather than a smear across eight hours — and if anyone later disputes a number, you can mount the same snapshot and check.
# zfs snapshot tank/scratch@xdu-$(date +%F) # mount -t zfs -o ro tank/scratch@xdu-$(date +%F) /mnt/xdu-snap # xdu /mnt/xdu-snap -o /index/scratch -j 32 # umount /mnt/xdu-snap && zfs destroy tank/scratch@xdu-$(date +%F)
Where this is going, and roughly when.
-o s3://bucket/prefix. Indexes on local disk are tied to the machine that
built them; the Hive layout maps onto object-store key prefixes with no change, so an
index can be built once centrally and read from anywhere by any tool.A central index built as root describes every tenant on the filesystem. Handing that to the people it describes is the hardest problem on this roadmap, and the answer is not a cleverer query — it is a service.
stat for
themselves. The scoping cannot live in the reader: DuckDB runs inside the calling process
with the calling user's privileges, so a filter a client applies is a convenience for
display and never a boundary — anyone who can open the files reads every row in them.
Two invariants replace what would otherwise be a page of configuration. An index
you built yourself inherits the permissions you had when you crawled it, and needs
no scoping machinery at all; that is what ships today. A shared index is only ever
read through the service, and needs none either.xdu-apixdu-loginxdu-find, xdu-view and
xdu-rm reach a served tree under exactly the invocation you would have used
locally. The token lands in your config directory, refreshes itself, and the readers carry
no authentication code of their own — so a query from a batch script works and
nothing ever prompts halfway through one. xdu-login --status reports who the
service thinks you are and how much of the tree that resolves to, which is the first
question worth asking when an expected row is missing.xdu-webxdu-view, compiled to WebAssembly: the same list and tree views, the same
filters and sorts, in a progressive web app. Against a shared index it authenticates and
queries through xdu-api, which scopes every answer to what you could see
yourself; against an index file you already hold it runs entirely in the browser, with no
service and nothing to configure. A browser has no Unix identity to scope with —
which is precisely why the service exists, rather than the page reaching object storage
on its own.Four supporting pieces carry that architecture and are tracked as their own
roadmap entries † v0.9–v1.0. The crawl computes
each row's visibility from its whole directory chain rather than the file's own mode — a
0600 file under world-readable directories is visible to du, and a
world-readable file under one private ancestor is not. The service speaks a typed operation set
rather than SQL, because SQL over a served index reaches the filesystem and the network. The
readers route to a service by path prefix. And the service refuses to start against a store that
anyone but itself can read.
xdu-cp and xdu-mvxdu-rm proved the pattern: select a set with an index query, act on exactly
that set, with a dry run, a confirmation, a --safe re-stat and
deterministic ordering under --limit. That path becomes one shared
select-and-act engine behind rm, cp and mv instead
of three copies of the same main. “Move everything untouched in two
years to cold storage” becomes one command that can be reviewed before it
runs.xdu-tar slice archivesxdu-find | xargs tar inherits every
xargs failure mode and assembles ten thousand appends where tape wants a
single stream.--newer-than filters on access time, the wrong clock for a
modification question — and becomes a measured difference, which is exactly the
selection xdu-tar wants for an incremental slice.--type video --min-size 1G. MIME detection recorded at crawl time, so
“every video file over a gigabyte” is an index query rather than a second
walk of the tree.xdu off a terminal and every diagnostic goes to stderr as one
timestamped, severity-tagged line: the invocation and its arguments, each partition's
finish with file and byte counts, every tolerated warning, the completion-marker verdict
and a final summary — so the log alone explains the exit status. Stdout stays clean
and pipeable, and the interactive display on a TTY is untouched. A 3 a.m. cron failure
used to leave a log that said what finished without saying when, or how badly.scanning….
Throughput was fine; the report was misleading, because each bar was owned by one
driver for life while the pool work-steals across all of them. Every line now carries its
own evidence of life, with labels that match the threading model — because on shared
scratch, “skewed” versus “stuck” is the difference between letting
a job run and killing it.scanning… with no
elapsed time, and quiet-for-seconds reads identically to quiet-for-an-hour. The fix is a
yield-independent refresh that leaves the single-pool work-stealing walk alone.hpccm/, so they cannot drift apart. A two-stage build installs the
published release binaries, man pages and completions — no Rust toolchain, so it
finishes in seconds — with the download verified against the published
SHA256SUMS in a throwaway first stage and the architecture resolved from
uname -m, so one spec serves x86-64 and arm64. They track release tags rather
than the working tree.manylinux_2_28-class builder floor, guarded in CI
so it cannot float back with the next toolchain bump. Until then the container specs above
are the way onto those hosts.apt install xdu, dnf install xdu. The tarball layout and
install.sh already define the exact file map these would ship.Smaller work, mostly invisible from the outside — except the first three, which are not.
| Item | What changes | Release |
|---|---|---|
| Clean exits on closed pipes | xdu-find … | head -1 succeeds instead of failing with
Broken pipe (os error 32), while a genuinely full disk stays loud. |
v0.6.1 |
| Panic-safe terminal restore | A drop guard and panic hook in xdu-view, plus multibyte-safe name
truncation, so neither a fault nor an awkward filename can leave a terminal
wedged. |
v0.6.1 |
Man pages that survive groff |
No soft hyphen inserted into the middle of OUTDIR/.xdu-complete, and a
rendering gate that uses the same renderer users do. |
v0.6.1 |
| Per-partition attestation | A --partition-scoped run stops rewriting the whole index's completion
marker from its own statistics, and so stops retiring warnings it knows nothing
about. |
v0.6.2 |
| Whole-index reconciliation | A partition whose source directory is gone is retired on re-index, instead of answering queries with rows for files that no longer exist. | v0.6.2 |
| Prune guard for unreadable partitions | finalize declines to prune chunks it could not read, so a permission
change or a stale mount can no longer cost you rows you already had. |
v0.6.2 |
Man-page gate from doc/*.scd |
The asserted page list is derived from the sources that are rendered, so a fifth man page cannot ship unchecked. | v0.6.x |
| Benchmark baseline guard | bench/run.sh baseline can no longer overwrite the committed reference it
is being compared against — the one file in bench/results/ that is
not reproducible on demand. |
v0.6.x |
| CI that binds | A branch ruleset with required checks and no direct pushes, plus a scheduled canary for the container image, so a red gate stops a merge instead of decorating one. | v0.6.x |
| Library extraction | Validated escaping on the index-glob seam, one count formatter rather than two, and
the TUI's pure helpers — strip_ansi above all, which is
load-bearing for terminal safety — lifted out of a 2,500-line binary into
lib where they can be tested. |
v0.6.x |
| The v1.0 checkpoint | A written motivation-and-architecture piece, a project identity, a contribution guide and a maintenance plan. A v1.0 is as much about explaining a project as shipping code. | v1.0 |
Delivered in the 0.5 series and present in the binaries described above: the
eight-column index — owner, group, mode, mtime and ctime alongside path, size and atime;
on-disk index format versioning, which readers refuse to guess at; glob-by-default path
matching, with --regex to opt back in; -V/--version on
every binary; the in-list file preview overlay in xdu-view; and an RPM spec.
Before that: the crawler and its shared pool, all four binaries, the release tarball,
install.sh, man pages and shell completions.