Datagator
Forensic-grade indexing and deduplication for filesystems, archives, containers and raw devices — every file, stream and path segment cryptographically fingerprinted.
MD5 + SHA256
Identity for every file, stream and path
Immutable
Time-sealed, audit-ready snapshots
No cloud
100% local, open output formats
You cannot prove what you cannot fingerprint.
Storage sprawls across filesystems, archives, containers, snapshots and raw disk images — trillions of files where walking the tree is no longer an option. Standard tools track paths, not truth: they can tell you a name existed, not that the bytes are what they were.
Without a cryptographic identity for every file, stream and path segment, and an immutable history to compare against, inventory drifts, deduplication is approximate, and compliance and provenance become an argument rather than a proof.
Store truth, not just paths.
01
Cryptographic identity
Every file, stream, directory and path segment gets an MD5 + SHA256 identity, with raw and canonical forms both preserved. Paths are Unicode-normalised and control-char-safe, so comparisons hold across encoding and case quirks that break naïve tools.
02
Dedup and composite scanning
Identical content is detected across nested archives, volumes and snapshots; scanning recurses into .zip, .iso, .dmg and .vhd, even inside disk images. It is filesystem-agnostic — NTFS, ext4, APFS, HFS+, FAT, ZFS and raw images all ingest the same way.
03
Immutable, audit-ready history
Time-stamped, signed snapshots carry hash-and-path provenance chains and ACL lineage across NTFS, APFS, ext4 and NFSv4 — the substrate for 21 CFR Part 11 and GDPR audit trails, SBOM generation and AI dataset lineage, with nothing offshored to a cloud.
Names are not paths.
“We do not store files — we store truth. Names are not paths — they are sequences of intention.” — Datagator Design Memo #001.
The uses follow from that stance: chain-of-custody for digital forensics, backup with real deduplication and drift detection, courtroom-admissible diffs, container and firmware reverse engineering, and provenance for AI datasets under the EU AI Act.
Alpha, on real collections.
| Milestone | Status | Evidence |
|---|---|---|
| Cryptographic FS indexing (Core) | Running | Alpha, in use |
| Archive + disk-image scanning | Running | In use |
| Deduplication engine (Pro) | In progress | Chunk-level, early |
| Live dashboards & segment sync (Pro) | Next | — |
Figures on this page are targets and internal measurements at proof-of-concept stage. We share the methodology with anyone who wants to run it.
Harden Core, ship Pro.
Core — the open, test-driven indexing engine — hardens toward trillion-node collections on a design partner’s real data, while Pro adds real-time indexing, dashboards, segment sync and signed, exportable compliance reports.
Datagator spins out of Research ByIQ once it holds forensic fidelity at collection scale in production.