Datagator

Forensic-grade indexing and deduplication for filesystems, archives, containers and raw devices — every file, stream and path segment cryptographically fingerprinted.


Domain

Forensic data

Stage

Alpha

Started

2025

Products

Datagator Core · Pro


MD5 + SHA256

Identity for every file, stream and path

Immutable

Time-sealed, audit-ready snapshots

No cloud

100% local, open output formats

The problem

You cannot prove what you cannot fingerprint.

Storage sprawls across filesystems, archives, containers, snapshots and raw disk images — trillions of files where walking the tree is no longer an option. Standard tools track paths, not truth: they can tell you a name existed, not that the bytes are what they were.

Without a cryptographic identity for every file, stream and path segment, and an immutable history to compare against, inventory drifts, deduplication is approximate, and compliance and provenance become an argument rather than a proof.


The approach

Store truth, not just paths.

01

Cryptographic identity

Every file, stream, directory and path segment gets an MD5 + SHA256 identity, with raw and canonical forms both preserved. Paths are Unicode-normalised and control-char-safe, so comparisons hold across encoding and case quirks that break naïve tools.


02

Dedup and composite scanning

Identical content is detected across nested archives, volumes and snapshots; scanning recurses into .zip, .iso, .dmg and .vhd, even inside disk images. It is filesystem-agnostic — NTFS, ext4, APFS, HFS+, FAT, ZFS and raw images all ingest the same way.


03

Immutable, audit-ready history

Time-stamped, signed snapshots carry hash-and-path provenance chains and ACL lineage across NTFS, APFS, ext4 and NFSv4 — the substrate for 21 CFR Part 11 and GDPR audit trails, SBOM generation and AI dataset lineage, with nothing offshored to a cloud.


Why it matters

Names are not paths.

“We do not store files — we store truth. Names are not paths — they are sequences of intention.” — Datagator Design Memo #001.

The uses follow from that stance: chain-of-custody for digital forensics, backup with real deduplication and drift detection, courtroom-admissible diffs, container and firmware reverse engineering, and provenance for AI datasets under the EU AI Act.


Where it stands

Alpha, on real collections.

MilestoneStatusEvidence
Cryptographic FS indexing (Core)RunningAlpha, in use
Archive + disk-image scanningRunningIn use
Deduplication engine (Pro)In progressChunk-level, early
Live dashboards & segment sync (Pro)Next

Figures on this page are targets and internal measurements at proof-of-concept stage. We share the methodology with anyone who wants to run it.


What's next

Harden Core, ship Pro.

Core — the open, test-driven indexing engine — hardens toward trillion-node collections on a design partner’s real data, while Pro adds real-time indexing, dashboards, segment sync and signed, exportable compliance reports.

Datagator spins out of Research ByIQ once it holds forensic fidelity at collection scale in production.