Solutions · Data provenance & compliance

Know where every training file came from

Every dataset ships with a Data ID Card: where each file came from, what changed it, who labeled it, and under what terms. Recorded as the work happens, not reconstructed for an audit.

A hand turning pages in a thick tabbed records binder, with a checklist and fountain pen beside it and a laptop behind

The problem

Why this is hard to do well

01

Regulators are asking

AI rules and procurement reviews increasingly ask what a model was trained on and whether you had the right to use it.

02

Long vendor chains

Data passes through collectors, tools, and labeling vendors, and each hand-off loses the paper trail.

03

Rights and consent

Licenses, consent, and usage limits have to follow each file, not sit in a contract folder.

How MLtwist does it

Platform and people, together

Our platform does the repeatable work; our team, yours, or both handle the judgment calls. Every version ships with a Data ID Card.

The Data ID Card
Origin, transformations, people and vendors, and compliance recorded for every file.
Versioned deliveries
Each delivery is a version you can diff, roll back, and cite.
Your data stays yours
Files stay in your GCS, S3, or Azure bucket, encrypted at rest, with role-based access.
Built for public sector
Chain of custody for regulated programs, and purchasing through Carahsoft and Google Cloud Marketplace.

Talk to us about data provenance & compliance

Tell us the data type, volume, and timeline. We'll scope it with our team, yours, or both — and deliver it versioned, in your format.

Also available through Carahsoft and Google Cloud Marketplace.