Sandia National Laboratories helps TSA develop machine learning algorithms that detect threats, such as weapons, in X-ray and body-scanner images at airport checkpoints. Under TSA's Open Architecture strategy, Sandia prepares datasets that many algorithm vendors train on, so errors in the data spread to every vendor.
The problem
Getting screening data ready meant collecting scans in controlled environments, annotating threat objects, merging metadata such as body type and item placement, running quality checks, and distributing the results to vendors. Sandia expected a simple process. In practice it found more than 75 places where errors could enter, and data preparation took up to eight weeks per batch.
What MLtwist did
- Automated cleaning and transformation. Ingested raw 3D scans and structured them for labeling.
- Multimodal labeling. AI-assisted workflows with human review across images, metadata, and 3D scan formats.
- Packaging. Labeled data written out as JSON, ready for detection-model training.
- Tracking. Every step versioned and traceable, to meet strict security requirements.
Because the pipeline was modular, Sandia could change it when new issues appeared instead of rebuilding it.
Results
- Data preparation dropped from eight weeks to three, a 60% reduction.
- More consistent, higher-quality labeled data for new threat detection models.
- The pilot, completed in May 2024, led to a $590K TSA contract.
Watch the webinar
Carahsoft hosted Sandia's Andrew Cox (R&D Systems Analyst), MLtwist CEO David Smith, and Google Public Sector's Steven Boesel to walk through the pipeline, what went wrong before it, and how the data stays secure. Watch the recording on Vimeo.
Screening data in this program uses the DICOS (Digital Imaging and Communications in Security) standard. MLtwist delivers in whatever format a program requires.