Stanford has used MLtwist since 2023, when Stanford's Institute for Human-Centered Artificial Intelligence (HAI) selected MLtwist as a technology provider. Since then MLtwist has supported two published studies from the Stanford NLP Group.
Sindhi: a first dataset for a low-resource language
Sindhi is spoken by about 40 million people in India and Pakistan but has few labeled datasets or pretrained embeddings. The researchers set out to build its first modern NLP foundation using the Universal Dependencies framework.
- Native-speaker annotators. MLtwist recruited two native Sindhi speakers from its partner network. Both are co-authors on the paper.
- Tool setup. MLtwist prepared data and set up projects in Datasaur and Kili so the labeling team worked consistently.
- Automated QA. Quality checks drove the human-in-the-loop revision rounds.
- Post-processing. Finished labels were pulled and processed for the researchers automatically.
The study analyzed about 6,000 Sindhi sentences and released a dataset in Universal Dependencies versions 2.16 and 2.17.
Global Englishes: do NER models work for non-native speakers?
A second study asked whether named entity recognizers trained on native-speaker English perform as well on English written by proficient non-native speakers. They don't: differences as basic as how names are structured are enough to hurt performance. MLtwist preprocessed the data, set it up in Datasaur for the labeling team, ran automated quality control through the revision rounds, and post-processed the results.
The Data ID Card
For both studies, researchers used MLtwist's Data ID Card to audit where the data came from and how it moved: which companies and tools transformed it, and how each met ethical and security requirements. For published research, that record is part of the method.