← All case studies

Stanford University

Research-grade language data for Stanford NLP

~6,000

Sindhi sentences prepared for Universal Dependencies v2.16 and v2.17

Research group
Stanford NLP Group
Data
Text: NER and dependency parsing
Tools
Datasaur and Kili, run by MLtwist
Record
Data ID Card per dataset

Stanford has used MLtwist since 2023, when Stanford's Institute for Human-Centered Artificial Intelligence (HAI) selected MLtwist as a technology provider. Since then MLtwist has supported two published studies from the Stanford NLP Group.

Sindhi: a first dataset for a low-resource language

Sindhi is spoken by about 40 million people in India and Pakistan but has few labeled datasets or pretrained embeddings. The researchers set out to build its first modern NLP foundation using the Universal Dependencies framework.

  • Native-speaker annotators. MLtwist recruited two native Sindhi speakers from its partner network. Both are co-authors on the paper.
  • Tool setup. MLtwist prepared data and set up projects in Datasaur and Kili so the labeling team worked consistently.
  • Automated QA. Quality checks drove the human-in-the-loop revision rounds.
  • Post-processing. Finished labels were pulled and processed for the researchers automatically.

The study analyzed about 6,000 Sindhi sentences and released a dataset in Universal Dependencies versions 2.16 and 2.17.

Global Englishes: do NER models work for non-native speakers?

A second study asked whether named entity recognizers trained on native-speaker English perform as well on English written by proficient non-native speakers. They don't: differences as basic as how names are structured are enough to hurt performance. MLtwist preprocessed the data, set it up in Datasaur for the labeling team, ran automated quality control through the revision rounds, and post-processed the results.

The Data ID Card

For both studies, researchers used MLtwist's Data ID Card to audit where the data came from and how it moved: which companies and tools transformed it, and how each met ethical and security requirements. For published research, that record is part of the method.

“The amount of published content is growing exponentially and it is increasingly important to understand not just the content of the data, but also the origin of the data.”
Audrey Smith Chief Operating Officer, MLtwist

Get a human-verified training set

Tell us what you're labeling. We'll scope it with our team, yours, or both — and deliver it versioned, in your format.

Also available through Carahsoft and Google Cloud Marketplace.