MASSIVE Multilingual NLU Dataset
Source, license and coverage
Supplier documentation. These claims are separate from the automated sample score. A listing edit date is not a data freshness date.
- License
- cc-by-4.0
- Source / creator
- AmazonScience/massive
- Collection method
- Amazon Science started from the English SLURP voice-assistant dataset and engaged professional native-speaker translators to localize each utterance into 50 additional languages, preserving intent labels and re-annotating slots in the target language. Translators used either direct translation or transcreation (cultural adaptation) per slot, and the dataset records which method was used. Multiple judgments per item were collected for quality control. Train/dev/test partitions are parallel across all 51 languages.
- Coverage start
- Not documented
- Coverage end
- Not documented
- Data last updated
- Not documented
- Update schedule
- Not documented
Utterances are short, single-turn voice-assistant queries — coverage of long-form or multi-turn dialogue is out of scope. The intent/slot schema reflects a 2020-era smart-speaker product surface (alarms, music, IoT, weather) and may not generalize to newer assistant capabilities. Some low-resource languages have higher translation-noise rates flagged in the source paper. Locale coverage is uneven for dialectal variation (e.g., only zh-CN and zh-TW for Chinese). Buyers should validate label consistency empirically for any specific locale before fine-tuning.
Sample structure score: 100 / 100
This automated check describes the inspected sample, not factual accuracy, legal rights, representativeness, or the quality of the entire dataset. It is not a customer rating.
Assessed 10 sample records (JSON) on 2026-10-09. All records in the provided sample were checked.
| Check | Points | Evidence |
|---|---|---|
| Populated cells | 50 / 50 | 100 of 100 top-level cells contain a value. Null, absent and blank values count as missing; zero and false count as populated. |
| Consistent value types | 30 / 30 | 100 of 100 populated cells match their column's most common observed type. Types are inferred, not checked against real-world truth. |
| Consistent record shape | 20 / 20 | 10 of 10 records have the expected fields. CSV/TSV use the header width; JSON uses the union of observed keys. |
Field-level findings and improvements
Check missing cells and mixed types below. Document intentional missing values or mixed types in your field descriptions. Do not fill legitimate unknowns with invented values just to increase this score.
| Field | Missing cells | Most common type | Other populated types |
|---|---|---|---|
| id | 0 / 10 | number | 0 / 10 |
| locale | 0 / 10 | string | 0 / 10 |
| partition | 0 / 10 | string | 0 / 10 |
| scenario | 0 / 10 | number | 0 / 10 |
| intent | 0 / 10 | number | 0 / 10 |
| utt | 0 / 10 | string | 0 / 10 |
| annot_utt | 0 / 10 | string | 0 / 10 |
| worker_id | 0 / 10 | number | 0 / 10 |
| slot_method | 0 / 10 | object | 0 / 10 |
| judgments | 0 / 10 | object | 0 / 10 |
About this data
Parallel natural language understanding benchmark with utterances annotated across 60 intents and 55 slot types, covering 51 languages. Localized from voice assistant interactions.
Retrieve with your agent or Python
Create an account and configure DATABAZAAR_API_KEY. This example retrieves free or already purchased data; it never makes a purchase. For a multi-file dataset, choose a file index from the manifest.
Download the Python examplepython3 retrieve-dataset.py 587b3f62-51c8-4ff7-8a79-4895cf3c00aa --output dataset.bin
Full supplier documentation
Schema
| Name | Type | Description |
|---|---|---|
| id | VARCHAR | Unique utterance identifier. |
| locale | VARCHAR | BCP 47 language-region code (e.g., en-US, ja-JP). |
| partition | VARCHAR | Dataset split: train, dev, or test. |
| scenario | BIGINT | Numeric identifier for high-level domain (e.g., alarm, music, weather). |
| intent | BIGINT | Numeric identifier for one of 60 intent classes (e.g., alarm_set). |
| utt | VARCHAR | Localized natural language utterance text. |
| annot_utt | VARCHAR | Utterance with inline slot annotations in [slot_type : value] format. |
| worker_id | VARCHAR | Anonymized identifier of the translator/annotator. |
| slot_method | STRUCT(slot VARCHAR[], "method" VARCHAR[]) | Per-slot localization method (translation, transcreation, etc.) paired with slot names. |
| judgments | STRUCT(worker_id VARCHAR[], intent_score TINYINT[], slots_score TINYINT[], grammar_score TINYINT[], spelling_score TINYINT[], language_identification VARCHAR[]) | Quality review scores (intent, slots, grammar, spelling) and language ID from multiple reviewers. |
Sample Data
Preview a sample of the data before downloading.
Public sample only. Sign in to retrieve the full dataset, including free datasets.
For AI Agents
# 1. Add to your agent's MCP config (claude_desktop_config.json or similar):
{
"mcpServers": {
"databazaar": { "command": "npx", "args": ["databazaar-mcp"] }
}
}
# 2. Your agent can then call:
search_datasets({ query: "MASSIVE Multilingual NLU Datas" })
// Found: 587b3f62-51c8-4ff7-8a79-4895cf3c00aa
get_download_url({ dataset_id: "587b3f62-51c8-4ff7-8a79-4895cf3c00aa" }) // free — sign in with MCP OAuth first# Free dataset — sign in or use your account API key: curl https://api.databazaar.io/datasets/587b3f62-51c8-4ff7-8a79-4895cf3c00aa/download-url -H "Authorization: Bearer $DATABAZAAR_API_KEY"