Nemotron Safety Guard Dataset v3 Multilingual
Source, license and coverage
Supplier documentation. These claims are separate from the automated sample score. A listing edit date is not a data freshness date.
- License
- cc-by-4.0
- Source / creator
- nvidia/Nemotron-Safety-Guard-Dataset-v3
- Collection method
- The dataset was primarily synthetically generated using NVIDIA's CultureGuard pipeline, which produces culturally and linguistically adapted safety training data across 12 languages. Prompts and responses are labeled against a safety taxonomy covering toxicity, harmful content, and other unsafe categories. See the linked arXiv paper and dataset card for full pipeline details, including translation, cultural adaptation, and quality filtering steps.
- Coverage start
- Not documented
- Coverage end
- Not documented
- Data last updated
- Not documented
- Update schedule
- Not documented
As a primarily synthetic dataset, distributions may not reflect real-world user prompts; coverage across the 12 languages is not guaranteed to be uniform in either volume or category coverage. Safety taxonomies are subjective and culture-specific — labels reflect NVIDIA/CultureGuard's taxonomy choices. Buyers training guard models should validate empirically on their target deployment distribution, and audit for any leakage against public safety benchmarks (e.g., ToxicChat, BeaverTails, XSTest). The supplier describes synthetic or modeled records. These should not be treated as verified real-world observations.
Sample structure score: 82.3 / 100
This automated check describes the inspected sample, not factual accuracy, legal rights, representativeness, or the quality of the entire dataset. It is not a customer rating.
Assessed 10 sample records (JSON) on 2026-10-09. All records in the provided sample were checked.
| Check | Points | Evidence |
|---|---|---|
| Populated cells | 32.3 / 50 | 71 of 110 top-level cells contain a value. Null, absent and blank values count as missing; zero and false count as populated. |
| Consistent value types | 30 / 30 | 71 of 71 populated cells match their column's most common observed type. Types are inferred, not checked against real-world truth. |
| Consistent record shape | 20 / 20 | 10 of 10 records have the expected fields. CSV/TSV use the header width; JSON uses the union of observed keys. |
Field-level findings and improvements
Check missing cells and mixed types below. Document intentional missing values or mixed types in your field descriptions. Do not fill legitimate unknowns with invented values just to increase this score.
| Field | Missing cells | Most common type | Other populated types |
|---|---|---|---|
| id | 0 / 10 | string | 0 / 10 |
| prompt | 0 / 10 | string | 0 / 10 |
| response | 8 / 10 | string | 0 / 2 |
| prompt_label | 0 / 10 | string | 0 / 10 |
| response_label | 8 / 10 | string | 0 / 2 |
| violated_categories | 6 / 10 | string | 0 / 4 |
| prompt_label_source | 0 / 10 | string | 0 / 10 |
| response_label_source | 8 / 10 | string | 0 / 2 |
| tag | 0 / 10 | string | 0 / 10 |
| language | 0 / 10 | string | 0 / 10 |
| reconstruction_id_if_redacted | 9 / 10 | number | 0 / 1 |
About this data
Multilingual safety dataset across 12 languages for training LLM safety guard models, generated via the CultureGuard pipeline.
Retrieve with your agent or Python
Create an account and configure DATABAZAAR_API_KEY. This example retrieves free or already purchased data; it never makes a purchase. For a multi-file dataset, choose a file index from the manifest.
Download the Python examplepython3 retrieve-dataset.py 5a12abd4-c2ed-4eb0-9cab-11a84eb5fc2c --output dataset.bin
Full supplier documentation
Schema
| Name | Type | Description |
|---|---|---|
| id | VARCHAR | Unique identifier for the sample (MD5 hash format). |
| prompt | VARCHAR | User input text to be evaluated for safety across 12 languages. |
| response | VARCHAR | Model or assistant response paired with the prompt; null if not provided. |
| prompt_label | VARCHAR | Safety classification of prompt: safe or unsafe with optional violation category. |
| response_label | VARCHAR | Safety classification of response: safe, unsafe, or empty if not applicable. |
| violated_categories | VARCHAR | Comma-separated safety taxonomy categories breached (e.g., Profanity, Violence); empty if none. |
| prompt_label_source | VARCHAR | Annotation source for prompt label: human or automated method. |
| response_label_source | VARCHAR | Annotation source for response label: human, automated, or null if not labeled. |
| tag | VARCHAR | Dataset partition or content type tag (e.g., generic, adversarial). |
| language | VARCHAR | ISO 639-1 language code (ar, de, en, es, fr, hi, it, ja, ko, nl, th, zh). |
| reconstruction_id_if_redacted | DOUBLE | Row ID of original unredacted sample if this prompt was redacted; null otherwise. |
Sample Data
Preview a sample of the data before downloading.
Public sample only. Sign in to retrieve the full dataset, including free datasets.
For AI Agents
# 1. Add to your agent's MCP config (claude_desktop_config.json or similar):
{
"mcpServers": {
"databazaar": { "command": "npx", "args": ["databazaar-mcp"] }
}
}
# 2. Your agent can then call:
search_datasets({ query: "Nemotron Safety Guard Dataset " })
// Found: 5a12abd4-c2ed-4eb0-9cab-11a84eb5fc2c
get_download_url({ dataset_id: "5a12abd4-c2ed-4eb0-9cab-11a84eb5fc2c" }) // free — sign in with MCP OAuth first# Free dataset — sign in or use your account API key: curl https://api.databazaar.io/datasets/5a12abd4-c2ed-4eb0-9cab-11a84eb5fc2c/download-url -H "Authorization: Bearer $DATABAZAAR_API_KEY"