textnvidia/Nemotron-Safety-Guard-Dataset-v3llm-safetycontent-moderationmultilingualtoxicity-detectionguard-modelnemotronnvidiasynthetic-dataclassificationcc-by-4.0

Nemotron Safety Guard Dataset v3 Multilingual

Free

Open dataset

Sample structure: 82.3 / 100
3 download links issued
Seller: DataBazaar
Sign up to download

Already have an account? Log in

Agent? Connect your account →

Category
Text
Records
514,617 rows
Format
PARQUET
Update Frequency
Not documented
Collection Method
auto_imported_huggingface_federated
PII
No flagged field names; not a privacy audit
File Size
~242.01 MB
Download links issued
3

Source, license and coverage

Supplier documentation. These claims are separate from the automated sample score. A listing edit date is not a data freshness date.

License
cc-by-4.0
Source / creator
nvidia/Nemotron-Safety-Guard-Dataset-v3
Collection method
The dataset was primarily synthetically generated using NVIDIA's CultureGuard pipeline, which produces culturally and linguistically adapted safety training data across 12 languages. Prompts and responses are labeled against a safety taxonomy covering toxicity, harmful content, and other unsafe categories. See the linked arXiv paper and dataset card for full pipeline details, including translation, cultural adaptation, and quality filtering steps.
Coverage start
Not documented
Coverage end
Not documented
Data last updated
Not documented
Update schedule
Not documented

Source documentation ↗

License terms ↗

As a primarily synthetic dataset, distributions may not reflect real-world user prompts; coverage across the 12 languages is not guaranteed to be uniform in either volume or category coverage. Safety taxonomies are subjective and culture-specific — labels reflect NVIDIA/CultureGuard's taxonomy choices. Buyers training guard models should validate empirically on their target deployment distribution, and audit for any leakage against public safety benchmarks (e.g., ToxicChat, BeaverTails, XSTest). The supplier describes synthetic or modeled records. These should not be treated as verified real-world observations.

Sample structure score: 82.3 / 100

This automated check describes the inspected sample, not factual accuracy, legal rights, representativeness, or the quality of the entire dataset. It is not a customer rating.

Assessed 10 sample records (JSON) on 2026-10-09. All records in the provided sample were checked.

CheckPointsEvidence
Populated cells32.3 / 5071 of 110 top-level cells contain a value. Null, absent and blank values count as missing; zero and false count as populated.
Consistent value types30 / 3071 of 71 populated cells match their column's most common observed type. Types are inferred, not checked against real-world truth.
Consistent record shape20 / 2010 of 10 records have the expected fields. CSV/TSV use the header width; JSON uses the union of observed keys.
Field-level findings and improvements

Check missing cells and mixed types below. Document intentional missing values or mixed types in your field descriptions. Do not fill legitimate unknowns with invented values just to increase this score.

FieldMissing cellsMost common typeOther populated types
id0 / 10string0 / 10
prompt0 / 10string0 / 10
response8 / 10string0 / 2
prompt_label0 / 10string0 / 10
response_label8 / 10string0 / 2
violated_categories6 / 10string0 / 4
prompt_label_source0 / 10string0 / 10
response_label_source8 / 10string0 / 2
tag0 / 10string0 / 10
language0 / 10string0 / 10
reconstruction_id_if_redacted9 / 10number0 / 1
How the score is calculated, its limitations, and how to correct an assessment →

About this data

Multilingual safety dataset across 12 languages for training LLM safety guard models, generated via the CultureGuard pipeline.

Retrieve with your agent or Python

Create an account and configure DATABAZAAR_API_KEY. This example retrieves free or already purchased data; it never makes a purchase. For a multi-file dataset, choose a file index from the manifest.

Download the Python example
python3 retrieve-dataset.py 5a12abd4-c2ed-4eb0-9cab-11a84eb5fc2c --output dataset.bin
Full supplier documentation
## Overview The Nemotron-Safety-Guard-Dataset-v3 (formerly Nemotron-Content-Safety-Dataset-Multilingual-v1) is a large, high-quality safety dataset designed for training multilingual LLM safety guard models. It contains approximately 514,617 samples spanning 12 languages: English, Arabic, German, Spanish, French, Hindi, Italian, Japanese, Korean, Dutch, Thai, and Mandarin Chinese. The data is JSON-formatted and primarily synthetically generated via NVIDIA's CultureGuard pipeline for content moderation, toxicity detection, and safety classification training. ## Schema - prompt — string — user prompt to be evaluated for safety - response — string — model or assistant response paired with the prompt - prompt_label — string/categorical — safety label assigned to the prompt (safe/unsafe + category) - response_label — string/categorical — safety label assigned to the response - language — string — ISO language code (one of 12 supported languages) - violated_categories — list[string] — safety taxonomy categories violated (if any) - (+ additional metadata columns; see HF dataset card) ## Sources - nvidia/Nemotron-Safety-Guard-Dataset-v3 on Hugging Face — https://huggingface.co/datasets/nvidia/Nemotron-Safety-Guard-Dataset-v3 — License: CC-BY-4.0 - Associated paper: CultureGuard (arXiv:2508.01710) ## Methodology The dataset was primarily synthetically generated using NVIDIA's CultureGuard pipeline, which produces culturally and linguistically adapted safety training data across 12 languages. Prompts and responses are labeled against a safety taxonomy covering toxicity, harmful content, and other unsafe categories. See the linked arXiv paper and dataset card for full pipeline details, including translation, cultural adaptation, and quality filtering steps. ## Known gaps & limitations As a primarily synthetic dataset, distributions may not reflect real-world user prompts; coverage across the 12 languages is not guaranteed to be uniform in either volume or category coverage. Safety taxonomies are subjective and culture-specific — labels reflect NVIDIA/CultureGuard's taxonomy choices. Buyers training guard models should validate empirically on their target deployment distribution, and audit for any leakage against public safety benchmarks (e.g., ToxicChat, BeaverTails, XSTest). ## Intended use & out-of-scope - IS for: training and fine-tuning multilingual LLM safety/guard classifiers, content moderation evaluation, toxicity detection research, multilingual safety alignment. - NOT for: use as a stand-alone safety benchmark without holdout validation; not deduplicated against common safety eval suites — leakage risk if used naively for benchmark training; synthetic generation means it should not be treated as ground truth for real-world user behavior. _Federated dataset: 3 parquet shards, 242.0 MB total. Queries and downloads stream through the DataBazaar API._ Original supplier listing: Nemotron Safety Guard Dataset v3 (Multilingual LLM Safety, 12 Languages) NVIDIA's 514K-sample multilingual safety dataset for training LLM safety guard models across 12 languages, generated via the CultureGuard pipeline. CC-BY-4.0.

Schema

NameTypeDescription
idVARCHARUnique identifier for the sample (MD5 hash format).
promptVARCHARUser input text to be evaluated for safety across 12 languages.
responseVARCHARModel or assistant response paired with the prompt; null if not provided.
prompt_labelVARCHARSafety classification of prompt: safe or unsafe with optional violation category.
response_labelVARCHARSafety classification of response: safe, unsafe, or empty if not applicable.
violated_categoriesVARCHARComma-separated safety taxonomy categories breached (e.g., Profanity, Violence); empty if none.
prompt_label_sourceVARCHARAnnotation source for prompt label: human or automated method.
response_label_sourceVARCHARAnnotation source for response label: human, automated, or null if not labeled.
tagVARCHARDataset partition or content type tag (e.g., generic, adversarial).
languageVARCHARISO 639-1 language code (ar, de, en, es, fr, hi, it, ja, ko, nl, th, zh).
reconstruction_id_if_redactedDOUBLERow ID of original unredacted sample if this prompt was redacted; null otherwise.

Sample Data

Preview a sample of the data before downloading.

Public sample only. Sign in to retrieve the full dataset, including free datasets.

For AI Agents

Via MCP Server
# 1. Add to your agent's MCP config (claude_desktop_config.json or similar):
{
  "mcpServers": {
    "databazaar": { "command": "npx", "args": ["databazaar-mcp"] }
  }
}

# 2. Your agent can then call:
search_datasets({ query: "Nemotron Safety Guard Dataset " })
// Found: 5a12abd4-c2ed-4eb0-9cab-11a84eb5fc2c
get_download_url({ dataset_id: "5a12abd4-c2ed-4eb0-9cab-11a84eb5fc2c" })  // free — sign in with MCP OAuth first
Via REST API
# Free dataset — sign in or use your account API key:
curl https://api.databazaar.io/datasets/5a12abd4-c2ed-4eb0-9cab-11a84eb5fc2c/download-url -H "Authorization: Bearer $DATABAZAAR_API_KEY"