Nemotron Content Safety Audio Dataset
Source, license and coverage
Supplier documentation. These claims are separate from the automated sample score. A listing edit date is not a data freshness date.
- License
- cc-by-4.0
- Source / creator
- nvidia/Nemotron-Content-Safety-Audio-Dataset
- Collection method
- Nvidia constructed the dataset by taking the test split of the Aegis 2.0 content-safety prompts (expert- and human-annotated adversarial and benign prompts spanning 23 violation categories) and synthesizing spoken-audio versions of each prompt via text-to-speech. The resulting audio files are aligned 1:1 with the source text records, preserving original safety labels and taxonomy assignments so that audio-modality safety classifiers can be evaluated against the same ground truth as the text-modality benchmark.
- Coverage start
- Not documented
- Coverage end
- Not documented
- Data last updated
- Not documented
- Update schedule
- Not documented
All prompts are English-only, so the dataset does not support multilingual audio safety evaluation. Audio is TTS-synthesized rather than recorded from natural speakers, so it under-represents accent, prosody, background noise, and disfluency seen in real-world audio attacks. The 1,928 file size is suitable for evaluation but small for training. Distribution of violation categories inherits any class imbalance present in the Aegis 2.0 test set. As a safety dataset it contains adversarial and harmful textual content by design.
Sample structure score: 91.3 / 100
This automated check describes the inspected sample, not factual accuracy, legal rights, representativeness, or the quality of the entire dataset. It is not a customer rating.
Assessed 10 sample records (JSON) on 2026-10-09. All records in the provided sample were checked.
| Check | Points | Evidence |
|---|---|---|
| Populated cells | 41.3 / 50 | 99 of 120 top-level cells contain a value. Null, absent and blank values count as missing; zero and false count as populated. |
| Consistent value types | 30 / 30 | 99 of 99 populated cells match their column's most common observed type. Types are inferred, not checked against real-world truth. |
| Consistent record shape | 20 / 20 | 10 of 10 records have the expected fields. CSV/TSV use the header width; JSON uses the union of observed keys. |
Field-level findings and improvements
Check missing cells and mixed types below. Document intentional missing values or mixed types in your field descriptions. Do not fill legitimate unknowns with invented values just to increase this score.
| Field | Missing cells | Most common type | Other populated types |
|---|---|---|---|
| id | 0 / 10 | string | 0 / 10 |
| response | 6 / 10 | string | 0 / 4 |
| prompt_label | 0 / 10 | string | 0 / 10 |
| response_label | 6 / 10 | string | 0 / 4 |
| violated_categories | 3 / 10 | string | 0 / 7 |
| prompt_label_source | 0 / 10 | string | 0 / 10 |
| response_label_source | 6 / 10 | string | 0 / 4 |
| prompt | 0 / 10 | string | 0 / 10 |
| audio_filename | 0 / 10 | string | 0 / 10 |
| audio_duration_seconds | 0 / 10 | number | 0 / 10 |
| speaker_name | 0 / 10 | string | 0 / 10 |
| speaker_native_language | 0 / 10 | string | 0 / 10 |
About this data
English audio files covering 23 violation categories designed for evaluating multimodal content-safety guardrails. Extends NVIDIA's Aegis 2.0 benchmark into the audio modality with adversarial and safety-critical examples.
Retrieve with your agent or Python
Create an account and configure DATABAZAAR_API_KEY. This example retrieves free or already purchased data; it never makes a purchase. For a multi-file dataset, choose a file index from the manifest.
Download the Python examplepython3 retrieve-dataset.py 834e6916-fc56-4b24-b397-3d7658d73d8c --output dataset.bin
Full supplier documentation
Schema
| Name | Type | Description |
|---|---|---|
| id | VARCHAR | Unique identifier for the prompt-response pair. |
| response | VARCHAR | LLM-generated text response to the prompt. |
| prompt_label | VARCHAR | Safety classification of the prompt: safe or unsafe. |
| response_label | VARCHAR | Safety classification of the response: safe or unsafe. |
| violated_categories | VARCHAR | Comma-separated list of violated safety categories from Aegis 2.0 taxonomy. |
| prompt_label_source | VARCHAR | Annotation source for prompt label: human or llm_jury. |
| response_label_source | VARCHAR | Annotation source for response label: human or llm_jury. |
| prompt | VARCHAR | Original English text prompt. |
| audio_filename | VARCHAR | Filename of the spoken-prompt audio file (WAV format). |
| audio_duration_seconds | FLOAT | Length of the audio file in seconds. |
| speaker_name | VARCHAR | TTS voice identifier or speaker name. |
| speaker_native_language | VARCHAR | Native language of the voice model or speaker. |
Sample Data
Preview a sample of the data before downloading.
Public sample only. Sign in to retrieve the full dataset, including free datasets.
For AI Agents
# 1. Add to your agent's MCP config (claude_desktop_config.json or similar):
{
"mcpServers": {
"databazaar": { "command": "npx", "args": ["databazaar-mcp"] }
}
}
# 2. Your agent can then call:
search_datasets({ query: "Nemotron Content Safety Audio " })
// Found: 834e6916-fc56-4b24-b397-3d7658d73d8c
get_download_url({ dataset_id: "834e6916-fc56-4b24-b397-3d7658d73d8c" }) // free — sign in with MCP OAuth first# Free dataset — sign in or use your account API key: curl https://api.databazaar.io/datasets/834e6916-fc56-4b24-b397-3d7658d73d8c/download-url -H "Authorization: Bearer $DATABAZAAR_API_KEY"