To search datasets programmatically: GET https://api.databazaar.io/datasets?query=your-search
Full API docs: https://api.databazaar.io/llms.txt
Agent discovery: https://databazaar.io/.well-known/agent.json
Browse Data
1–24 of 94OpenPII — 1.4M Multilingual PII Masking Examples
1,428,143 synthetic text examples with fine-grained PII annotations spanning 23 European languages. Each example pairs source and masked text with privacy masks and token-level labels, designed for training privacy-redaction and anonymization models.
OPUS-100 — English-Centric Parallel Corpus (100 Languages)
A large English-centric multilingual translation corpus covering 100 languages, where every training pair includes English on the source or target side. A standard resource for machine-translation research and multilingual model training.
Nemotron Cascade-2 — SFT Training Data (Math, Code & Chat)
~1.94M supervised fine-tuning examples used to train NVIDIA's Nemotron-Cascade-2 models, spanning math, code, and conversational prompts paired with model-generated responses and source/domain labels.
WideSearch — Agentic Broad Information-Seeking Benchmark
A 200-instance benchmark for evaluating LLM-driven agents on broad, wide-coverage information-seeking tasks. Each instance provides a query, structured evaluation criteria, and language metadata for measuring agent search performance.
Emotion — English Tweets Labeled by Emotion
English Twitter messages labeled with six basic emotions — anger, fear, joy, love, sadness, and surprise. A widely used benchmark for text-based emotion classification and sentiment research.
OpenMath GSM8K (Masked) — Grade-School Math Solutions
A masked version of the GSM8K grade-school math word problems (7,473 items), pairing each question with masked and full reference solutions plus the expected answer. Designed to support consistent synthetic generation of additional solutions.
MiniPile — 1M-Document Deduplicated Text Corpus
A compact 6 GB subset of the deduplicated Pile corpus (~1,010,500 documents), curated for data-efficient language-model pretraining and experimentation without the storage cost of the full Pile.
PII Masking Dataset — 300K Labeled Text Samples
225,000+ text samples annotated for personally identifiable information (PII) detection and masking, built to train and evaluate models that redact sensitive data from text. Each record pairs source and target text with privacy masks, span labels, and token-level annotations across multiple languages.
Stack Overflow Posts — 58M Questions & Answers (Markdown)
Every Stack Overflow post submitted before June 14, 2023 — roughly 58 million questions and answers (~35 GB) formatted as Markdown. Includes scores, tags, view counts, creation/edit timestamps, and full post metadata.
Waveform-5000 (UCI / OpenML)
Classic 3-class synthetic waveform classification benchmark with 5,000 instances and 40 numeric attributes (21 informative + 19 noise). Originally from Breiman et al. 1984, distributed via UCI and OpenML.
Nemotron Content Safety Audio Dataset (Aegis 2.0 Multimodal)
1,928 English audio files of adversarial and safety-critical prompts across 23 violation categories, extending Nvidia's Aegis 2.0 content-safety benchmark into the audio modality for multimodal guardrail evaluation.
Hindawi Arabic Books — Section-Level NLP Dataset
Cleaned, section-level Arabic text from Hindawi.org books spanning literature, philosophy, history, and science. 10K-100K rows in Parquet format, prepared for Arabic NLP training and research.
MInDS-14: Multilingual Spoken Intent Detection (14 Languages, e-Banking)
Spoken intent detection benchmark covering 14 e-banking intents across 14 language varieties. Audio + transcriptions in parquet format, ideal for speech understanding evals and multilingual ASR/NLU fine-tuning.
AG News - Topic Classification Benchmark
Classic 4-class news topic classification dataset (~127K articles across World, Sports, Business, Sci/Tech). Standard benchmark for text classification, fine-tuning, and NLP evals.
Go-Code-Large: 316K Go Source Code Samples
Large-scale corpus of 316,427 Go (Golang) source code samples in JSONL format. Curated for LLM pretraining, code generation fine-tuning, and static analysis research on cloud-native and backend systems.
TaskTrove — 750K+ Agentic Tasks for RL & SFT Training
Open-source collection of 750,000+ unique agentic tasks aggregated from 100+ sources including SWE-Smith, R2EGym, and SWE-Re-Bench. Apache-2.0 licensed, parquet format, designed for agent training and evaluation.
Bhasha SFT — 13M+ Multilingual Instruction-Response Pairs (Hindi, Bengali, Gujarati, English)
13M+ instruction-response pairs across Hindi, Bengali, Gujarati, and English for supervised fine-tuning of multilingual LLMs. Mix of human-annotated and synthetic data from open-source SFT collections, curated by Soket AI Labs.
SciCode Domain Code: 1.1B+ Lines of Scientific Code Across 178 Domains
Large-scale domain-specific code dataset (~115 GB, 1.1B+ lines) from GitHub covering biology, chemistry, materials science, physics, and 174 other scientific domains. Apache 2.0 licensed.
HPDv3 — Human Preference Dataset v3 (1.08M text-image pairs)
Wide-spectrum human preference dataset for text-to-image generation: 1.08M text-image pairs and 1.17M pairwise human preference annotations. MIT-licensed, used to train HPSv3 (ICCV 2025) reward models.
LongCoT: Long-Horizon Reasoning Benchmark
Benchmark for evaluating sustained long chain-of-thought reasoning across logic, computer science, chemistry, chess, and mathematics. Parquet format, MIT licensed.
UGMathBench: Undergraduate Math Reasoning Benchmark
5,062 undergraduate-level math problems across 16 subjects and 111 topics, with 10 answer types and 3 randomized versions each. Designed for evaluating LLM mathematical reasoning.
Common Crawl Creative Commons Fine (C5f) — Multilingual High-Quality Web Corpus
Filtered Creative Commons web corpus from Common Crawl, intersected with FineWeb/FineWeb-2 for quality. Multilingual (EN, DE, FR, NL, ES, IT, AF, FY), 10M-100M rows, Parquet format. Ideal for LLM pretraining and fine-tuning.
MASSIVE: Multilingual NLU Dataset (51 Languages, 1M+ Utterances)
Parallel multilingual NLU benchmark from Amazon Science with 1M+ utterances across 51 languages, annotated with 60 intents and 55 slot types. Built by localizing SLURP voice assistant interactions.
Salad-Data: LLM Safety & Jailbreak Evaluation Benchmark
21K+ safety questions for LLM red-teaming and jailbreak evaluation, aggregated from HH-RLHF, AdvBench, ToxicChat, GPTFuzzer, and GPT-3.5 self-instructed prompts. Apache-2.0 licensed.