To search datasets programmatically: GET https://api.databazaar.io/datasets?query=your-search

Full API docs: https://api.databazaar.io/llms.txt

Agent discovery: https://databazaar.io/.well-known/agent.json

Browse Data

1–24 of 93
Filters1
Newest
text
Freefixed price

OpenPII — 1.4M Multilingual PII Masking Examples

1,428,143 synthetic text examples with fine-grained PII annotations spanning 23 European languages. Each example pairs source and masked text with privacy masks and token-level labels, designed for training privacy-redaction and anonymization models.

1,428,143 rows·PARQUET·1 downloads
text
Freefixed price

OPUS-100 — English-Centric Parallel Corpus (100 Languages)

A large English-centric multilingual translation corpus covering 100 languages, where every training pair includes English on the source or target side. A standard resource for machine-translation research and multilingual model training.

55,057,504 rows·PARQUET·2 downloads
text
Freefixed price

Nemotron Cascade-2 — SFT Training Data (Math, Code & Chat)

~1.94M supervised fine-tuning examples used to train NVIDIA's Nemotron-Cascade-2 models, spanning math, code, and conversational prompts paired with model-generated responses and source/domain labels.

1,940,375 rows·PARQUET·2 downloads
text
Freefixed price

WideSearch — Agentic Broad Information-Seeking Benchmark

A 200-instance benchmark for evaluating LLM-driven agents on broad, wide-coverage information-seeking tasks. Each instance provides a query, structured evaluation criteria, and language metadata for measuring agent search performance.

200 rows·PARQUET·2 downloads
text
Freefixed price

Emotion — English Tweets Labeled by Emotion

English Twitter messages labeled with six basic emotions — anger, fear, joy, love, sadness, and surprise. A widely used benchmark for text-based emotion classification and sentiment research.

436,809 rows·PARQUET·2 downloads
text
Freefixed price

OpenMath GSM8K (Masked) — Grade-School Math Solutions

A masked version of the GSM8K grade-school math word problems (7,473 items), pairing each question with masked and full reference solutions plus the expected answer. Designed to support consistent synthetic generation of additional solutions.

7,473 rows·PARQUET·0 downloads
text
Freefixed price

MiniPile — 1M-Document Deduplicated Text Corpus

A compact 6 GB subset of the deduplicated Pile corpus (~1,010,500 documents), curated for data-efficient language-model pretraining and experimentation without the storage cost of the full Pile.

1,010,500 rows·PARQUET·2 downloads
text
Freefixed price

PII Masking Dataset — 300K Labeled Text Samples

225,000+ text samples annotated for personally identifiable information (PII) detection and masking, built to train and evaluate models that redact sensitive data from text. Each record pairs source and target text with privacy masks, span labels, and token-level annotations across multiple languages.

225,405 rows·PARQUET·2 downloads
text
Freefixed price

Stack Overflow Posts — 58M Questions & Answers (Markdown)

Every Stack Overflow post submitted before June 14, 2023 — roughly 58 million questions and answers (~35 GB) formatted as Markdown. Includes scores, tags, view counts, creation/edit timestamps, and full post metadata.

58,329,355 rows·PARQUET·2 downloads
text
Freefixed price

Waveform-5000 (UCI / OpenML)

Classic 3-class synthetic waveform classification benchmark with 5,000 instances and 40 numeric attributes (21 informative + 19 noise). Originally from Breiman et al. 1984, distributed via UCI and OpenML.

5,000 rows·PARQUET·3 downloads
text
Freefixed price

Nemotron Content Safety Audio Dataset (Aegis 2.0 Multimodal)

1,928 English audio files of adversarial and safety-critical prompts across 23 violation categories, extending Nvidia's Aegis 2.0 content-safety benchmark into the audio modality for multimodal guardrail evaluation.

1,928 rows·PARQUET·1 downloads
text
Freefixed price

Hindawi Arabic Books — Section-Level NLP Dataset

Cleaned, section-level Arabic text from Hindawi.org books spanning literature, philosophy, history, and science. 10K-100K rows in Parquet format, prepared for Arabic NLP training and research.

52,830 rows·PARQUET·2 downloads
text
Freefixed price

MInDS-14: Multilingual Spoken Intent Detection (14 Languages, e-Banking)

Spoken intent detection benchmark covering 14 e-banking intents across 14 language varieties. Audio + transcriptions in parquet format, ideal for speech understanding evals and multilingual ASR/NLU fine-tuning.

16,336 rows·PARQUET·1 downloads
text
Freefixed price

AG News - Topic Classification Benchmark

Classic 4-class news topic classification dataset (~127K articles across World, Sports, Business, Sci/Tech). Standard benchmark for text classification, fine-tuning, and NLP evals.

127,600 rows·PARQUET·1 downloads
text
Freefixed price

Go-Code-Large: 316K Go Source Code Samples

Large-scale corpus of 316,427 Go (Golang) source code samples in JSONL format. Curated for LLM pretraining, code generation fine-tuning, and static analysis research on cloud-native and backend systems.

316,427 rows·PARQUET·1 downloads
text
Freefixed price

TaskTrove — 750K+ Agentic Tasks for RL & SFT Training

Open-source collection of 750,000+ unique agentic tasks aggregated from 100+ sources including SWE-Smith, R2EGym, and SWE-Re-Bench. Apache-2.0 licensed, parquet format, designed for agent training and evaluation.

17,191 rows·PARQUET·1 downloads
text
Freefixed price

Bhasha SFT — 13M+ Multilingual Instruction-Response Pairs (Hindi, Bengali, Gujarati, English)

13M+ instruction-response pairs across Hindi, Bengali, Gujarati, and English for supervised fine-tuning of multilingual LLMs. Mix of human-annotated and synthetic data from open-source SFT collections, curated by Soket AI Labs.

18,139,035 rows·PARQUET·1 downloads
text
Freefixed price

SciCode Domain Code: 1.1B+ Lines of Scientific Code Across 178 Domains

Large-scale domain-specific code dataset (~115 GB, 1.1B+ lines) from GitHub covering biology, chemistry, materials science, physics, and 174 other scientific domains. Apache 2.0 licensed.

155,855 rows·PARQUET·1 downloads
text
Freefixed price

HPDv3 — Human Preference Dataset v3 (1.08M text-image pairs)

Wide-spectrum human preference dataset for text-to-image generation: 1.08M text-image pairs and 1.17M pairwise human preference annotations. MIT-licensed, used to train HPSv3 (ICCV 2025) reward models.

1,154,324 rows·PARQUET·2 downloads
text
Freefixed price

LongCoT: Long-Horizon Reasoning Benchmark

Benchmark for evaluating sustained long chain-of-thought reasoning across logic, computer science, chemistry, chess, and mathematics. Parquet format, MIT licensed.

5,004 rows·PARQUET·1 downloads
text
Freefixed price

UGMathBench: Undergraduate Math Reasoning Benchmark

5,062 undergraduate-level math problems across 16 subjects and 111 topics, with 10 answer types and 3 randomized versions each. Designed for evaluating LLM mathematical reasoning.

5,061 rows·PARQUET·1 downloads
text
Freefixed price

Common Crawl Creative Commons Fine (C5f) — Multilingual High-Quality Web Corpus

Filtered Creative Commons web corpus from Common Crawl, intersected with FineWeb/FineWeb-2 for quality. Multilingual (EN, DE, FR, NL, ES, IT, AF, FY), 10M-100M rows, Parquet format. Ideal for LLM pretraining and fine-tuning.

75,055,472 rows·PARQUET·2 downloads
text
Freefixed price

MASSIVE: Multilingual NLU Dataset (51 Languages, 1M+ Utterances)

Parallel multilingual NLU benchmark from Amazon Science with 1M+ utterances across 51 languages, annotated with 60 intents and 55 slot types. Built by localizing SLURP voice assistant interactions.

2,560,755 rows·PARQUET·1 downloads
text
Freefixed price

Salad-Data: LLM Safety & Jailbreak Evaluation Benchmark

21K+ safety questions for LLM red-teaming and jailbreak evaluation, aggregated from HH-RLHF, AdvBench, ToxicChat, GPTFuzzer, and GPT-3.5 self-instructed prompts. Apache-2.0 licensed.

30,358 rows·PARQUET·1 downloads