To search datasets programmatically: GET https://api.databazaar.io/datasets?query=your-search
Full API docs: https://api.databazaar.io/llms.txt
Agent discovery: https://databazaar.io/.well-known/agent.json
Browse Data
49–72 of 93HelpSteer: NVIDIA Helpfulness Preference Dataset
NVIDIA's open helpfulness dataset with multi-attribute human ratings (helpfulness, correctness, coherence, complexity, verbosity) for prompts and responses. Used to train SteerLM-aligned LLMs. CC-BY-4.0.
OpenThoughts-114k: Synthetic Reasoning Dataset
114k high-quality synthetic reasoning examples across math, science, code, and puzzles. Used to fine-tune OpenThinker-7B/32B models. Apache 2.0 licensed, parquet format.
ClimbMix: 400B-Token Pre-training Corpus (NVIDIA)
A 400-billion-token English pre-training corpus from NVIDIA, filtered and topic-mixed for efficient LLM pre-training with superior performance per token.
MATH: Competition Mathematics Problems with Step-by-Step Solutions
12,500 competition math problems (AMC 10/12, AIME, etc.) with full worked solutions. Standard benchmark for math reasoning, chain-of-thought training, and LLM evaluation. MIT licensed, parquet format.
AgentTrove: 1.7M Agentic Interaction Traces
Largest open-source collection of agentic interaction traces (1.7M rows from 219 source datasets) covering code repair, shell scripting, math, competitive programming, and computer-use tasks. Apache 2.0.
Sangraha — 251B Token Indic Language Pretraining Corpus (22 Languages)
Largest cleaned Indic-language pretraining dataset: 251B tokens across 22 Indian languages, curated from web sources, multilingual corpora, and large-scale translations. CC-BY-4.0, parquet format.
OpenAssistant Conversations (OASST1)
161,443 human-generated assistant conversation messages across 35 languages with 461,292 quality ratings — a foundational dataset for alignment, RLHF, and instruction-tuning research.
BEIR: MS MARCO Passage Retrieval Benchmark
MS MARCO subset of the BEIR heterogeneous IR benchmark — 8.8M passages and ~500K queries in standard retrieval layout. Parquet format, CC-BY-SA-4.0, English.
UltraFeedback — Large-Scale Fine-Grained Preference Dataset
64k prompts × 4 LLM responses (256k samples) with fine-grained GPT-4 preference annotations across instruction-following, truthfulness, honesty, and helpfulness. Canonical dataset for reward model and RLHF training.
Dolci Instruct SFT Mixture (Olmo 3 Training Data)
2.15M-sample multilingual instruction SFT mixture used to train AI2's Olmo 3 7B Instruct. Parquet format, ODC-BY licensed, covers 70+ languages and includes OpenThoughts 3 and other curated prompt sets.
WikiText (WikiText-2 & WikiText-103) Language Modeling Corpus
Canonical language modeling benchmark: 100M+ tokens from verified Good/Featured Wikipedia articles. Includes WikiText-2 and WikiText-103 in raw and tokenized variants. Parquet format, CC-BY-SA 3.0.
KILT: Knowledge-Intensive Language Tasks Benchmark
Facebook AI's KILT benchmark — 11 datasets across fact-checking, entity linking, slot filling, open-domain QA, and dialog generation, all grounded in a unified Wikipedia snapshot. MIT licensed, parquet format, 1M-10M examples.
FineWeb-Edu Translated (36 Languages, 960B+ Tokens)
Machine-translated FineWeb-Edu corpus covering 36.7M aligned documents across 36 European languages — 960B+ tokens for multilingual LLM pretraining and translation training.
People's Speech — 30,000+ Hours of Transcribed English Speech (MLCommons)
One of the world's largest open English ASR corpora: 30,000+ hours of transcribed speech under CC-BY/CC-BY-SA. Built by MLCommons for training and evaluating speech-to-text systems.
MS MARCO — Question Answering & Passage Ranking
Microsoft's MS MARCO dataset: 1M+ real Bing questions with human-generated answers and passage relevance judgments. Foundational benchmark for retrieval, RAG, and QA systems.
UltraChat — Large-Scale Multi-Round Dialogue Dataset
Open-source large-scale multi-round dialogue dataset (1.5M+ conversations) generated via dual ChatGPT Turbo API role-play. Widely used for instruction tuning and chat model fine-tuning. MIT licensed.
MATH-500 (HuggingFaceH4)
500-problem subset of the MATH benchmark used by OpenAI in 'Let's Verify Step by Step' — the canonical math reasoning eval for LLMs and reasoning models.
BigCodeBench — Code Generation Benchmark (Complete & Instruct)
1,140-task code generation benchmark from BigCode with both docstring-based completion and NL-instruction variants, 99% test coverage, Apache-2.0 licensed.
WildChat-1M: 1 Million Real Human-ChatGPT Conversations
1M real-world conversations between human users and ChatGPT (GPT-3.5/4), with demographics, timestamps, languages, and toxicity labels. Parquet format, ODC-BY licensed. Widely used for instruction tuning, eval, and alignment research.
HelpSteer3 — NVIDIA Human Preference & Feedback Dataset for RLHF/Reward Modeling
NVIDIA's open-source multilingual preference dataset for training reward models and aligning LLMs. CC-BY-4.0, 100K+ samples across 15 languages, used to train SOTA reward models on RM-Bench (85.5%) and JudgeBench (78.6%).
Alpaca Cleaned — Instruction Fine-Tuning Dataset
Cleaned version of Stanford's Alpaca instruction-following dataset (~52K examples). Fixes hallucinations, merged instructions, empty outputs, and other quality issues. CC-BY-4.0, ready for LLM fine-tuning.
OpenR1-Math-220k: Math Reasoning Traces from DeepSeek R1
220k math problems with verified DeepSeek R1 reasoning traces, sourced from NuminaMath 1.5. Apache-2.0 licensed dataset for training and evaluating mathematical reasoning models.
Nemotron Content Safety Dataset V2 (Aegis 2.0)
33,416 annotated human-LLM interactions for content safety classification across 12+ harm categories. Used for training and evaluating LLM safety guardrails like NeMo Guard.
Hermes Agent Reasoning Traces (Kimi-K2.5 & GLM-5.1)
14,700+ multi-turn agent tool-calling trajectories with step-by-step reasoning traces and real tool execution results, generated via the Hermes Agent harness from Kimi-K2.5 and GLM-5.1 models. Apache-2.0.