To search datasets programmatically: GET https://api.databazaar.io/datasets?query=your-search

Full API docs: https://api.databazaar.io/llms.txt

Agent discovery: https://databazaar.io/.well-known/agent.json

Browse Data

49–72 of 94
Filters
Newest
text
Freefixed price

HelpSteer: NVIDIA Helpfulness Preference Dataset

NVIDIA's open helpfulness dataset with multi-attribute human ratings (helpfulness, correctness, coherence, complexity, verbosity) for prompts and responses. Used to train SteerLM-aligned LLMs. CC-BY-4.0.

37,120 rows·PARQUET·1 downloads
text
Freefixed price

OpenThoughts-114k: Synthetic Reasoning Dataset

114k high-quality synthetic reasoning examples across math, science, code, and puzzles. Used to fine-tune OpenThinker-7B/32B models. Apache 2.0 licensed, parquet format.

227,914 rows·PARQUET·3 downloads
text
Freefixed price

ClimbMix: 400B-Token Pre-training Corpus (NVIDIA)

A 400-billion-token English pre-training corpus from NVIDIA, filtered and topic-mixed for efficient LLM pre-training with superior performance per token.

1,794,054 rows·PARQUET·2 downloads
text
Freefixed price

MATH: Competition Mathematics Problems with Step-by-Step Solutions

12,500 competition math problems (AMC 10/12, AIME, etc.) with full worked solutions. Standard benchmark for math reasoning, chain-of-thought training, and LLM evaluation. MIT licensed, parquet format.

12,500 rows·PARQUET·1 downloads
text
Freefixed price

AgentTrove: 1.7M Agentic Interaction Traces

Largest open-source collection of agentic interaction traces (1.7M rows from 219 source datasets) covering code repair, shell scripting, math, competitive programming, and computer-use tasks. Apache 2.0.

1,696,847 rows·PARQUET·1 downloads
text
Freefixed price

Sangraha — 251B Token Indic Language Pretraining Corpus (22 Languages)

Largest cleaned Indic-language pretraining dataset: 251B tokens across 22 Indian languages, curated from web sources, multilingual corpora, and large-scale translations. CC-BY-4.0, parquet format.

177,355,164 rows·PARQUET·1 downloads
text
Freefixed price

OpenAssistant Conversations (OASST1)

161,443 human-generated assistant conversation messages across 35 languages with 461,292 quality ratings — a foundational dataset for alignment, RLHF, and instruction-tuning research.

88,838 rows·PARQUET·0 downloads
text
Freefixed price

BEIR: MS MARCO Passage Retrieval Benchmark

MS MARCO subset of the BEIR heterogeneous IR benchmark — 8.8M passages and ~500K queries in standard retrieval layout. Parquet format, CC-BY-SA-4.0, English.

9,351,785 rows·PARQUET·1 downloads
text
Freefixed price

UltraFeedback — Large-Scale Fine-Grained Preference Dataset

64k prompts × 4 LLM responses (256k samples) with fine-grained GPT-4 preference annotations across instruction-following, truthfulness, honesty, and helpfulness. Canonical dataset for reward model and RLHF training.

63,967 rows·PARQUET·1 downloads
text
Freefixed price

Dolci Instruct SFT Mixture (Olmo 3 Training Data)

2.15M-sample multilingual instruction SFT mixture used to train AI2's Olmo 3 7B Instruct. Parquet format, ODC-BY licensed, covers 70+ languages and includes OpenThoughts 3 and other curated prompt sets.

2,152,112 rows·PARQUET·1 downloads
text
Freefixed price

WikiText (WikiText-2 & WikiText-103) Language Modeling Corpus

Canonical language modeling benchmark: 100M+ tokens from verified Good/Featured Wikipedia articles. Includes WikiText-2 and WikiText-103 in raw and tokenized variants. Parquet format, CC-BY-SA 3.0.

3,708,608 rows·PARQUET·0 downloads
text
Freefixed price

KILT: Knowledge-Intensive Language Tasks Benchmark

Facebook AI's KILT benchmark — 11 datasets across fact-checking, entity linking, slot filling, open-domain QA, and dialog generation, all grounded in a unified Wikipedia snapshot. MIT licensed, parquet format, 1M-10M examples.

3,231,786 rows·PARQUET·1 downloads
text
Freefixed price

FineWeb-Edu Translated (36 Languages, 960B+ Tokens)

Machine-translated FineWeb-Edu corpus covering 36.7M aligned documents across 36 European languages — 960B+ tokens for multilingual LLM pretraining and translation training.

1,999,563,091 rows·PARQUET·2 downloads
text
Freefixed price

People's Speech — 30,000+ Hours of Transcribed English Speech (MLCommons)

One of the world's largest open English ASR corpora: 30,000+ hours of transcribed speech under CC-BY/CC-BY-SA. Built by MLCommons for training and evaluating speech-to-text systems.

8,051,212 rows·PARQUET·1 downloads
text
Freefixed price

MS MARCO — Question Answering & Passage Ranking

Microsoft's MS MARCO dataset: 1M+ real Bing questions with human-generated answers and passage relevance judgments. Foundational benchmark for retrieval, RAG, and QA systems.

1,112,939 rows·PARQUET·0 downloads
text
Freefixed price

UltraChat — Large-Scale Multi-Round Dialogue Dataset

Open-source large-scale multi-round dialogue dataset (1.5M+ conversations) generated via dual ChatGPT Turbo API role-play. Widely used for instruction tuning and chat model fine-tuning. MIT licensed.

773,913 rows·PARQUET·0 downloads
text
Freefixed price

MATH-500 (HuggingFaceH4)

500-problem subset of the MATH benchmark used by OpenAI in 'Let's Verify Step by Step' — the canonical math reasoning eval for LLMs and reasoning models.

500 rows·PARQUET·0 downloads
text
Freefixed price

BigCodeBench — Code Generation Benchmark (Complete & Instruct)

1,140-task code generation benchmark from BigCode with both docstring-based completion and NL-instruction variants, 99% test coverage, Apache-2.0 licensed.

5,700 rows·PARQUET·1 downloads
text
Freefixed price

WildChat-1M: 1 Million Real Human-ChatGPT Conversations

1M real-world conversations between human users and ChatGPT (GPT-3.5/4), with demographics, timestamps, languages, and toxicity labels. Parquet format, ODC-BY licensed. Widely used for instruction tuning, eval, and alignment research.

837,989 rows·PARQUET·1 downloads
text
Freefixed price

HelpSteer3 — NVIDIA Human Preference & Feedback Dataset for RLHF/Reward Modeling

NVIDIA's open-source multilingual preference dataset for training reward models and aligning LLMs. CC-BY-4.0, 100K+ samples across 15 languages, used to train SOTA reward models on RM-Bench (85.5%) and JudgeBench (78.6%).

132,937 rows·PARQUET·1 downloads
text
Freefixed price

Alpaca Cleaned — Instruction Fine-Tuning Dataset

Cleaned version of Stanford's Alpaca instruction-following dataset (~52K examples). Fixes hallucinations, merged instructions, empty outputs, and other quality issues. CC-BY-4.0, ready for LLM fine-tuning.

51,760 rows·PARQUET·1 downloads
text
Freefixed price

OpenR1-Math-220k: Math Reasoning Traces from DeepSeek R1

220k math problems with verified DeepSeek R1 reasoning traces, sourced from NuminaMath 1.5. Apache-2.0 licensed dataset for training and evaluating mathematical reasoning models.

450,258 rows·PARQUET·2 downloads
text
Freefixed price

Nemotron Content Safety Dataset V2 (Aegis 2.0)

33,416 annotated human-LLM interactions for content safety classification across 12+ harm categories. Used for training and evaluating LLM safety guardrails like NeMo Guard.

33,416 rows·PARQUET·1 downloads
text
Freefixed price

Hermes Agent Reasoning Traces (Kimi-K2.5 & GLM-5.1)

14,700+ multi-turn agent tool-calling trajectories with step-by-step reasoning traces and real tool execution results, generated via the Hermes Agent harness from Kimi-K2.5 and GLM-5.1 models. Apache-2.0.

14,701 rows·PARQUET·1 downloads