textnvidia/Nemotron-Math-v2mathematicsreasoningllm-trainingdistillationlong-contexttool-usenvidiafine-tuningsynthetic-datacot

Nemotron-Math-v2 Mathematical Reasoning Trajectories

Free

Open dataset

Sample structure: 84.8 / 100
2 download links issued
Seller: DataBazaar
Sign up to download

Already have an account? Log in

Agent? Connect your account →

Category
Text
Records
7,085,839 rows
Format
PARQUET
Update Frequency
Not documented
Collection Method
auto_imported_huggingface_federated
PII
No flagged field names; not a privacy audit
File Size
~47995.71 MB
Download links issued
2

Source, license and coverage

Supplier documentation. These claims are separate from the automated sample score. A listing edit date is not a data freshness date.

License
cc-by-4.0,cc-by-sa-4.0
Source / creator
nvidia/Nemotron-Math-v2
Collection method
NVIDIA curated ~347K mathematical problems from a range of public math sources and generated multiple reasoning trajectories per problem using strong teacher models under multi-mode supervision (e.g., chain-of-thought and tool-augmented modes). Trajectories are intended to support long-context distillation into smaller student models. Refer to the accompanying paper and NeMo-Skills documentation for full collection, filtering, and quality-control procedures.
Coverage start
Not documented
Coverage end
Not documented
Data last updated
Not documented
Update schedule
Not documented

Source documentation ↗

The dataset is English-only and focused on mathematical reasoning, so it does not generalize to other domains or languages. Trajectories are synthetic (model-generated) and may contain reasoning errors, hallucinated steps, or biases inherited from the teacher model. Problem coverage overlaps with public math corpora (e.g., MATH, GSM8K-style data), creating potential leakage risk if used to train models evaluated on standard math benchmarks. Source does not exhaustively document deduplication against common eval suites; buyers should validate empirically before using for benchmark-adjacent training.

Sample structure score: 84.8 / 100

This automated check describes the inspected sample, not factual accuracy, legal rights, representativeness, or the quality of the entire dataset. It is not a customer rating.

Assessed 10 sample records (JSON) on 2026-10-09. All records in the provided sample were checked.

CheckPointsEvidence
Populated cells35.7 / 50100 of 140 top-level cells contain a value. Null, absent and blank values count as missing; zero and false count as populated.
Consistent value types29.1 / 3097 of 100 populated cells match their column's most common observed type. Types are inferred, not checked against real-world truth.
Consistent record shape20 / 2010 of 10 records have the expected fields. CSV/TSV use the header width; JSON uses the union of observed keys.
Field-level findings and improvements

Check missing cells and mixed types below. Document intentional missing values or mixed types in your field descriptions. Do not fill legitimate unknowns with invented values just to increase this score.

FieldMissing cellsMost common typeOther populated types
uuid0 / 10string0 / 10
expected_answer0 / 10string3 / 10
problem0 / 10string0 / 10
original_expected_answer10 / 10unknown0 / 0
changed_answer_to_majority0 / 10boolean0 / 10
data_source0 / 10string0 / 10
messages0 / 10object0 / 10
used_in0 / 10object0 / 10
metadata0 / 10object0 / 10
license0 / 10string0 / 10
tools0 / 10object0 / 10
url10 / 10unknown0 / 0
user_name10 / 10unknown0 / 0
user_url10 / 10unknown0 / 0
How the score is calculated, its limitations, and how to correct an assessment →

About this data

Collection of 347K mathematical problems paired with 7M model-generated reasoning trajectories for training mathematical reasoning. Includes long-context reasoning, tool-use integration, and multi-mode supervision signals.

Retrieve with your agent or Python

Create an account and configure DATABAZAAR_API_KEY. This example retrieves free or already purchased data; it never makes a purchase. For a multi-file dataset, choose a file index from the manifest.

Download the Python example
python3 retrieve-dataset.py 17600013-74d6-45d5-b618-c2f8b0afe9c8 --output dataset.bin
Full supplier documentation
## Overview Nemotron-Math-v2 is a large-scale mathematical reasoning dataset released by NVIDIA, containing approximately 347,000 high-quality mathematical problems paired with roughly 7 million model-generated reasoning trajectories. It accompanies the paper "Nemotron-Math: Efficient Long-Context Distillation of Mathematical Reasoning from Multi-Mode Supervision" (arXiv:2512.15489) and is distributed in Parquet format. The dataset supports long-context reasoning, tool use, and multi-mode supervision for training and distilling math-capable LLMs. ## Schema - problem — string — the mathematical problem statement - solution / trajectory — string — model-generated reasoning chain (often long-context) - answer — string — final answer (where applicable) - mode — string — supervision mode (e.g., CoT, tool-use) - source — string — origin or problem family - metadata — struct — additional per-record annotations - +additional columns documented on the source dataset card ## Sources - nvidia/Nemotron-Math-v2 on HuggingFace — https://huggingface.co/datasets/nvidia/Nemotron-Math-v2 — licensed CC-BY-4.0 and CC-BY-SA-4.0 - Companion paper: arXiv:2512.15489 - Code: NeMo-Skills (NVIDIA) ## Methodology NVIDIA curated ~347K mathematical problems from a range of public math sources and generated multiple reasoning trajectories per problem using strong teacher models under multi-mode supervision (e.g., chain-of-thought and tool-augmented modes). Trajectories are intended to support long-context distillation into smaller student models. Refer to the accompanying paper and NeMo-Skills documentation for full collection, filtering, and quality-control procedures. ## Known gaps & limitations The dataset is English-only and focused on mathematical reasoning, so it does not generalize to other domains or languages. Trajectories are synthetic (model-generated) and may contain reasoning errors, hallucinated steps, or biases inherited from the teacher model. Problem coverage overlaps with public math corpora (e.g., MATH, GSM8K-style data), creating potential leakage risk if used to train models evaluated on standard math benchmarks. Source does not exhaustively document deduplication against common eval suites; buyers should validate empirically before using for benchmark-adjacent training. ## Intended use & out-of-scope - Intended: fine-tuning and distilling mathematical reasoning into LLMs, long-context reasoning research, tool-use training, supervised trajectory learning, agent math evals. - Out-of-scope: not deduplicated against common math eval suites (MATH, GSM8K, AIME) — leakage risk for benchmark training; not suitable for non-math reasoning, non-English math, or as ground-truth answer key without verification. _Federated dataset: 5 parquet shards, 46.87 GB total. Queries and downloads stream through the DataBazaar API._ Original supplier listing: Nemotron-Math-v2: Mathematical Reasoning Trajectories (NVIDIA) NVIDIA's 347K math problems with 7M model-generated reasoning trajectories for distilling mathematical reasoning. Long-context, tool-use, multi-mode supervision. CC-BY-4.0/CC-BY-SA-4.0.

Schema

NameTypeDescription
uuidVARCHARUnique identifier (UUID v4 format) for the record.
expected_answerVARCHARFinal numerical or symbolic answer to the mathematical problem.
problemVARCHARMathematical problem statement in LaTeX or plain text format.
original_expected_answerVARCHARInitial expected answer before any corrections or majority voting.
changed_answer_to_majorityBOOLEANBoolean flag indicating if answer was updated to match majority model consensus.
data_sourceVARCHAROrigin dataset or problem collection (e.g., aops, competition, textbook).
messagesSTRUCT("role" VARCHAR, "content" VARCHAR, reasoning_content VARCHAR, tool_calls STRUCT(id VARCHAR, "type" VARCHAR, "function" STRUCT("name" VARCHAR, arguments VARCHAR))[], tool_call_id VARCHAR, "name" VARCHAR)[]Array of conversational turns with role, content, reasoning traces, and tool invocations.
used_inVARCHAR[]Array of dataset splits or benchmarks this record appears in.
metadataSTRUCT(reason_low_with_tool STRUCT(count BIGINT, pass BIGINT, accuracy DOUBLE), reason_low_no_tool STRUCT(count BIGINT, pass BIGINT, accuracy DOUBLE), reason_medium_with_tool STRUCT(count BIGINT, pass BIGINT, accuracy DOUBLE), reason_medium_no_tool STRUCT(count BIGINT, pass BIGINT, accuracy DOUBLE), reason_high_with_tool STRUCT(count BIGINT, pass BIGINT, accuracy DOUBLE), reason_high_no_tool STRUCT(count BIGINT, pass BIGINT, accuracy DOUBLE))Nested counts and accuracy metrics by reasoning difficulty level and tool availability.
licenseVARCHARLicense identifier governing dataset usage (e.g., CC-BY-4.0, CC-BY-SA-4.0).
toolsSTRUCT("type" VARCHAR, "function" STRUCT("name" VARCHAR, description VARCHAR, parameters STRUCT("type" VARCHAR, properties STRUCT(code STRUCT("type" VARCHAR, description VARCHAR)), required VARCHAR[])))[]Array of available function definitions with parameters for tool-use trajectories.
urlVARCHARSource URL or reference link for the problem.
user_nameVARCHARUsername or author identifier of problem contributor or source.
user_urlVARCHARProfile or homepage URL of the user who contributed the problem.

Sample Data

Preview a sample of the data before downloading.

Public sample only. Sign in to retrieve the full dataset, including free datasets.

For AI Agents

Via MCP Server
# 1. Add to your agent's MCP config (claude_desktop_config.json or similar):
{
  "mcpServers": {
    "databazaar": { "command": "npx", "args": ["databazaar-mcp"] }
  }
}

# 2. Your agent can then call:
search_datasets({ query: "Nemotron-Math-v2 Mathematical " })
// Found: 17600013-74d6-45d5-b618-c2f8b0afe9c8
get_download_url({ dataset_id: "17600013-74d6-45d5-b618-c2f8b0afe9c8" })  // free — sign in with MCP OAuth first
Via REST API
# Free dataset — sign in or use your account API key:
curl https://api.databazaar.io/datasets/17600013-74d6-45d5-b618-c2f8b0afe9c8/download-url -H "Authorization: Bearer $DATABAZAAR_API_KEY"