textHuggingFaceH4/MATH-500mathreasoningbenchmarkllm-evalmathematicsprm800kopenaitext-generation

MATH-500 Benchmark

Free

Open dataset

Sample structure: 97.5 / 100
3 download links issued
Seller: DataBazaar
Sign up to download

Already have an account? Log in

Agent? Connect your account →

Category
Text
Records
500 rows
Format
PARQUET
Update Frequency
Not documented
Collection Method
uploaded
PII
No flagged field names; not a privacy audit
File Size
~0.2 MB
Download links issued
3

Source, license and coverage

Supplier documentation. These claims are separate from the automated sample score. A listing edit date is not a data freshness date.

License
Not documented — confirm reuse terms with the seller
Source / creator
HuggingFaceH4/MATH-500
Collection method
Problems were originally scraped from US high-school math competitions (AMC, AIME, etc.) by Hendrycks et al. and annotated with step-by-step solutions in LaTeX. OpenAI then sampled 500 problems from the MATH test split to form a stable, smaller evaluation subset for their process-reward-model work; the selection is published in PRM800K. HuggingFace H4 mirrors that exact subset as JSON with no transformations beyond format conversion.
Coverage start
Not documented
Coverage end
Not documented
Data last updated
Not documented
Update schedule
Not documented

Source documentation ↗

This is a small evaluation set (n=500), not a training corpus. It is widely used and almost certainly contaminated in the pretraining data of most frontier LLMs, so scores should be interpreted as upper-bound capability indicators rather than held-out generalization. Coverage is English-only, US competition-math style, and skews toward symbolic manipulation rather than applied or proof-based math. Answers are extracted from \boxed{} expressions, so grading requires careful normalization (equivalent forms, LaTeX whitespace). The source does not document selection bias in how OpenAI chose the 500.

Sample structure score: 97.5 / 100

This automated check describes the inspected sample, not factual accuracy, legal rights, representativeness, or the quality of the entire dataset. It is not a customer rating.

Assessed 10 sample records (JSON) on 2026-10-09. All records in the provided sample were checked.

CheckPointsEvidence
Populated cells50 / 5060 of 60 top-level cells contain a value. Null, absent and blank values count as missing; zero and false count as populated.
Consistent value types27.5 / 3055 of 60 populated cells match their column's most common observed type. Types are inferred, not checked against real-world truth.
Consistent record shape20 / 2010 of 10 records have the expected fields. CSV/TSV use the header width; JSON uses the union of observed keys.
Field-level findings and improvements

Check missing cells and mixed types below. Document intentional missing values or mixed types in your field descriptions. Do not fill legitimate unknowns with invented values just to increase this score.

FieldMissing cellsMost common typeOther populated types
problem0 / 10string0 / 10
solution0 / 10string0 / 10
answer0 / 10number5 / 10
subject0 / 10string0 / 10
level0 / 10number0 / 10
unique_id0 / 10string0 / 10
How the score is calculated, its limitations, and how to correct an assessment →

About this data

500-problem subset of the MATH benchmark used in mathematical reasoning evaluation for language models and reasoning systems. Covers diverse mathematics domains and difficulty levels.

Retrieve with your agent or Python

Create an account and configure DATABAZAAR_API_KEY. This example retrieves free or already purchased data; it never makes a purchase. For a multi-file dataset, choose a file index from the manifest.

Download the Python example
python3 retrieve-dataset.py 9c009b4d-3ecb-4b9f-9f06-6f61b3a36a9a --output dataset.bin
Full supplier documentation
## Overview MATH-500 is a 500-problem subset of the Hendrycks MATH benchmark, curated by OpenAI for the 'Let's Verify Step by Step' paper and redistributed by HuggingFace H4. It is the de-facto standard evaluation set for measuring mathematical reasoning capability in modern LLMs and reasoning models (o1, R1, etc.). The dataset is small (500 rows), text-only English, JSON-formatted, and covers competition-level problems across algebra, geometry, number theory, counting & probability, intermediate algebra, prealgebra, and precalculus, each with difficulty levels 1–5 and a ground-truth solution. ## Schema - problem — string — the math problem statement (LaTeX-formatted) - solution — string — full worked solution with final answer in \boxed{} - answer — string — extracted final answer - subject — string — topic category (e.g. Algebra, Geometry, Number Theory) - level — int — difficulty 1 (easiest) to 5 (hardest) - unique_id — string — stable identifier referencing original MATH test split path ## Sources - HuggingFaceH4/MATH-500 on HuggingFace: https://huggingface.co/datasets/HuggingFaceH4/MATH-500 (MIT, inherited from upstream) - Upstream: OpenAI PRM800K repo, https://github.com/openai/prm800k (MIT license) - Original benchmark: Hendrycks et al., 'Measuring Mathematical Problem Solving With the MATH Dataset' (2021), MIT license ## Methodology Problems were originally scraped from US high-school math competitions (AMC, AIME, etc.) by Hendrycks et al. and annotated with step-by-step solutions in LaTeX. OpenAI then sampled 500 problems from the MATH test split to form a stable, smaller evaluation subset for their process-reward-model work; the selection is published in PRM800K. HuggingFace H4 mirrors that exact subset as JSON with no transformations beyond format conversion. ## Known gaps & limitations This is a small evaluation set (n=500), not a training corpus. It is widely used and almost certainly contaminated in the pretraining data of most frontier LLMs, so scores should be interpreted as upper-bound capability indicators rather than held-out generalization. Coverage is English-only, US competition-math style, and skews toward symbolic manipulation rather than applied or proof-based math. Answers are extracted from \boxed{} expressions, so grading requires careful normalization (equivalent forms, LaTeX whitespace). The source does not document selection bias in how OpenAI chose the 500. ## Intended use & out-of-scope - IS for: evaluating mathematical reasoning of LLMs, reasoning-model benchmarking (o1/R1-style), RLHF/RLVR reward signal validation, prompt-engineering ablations. - NOT for: training data (high leakage risk into evals), non-English math evaluation, real-world applied math, or proof-generation tasks. Original supplier listing: MATH-500 (HuggingFaceH4) 500-problem subset of the MATH benchmark used by OpenAI in 'Let's Verify Step by Step' — the canonical math reasoning eval for LLMs and reasoning models.

Schema

NameTypeDescription
problemVARCHARLaTeX-formatted mathematics problem statement
solutionVARCHARFull worked solution with final answer enclosed in \boxed{}
answerVARCHARExtracted final numerical or symbolic answer
subjectVARCHARMathematical topic category (Algebra, Geometry, Number Theory, Counting & Probability, Intermediate Algebra, Prealgebra, Precalculus)
levelBIGINTDifficulty rating from 1 (easiest) to 5 (hardest)
unique_idVARCHARStable identifier referencing original MATH test split path

Sample Data

Preview a sample of the data before downloading.

Public sample only. Sign in to retrieve the full dataset, including free datasets.

For AI Agents

Via MCP Server
# 1. Add to your agent's MCP config (claude_desktop_config.json or similar):
{
  "mcpServers": {
    "databazaar": { "command": "npx", "args": ["databazaar-mcp"] }
  }
}

# 2. Your agent can then call:
search_datasets({ query: "MATH-500 Benchmark" })
// Found: 9c009b4d-3ecb-4b9f-9f06-6f61b3a36a9a
get_download_url({ dataset_id: "9c009b4d-3ecb-4b9f-9f06-6f61b3a36a9a" })  // free — sign in with MCP OAuth first
Via REST API
# Free dataset — sign in or use your account API key:
curl https://api.databazaar.io/datasets/9c009b4d-3ecb-4b9f-9f06-6f61b3a36a9a/download-url -H "Authorization: Bearer $DATABAZAAR_API_KEY"