textopenai/gsm8kmathreasoningbenchmarkllm-evalchain-of-thoughtword-problemsopenaifine-tuning

GSM8K Grade School Math Word Problems

Free

Open dataset

Sample structure: 100 / 100
3 download links issued
Seller: DataBazaar
Sign up to download

Already have an account? Log in

Agent? Connect your account →

Category
Text
Records
17,584 rows
Format
PARQUET
Update Frequency
Not documented
Collection Method
auto_imported_huggingface_federated
PII
No flagged field names; not a privacy audit
File Size
~5.62 MB
Download links issued
3

Source, license and coverage

Supplier documentation. These claims are separate from the automated sample score. A listing edit date is not a data freshness date.

License
mit
Source / creator
openai/gsm8k
Collection method
Problems were written by human contractors (crowdsourced) hired by OpenAI, with explicit instructions to produce linguistically diverse problems solvable by a bright middle-schooler. Each problem includes a natural-language solution that walks through the calculation steps. Calculator annotations (`<<calc>>`) appear inline in solutions to mark arithmetic operations. The final numeric answer is delimited by `####` for programmatic grading.
Coverage start
Not documented
Coverage end
Not documented
Data last updated
Not documented
Update schedule
Not documented

Source documentation ↗

License terms ↗

- English-only; no multilingual coverage. - Grade-school difficulty only — does not probe algebra, geometry, or competition-level math (see MATH, AIME, etc. for harder tiers). - Heavy contamination risk: GSM8K is one of the most widely leaked benchmarks; it appears in pretraining corpora and instruction-tuning mixes. Reported scores on frontier models should be interpreted with leakage caveats. - Errata: a small number of problems have been flagged by the community as ambiguous or incorrectly labeled; OpenAI has not published an official errata. - Source does not document demographic or topical biases in problem authoring.

Sample structure score: 100 / 100

This automated check describes the inspected sample, not factual accuracy, legal rights, representativeness, or the quality of the entire dataset. It is not a customer rating.

Assessed 10 sample records (JSON) on 2026-10-09. All records in the provided sample were checked.

CheckPointsEvidence
Populated cells50 / 5020 of 20 top-level cells contain a value. Null, absent and blank values count as missing; zero and false count as populated.
Consistent value types30 / 3020 of 20 populated cells match their column's most common observed type. Types are inferred, not checked against real-world truth.
Consistent record shape20 / 2010 of 10 records have the expected fields. CSV/TSV use the header width; JSON uses the union of observed keys.
Field-level findings and improvements

Check missing cells and mixed types below. Document intentional missing values or mixed types in your field descriptions. Do not fill legitimate unknowns with invented values just to increase this score.

FieldMissing cellsMost common typeOther populated types
question0 / 10string0 / 10
answer0 / 10string0 / 10
How the score is calculated, its limitations, and how to correct an assessment →

About this data

Linguistically diverse grade-school math word problems with step-by-step solutions, used as a standard benchmark for multi-step arithmetic reasoning in language models.

Retrieve with your agent or Python

Create an account and configure DATABAZAAR_API_KEY. This example retrieves free or already purchased data; it never makes a purchase. For a multi-file dataset, choose a file index from the manifest.

Download the Python example
python3 retrieve-dataset.py 50b6229e-be6d-460b-aeb8-de672b933135 --output dataset.bin
Full supplier documentation
## Overview GSM8K (Grade School Math 8K) is a benchmark of ~8,500 high-quality, linguistically diverse grade-school math word problems created by OpenAI to evaluate multi-step arithmetic reasoning. Each problem requires 2–8 reasoning steps and is solved using elementary operations (+, −, ×, ÷). Released in 2021 alongside the paper "Training Verifiers to Solve Math Word Problems" (arXiv:2110.14168). Split into ~7.5K train and ~1.3K test examples. Format: Parquet. Language: English only. ## Schema - `question` — string — natural-language math word problem - `answer` — string — full chain-of-thought solution ending with `#### <final_numeric_answer>` marker for easy extraction Two configurations available: `main` (standard) and `socratic` (solutions rewritten as Socratic-style guiding questions). ## Sources - HuggingFace: https://huggingface.co/datasets/openai/gsm8k — License: MIT - Paper: Cobbe et al., "Training Verifiers to Solve Math Word Problems," arXiv:2110.14168 - Original repo: https://github.com/openai/grade-school-math ## Methodology Problems were written by human contractors (crowdsourced) hired by OpenAI, with explicit instructions to produce linguistically diverse problems solvable by a bright middle-schooler. Each problem includes a natural-language solution that walks through the calculation steps. Calculator annotations (`<<calc>>`) appear inline in solutions to mark arithmetic operations. The final numeric answer is delimited by `####` for programmatic grading. ## Known gaps & limitations - English-only; no multilingual coverage. - Grade-school difficulty only — does not probe algebra, geometry, or competition-level math (see MATH, AIME, etc. for harder tiers). - Heavy contamination risk: GSM8K is one of the most widely leaked benchmarks; it appears in pretraining corpora and instruction-tuning mixes. Reported scores on frontier models should be interpreted with leakage caveats. - Errata: a small number of problems have been flagged by the community as ambiguous or incorrectly labeled; OpenAI has not published an official errata. - Source does not document demographic or topical biases in problem authoring. ## Intended use & out-of-scope - IS for: evaluating arithmetic/chain-of-thought reasoning, few-shot prompting research, fine-tuning datasets for math reasoning, verifier training, RLHF reward modeling on math. - NOT for: training models you then plan to evaluate on GSM8K itself (leakage); not a measure of advanced mathematical ability; not deduplicated against common pretraining corpora. _Federated dataset: 4 parquet shards, 5.6 MB total. Queries and downloads stream through the DataBazaar API._ Original supplier listing: GSM8K — Grade School Math Word Problems (OpenAI) 8.5K linguistically diverse grade-school math word problems with step-by-step solutions. Standard benchmark for multi-step arithmetic reasoning in LLMs.

Schema

NameTypeDescription
questionVARCHARNatural-language grade-school math word problem requiring 2–8 arithmetic reasoning steps.
answerVARCHARChain-of-thought solution with intermediate calculations and final answer marked by `#### <number>` delimiter.

Sample Data

Preview a sample of the data before downloading.

Public sample only. Sign in to retrieve the full dataset, including free datasets.

For AI Agents

Via MCP Server
# 1. Add to your agent's MCP config (claude_desktop_config.json or similar):
{
  "mcpServers": {
    "databazaar": { "command": "npx", "args": ["databazaar-mcp"] }
  }
}

# 2. Your agent can then call:
search_datasets({ query: "GSM8K Grade School Math Word P" })
// Found: 50b6229e-be6d-460b-aeb8-de672b933135
get_download_url({ dataset_id: "50b6229e-be6d-460b-aeb8-de672b933135" })  // free — sign in with MCP OAuth first
Via REST API
# Free dataset — sign in or use your account API key:
curl https://api.databazaar.io/datasets/50b6229e-be6d-460b-aeb8-de672b933135/download-url -H "Authorization: Bearer $DATABAZAAR_API_KEY"