textSWE-bench/SWE-benchswe-benchcode-generationagentsbenchmarkgithub-issuespythonevaluationllm-evalcoding-agentsfine-tuning

SWE-bench GitHub Issue Resolution Benchmark

Free

Open dataset

Sample structure: 98.8 / 100
3 download links issued
Seller: DataBazaar
Sign up to download

Already have an account? Log in

Agent? Connect your account →

Category
Text
Records
21,527 rows
Format
PARQUET
Update Frequency
Not documented
Collection Method
auto_imported_huggingface_federated
PII
No flagged field names; not a privacy audit
File Size
~114.37 MB
Download links issued
3

Source, license and coverage

Supplier documentation. These claims are separate from the automated sample score. A listing edit date is not a data freshness date.

License
Not documented — confirm reuse terms with the seller
Source / creator
SWE-bench/SWE-bench
Collection method
The authors mined merged pull requests from 12 widely-used Python repositories on GitHub, filtering to PRs that (a) closed an issue, (b) modified source code, and (c) added or modified unit tests. For each candidate, they reconstructed the repository state at the base commit, applied the test patch, and verified that the gold code patch caused a specified set of tests to transition from failing to passing (FAIL_TO_PASS) while not breaking previously passing tests (PASS_TO_PASS). Each instance is bundled with metadata sufficient to reproduce a Python environment for execution-based evaluation. No content rewriting is applied; problem_statements are the raw issue text.
Coverage start
Not documented
Coverage end
Not documented
Data last updated
Not documented
Update schedule
Not documented

Source documentation ↗

- Python-only and limited to 12 repositories — does not represent broader software engineering work (no C/C++, JS, Go, Java, etc.). - Heavy weighting toward scientific/web Python ecosystem; agent performance on these may not generalize. - Known contamination concerns: many frontier LLMs were trained on the underlying public GitHub repos, so issues and their fixes may appear in pretraining data — leakage risk for any model with a training cutoff after the PR merge date. - This base split contains only problem_statement and metadata; for execution you typically need SWE-bench_Lite, SWE-bench_Verified, or the companion Docker images / harness from the upstream repo. - Hints_text may contain spoilers (maintainer-written fix suggestions) that make some tasks easier than a realistic from-scratch issue triage. - Time coverage: PRs span the historical lifetime of the included repos through ~2023; not refreshed with newer issues.

Sample structure score: 98.8 / 100

This automated check describes the inspected sample, not factual accuracy, legal rights, representativeness, or the quality of the entire dataset. It is not a customer rating.

Assessed 10 sample records (JSON) on 2026-10-09. All records in the provided sample were checked.

CheckPointsEvidence
Populated cells48.8 / 50117 of 120 top-level cells contain a value. Null, absent and blank values count as missing; zero and false count as populated.
Consistent value types30 / 30117 of 117 populated cells match their column's most common observed type. Types are inferred, not checked against real-world truth.
Consistent record shape20 / 2010 of 10 records have the expected fields. CSV/TSV use the header width; JSON uses the union of observed keys.
Field-level findings and improvements

Check missing cells and mixed types below. Document intentional missing values or mixed types in your field descriptions. Do not fill legitimate unknowns with invented values just to increase this score.

FieldMissing cellsMost common typeOther populated types
repo0 / 10string0 / 10
instance_id0 / 10string0 / 10
base_commit0 / 10string0 / 10
patch0 / 10string0 / 10
test_patch0 / 10string0 / 10
problem_statement0 / 10string0 / 10
hints_text3 / 10string0 / 7
created_at0 / 10string0 / 10
version0 / 10number0 / 10
FAIL_TO_PASS0 / 10string0 / 10
PASS_TO_PASS0 / 10string0 / 10
environment_setup_commit0 / 10string0 / 10
How the score is calculated, its limitations, and how to correct an assessment →

About this data

Real-world issue-pull request pairs from 12 popular Python repositories, with resolution verified by post-PR unit tests. Designed to evaluate language model and agent performance on practical software engineering tasks.

Retrieve with your agent or Python

Create an account and configure DATABAZAAR_API_KEY. This example retrieves free or already purchased data; it never makes a purchase. For a multi-file dataset, choose a file index from the manifest.

Download the Python example
python3 retrieve-dataset.py c8dddd6d-99ec-4934-9396-43455b088bd8 --output dataset.bin
Full supplier documentation
## Overview SWE-bench is the canonical benchmark for evaluating language models and agents on resolving real-world GitHub issues. It contains 2,294 Issue–Pull Request pairs collected from 12 popular Python repositories (e.g., django, sympy, scikit-learn, matplotlib, flask, requests, pytest, sphinx, astropy, xarray, pylint, pyodide). Each instance includes the issue text, the target repository state, and the unit tests that validate a correct fix. Evaluation uses post-PR test behavior as the reference solution. Format: Parquet, ~10K–100K rows across splits (dev/test), text modality. ## Schema - instance_id — string — unique identifier for the issue/PR pair - repo — string — owner/name of the source GitHub repository - base_commit — string — commit SHA of the repo state before the fix - patch — string — the gold patch (PR diff) that resolves the issue - test_patch — string — diff containing the unit tests used for verification - problem_statement — string — the issue text describing the bug or feature request - hints_text — string — additional discussion/comments from the issue thread - created_at — string — issue creation timestamp - version — string — repository version tag for environment setup - FAIL_TO_PASS — string (JSON list) — tests expected to flip from fail→pass after the fix - PASS_TO_PASS — string (JSON list) — tests that should continue passing - environment_setup_commit — string — commit used to install the testing environment ## Sources - HuggingFace: https://huggingface.co/datasets/SWE-bench/SWE-bench — license: CC BY 4.0 (per the dataset card / project repo) - Paper: "SWE-bench: Can Language Models Resolve Real-World GitHub Issues?" (Jimenez et al., arXiv:2310.06770) - Upstream project: https://github.com/princeton-nlp/SWE-bench ## Methodology The authors mined merged pull requests from 12 widely-used Python repositories on GitHub, filtering to PRs that (a) closed an issue, (b) modified source code, and (c) added or modified unit tests. For each candidate, they reconstructed the repository state at the base commit, applied the test patch, and verified that the gold code patch caused a specified set of tests to transition from failing to passing (FAIL_TO_PASS) while not breaking previously passing tests (PASS_TO_PASS). Each instance is bundled with metadata sufficient to reproduce a Python environment for execution-based evaluation. No content rewriting is applied; problem_statements are the raw issue text. ## Known gaps & limitations - Python-only and limited to 12 repositories — does not represent broader software engineering work (no C/C++, JS, Go, Java, etc.). - Heavy weighting toward scientific/web Python ecosystem; agent performance on these may not generalize. - Known contamination concerns: many frontier LLMs were trained on the underlying public GitHub repos, so issues and their fixes may appear in pretraining data — leakage risk for any model with a training cutoff after the PR merge date. - This base split contains only problem_statement and metadata; for execution you typically need SWE-bench_Lite, SWE-bench_Verified, or the companion Docker images / harness from the upstream repo. - Hints_text may contain spoilers (maintainer-written fix suggestions) that make some tasks easier than a realistic from-scratch issue triage. - Time coverage: PRs span the historical lifetime of the included repos through ~2023; not refreshed with newer issues. ## Intended use & out-of-scope - Intended: evaluating coding agents, LLM patch generation, retrieval-augmented code fix systems, agent planning/tool-use benchmarks, and SFT/RL fine-tuning on agentic coding trajectories. - Out-of-scope: training a model and then reporting numbers on this same split without decontamination — leakage against GitHub-trained LLMs is well-documented. For leaderboard-quality eval use SWE-bench_Verified. _Federated dataset: 3 parquet shards, 114.4 MB total. Queries and downloads stream through the DataBazaar API._ _PII signals: email×8 present in the sample. Common in public datasets (papers, logs) but worth knowing before joining with private data._ Original supplier listing: SWE-bench: Real-World GitHub Issue Resolution Benchmark 2,294 Issue-PR pairs from 12 popular Python repos for evaluating LLM/agent ability to resolve real GitHub issues. Verified via post-PR unit tests. Canonical agent coding benchmark.

Schema

NameTypeDescription
repoVARCHARGitHub repository identifier in owner/name format
instance_idVARCHARUnique identifier for the issue-PR pair
base_commitVARCHARGit commit SHA of repository state before fix
patchVARCHARUnified diff of the gold-standard fix (PR changes)
test_patchVARCHARUnified diff containing unit tests for verification
problem_statementVARCHARFull text of the GitHub issue describing the bug or feature request
hints_textVARCHARAdditional discussion comments from the issue thread
created_atVARCHARISO 8601 timestamp when the issue was created
versionVARCHARRepository version tag or release identifier
FAIL_TO_PASSVARCHARTest cases that fail before fix and pass after fix
PASS_TO_PASSVARCHARTest cases that pass both before and after fix
environment_setup_commitVARCHARCommit SHA used to set up the test environment

Sample Data

Preview a sample of the data before downloading.

Public sample only. Sign in to retrieve the full dataset, including free datasets.

For AI Agents

Via MCP Server
# 1. Add to your agent's MCP config (claude_desktop_config.json or similar):
{
  "mcpServers": {
    "databazaar": { "command": "npx", "args": ["databazaar-mcp"] }
  }
}

# 2. Your agent can then call:
search_datasets({ query: "SWE-bench GitHub Issue Resolut" })
// Found: c8dddd6d-99ec-4934-9396-43455b088bd8
get_download_url({ dataset_id: "c8dddd6d-99ec-4934-9396-43455b088bd8" })  // free — sign in with MCP OAuth first
Via REST API
# Free dataset — sign in or use your account API key:
curl https://api.databazaar.io/datasets/c8dddd6d-99ec-4934-9396-43455b088bd8/download-url -H "Authorization: Bearer $DATABAZAAR_API_KEY"