SWE-bench GitHub Issue Resolution Benchmark
Source, license and coverage
Supplier documentation. These claims are separate from the automated sample score. A listing edit date is not a data freshness date.
- License
- Not documented — confirm reuse terms with the seller
- Source / creator
- SWE-bench/SWE-bench
- Collection method
- The authors mined merged pull requests from 12 widely-used Python repositories on GitHub, filtering to PRs that (a) closed an issue, (b) modified source code, and (c) added or modified unit tests. For each candidate, they reconstructed the repository state at the base commit, applied the test patch, and verified that the gold code patch caused a specified set of tests to transition from failing to passing (FAIL_TO_PASS) while not breaking previously passing tests (PASS_TO_PASS). Each instance is bundled with metadata sufficient to reproduce a Python environment for execution-based evaluation. No content rewriting is applied; problem_statements are the raw issue text.
- Coverage start
- Not documented
- Coverage end
- Not documented
- Data last updated
- Not documented
- Update schedule
- Not documented
- Python-only and limited to 12 repositories — does not represent broader software engineering work (no C/C++, JS, Go, Java, etc.). - Heavy weighting toward scientific/web Python ecosystem; agent performance on these may not generalize. - Known contamination concerns: many frontier LLMs were trained on the underlying public GitHub repos, so issues and their fixes may appear in pretraining data — leakage risk for any model with a training cutoff after the PR merge date. - This base split contains only problem_statement and metadata; for execution you typically need SWE-bench_Lite, SWE-bench_Verified, or the companion Docker images / harness from the upstream repo. - Hints_text may contain spoilers (maintainer-written fix suggestions) that make some tasks easier than a realistic from-scratch issue triage. - Time coverage: PRs span the historical lifetime of the included repos through ~2023; not refreshed with newer issues.
Sample structure score: 98.8 / 100
This automated check describes the inspected sample, not factual accuracy, legal rights, representativeness, or the quality of the entire dataset. It is not a customer rating.
Assessed 10 sample records (JSON) on 2026-10-09. All records in the provided sample were checked.
| Check | Points | Evidence |
|---|---|---|
| Populated cells | 48.8 / 50 | 117 of 120 top-level cells contain a value. Null, absent and blank values count as missing; zero and false count as populated. |
| Consistent value types | 30 / 30 | 117 of 117 populated cells match their column's most common observed type. Types are inferred, not checked against real-world truth. |
| Consistent record shape | 20 / 20 | 10 of 10 records have the expected fields. CSV/TSV use the header width; JSON uses the union of observed keys. |
Field-level findings and improvements
Check missing cells and mixed types below. Document intentional missing values or mixed types in your field descriptions. Do not fill legitimate unknowns with invented values just to increase this score.
| Field | Missing cells | Most common type | Other populated types |
|---|---|---|---|
| repo | 0 / 10 | string | 0 / 10 |
| instance_id | 0 / 10 | string | 0 / 10 |
| base_commit | 0 / 10 | string | 0 / 10 |
| patch | 0 / 10 | string | 0 / 10 |
| test_patch | 0 / 10 | string | 0 / 10 |
| problem_statement | 0 / 10 | string | 0 / 10 |
| hints_text | 3 / 10 | string | 0 / 7 |
| created_at | 0 / 10 | string | 0 / 10 |
| version | 0 / 10 | number | 0 / 10 |
| FAIL_TO_PASS | 0 / 10 | string | 0 / 10 |
| PASS_TO_PASS | 0 / 10 | string | 0 / 10 |
| environment_setup_commit | 0 / 10 | string | 0 / 10 |
About this data
Real-world issue-pull request pairs from 12 popular Python repositories, with resolution verified by post-PR unit tests. Designed to evaluate language model and agent performance on practical software engineering tasks.
Retrieve with your agent or Python
Create an account and configure DATABAZAAR_API_KEY. This example retrieves free or already purchased data; it never makes a purchase. For a multi-file dataset, choose a file index from the manifest.
Download the Python examplepython3 retrieve-dataset.py c8dddd6d-99ec-4934-9396-43455b088bd8 --output dataset.bin
Full supplier documentation
Schema
| Name | Type | Description |
|---|---|---|
| repo | VARCHAR | GitHub repository identifier in owner/name format |
| instance_id | VARCHAR | Unique identifier for the issue-PR pair |
| base_commit | VARCHAR | Git commit SHA of repository state before fix |
| patch | VARCHAR | Unified diff of the gold-standard fix (PR changes) |
| test_patch | VARCHAR | Unified diff containing unit tests for verification |
| problem_statement | VARCHAR | Full text of the GitHub issue describing the bug or feature request |
| hints_text | VARCHAR | Additional discussion comments from the issue thread |
| created_at | VARCHAR | ISO 8601 timestamp when the issue was created |
| version | VARCHAR | Repository version tag or release identifier |
| FAIL_TO_PASS | VARCHAR | Test cases that fail before fix and pass after fix |
| PASS_TO_PASS | VARCHAR | Test cases that pass both before and after fix |
| environment_setup_commit | VARCHAR | Commit SHA used to set up the test environment |
Sample Data
Preview a sample of the data before downloading.
Public sample only. Sign in to retrieve the full dataset, including free datasets.
For AI Agents
# 1. Add to your agent's MCP config (claude_desktop_config.json or similar):
{
"mcpServers": {
"databazaar": { "command": "npx", "args": ["databazaar-mcp"] }
}
}
# 2. Your agent can then call:
search_datasets({ query: "SWE-bench GitHub Issue Resolut" })
// Found: c8dddd6d-99ec-4934-9396-43455b088bd8
get_download_url({ dataset_id: "c8dddd6d-99ec-4934-9396-43455b088bd8" }) // free — sign in with MCP OAuth first# Free dataset — sign in or use your account API key: curl https://api.databazaar.io/datasets/c8dddd6d-99ec-4934-9396-43455b088bd8/download-url -H "Authorization: Bearer $DATABAZAAR_API_KEY"