textmicrosoft/ms_marcoquestion-answeringpassage-rankingretrievalragir-benchmarkmsmarcobingenglish

MS MARCO Question Answering & Passage Ranking

Free

Open dataset

Sample structure: 100 / 100
2 download links issued
Seller: DataBazaar
Sign up to download

Already have an account? Log in

Agent? Connect your account →

Category
Text
Records
1,112,939 rows
Format
PARQUET
Update Frequency
Not documented
Collection Method
auto_imported_huggingface_federated
PII
No flagged field names; not a privacy audit
File Size
~2215.43 MB
Download links issued
2

Source, license and coverage

Supplier documentation. These claims are separate from the automated sample score. A listing edit date is not a data freshness date.

License
Not documented — confirm reuse terms with the seller
Source / creator
microsoft/ms_marco
Collection method
Questions were sampled from anonymized Bing search query logs and filtered to those expressing information need. For each query, ten passages were retrieved from the web index, and human judges wrote a natural-language answer based on those passages while also labeling which passage(s) they used (is_selected). Subsequent releases added well-formed answer rewrites and expanded the original 100K v1 set to ~1M questions in v2.1.
Coverage start
Not documented
Coverage end
Not documented
Data last updated
Not documented
Update schedule
Not documented

Source documentation ↗

English-only; queries reflect Bing user distribution and time period (mid-2010s), so topical coverage skews to that era's web. Answers are extractive/abstractive from web passages and may contain factual errors inherited from sources. Not all rows have well-formed answers. The dataset has been extensively used for benchmark training and is known to overlap with many downstream IR evaluations — leakage risk is real. Microsoft's terms restrict use to non-commercial research in some interpretations; buyers should review the official MS MARCO terms at microsoft.github.io/msmarco for their use case.

Sample structure score: 100 / 100

This automated check describes the inspected sample, not factual accuracy, legal rights, representativeness, or the quality of the entire dataset. It is not a customer rating.

Assessed 10 sample records (JSON) on 2026-10-09. All records in the provided sample were checked.

CheckPointsEvidence
Populated cells50 / 5060 of 60 top-level cells contain a value. Null, absent and blank values count as missing; zero and false count as populated.
Consistent value types30 / 3060 of 60 populated cells match their column's most common observed type. Types are inferred, not checked against real-world truth.
Consistent record shape20 / 2010 of 10 records have the expected fields. CSV/TSV use the header width; JSON uses the union of observed keys.
Field-level findings and improvements

Check missing cells and mixed types below. Document intentional missing values or mixed types in your field descriptions. Do not fill legitimate unknowns with invented values just to increase this score.

FieldMissing cellsMost common typeOther populated types
answers0 / 10object0 / 10
passages0 / 10object0 / 10
query0 / 10string0 / 10
query_id0 / 10number0 / 10
query_type0 / 10string0 / 10
wellFormedAnswers0 / 10object0 / 10
How the score is calculated, its limitations, and how to correct an assessment →

About this data

Bing search questions with human-generated answers and passage relevance judgments for retrieval and question answering. Supplier notes flag possible noncommercial research restrictions; review the official source terms for your use case.

Retrieve with your agent or Python

Create an account and configure DATABAZAAR_API_KEY. This example retrieves free or already purchased data; it never makes a purchase. For a multi-file dataset, choose a file index from the manifest.

Download the Python example
python3 retrieve-dataset.py 28fcfdda-bc59-4d35-afe1-64d077455e84 --output dataset.bin
Full supplier documentation
## Overview MS MARCO (Microsoft MAchine Reading COmprehension) is a large-scale collection for deep learning in search, built from real anonymized Bing user queries. The QA configuration contains roughly 1M questions paired with human-generated answers and ~10 candidate passages each (with relevance labels). Format is Parquet (text modality), English only. Total size is in the 1M–10M row category across splits. ## Schema - query_id — int — unique identifier for the question - query — string — the natural-language question from Bing logs - query_type — string — category (e.g., DESCRIPTION, NUMERIC, ENTITY, LOCATION, PERSON) - passages — sequence — list of candidate passages, each with: passage_text (string), url (string), is_selected (int, relevance label) - answers — sequence[string] — human-written answer(s) to the query - wellFormedAnswers — sequence[string] — rewritten well-formed answers (subset of rows) ## Sources - HuggingFace: https://huggingface.co/datasets/microsoft/ms_marco — license: MIT (per Microsoft's MS MARCO terms of use / dataset card) - Original paper: Bajaj et al., "MS MARCO: A Human Generated MAchine Reading COmprehension Dataset" (arXiv:1611.09268) ## Methodology Questions were sampled from anonymized Bing search query logs and filtered to those expressing information need. For each query, ten passages were retrieved from the web index, and human judges wrote a natural-language answer based on those passages while also labeling which passage(s) they used (is_selected). Subsequent releases added well-formed answer rewrites and expanded the original 100K v1 set to ~1M questions in v2.1. ## Known gaps & limitations English-only; queries reflect Bing user distribution and time period (mid-2010s), so topical coverage skews to that era's web. Answers are extractive/abstractive from web passages and may contain factual errors inherited from sources. Not all rows have well-formed answers. The dataset has been extensively used for benchmark training and is known to overlap with many downstream IR evaluations — leakage risk is real. Microsoft's terms restrict use to non-commercial research in some interpretations; buyers should review the official MS MARCO terms at microsoft.github.io/msmarco for their use case. ## Intended use & out-of-scope - IS for: training and evaluating passage retrievers (BM25 baselines, dense retrievers like DPR/ColBERT), RAG pipelines, open-domain QA models, and reranker fine-tuning. - NOT for: training models that will be evaluated on BEIR, TREC-DL, or other MS MARCO–derived benchmarks without careful decontamination; not a source of ground-truth world knowledge (answers reflect 2016-era web). _Federated dataset: 12 parquet shards, 2.16 GB total. Queries and downloads stream through the DataBazaar API._ Original supplier listing: MS MARCO — Question Answering & Passage Ranking Microsoft's MS MARCO dataset: 1M+ real Bing questions with human-generated answers and passage relevance judgments. Foundational benchmark for retrieval, RAG, and QA systems.

Schema

NameTypeDescription
answersVARCHAR[]Human-written natural-language answer(s) to the query.
passagesSTRUCT(is_selected INTEGER[], passage_text VARCHAR[], url VARCHAR[])List of ~10 candidate passages with text, source URL, and binary relevance label (1=selected, 0=not selected).
queryVARCHARNatural-language question from anonymized Bing user logs.
query_idINTEGERUnique integer identifier for the question.
query_typeVARCHARQuestion category: DESCRIPTION, NUMERIC, ENTITY, LOCATION, or PERSON.
wellFormedAnswersVARCHAR[]Rewritten well-formed answers (present on subset of rows).

Sample Data

Preview a sample of the data before downloading.

Public sample only. Sign in to retrieve the full dataset, including free datasets.

For AI Agents

Via MCP Server
# 1. Add to your agent's MCP config (claude_desktop_config.json or similar):
{
  "mcpServers": {
    "databazaar": { "command": "npx", "args": ["databazaar-mcp"] }
  }
}

# 2. Your agent can then call:
search_datasets({ query: "MS MARCO Question Answering & " })
// Found: 28fcfdda-bc59-4d35-afe1-64d077455e84
get_download_url({ dataset_id: "28fcfdda-bc59-4d35-afe1-64d077455e84" })  // free — sign in with MCP OAuth first
Via REST API
# Free dataset — sign in or use your account API key:
curl https://api.databazaar.io/datasets/28fcfdda-bc59-4d35-afe1-64d077455e84/download-url -H "Authorization: Bearer $DATABAZAAR_API_KEY"