textBeIR/msmarcoinformation-retrievalbeirmsmarcopassage-retrievalragbenchmarkenglishdense-retrievalevaluationnlp

BEIR MS MARCO Passage Retrieval Benchmark

Free

Open dataset

Sample structure: 83.3 / 100
2 download links issued
Seller: DataBazaar
Sign up to download

Already have an account? Log in

Agent? Connect your account →

Category
Text
Records
9,351,785 rows
Format
PARQUET
Update Frequency
Not documented
Collection Method
auto_imported_huggingface_federated
PII
No flagged field names; not a privacy audit
File Size
~1571.67 MB
Download links issued
2

Source, license and coverage

Supplier documentation. These claims are separate from the automated sample score. A listing edit date is not a data freshness date.

License
cc-by-sa-4.0
Source / creator
BeIR/msmarco
Collection method
MS MARCO was originally constructed by Microsoft from anonymized Bing query logs, with human annotators identifying relevant passages from web documents that answered each query. BEIR repackages MS MARCO into a uniform JSON/Parquet retrieval format (corpus, queries, qrels) to enable apples-to-apples comparison across its 18 IR tasks. No additional content filtering or augmentation is applied by BEIR beyond schema normalization.
Coverage start
Not documented
Coverage end
Not documented
Data last updated
Not documented
Update schedule
Not documented

Source documentation ↗

License terms ↗

- English-only; not suitable for multilingual retrieval evaluation. - Queries reflect Bing user search patterns circa 2018, which skews toward informational web search and may underrepresent conversational, multi-hop, or domain-specialist queries. - Many passages lack titles (`title` field often empty). - Sparse relevance judgments: most queries have only 1 judged relevant passage, leading to known underestimation of true recall. - Original Bing log collection date is not precisely documented; treat the dataset as a static 2018-era snapshot.

Sample structure score: 83.3 / 100

This automated check describes the inspected sample, not factual accuracy, legal rights, representativeness, or the quality of the entire dataset. It is not a customer rating.

Assessed 10 sample records (JSON) on 2026-10-09. All records in the provided sample were checked.

CheckPointsEvidence
Populated cells33.3 / 5020 of 30 top-level cells contain a value. Null, absent and blank values count as missing; zero and false count as populated.
Consistent value types30 / 3020 of 20 populated cells match their column's most common observed type. Types are inferred, not checked against real-world truth.
Consistent record shape20 / 2010 of 10 records have the expected fields. CSV/TSV use the header width; JSON uses the union of observed keys.
Field-level findings and improvements

Check missing cells and mixed types below. Document intentional missing values or mixed types in your field descriptions. Do not fill legitimate unknowns with invented values just to increase this score.

FieldMissing cellsMost common typeOther populated types
_id0 / 10number0 / 10
title10 / 10unknown0 / 0
text0 / 10string0 / 10
How the score is calculated, its limitations, and how to correct an assessment →

About this data

MS MARCO subset of the BEIR heterogeneous IR benchmark in standard retrieval layout. English-language passages and queries for passage retrieval evaluation.

Retrieve with your agent or Python

Create an account and configure DATABAZAAR_API_KEY. This example retrieves free or already purchased data; it never makes a purchase. For a multi-file dataset, choose a file index from the manifest.

Download the Python example
python3 retrieve-dataset.py 8f53147d-be88-4e77-b967-f3c6df1a0458 --output dataset.bin
Full supplier documentation
## Overview The MS MARCO subset of the BEIR (Benchmarking IR) benchmark — a large-scale English passage retrieval dataset widely used to evaluate dense retrievers, sparse retrievers, and RAG systems. Contains approximately 8.8M passages in the corpus and ~500K natural-language queries derived from anonymized Bing search logs, distributed in Parquet using the standard BEIR retrieval layout. English only. ## Schema - `_id` — string — unique identifier for a document or query - `title` — string — title field (often empty for MS MARCO passages) - `text` — string — passage text (corpus) or query text (queries) The dataset is split into two configurations: `corpus` (one row per passage) and `queries` (one row per query). Relevance judgments (qrels) are distributed via the companion BeIR/msmarco-qrels dataset. ## Sources - BeIR/msmarco on HuggingFace — https://huggingface.co/datasets/BeIR/msmarco — license: CC-BY-SA-4.0 - Original MS MARCO dataset (Microsoft) — https://microsoft.github.io/msmarco/ - BEIR paper (Thakur et al., NeurIPS 2021) — https://arxiv.org/abs/2104.08663 ## Methodology MS MARCO was originally constructed by Microsoft from anonymized Bing query logs, with human annotators identifying relevant passages from web documents that answered each query. BEIR repackages MS MARCO into a uniform JSON/Parquet retrieval format (corpus, queries, qrels) to enable apples-to-apples comparison across its 18 IR tasks. No additional content filtering or augmentation is applied by BEIR beyond schema normalization. ## Known gaps & limitations - English-only; not suitable for multilingual retrieval evaluation. - Queries reflect Bing user search patterns circa 2018, which skews toward informational web search and may underrepresent conversational, multi-hop, or domain-specialist queries. - Many passages lack titles (`title` field often empty). - Sparse relevance judgments: most queries have only 1 judged relevant passage, leading to known underestimation of true recall. - Original Bing log collection date is not precisely documented; treat the dataset as a static 2018-era snapshot. ## Intended use & out-of-scope - Intended for: training and evaluating dense/sparse retrievers, reranker fine-tuning, RAG component benchmarking, zero-shot retrieval research, and as a base corpus for synthetic query generation. - Out-of-scope: training models that will be evaluated on BEIR itself (severe leakage — MS MARCO is part of BEIR), production search over current web content (data is stale), and any non-English IR use case. _Federated dataset: 2 parquet shards, 1.53 GB total. Queries and downloads stream through the DataBazaar API._ Original supplier listing: BEIR: MS MARCO Passage Retrieval Benchmark MS MARCO subset of the BEIR heterogeneous IR benchmark — 8.8M passages and ~500K queries in standard retrieval layout. Parquet format, CC-BY-SA-4.0, English.

Schema

NameTypeDescription
_idVARCHARUnique identifier for a passage (corpus) or query (queries split).
titleVARCHARPassage title field; typically empty for MS MARCO passages.
textVARCHARPassage text content (corpus split) or natural-language query text (queries split).

Sample Data

Preview a sample of the data before downloading.

Public sample only. Sign in to retrieve the full dataset, including free datasets.

For AI Agents

Via MCP Server
# 1. Add to your agent's MCP config (claude_desktop_config.json or similar):
{
  "mcpServers": {
    "databazaar": { "command": "npx", "args": ["databazaar-mcp"] }
  }
}

# 2. Your agent can then call:
search_datasets({ query: "BEIR MS MARCO Passage Retrieva" })
// Found: 8f53147d-be88-4e77-b967-f3c6df1a0458
get_download_url({ dataset_id: "8f53147d-be88-4e77-b967-f3c6df1a0458" })  // free — sign in with MCP OAuth first
Via REST API
# Free dataset — sign in or use your account API key:
curl https://api.databazaar.io/datasets/8f53147d-be88-4e77-b967-f3c6df1a0458/download-url -H "Authorization: Bearer $DATABAZAAR_API_KEY"