textfacebook/kilt_tasksnlpbenchmarkquestion-answeringfact-checkingentity-linkingragwikipediaretrievalevaluationkilt

KILT Knowledge-Intensive Language Tasks Benchmark

Free

Open dataset

Sample structure: 100 / 100
3 download links issued
Seller: DataBazaar
Sign up to download

Already have an account? Log in

Agent? Connect your account →

Category
Text
Records
3,231,786 rows
Format
PARQUET
Update Frequency
Not documented
Collection Method
auto_imported_huggingface_federated
PII
No flagged field names; not a privacy audit
File Size
~1001.61 MB
Download links issued
3

Source, license and coverage

Supplier documentation. These claims are separate from the automated sample score. A listing edit date is not a data freshness date.

License
mit
Source / creator
facebook/kilt_tasks
Collection method
KILT's authors took 11 pre-existing English NLP datasets and re-grounded all of them against a single 2019 Wikipedia snapshot, mapping each example's supporting evidence to specific Wikipedia pages, paragraphs, and character spans. This normalization allows the same retriever and reader to be evaluated across all tasks under identical knowledge-source assumptions, and supports multi-task and transfer learning. Annotations are a mix of crowdsourced, found, and machine-generated depending on the upstream source dataset.
Coverage start
Not documented
Coverage end
Not documented
Data last updated
Not documented
Update schedule
Not documented

Source documentation ↗

License terms ↗

- English-only (monolingual); not suitable for multilingual evaluation. - Grounded on a 2019 Wikipedia snapshot — temporal drift for any question whose answer has changed since. - Upstream datasets have heterogeneous annotation quality; e.g. TriviaQA is distantly supervised, Wizard of Wikipedia is dialog with crowdsourced grounding. - The `triviaqa_support_only` config is a subset (only examples with KILT-mapped provenance), not full TriviaQA. - Heavy overlap with widely-used eval suites (NQ, HotpotQA, TriviaQA, FEVER) — significant leakage risk if used as training data for models later benchmarked on these. - Provenance spans are mapped automatically and may have alignment errors at character-offset granularity.

Sample structure score: 100 / 100

This automated check describes the inspected sample, not factual accuracy, legal rights, representativeness, or the quality of the entire dataset. It is not a customer rating.

Assessed 10 sample records (JSON) on 2026-10-09. All records in the provided sample were checked.

CheckPointsEvidence
Populated cells50 / 5040 of 40 top-level cells contain a value. Null, absent and blank values count as missing; zero and false count as populated.
Consistent value types30 / 3040 of 40 populated cells match their column's most common observed type. Types are inferred, not checked against real-world truth.
Consistent record shape20 / 2010 of 10 records have the expected fields. CSV/TSV use the header width; JSON uses the union of observed keys.
Field-level findings and improvements

Check missing cells and mixed types below. Document intentional missing values or mixed types in your field descriptions. Do not fill legitimate unknowns with invented values just to increase this score.

FieldMissing cellsMost common typeOther populated types
id0 / 10string0 / 10
input0 / 10string0 / 10
meta0 / 10object0 / 10
output0 / 10object0 / 10
How the score is calculated, its limitations, and how to correct an assessment →

About this data

Eleven datasets spanning fact-checking, entity linking, slot filling, open-domain question answering, and dialog generation, all grounded in a unified Wikipedia snapshot. Covers approximately 3.2 million examples across multiple task types.

Retrieve with your agent or Python

Create an account and configure DATABAZAAR_API_KEY. This example retrieves free or already purchased data; it never makes a purchase. For a multi-file dataset, choose a file index from the manifest.

Download the Python example
python3 retrieve-dataset.py 8997fb69-ff6d-48af-926e-b3840702fd18 --output dataset.bin
Full supplier documentation
## Overview KILT (Knowledge-Intensive Language Tasks) is a unified benchmark from Facebook AI Research consolidating 11 datasets across 5 task families (fact-checking, entity linking, slot filling, open-domain QA, dialog generation) all grounded in a single pre-processed Wikipedia dump. The dataset contains between 1M and 10M examples in parquet format, English-only. Original release accompanies the EMNLP 2020 paper (arXiv:2009.02252); the HF mirror was last updated January 2024. ## Schema Schema varies per subtask config, but core fields shared across configs include: - `id` — string — unique example identifier - `input` — string — the query, claim, or dialog context - `output` — list of structs — gold answers/labels, each with: - `answer` — string — gold answer text (for QA/slot filling) or label (for fact-checking) - `provenance` — list of structs — supporting Wikipedia evidence with `wikipedia_id`, `title`, `section`, `start_paragraph_id`, `end_paragraph_id`, `start_character`, `end_character`, `bleu_score`, `meta` - `meta` — struct — task-specific metadata (e.g. left/right context for entity linking, template for slot filling) Subtask configs include: `aidayago2`, `fever`, `hotpotqa`, `nq`, `trex`, `triviaqa_support_only`, `structured_zeroshot`, `wned`, `cweb`, `wow`. ## Sources - HuggingFace: https://huggingface.co/datasets/facebook/kilt_tasks — MIT license - Original paper: Petroni et al., "KILT: a Benchmark for Knowledge Intensive Language Tasks", NAACL 2021 (arXiv:2009.02252) - Source datasets (extended): Natural Questions, AIDA-YAGO, FEVER, HotpotQA, T-REx, TriviaQA, Wizard of Wikipedia, WNED-CWEB, WNED-WIKI, Zero-Shot RE ## Methodology KILT's authors took 11 pre-existing English NLP datasets and re-grounded all of them against a single 2019 Wikipedia snapshot, mapping each example's supporting evidence to specific Wikipedia pages, paragraphs, and character spans. This normalization allows the same retriever and reader to be evaluated across all tasks under identical knowledge-source assumptions, and supports multi-task and transfer learning. Annotations are a mix of crowdsourced, found, and machine-generated depending on the upstream source dataset. ## Known gaps & limitations - English-only (monolingual); not suitable for multilingual evaluation. - Grounded on a 2019 Wikipedia snapshot — temporal drift for any question whose answer has changed since. - Upstream datasets have heterogeneous annotation quality; e.g. TriviaQA is distantly supervised, Wizard of Wikipedia is dialog with crowdsourced grounding. - The `triviaqa_support_only` config is a subset (only examples with KILT-mapped provenance), not full TriviaQA. - Heavy overlap with widely-used eval suites (NQ, HotpotQA, TriviaQA, FEVER) — significant leakage risk if used as training data for models later benchmarked on these. - Provenance spans are mapped automatically and may have alignment errors at character-offset granularity. ## Intended use & out-of-scope - IS for: benchmarking retrieval-augmented generation (RAG) and knowledge-intensive LM systems, multi-task evaluation across QA / fact-checking / entity linking / slot filling / dialog, training retrievers on a unified knowledge source. - NOT for: training models that will later be evaluated on NQ/HotpotQA/TriviaQA/FEVER (leakage); multilingual or non-English evaluation; tasks requiring post-2019 world knowledge. _Federated dataset: 34 parquet shards, 1001.6 MB total. Queries and downloads stream through the DataBazaar API._ Original supplier listing: KILT: Knowledge-Intensive Language Tasks Benchmark Facebook AI's KILT benchmark — 11 datasets across fact-checking, entity linking, slot filling, open-domain QA, and dialog generation, all grounded in a unified Wikipedia snapshot. MIT licensed, parquet format, 1M-10M examples.

Schema

NameTypeDescription
idVARCHARUnique example identifier string
inputVARCHARQuery, claim, or dialog context text
metaSTRUCT(left_context VARCHAR, mention VARCHAR, right_context VARCHAR, partial_evidence STRUCT(start_paragraph_id INTEGER, end_paragraph_id INTEGER, title VARCHAR, section VARCHAR, wikipedia_id VARCHAR, meta STRUCT(evidence_span VARCHAR[]))[], obj_surface VARCHAR[], sub_surface VARCHAR[], subj_aliases VARCHAR[], template_questions VARCHAR[])Task-specific metadata including entity mention context, surface forms, aliases, template questions, and partial Wikipedia evidence references
outputSTRUCT(answer VARCHAR, meta STRUCT(score INTEGER), provenance STRUCT(bleu_score FLOAT, start_character INTEGER, start_paragraph_id INTEGER, end_character INTEGER, end_paragraph_id INTEGER, meta STRUCT(fever_page_id VARCHAR, fever_sentence_id INTEGER, annotation_id VARCHAR, yes_no_answer VARCHAR, evidence_span VARCHAR[]), section VARCHAR, title VARCHAR, wikipedia_id VARCHAR)[])[]List of gold answers with text and supporting Wikipedia provenance (document ID, section, character spans, evidence metrics)

Sample Data

Preview a sample of the data before downloading.

Public sample only. Sign in to retrieve the full dataset, including free datasets.

For AI Agents

Via MCP Server
# 1. Add to your agent's MCP config (claude_desktop_config.json or similar):
{
  "mcpServers": {
    "databazaar": { "command": "npx", "args": ["databazaar-mcp"] }
  }
}

# 2. Your agent can then call:
search_datasets({ query: "KILT Knowledge-Intensive Langu" })
// Found: 8997fb69-ff6d-48af-926e-b3840702fd18
get_download_url({ dataset_id: "8997fb69-ff6d-48af-926e-b3840702fd18" })  // free — sign in with MCP OAuth first
Via REST API
# Free dataset — sign in or use your account API key:
curl https://api.databazaar.io/datasets/8997fb69-ff6d-48af-926e-b3840702fd18/download-url -H "Authorization: Bearer $DATABAZAAR_API_KEY"