KILT Knowledge-Intensive Language Tasks Benchmark
Source, license and coverage
Supplier documentation. These claims are separate from the automated sample score. A listing edit date is not a data freshness date.
- License
- mit
- Source / creator
- facebook/kilt_tasks
- Collection method
- KILT's authors took 11 pre-existing English NLP datasets and re-grounded all of them against a single 2019 Wikipedia snapshot, mapping each example's supporting evidence to specific Wikipedia pages, paragraphs, and character spans. This normalization allows the same retriever and reader to be evaluated across all tasks under identical knowledge-source assumptions, and supports multi-task and transfer learning. Annotations are a mix of crowdsourced, found, and machine-generated depending on the upstream source dataset.
- Coverage start
- Not documented
- Coverage end
- Not documented
- Data last updated
- Not documented
- Update schedule
- Not documented
- English-only (monolingual); not suitable for multilingual evaluation. - Grounded on a 2019 Wikipedia snapshot — temporal drift for any question whose answer has changed since. - Upstream datasets have heterogeneous annotation quality; e.g. TriviaQA is distantly supervised, Wizard of Wikipedia is dialog with crowdsourced grounding. - The `triviaqa_support_only` config is a subset (only examples with KILT-mapped provenance), not full TriviaQA. - Heavy overlap with widely-used eval suites (NQ, HotpotQA, TriviaQA, FEVER) — significant leakage risk if used as training data for models later benchmarked on these. - Provenance spans are mapped automatically and may have alignment errors at character-offset granularity.
Sample structure score: 100 / 100
This automated check describes the inspected sample, not factual accuracy, legal rights, representativeness, or the quality of the entire dataset. It is not a customer rating.
Assessed 10 sample records (JSON) on 2026-10-09. All records in the provided sample were checked.
| Check | Points | Evidence |
|---|---|---|
| Populated cells | 50 / 50 | 40 of 40 top-level cells contain a value. Null, absent and blank values count as missing; zero and false count as populated. |
| Consistent value types | 30 / 30 | 40 of 40 populated cells match their column's most common observed type. Types are inferred, not checked against real-world truth. |
| Consistent record shape | 20 / 20 | 10 of 10 records have the expected fields. CSV/TSV use the header width; JSON uses the union of observed keys. |
Field-level findings and improvements
Check missing cells and mixed types below. Document intentional missing values or mixed types in your field descriptions. Do not fill legitimate unknowns with invented values just to increase this score.
| Field | Missing cells | Most common type | Other populated types |
|---|---|---|---|
| id | 0 / 10 | string | 0 / 10 |
| input | 0 / 10 | string | 0 / 10 |
| meta | 0 / 10 | object | 0 / 10 |
| output | 0 / 10 | object | 0 / 10 |
About this data
Eleven datasets spanning fact-checking, entity linking, slot filling, open-domain question answering, and dialog generation, all grounded in a unified Wikipedia snapshot. Covers approximately 3.2 million examples across multiple task types.
Retrieve with your agent or Python
Create an account and configure DATABAZAAR_API_KEY. This example retrieves free or already purchased data; it never makes a purchase. For a multi-file dataset, choose a file index from the manifest.
Download the Python examplepython3 retrieve-dataset.py 8997fb69-ff6d-48af-926e-b3840702fd18 --output dataset.bin
Full supplier documentation
Schema
| Name | Type | Description |
|---|---|---|
| id | VARCHAR | Unique example identifier string |
| input | VARCHAR | Query, claim, or dialog context text |
| meta | STRUCT(left_context VARCHAR, mention VARCHAR, right_context VARCHAR, partial_evidence STRUCT(start_paragraph_id INTEGER, end_paragraph_id INTEGER, title VARCHAR, section VARCHAR, wikipedia_id VARCHAR, meta STRUCT(evidence_span VARCHAR[]))[], obj_surface VARCHAR[], sub_surface VARCHAR[], subj_aliases VARCHAR[], template_questions VARCHAR[]) | Task-specific metadata including entity mention context, surface forms, aliases, template questions, and partial Wikipedia evidence references |
| output | STRUCT(answer VARCHAR, meta STRUCT(score INTEGER), provenance STRUCT(bleu_score FLOAT, start_character INTEGER, start_paragraph_id INTEGER, end_character INTEGER, end_paragraph_id INTEGER, meta STRUCT(fever_page_id VARCHAR, fever_sentence_id INTEGER, annotation_id VARCHAR, yes_no_answer VARCHAR, evidence_span VARCHAR[]), section VARCHAR, title VARCHAR, wikipedia_id VARCHAR)[])[] | List of gold answers with text and supporting Wikipedia provenance (document ID, section, character spans, evidence metrics) |
Sample Data
Preview a sample of the data before downloading.
Public sample only. Sign in to retrieve the full dataset, including free datasets.
For AI Agents
# 1. Add to your agent's MCP config (claude_desktop_config.json or similar):
{
"mcpServers": {
"databazaar": { "command": "npx", "args": ["databazaar-mcp"] }
}
}
# 2. Your agent can then call:
search_datasets({ query: "KILT Knowledge-Intensive Langu" })
// Found: 8997fb69-ff6d-48af-926e-b3840702fd18
get_download_url({ dataset_id: "8997fb69-ff6d-48af-926e-b3840702fd18" }) // free — sign in with MCP OAuth first# Free dataset — sign in or use your account API key: curl https://api.databazaar.io/datasets/8997fb69-ff6d-48af-926e-b3840702fd18/download-url -H "Authorization: Bearer $DATABAZAAR_API_KEY"