textHuggingFaceFW/fineweb-2pretrainingmultilingualllmweb-crawlcommoncrawlhuggingfaceodc-byfineweb

FineWeb2 Multilingual Web Corpus

Free

Open dataset

Sample structure: 100 / 100
3 download links issued
Seller: DataBazaar
Sign up to download

Already have an account? Log in

Agent? Connect your account →

Category
Text
Records
4,484,929,995 rows
Format
PARQUET
Update Frequency
Not documented
Collection Method
auto_imported_huggingface_federated
PII
No flagged field names; not a privacy audit
File Size
~8263788.18 MB
Download links issued
3

Source, license and coverage

Supplier documentation. These claims are separate from the automated sample score. A listing edit date is not a data freshness date.

License
odc-by
Source / creator
HuggingFaceFW/fineweb-2
Collection method
FineWeb2 starts from raw CommonCrawl WARC files, applies language identification (GlotLID and similar classifiers), then performs per-language quality filtering, MinHash near-deduplication, and PII redaction. Processing decisions were guided by hundreds of ablation experiments on 9 diverse pilot languages, and the full pipeline (datatrove) is open source and reproducible. Each language partition is shipped as parquet shards with metadata columns preserved for downstream filtering.
Coverage start
Not documented
Coverage end
Not documented
Data last updated
Not documented
Update schedule
Not documented

Source documentation ↗

License terms ↗

Quality varies dramatically across the 1000+ languages — high-resource languages (English, Mandarin, Spanish, Hindi) have far more data and stricter filter validation than low-resource languages where ablation studies were not run. Language identification errors are non-trivial for closely related languages and code-switched content. Web-sourced content includes biases toward commercial, technical, and Western perspectives; some toxic, NSFW, or factually incorrect material remains despite filtering. Temporal coverage is tied to CommonCrawl snapshots and skews recent. Not aligned to any specific eval benchmark — leakage risk if used naively for evaluation training.

Sample structure score: 100 / 100

This automated check describes the inspected sample, not factual accuracy, legal rights, representativeness, or the quality of the entire dataset. It is not a customer rating.

Assessed 1 sample records (JSON) on 2026-10-09. All records in the provided sample were checked.

CheckPointsEvidence
Populated cells50 / 5011 of 11 top-level cells contain a value. Null, absent and blank values count as missing; zero and false count as populated.
Consistent value types30 / 3011 of 11 populated cells match their column's most common observed type. Types are inferred, not checked against real-world truth.
Consistent record shape20 / 201 of 1 records have the expected fields. CSV/TSV use the header width; JSON uses the union of observed keys.
Field-level findings and improvements

Check missing cells and mixed types below. Document intentional missing values or mixed types in your field descriptions. Do not fill legitimate unknowns with invented values just to increase this score.

FieldMissing cellsMost common typeOther populated types
text0 / 1string0 / 1
id0 / 1string0 / 1
dump0 / 1string0 / 1
url0 / 1string0 / 1
date0 / 1string0 / 1
file_path0 / 1string0 / 1
language0 / 1string0 / 1
language_score0 / 1number0 / 1
language_script0 / 1string0 / 1
minhash_cluster_size0 / 1number0 / 1
top_langs0 / 1string0 / 1
How the score is calculated, its limitations, and how to correct an assessment →

About this data

Filtered web text for language model pretraining across 1000+ languages. Second iteration of HuggingFace's FineWeb dataset, validated through ablation experiments.

Retrieve with your agent or Python

Create an account and configure DATABAZAAR_API_KEY. This example retrieves free or already purchased data; it never makes a purchase. For a multi-file dataset, choose a file index from the manifest.

Download the Python example
python3 retrieve-dataset.py bb3745f0-b140-4136-a72f-e290a474dad7 --output dataset.bin
Full supplier documentation
## Overview FineWeb2 is a massive multilingual web pretraining corpus from HuggingFace covering over 1,000 languages, derived from CommonCrawl snapshots and processed with language-specific filtering and deduplication pipelines. The dataset contains billions of documents (1B<n<10B rows), distributed as parquet shards organized by language. It is the successor to the popular English-focused FineWeb dataset and is intended as a foundation for training multilingual LLMs. ## Schema - `text` — string — main document text content - `id` — string — unique document identifier - `dump` — string — CommonCrawl dump identifier (e.g., CC-MAIN-2024-XX) - `url` — string — source URL of the document - `date` — string — crawl timestamp - `file_path` — string — path within the CommonCrawl WARC archive - `language` — string — detected language code (ISO 639-3) - `language_score` — float — confidence of language detection - `language_script` — string — detected writing script - `minhash_cluster_size` — int — deduplication cluster size - `top_langs` — string — JSON of top detected language candidates ## Sources - HuggingFaceFW/fineweb-2 on HuggingFace: https://huggingface.co/datasets/HuggingFaceFW/fineweb-2 — License: ODC-By 1.0 - Upstream: CommonCrawl (https://commoncrawl.org) — also ODC-By aligned - Reference papers: arXiv:2506.20920 (FineWeb2), arXiv:2406.17557 (FineWeb), arXiv:2109.07445 (OSCAR-related methodology) ## Methodology FineWeb2 starts from raw CommonCrawl WARC files, applies language identification (GlotLID and similar classifiers), then performs per-language quality filtering, MinHash near-deduplication, and PII redaction. Processing decisions were guided by hundreds of ablation experiments on 9 diverse pilot languages, and the full pipeline (datatrove) is open source and reproducible. Each language partition is shipped as parquet shards with metadata columns preserved for downstream filtering. ## Known gaps & limitations Quality varies dramatically across the 1000+ languages — high-resource languages (English, Mandarin, Spanish, Hindi) have far more data and stricter filter validation than low-resource languages where ablation studies were not run. Language identification errors are non-trivial for closely related languages and code-switched content. Web-sourced content includes biases toward commercial, technical, and Western perspectives; some toxic, NSFW, or factually incorrect material remains despite filtering. Temporal coverage is tied to CommonCrawl snapshots and skews recent. Not aligned to any specific eval benchmark — leakage risk if used naively for evaluation training. ## Intended use & out-of-scope - **Intended for:** Pretraining and continued pretraining of multilingual LLMs, language-specific model training (especially for low-resource languages), tokenizer training, linguistic research on web text distributions. - **Not for:** Production use without further filtering for the user's safety requirements; benchmark evaluation (not deduplicated against common eval suites — leakage risk); authoritative source of facts; assumed-clean PII-free corpus despite redaction efforts. _Federated dataset: 4,569 parquet shards, 8070.11 GB total. Queries and downloads stream through the DataBazaar API._ _PII signals: cc_shape×5 (Luhn-valid: 0) present in the sample. Common in public datasets (papers, logs) but worth knowing before joining with private data._ Original supplier listing: FineWeb2: Multilingual Web Pretraining Corpus (1000+ Languages) Second iteration of HuggingFace's FineWeb dataset, covering 1000+ languages with high-quality filtered web text for LLM pretraining. ODC-By 1.0 licensed, validated through hundreds of ablation experiments.

Schema

NameTypeDescription
textVARCHARDocument text content in detected language
idVARCHARUnique document identifier (UUID URN format)
dumpVARCHARCommonCrawl dump identifier (CC-MAIN-YYYY-XX format)
urlVARCHARSource URL of the document
dateVARCHARISO 8601 crawl timestamp
file_pathVARCHARS3 path to document in CommonCrawl WARC archive
languageVARCHARISO 639-3 language code
language_scoreDOUBLELanguage detection confidence score (0.0–1.0)
language_scriptVARCHARISO 15924 script code for detected writing system
minhash_cluster_sizeBIGINTNumber of documents in deduplication cluster
top_langsVARCHARJSON object of top language candidates with confidence scores

Sample Data

Preview a sample of the data before downloading.

Public sample only. Sign in to retrieve the full dataset, including free datasets.

For AI Agents

Via MCP Server
# 1. Add to your agent's MCP config (claude_desktop_config.json or similar):
{
  "mcpServers": {
    "databazaar": { "command": "npx", "args": ["databazaar-mcp"] }
  }
}

# 2. Your agent can then call:
search_datasets({ query: "FineWeb2 Multilingual Web Corp" })
// Found: bb3745f0-b140-4136-a72f-e290a474dad7
get_download_url({ dataset_id: "bb3745f0-b140-4136-a72f-e290a474dad7" })  // free — sign in with MCP OAuth first
Via REST API
# Free dataset — sign in or use your account API key:
curl https://api.databazaar.io/datasets/bb3745f0-b140-4136-a72f-e290a474dad7/download-url -H "Authorization: Bearer $DATABAZAAR_API_KEY"