textPleIAs/common_corpusllm-pretrainingmultilingualopen-licensetext-corpusparquetbooksscientificlegalcoderag

Common Corpus Multilingual Text

Free

Open dataset

Sample structure: 98.8 / 100
3 download links issued
Seller: DataBazaar
Sign up to download

Already have an account? Log in

Agent? Connect your account →

Category
Text
Records
69,907 rows
Format
PARQUET
Update Frequency
Not documented
Collection Method
uploaded
PII
No flagged field names; not a privacy audit
File Size
~410.04 MB
Download links issued
3

Source, license and coverage

Supplier documentation. These claims are separate from the automated sample score. A listing edit date is not a data freshness date.

License
Per-document public domain or open license
Source / creator
PleIAs/common_corpus
Collection method
PleIAs aggregated text from sources whose licenses permit redistribution and downstream use — primarily public domain works, government documents, openly-licensed scientific publications, and permissively-licensed code. Per-document license metadata is retained so downstream users can filter by license type. The corpus has been normalized to a uniform Parquet schema, with language tagging, deduplication, and quality filtering applied by the curators. See the full paper for methodology details on filtering, OCR cleanup, and provenance tracking.
Coverage start
Not documented
Coverage end
Not documented
Data last updated
Not documented
Update schedule
Not documented

Source documentation ↗

License terms ↗

- Language balance is heavily skewed toward English and major European languages; lower-resource languages are underrepresented. - Significant portions are derived from older public-domain texts (pre-1928), which introduces archaic language, OCR artifacts, and outdated factual content. - OCR quality varies across cultural heritage sources; some documents contain recognition errors. - Per-document license metadata is retained but downstream users are responsible for honoring source-level license terms (e.g. attribution for CC-BY components). - Source does not exhaustively document deduplication against common evaluation benchmarks; buyers should validate empirically for eval contamination. The source records rights per document. No single blanket license covers the entire collection.

Sample structure score: 98.8 / 100

This automated check describes the inspected sample, not factual accuracy, legal rights, representativeness, or the quality of the entire dataset. It is not a customer rating.

Assessed 10 sample records (JSON) on 2026-10-09. All records in the provided sample were checked.

CheckPointsEvidence
Populated cells48.8 / 50127 of 130 top-level cells contain a value. Null, absent and blank values count as missing; zero and false count as populated.
Consistent value types30 / 30127 of 127 populated cells match their column's most common observed type. Types are inferred, not checked against real-world truth.
Consistent record shape20 / 2010 of 10 records have the expected fields. CSV/TSV use the header width; JSON uses the union of observed keys.
Field-level findings and improvements

Check missing cells and mixed types below. Document intentional missing values or mixed types in your field descriptions. Do not fill legitimate unknowns with invented values just to increase this score.

FieldMissing cellsMost common typeOther populated types
identifier0 / 10string0 / 10
collection0 / 10string0 / 10
open_type0 / 10string0 / 10
curator0 / 10string0 / 10
license0 / 10string0 / 10
date3 / 10number0 / 7
title0 / 10string0 / 10
creator0 / 10string0 / 10
language0 / 10string0 / 10
language_type0 / 10string0 / 10
word_count0 / 10number0 / 10
token_count0 / 10number0 / 10
text0 / 10string0 / 10
How the score is calculated, its limitations, and how to correct an assessment →

About this data

Open and permissibly-licensed text dataset across books, newspapers, scientific articles, legal documents, code, and other sources in 13+ languages, designed for LLM pretraining.

Retrieve with your agent or Python

Create an account and configure DATABAZAAR_API_KEY. This example retrieves free or already purchased data; it never makes a purchase. For a multi-file dataset, choose a file index from the manifest.

Download the Python example
python3 retrieve-dataset.py 4252ca2b-9cc3-4573-9616-e393bdaffab7 --output dataset.bin
Full supplier documentation
## Overview Common Corpus is the largest openly-licensed text corpus available, comprising approximately 2.27 trillion tokens (2,267,302,720,836) of text drawn from books, newspapers, scientific articles, government and legal documents, source code, and other permissibly-licensed sources. The dataset is distributed in Parquet format and spans 13+ languages including English, French, German, Chinese, Italian, Spanish, Japanese, Polish, Latin, Dutch, Russian, Arabic, and Korean. It was assembled by PleIAs in collaboration with partners and is intended as a fully-open alternative to mixed-license web crawl corpora for LLM pretraining. ## Schema - `text` — string — the document or passage content - `identifier` — string — source/document ID - `collection` — string — sub-corpus the document belongs to (e.g. cultural heritage, scientific, legal, code) - `language` — string — ISO language code - `license` — string — original license of the source document - `date` — string/int — publication or source date where available - `source` — string — upstream provider (e.g. Wikisource, PubMed, government archive) - `token_count` — int — approximate token count - +additional provenance and quality columns per sub-collection ## Sources - PleIAs/common_corpus on Hugging Face — https://huggingface.co/datasets/PleIAs/common_corpus — openly/permissibly licensed (per-document licenses retained; aggregate is open) - Reference paper: arXiv:2410.22587 (ICLR 2026 oral) - Underlying sources include public domain books, openly-licensed scientific literature (e.g. PubMed Central OA subset), government/legal documents, permissively-licensed code, and openly-licensed news. ## Methodology PleIAs aggregated text from sources whose licenses permit redistribution and downstream use — primarily public domain works, government documents, openly-licensed scientific publications, and permissively-licensed code. Per-document license metadata is retained so downstream users can filter by license type. The corpus has been normalized to a uniform Parquet schema, with language tagging, deduplication, and quality filtering applied by the curators. See the full paper for methodology details on filtering, OCR cleanup, and provenance tracking. ## Known gaps & limitations - Language balance is heavily skewed toward English and major European languages; lower-resource languages are underrepresented. - Significant portions are derived from older public-domain texts (pre-1928), which introduces archaic language, OCR artifacts, and outdated factual content. - OCR quality varies across cultural heritage sources; some documents contain recognition errors. - Per-document license metadata is retained but downstream users are responsible for honoring source-level license terms (e.g. attribution for CC-BY components). - Source does not exhaustively document deduplication against common evaluation benchmarks; buyers should validate empirically for eval contamination. ## Intended use & out-of-scope - **Intended:** LLM pretraining and continued pretraining where license cleanliness matters; research on open-data language modeling; multilingual model development; RAG corpora needing permissive provenance. - **Out-of-scope:** Not deduplicated against common eval suites (MMLU, HellaSwag, etc.) — leakage risk for benchmark training. Not a current-events corpus — heavy historical/public-domain skew makes it unsuitable as a sole source for recency-sensitive applications. _PII signals: cc_shape×2 (Luhn-valid: 0), email×1, us_phone×1 present in the sample. Common in public datasets (papers, logs) but worth knowing before joining with private data._ Original supplier listing: Common Corpus — 2.27T Tokens of Permissibly-Licensed Text Largest open and permissibly-licensed text dataset: 2.27 trillion tokens across books, newspapers, scientific articles, legal docs, code, and more in 13+ languages. Built by PleIAs for LLM pretraining.

Schema

NameTypeDescription
identifierVARCHARUnique hash or ID for the document/passage
collectionVARCHARSub-corpus category (e.g., French Open Data, scientific, legal, code)
open_typeVARCHARClassification of openness (e.g., Open Government, Creative Commons)
curatorVARCHAROrganization responsible for curation (e.g., Pleias)
licenseVARCHARPermissive license type of the source document
dateBIGINTPublication or source date as Unix timestamp or null if unavailable
titleVARCHARDocument title or filename
creatorVARCHARAuthor, institution, or upstream source
languageVARCHARHuman language name
language_typeVARCHARMedium of text (Written or Spoken)
word_countBIGINTNumber of words in the document
token_countBIGINTApproximate token count using standard tokenization
textVARCHARFull document or passage content

Sample Data

Preview a sample of the data before downloading.

Public sample only. Sign in to retrieve the full dataset, including free datasets.

For AI Agents

Via MCP Server
# 1. Add to your agent's MCP config (claude_desktop_config.json or similar):
{
  "mcpServers": {
    "databazaar": { "command": "npx", "args": ["databazaar-mcp"] }
  }
}

# 2. Your agent can then call:
search_datasets({ query: "Common Corpus Multilingual Tex" })
// Found: 4252ca2b-9cc3-4573-9616-e393bdaffab7
get_download_url({ dataset_id: "4252ca2b-9cc3-4573-9616-e393bdaffab7" })  // free — sign in with MCP OAuth first
Via REST API
# Free dataset — sign in or use your account API key:
curl https://api.databazaar.io/datasets/4252ca2b-9cc3-4573-9616-e393bdaffab7/download-url -H "Authorization: Bearer $DATABAZAAR_API_KEY"