Common Corpus Multilingual Text
Source, license and coverage
Supplier documentation. These claims are separate from the automated sample score. A listing edit date is not a data freshness date.
- License
- Per-document public domain or open license
- Source / creator
- PleIAs/common_corpus
- Collection method
- PleIAs aggregated text from sources whose licenses permit redistribution and downstream use — primarily public domain works, government documents, openly-licensed scientific publications, and permissively-licensed code. Per-document license metadata is retained so downstream users can filter by license type. The corpus has been normalized to a uniform Parquet schema, with language tagging, deduplication, and quality filtering applied by the curators. See the full paper for methodology details on filtering, OCR cleanup, and provenance tracking.
- Coverage start
- Not documented
- Coverage end
- Not documented
- Data last updated
- Not documented
- Update schedule
- Not documented
- Language balance is heavily skewed toward English and major European languages; lower-resource languages are underrepresented. - Significant portions are derived from older public-domain texts (pre-1928), which introduces archaic language, OCR artifacts, and outdated factual content. - OCR quality varies across cultural heritage sources; some documents contain recognition errors. - Per-document license metadata is retained but downstream users are responsible for honoring source-level license terms (e.g. attribution for CC-BY components). - Source does not exhaustively document deduplication against common evaluation benchmarks; buyers should validate empirically for eval contamination. The source records rights per document. No single blanket license covers the entire collection.
Sample structure score: 98.8 / 100
This automated check describes the inspected sample, not factual accuracy, legal rights, representativeness, or the quality of the entire dataset. It is not a customer rating.
Assessed 10 sample records (JSON) on 2026-10-09. All records in the provided sample were checked.
| Check | Points | Evidence |
|---|---|---|
| Populated cells | 48.8 / 50 | 127 of 130 top-level cells contain a value. Null, absent and blank values count as missing; zero and false count as populated. |
| Consistent value types | 30 / 30 | 127 of 127 populated cells match their column's most common observed type. Types are inferred, not checked against real-world truth. |
| Consistent record shape | 20 / 20 | 10 of 10 records have the expected fields. CSV/TSV use the header width; JSON uses the union of observed keys. |
Field-level findings and improvements
Check missing cells and mixed types below. Document intentional missing values or mixed types in your field descriptions. Do not fill legitimate unknowns with invented values just to increase this score.
| Field | Missing cells | Most common type | Other populated types |
|---|---|---|---|
| identifier | 0 / 10 | string | 0 / 10 |
| collection | 0 / 10 | string | 0 / 10 |
| open_type | 0 / 10 | string | 0 / 10 |
| curator | 0 / 10 | string | 0 / 10 |
| license | 0 / 10 | string | 0 / 10 |
| date | 3 / 10 | number | 0 / 7 |
| title | 0 / 10 | string | 0 / 10 |
| creator | 0 / 10 | string | 0 / 10 |
| language | 0 / 10 | string | 0 / 10 |
| language_type | 0 / 10 | string | 0 / 10 |
| word_count | 0 / 10 | number | 0 / 10 |
| token_count | 0 / 10 | number | 0 / 10 |
| text | 0 / 10 | string | 0 / 10 |
About this data
Open and permissibly-licensed text dataset across books, newspapers, scientific articles, legal documents, code, and other sources in 13+ languages, designed for LLM pretraining.
Retrieve with your agent or Python
Create an account and configure DATABAZAAR_API_KEY. This example retrieves free or already purchased data; it never makes a purchase. For a multi-file dataset, choose a file index from the manifest.
Download the Python examplepython3 retrieve-dataset.py 4252ca2b-9cc3-4573-9616-e393bdaffab7 --output dataset.bin
Full supplier documentation
Schema
| Name | Type | Description |
|---|---|---|
| identifier | VARCHAR | Unique hash or ID for the document/passage |
| collection | VARCHAR | Sub-corpus category (e.g., French Open Data, scientific, legal, code) |
| open_type | VARCHAR | Classification of openness (e.g., Open Government, Creative Commons) |
| curator | VARCHAR | Organization responsible for curation (e.g., Pleias) |
| license | VARCHAR | Permissive license type of the source document |
| date | BIGINT | Publication or source date as Unix timestamp or null if unavailable |
| title | VARCHAR | Document title or filename |
| creator | VARCHAR | Author, institution, or upstream source |
| language | VARCHAR | Human language name |
| language_type | VARCHAR | Medium of text (Written or Spoken) |
| word_count | BIGINT | Number of words in the document |
| token_count | BIGINT | Approximate token count using standard tokenization |
| text | VARCHAR | Full document or passage content |
Sample Data
Preview a sample of the data before downloading.
Public sample only. Sign in to retrieve the full dataset, including free datasets.
For AI Agents
# 1. Add to your agent's MCP config (claude_desktop_config.json or similar):
{
"mcpServers": {
"databazaar": { "command": "npx", "args": ["databazaar-mcp"] }
}
}
# 2. Your agent can then call:
search_datasets({ query: "Common Corpus Multilingual Tex" })
// Found: 4252ca2b-9cc3-4573-9616-e393bdaffab7
get_download_url({ dataset_id: "4252ca2b-9cc3-4573-9616-e393bdaffab7" }) // free — sign in with MCP OAuth first# Free dataset — sign in or use your account API key: curl https://api.databazaar.io/datasets/4252ca2b-9cc3-4573-9616-e393bdaffab7/download-url -H "Authorization: Bearer $DATABAZAAR_API_KEY"