imagesajimeno/PubTabNettable-recognitiondocument-aiocrcomputer-visionhtmlpubmedwebdatasetimage-to-textlayout-analysisbenchmark

PubTabNet Table Images with HTML Annotations

Free

Open dataset

Sample structure: 100 / 100
5 download links issued
Seller: DataBazaar
Sign up to download

Already have an account? Log in

Agent? Connect your account →

Category
Images
Records
223,100 rows
Format
PARQUET
Update Frequency
Not documented
Collection Method
auto_imported_huggingface_federated
PII
No flagged field names; not a privacy audit
File Size
~4747.52 MB
Download links issued
5

Source, license and coverage

Supplier documentation. These claims are separate from the automated sample score. A listing edit date is not a data freshness date.

License
cdla-sharing-1.0
Source / creator
ajimeno/PubTabNet
Collection method
Table regions were identified by aligning the PDF rendering of each article with its corresponding XML representation in the PubMed Central Open Access Subset. Where the PDF layout could be reliably matched to XML-defined table content, the table region was cropped as an image and paired with HTML reconstructed from the XML structure, including cell-level text and span information. This HuggingFace mirror repackages the original IBM release into the WebDataset format for streamable consumption.
Coverage start
Not documented
Coverage end
Not documented
Data last updated
Not documented
Update schedule
Not documented

Source documentation ↗

License terms ↗

Content is restricted to scientific/biomedical publications, so visual style, terminology, and table conventions are skewed toward that domain — generalization to financial, legal, or handwritten tables is not guaranteed. Annotations are derived from PDF/XML alignment heuristics and may contain misaligned cells or imperfect HTML for complex tables (deeply nested headers, rotated text, multi-page tables). Cell content text is the XML-extracted string, not OCR — it may differ from what a model would read off pixels. Source does not enumerate per-example quality metadata; buyers should validate empirically on their target distribution.

Sample structure score: 100 / 100

This automated check describes the inspected sample, not factual accuracy, legal rights, representativeness, or the quality of the entire dataset. It is not a customer rating.

Assessed 10 sample records (JSON) on 2026-10-09. All records in the provided sample were checked.

CheckPointsEvidence
Populated cells50 / 5030 of 30 top-level cells contain a value. Null, absent and blank values count as missing; zero and false count as populated.
Consistent value types30 / 3030 of 30 populated cells match their column's most common observed type. Types are inferred, not checked against real-world truth.
Consistent record shape20 / 2010 of 10 records have the expected fields. CSV/TSV use the header width; JSON uses the union of observed keys.
Field-level findings and improvements

Check missing cells and mixed types below. Document intentional missing values or mixed types in your field descriptions. Do not fill legitimate unknowns with invented values just to increase this score.

FieldMissing cellsMost common typeOther populated types
png0 / 10object0 / 10
__key__0 / 10string0 / 10
__url__0 / 10string0 / 10
How the score is calculated, its limitations, and how to correct an assessment →

About this data

Scientific table images extracted from PubMed Central Open Access articles, paired with HTML structure annotations. Standard benchmark for image-based table recognition and document AI.

Retrieve with your agent or Python

Create an account and configure DATABAZAAR_API_KEY. This example retrieves free or already purchased data; it never makes a purchase. For a multi-file dataset, choose a file index from the manifest.

Download the Python example
python3 retrieve-dataset.py e82724e4-9b1e-47f5-bf62-c21e5982a2b2 --output dataset.bin
Full supplier documentation
## Overview PubTabNet is a large-scale dataset for image-based table recognition containing 568,000+ images of tabular data, each annotated with the corresponding HTML representation describing both table structure and cell content. The tables were extracted from scientific publications in the PubMed Central Open Access Subset (commercial use collection). The dataset is distributed in WebDataset format and combines image and text modalities, making it suitable for training table structure recognition, OCR, and document understanding models. ## Schema - `image` — image (PNG) — rendered table region cropped from the source PDF - `html` / `structure` — text — HTML annotation encoding table structure (rows, columns, spans) and cell text - `filename` — string — original identifier tying back to the source PMC article - `split` — string — train/val/test partition (per original PubTabNet release) ## Sources - HuggingFace: https://huggingface.co/datasets/ajimeno/PubTabNet — license: CDLA-Sharing-1.0 - Original dataset: IBM Research PubTabNet (https://github.com/ibm-aur-nlp/PubTabNet) - Underlying corpus: PubMed Central Open Access Subset (commercial use collection) - Reference: Zhong et al., "Image-based table recognition: data, model, and evaluation" (arXiv:1911.10683) ## Methodology Table regions were identified by aligning the PDF rendering of each article with its corresponding XML representation in the PubMed Central Open Access Subset. Where the PDF layout could be reliably matched to XML-defined table content, the table region was cropped as an image and paired with HTML reconstructed from the XML structure, including cell-level text and span information. This HuggingFace mirror repackages the original IBM release into the WebDataset format for streamable consumption. ## Known gaps & limitations Content is restricted to scientific/biomedical publications, so visual style, terminology, and table conventions are skewed toward that domain — generalization to financial, legal, or handwritten tables is not guaranteed. Annotations are derived from PDF/XML alignment heuristics and may contain misaligned cells or imperfect HTML for complex tables (deeply nested headers, rotated text, multi-page tables). Cell content text is the XML-extracted string, not OCR — it may differ from what a model would read off pixels. Source does not enumerate per-example quality metadata; buyers should validate empirically on their target distribution. ## Intended use & out-of-scope - IS for: training and evaluating table structure recognition models, document AI / layout models, image-to-HTML sequence models, OCR pipelines that need table-aware post-processing, and RAG over scientific documents. - NOT for: claims of coverage outside biomedical literature; benchmark training without checking leakage against PubTabNet's official val/test splits; tasks needing handwritten or non-English tables. ## License CDLA-Sharing-1.0 (Community Data License Agreement – Sharing). Share-alike: redistribution and derivative datasets are permitted provided the same license terms are preserved. _Federated dataset: 10 parquet shards, 4.64 GB total. Queries and downloads stream through the DataBazaar API._ Original supplier listing: PubTabNet — 568K Table Images with HTML Annotations 568K+ scientific table images paired with HTML structure annotations, extracted from PubMed Central Open Access articles. Standard benchmark for image-based table recognition and document AI.

Schema

NameTypeDescription
pngSTRUCT(bytes BLOB, path VARCHAR)PNG image file as binary blob with embedded path reference to source table image
__key__VARCHARUnique identifier for the WebDataset sample linking image and HTML annotation
__url__VARCHARSource URL or reference path to the original table image in the dataset repository

Sample Data

Preview a sample of the data before downloading.

Public sample only. Sign in to retrieve the full dataset, including free datasets.

For AI Agents

Via MCP Server
# 1. Add to your agent's MCP config (claude_desktop_config.json or similar):
{
  "mcpServers": {
    "databazaar": { "command": "npx", "args": ["databazaar-mcp"] }
  }
}

# 2. Your agent can then call:
search_datasets({ query: "PubTabNet Table Images with HT" })
// Found: e82724e4-9b1e-47f5-bf62-c21e5982a2b2
get_download_url({ dataset_id: "e82724e4-9b1e-47f5-bf62-c21e5982a2b2" })  // free — sign in with MCP OAuth first
Via REST API
# Free dataset — sign in or use your account API key:
curl https://api.databazaar.io/datasets/e82724e4-9b1e-47f5-bf62-c21e5982a2b2/download-url -H "Authorization: Bearer $DATABAZAAR_API_KEY"