PubTabNet Table Images with HTML Annotations
Source, license and coverage
Supplier documentation. These claims are separate from the automated sample score. A listing edit date is not a data freshness date.
- License
- cdla-sharing-1.0
- Source / creator
- ajimeno/PubTabNet
- Collection method
- Table regions were identified by aligning the PDF rendering of each article with its corresponding XML representation in the PubMed Central Open Access Subset. Where the PDF layout could be reliably matched to XML-defined table content, the table region was cropped as an image and paired with HTML reconstructed from the XML structure, including cell-level text and span information. This HuggingFace mirror repackages the original IBM release into the WebDataset format for streamable consumption.
- Coverage start
- Not documented
- Coverage end
- Not documented
- Data last updated
- Not documented
- Update schedule
- Not documented
Content is restricted to scientific/biomedical publications, so visual style, terminology, and table conventions are skewed toward that domain — generalization to financial, legal, or handwritten tables is not guaranteed. Annotations are derived from PDF/XML alignment heuristics and may contain misaligned cells or imperfect HTML for complex tables (deeply nested headers, rotated text, multi-page tables). Cell content text is the XML-extracted string, not OCR — it may differ from what a model would read off pixels. Source does not enumerate per-example quality metadata; buyers should validate empirically on their target distribution.
Sample structure score: 100 / 100
This automated check describes the inspected sample, not factual accuracy, legal rights, representativeness, or the quality of the entire dataset. It is not a customer rating.
Assessed 10 sample records (JSON) on 2026-10-09. All records in the provided sample were checked.
| Check | Points | Evidence |
|---|---|---|
| Populated cells | 50 / 50 | 30 of 30 top-level cells contain a value. Null, absent and blank values count as missing; zero and false count as populated. |
| Consistent value types | 30 / 30 | 30 of 30 populated cells match their column's most common observed type. Types are inferred, not checked against real-world truth. |
| Consistent record shape | 20 / 20 | 10 of 10 records have the expected fields. CSV/TSV use the header width; JSON uses the union of observed keys. |
Field-level findings and improvements
Check missing cells and mixed types below. Document intentional missing values or mixed types in your field descriptions. Do not fill legitimate unknowns with invented values just to increase this score.
| Field | Missing cells | Most common type | Other populated types |
|---|---|---|---|
| png | 0 / 10 | object | 0 / 10 |
| __key__ | 0 / 10 | string | 0 / 10 |
| __url__ | 0 / 10 | string | 0 / 10 |
About this data
Scientific table images extracted from PubMed Central Open Access articles, paired with HTML structure annotations. Standard benchmark for image-based table recognition and document AI.
Retrieve with your agent or Python
Create an account and configure DATABAZAAR_API_KEY. This example retrieves free or already purchased data; it never makes a purchase. For a multi-file dataset, choose a file index from the manifest.
Download the Python examplepython3 retrieve-dataset.py e82724e4-9b1e-47f5-bf62-c21e5982a2b2 --output dataset.bin
Full supplier documentation
Schema
| Name | Type | Description |
|---|---|---|
| png | STRUCT(bytes BLOB, path VARCHAR) | PNG image file as binary blob with embedded path reference to source table image |
| __key__ | VARCHAR | Unique identifier for the WebDataset sample linking image and HTML annotation |
| __url__ | VARCHAR | Source URL or reference path to the original table image in the dataset repository |
Sample Data
Preview a sample of the data before downloading.
Public sample only. Sign in to retrieve the full dataset, including free datasets.
For AI Agents
# 1. Add to your agent's MCP config (claude_desktop_config.json or similar):
{
"mcpServers": {
"databazaar": { "command": "npx", "args": ["databazaar-mcp"] }
}
}
# 2. Your agent can then call:
search_datasets({ query: "PubTabNet Table Images with HT" })
// Found: e82724e4-9b1e-47f5-bf62-c21e5982a2b2
get_download_url({ dataset_id: "e82724e4-9b1e-47f5-bf62-c21e5982a2b2" }) // free — sign in with MCP OAuth first# Free dataset — sign in or use your account API key: curl https://api.databazaar.io/datasets/e82724e4-9b1e-47f5-bf62-c21e5982a2b2/download-url -H "Authorization: Bearer $DATABAZAAR_API_KEY"