TextCaps Image Captioning Benchmark
Source, license and coverage
Supplier documentation. These claims are separate from the automated sample score. A listing edit date is not a data freshness date.
- License
- Not documented — confirm reuse terms with the seller
- Source / creator
- lmms-lab/TextCaps
- Collection method
- The original TextCaps was constructed by selecting images from the Open Images dataset that contain text, then crowdsourcing multiple captions per image where annotators were instructed to describe the image including any text present. lmms-lab reformatted the dataset into parquet with embedded images and standardized fields for plug-and-play evaluation in their lmms-eval framework.
- Coverage start
- Not documented
- Coverage end
- Not documented
- Data last updated
- Not documented
- Update schedule
- Not documented
Embedded binary payloads are explicitly replaced with byte-length descriptors; the preview preserves accompanying text and metadata. English captions only; OCR difficulty varies and some text in images may be unreadable. Images are sourced from Open Images so biases of that collection (Western, web-scraped) carry over. The lmms-lab reformatting is for evaluation use — the train split formatting may differ from the original release. Source does not exhaustively document reformatting deltas; buyers should validate against the canonical TextCaps release if exact parity matters.
Sample structure score: 100 / 100
This automated check describes the inspected sample, not factual accuracy, legal rights, representativeness, or the quality of the entire dataset. It is not a customer rating.
Assessed 10 sample records (JSON) on 2026-10-09. All records in the provided sample were checked.
| Check | Points | Evidence |
|---|---|---|
| Populated cells | 50 / 50 | 150 of 150 top-level cells contain a value. Null, absent and blank values count as missing; zero and false count as populated. |
| Consistent value types | 30 / 30 | 150 of 150 populated cells match their column's most common observed type. Types are inferred, not checked against real-world truth. |
| Consistent record shape | 20 / 20 | 10 of 10 records have the expected fields. CSV/TSV use the header width; JSON uses the union of observed keys. |
Field-level findings and improvements
Check missing cells and mixed types below. Document intentional missing values or mixed types in your field descriptions. Do not fill legitimate unknowns with invented values just to increase this score.
| Field | Missing cells | Most common type | Other populated types |
|---|---|---|---|
| question_id | 0 / 10 | string | 0 / 10 |
| question | 0 / 10 | string | 0 / 10 |
| image | 0 / 10 | object | 0 / 10 |
| image_id | 0 / 10 | string | 0 / 10 |
| image_classes | 0 / 10 | object | 0 / 10 |
| flickr_original_url | 0 / 10 | string | 0 / 10 |
| flickr_300k_url | 0 / 10 | string | 0 / 10 |
| image_width | 0 / 10 | number | 0 / 10 |
| image_height | 0 / 10 | number | 0 / 10 |
| set_name | 0 / 10 | string | 0 / 10 |
| image_name | 0 / 10 | string | 0 / 10 |
| image_path | 0 / 10 | string | 0 / 10 |
| caption_id | 0 / 10 | object | 0 / 10 |
| caption_str | 0 / 10 | object | 0 / 10 |
| reference_strs | 0 / 10 | object | 0 / 10 |
About this data
Image captioning benchmark requiring optical character recognition and text reading within images. Formatted for multimodal model evaluation.
Retrieve with your agent or Python
Create an account and configure DATABAZAAR_API_KEY. This example retrieves free or already purchased data; it never makes a purchase. For a multi-file dataset, choose a file index from the manifest.
Download the Python examplepython3 retrieve-dataset.py 992a5d4b-efc7-40b0-ac8a-26c805d510c2 --output dataset.bin
Full supplier documentation
Schema
| Name | Type | Description |
|---|---|---|
| question_id | VARCHAR | Unique identifier for the image/question pair in the evaluation set |
| question | VARCHAR | Instruction prompt asking the model to generate a caption for the image |
| image | STRUCT(bytes BLOB, path VARCHAR) | Image data as binary blob with optional file path reference |
| image_id | VARCHAR | Unique identifier for the source image |
| image_classes | VARCHAR[] | List of semantic object class labels present in the image |
| flickr_original_url | VARCHAR | URL to original image on Flickr |
| flickr_300k_url | VARCHAR | URL to image from Flickr 300K subset |
| image_width | BIGINT | Image width in pixels |
| image_height | BIGINT | Image height in pixels |
| set_name | VARCHAR | Evaluation split designation (train/val/test) |
| image_name | VARCHAR | Filename of the image |
| image_path | VARCHAR | File path or identifier for image location |
| caption_id | BIGINT[] | List of numeric identifiers for captions associated with the image |
| caption_str | VARCHAR[] | List of reference captions describing the image including OCR text |
| reference_strs | VARCHAR[] | Multiple human-written reference captions for evaluation metrics |
Sample Data
Preview a sample of the data before downloading.
Public sample only. Sign in to retrieve the full dataset, including free datasets.
For AI Agents
# 1. Add to your agent's MCP config (claude_desktop_config.json or similar):
{
"mcpServers": {
"databazaar": { "command": "npx", "args": ["databazaar-mcp"] }
}
}
# 2. Your agent can then call:
search_datasets({ query: "TextCaps Image Captioning Benc" })
// Found: 992a5d4b-efc7-40b0-ac8a-26c805d510c2
get_download_url({ dataset_id: "992a5d4b-efc7-40b0-ac8a-26c805d510c2" }) // free — sign in with MCP OAuth first# Free dataset — sign in or use your account API key: curl https://api.databazaar.io/datasets/992a5d4b-efc7-40b0-ac8a-26c805d510c2/download-url -H "Authorization: Bearer $DATABAZAAR_API_KEY"