imageslmms-lab/TextCapsimage-captioningocrmultimodalvision-languagebenchmarklmms-evalevaluationvqa

TextCaps Image Captioning Benchmark

Free

Open dataset

Sample structure: 100 / 100
4 download links issued
Seller: DataBazaar
Sign up to download

Already have an account? Log in

Agent? Connect your account →

Category
Images
Records
28,408 rows
Format
PARQUET
Update Frequency
Not documented
Collection Method
auto_imported_huggingface_federated
PII
No flagged field names; not a privacy audit
File Size
~7689.6 MB
Download links issued
4

Source, license and coverage

Supplier documentation. These claims are separate from the automated sample score. A listing edit date is not a data freshness date.

License
Not documented — confirm reuse terms with the seller
Source / creator
lmms-lab/TextCaps
Collection method
The original TextCaps was constructed by selecting images from the Open Images dataset that contain text, then crowdsourcing multiple captions per image where annotators were instructed to describe the image including any text present. lmms-lab reformatted the dataset into parquet with embedded images and standardized fields for plug-and-play evaluation in their lmms-eval framework.
Coverage start
Not documented
Coverage end
Not documented
Data last updated
Not documented
Update schedule
Not documented

Source documentation ↗

Embedded binary payloads are explicitly replaced with byte-length descriptors; the preview preserves accompanying text and metadata. English captions only; OCR difficulty varies and some text in images may be unreadable. Images are sourced from Open Images so biases of that collection (Western, web-scraped) carry over. The lmms-lab reformatting is for evaluation use — the train split formatting may differ from the original release. Source does not exhaustively document reformatting deltas; buyers should validate against the canonical TextCaps release if exact parity matters.

Sample structure score: 100 / 100

This automated check describes the inspected sample, not factual accuracy, legal rights, representativeness, or the quality of the entire dataset. It is not a customer rating.

Assessed 10 sample records (JSON) on 2026-10-09. All records in the provided sample were checked.

CheckPointsEvidence
Populated cells50 / 50150 of 150 top-level cells contain a value. Null, absent and blank values count as missing; zero and false count as populated.
Consistent value types30 / 30150 of 150 populated cells match their column's most common observed type. Types are inferred, not checked against real-world truth.
Consistent record shape20 / 2010 of 10 records have the expected fields. CSV/TSV use the header width; JSON uses the union of observed keys.
Field-level findings and improvements

Check missing cells and mixed types below. Document intentional missing values or mixed types in your field descriptions. Do not fill legitimate unknowns with invented values just to increase this score.

FieldMissing cellsMost common typeOther populated types
question_id0 / 10string0 / 10
question0 / 10string0 / 10
image0 / 10object0 / 10
image_id0 / 10string0 / 10
image_classes0 / 10object0 / 10
flickr_original_url0 / 10string0 / 10
flickr_300k_url0 / 10string0 / 10
image_width0 / 10number0 / 10
image_height0 / 10number0 / 10
set_name0 / 10string0 / 10
image_name0 / 10string0 / 10
image_path0 / 10string0 / 10
caption_id0 / 10object0 / 10
caption_str0 / 10object0 / 10
reference_strs0 / 10object0 / 10
How the score is calculated, its limitations, and how to correct an assessment →

About this data

Image captioning benchmark requiring optical character recognition and text reading within images. Formatted for multimodal model evaluation.

Retrieve with your agent or Python

Create an account and configure DATABAZAAR_API_KEY. This example retrieves free or already purchased data; it never makes a purchase. For a multi-file dataset, choose a file index from the manifest.

Download the Python example
python3 retrieve-dataset.py 992a5d4b-efc7-40b0-ac8a-26c805d510c2 --output dataset.bin
Full supplier documentation
## Overview TextCaps is an image captioning dataset that specifically requires models to read and reason about text appearing in images. This is the lmms-lab formatted version, prepared for use in the lmms-eval evaluation pipeline for large multimodality models (LMMs). The dataset contains roughly 28K images with multiple reference captions each, distributed in parquet format with image and text modalities. Size category is 10K-100K rows. ## Schema - image — image (bytes) — the source image containing text to be read - image_id — string — unique identifier for the image - caption_str — string — reference caption(s) describing the image (including OCR'd text) - reference_strs — list[string] — multiple human reference captions for evaluation - ocr_tokens — list[string] — OCR tokens extracted from the image (where provided) - question_id / set_name — string — evaluation split metadata ## Sources - HuggingFace: https://huggingface.co/datasets/lmms-lab/TextCaps - Original TextCaps paper: Sidorov et al., ECCV 2020 — "TextCaps: a Dataset for Image Captioning with Reading Comprehension" - License: per source (TextCaps original is CC-BY 4.0; images sourced from Open Images) ## Methodology The original TextCaps was constructed by selecting images from the Open Images dataset that contain text, then crowdsourcing multiple captions per image where annotators were instructed to describe the image including any text present. lmms-lab reformatted the dataset into parquet with embedded images and standardized fields for plug-and-play evaluation in their lmms-eval framework. ## Known gaps & limitations English captions only; OCR difficulty varies and some text in images may be unreadable. Images are sourced from Open Images so biases of that collection (Western, web-scraped) carry over. The lmms-lab reformatting is for evaluation use — the train split formatting may differ from the original release. Source does not exhaustively document reformatting deltas; buyers should validate against the canonical TextCaps release if exact parity matters. ## Intended use & out-of-scope - IS for: evaluating multimodal/vision-language models on OCR-aware image captioning, benchmarking LMMs via lmms-eval, RAG over image+text content. - NOT for: training production captioning systems without deduplication against the eval set; this is a public benchmark with high contamination risk for any model trained on web data. _Federated dataset: 17 parquet shards, 7.51 GB total. Queries and downloads stream through the DataBazaar API._ Original supplier listing: TextCaps (lmms-eval formatted) Image captioning benchmark requiring OCR/text reading in images. Formatted by lmms-lab for one-click multimodal model evaluation. ~28K images with captions.

Schema

NameTypeDescription
question_idVARCHARUnique identifier for the image/question pair in the evaluation set
questionVARCHARInstruction prompt asking the model to generate a caption for the image
imageSTRUCT(bytes BLOB, path VARCHAR)Image data as binary blob with optional file path reference
image_idVARCHARUnique identifier for the source image
image_classesVARCHAR[]List of semantic object class labels present in the image
flickr_original_urlVARCHARURL to original image on Flickr
flickr_300k_urlVARCHARURL to image from Flickr 300K subset
image_widthBIGINTImage width in pixels
image_heightBIGINTImage height in pixels
set_nameVARCHAREvaluation split designation (train/val/test)
image_nameVARCHARFilename of the image
image_pathVARCHARFile path or identifier for image location
caption_idBIGINT[]List of numeric identifiers for captions associated with the image
caption_strVARCHAR[]List of reference captions describing the image including OCR text
reference_strsVARCHAR[]Multiple human-written reference captions for evaluation metrics

Sample Data

Preview a sample of the data before downloading.

Public sample only. Sign in to retrieve the full dataset, including free datasets.

For AI Agents

Via MCP Server
# 1. Add to your agent's MCP config (claude_desktop_config.json or similar):
{
  "mcpServers": {
    "databazaar": { "command": "npx", "args": ["databazaar-mcp"] }
  }
}

# 2. Your agent can then call:
search_datasets({ query: "TextCaps Image Captioning Benc" })
// Found: 992a5d4b-efc7-40b0-ac8a-26c805d510c2
get_download_url({ dataset_id: "992a5d4b-efc7-40b0-ac8a-26c805d510c2" })  // free — sign in with MCP OAuth first
Via REST API
# Free dataset — sign in or use your account API key:
curl https://api.databazaar.io/datasets/992a5d4b-efc7-40b0-ac8a-26c805d510c2/download-url -H "Authorization: Bearer $DATABAZAAR_API_KEY"