To search this category programmatically: GET https://api.databazaar.io/datasets?category=images
Full API docs: https://api.databazaar.io/llms.txt
Browse / images
Image & Visual Datasets
30 listingsVisual data trains and evaluates computer-vision models — classification, detection, segmentation, and multimodal LLMs. This category includes labeled photo collections, annotated image sets, and video-derived frames with structured metadata. Every listing is scanned for prohibited content before going live, and the free sample lets you inspect annotation quality first-hand.
Conceptual Captions — 5.3M Image–Caption Pairs
Millions of web images paired with natural-language captions, harvested and filtered from alt-text descriptions. A large-scale resource for training and evaluating image-captioning and vision-language models.
Vessel Detection — Labeled Satellite Image Patches
1,920 validated satellite image patches labeled for maritime vessel detection, exported from a boat-detection review workflow. Each patch ships with Hugging Face image-folder-compatible metadata for object-detection and classification training.
Global Wheat Full Semantic Organ Segmentation (GWFSS) v1.0
Labelled image dataset for semantic segmentation of wheat plant organs (canopy, leaves, stems, heads) across global field conditions. CC-BY-4.0, ~1K-10K images in parquet format.
RxRx3-core — Phenomics Microscopy Image Challenge
Recursion's RxRx3-core phenomics dataset: labeled microscopy images of 735 genetic knockouts and 1,674 small-molecule perturbations drawn from RxRx3, released as a benchmark for cellular image representation learning.
Pexels 568K Synthetic Captions (InternVL2-40B)
567,573 synthetic English captions for Pexels photos, generated with InternVL2-40B-AWQ and grounded with original tags. JSON format, ideal for text-to-image and image-to-text model training.
American Sign Language (ASL) Video Dataset — 108K Videos, 2,208 Words
108,618 ASL gesture videos covering 2,208 distinct words (≥30 videos per word). MIT-licensed, preprocessed for ML training and gesture recognition.
ShotBench: Cinematic Understanding Benchmark for VLMs
3,572 expert-level QA pairs over 3,049 images and 464 video clips from Oscar-nominated cinematography films, for evaluating cinematic understanding in vision-language models.
UAVIT-1M: UAV Visual Instruction Tuning Dataset (1M+)
Largest instruction-tuning dataset for low-altitude UAV visual understanding, with 1M+ samples across 11 image- and region-level tasks. CC-BY-4.0.
TextCaps (lmms-eval formatted)
Image captioning benchmark requiring OCR/text reading in images. Formatted by lmms-lab for one-click multimodal model evaluation. ~28K images with captions.
Vero-600K: Multi-Task Visual Reasoning RL Dataset
600K curated reinforcement learning samples from 59 datasets across 6 visual reasoning categories for training and evaluating vision-language models.
PubTabNet — 568K Table Images with HTML Annotations
568K+ scientific table images paired with HTML structure annotations, extracted from PubMed Central Open Access articles. Standard benchmark for image-based table recognition and document AI.
EO-Data-1.5M: Interleaved Vision-Text-Action Dataset for Embodied AI
1.5M-sample interleaved vision-language-action dataset for embodied AI and robot learning, emphasizing temporal dynamics and causal dependencies across modalities. Apache-2.0 licensed, parquet format.
FLUX-Reason-6M: 6M Reasoning-Focused Text-to-Image Dataset
6 million high-quality images synthesized by FLUX.1-dev with 20M bilingual (English/Chinese) descriptions, engineered to instill complex reasoning capabilities in text-to-image generative models.
Text-to-Image 2M: Curated Text-Image Pair Training Dataset
~2M curated high-quality text-image pairs in webdataset format, designed for fine-tuning text-to-image generative models. MIT licensed.
TreeOfLife-10M: 10M Biological Organism Images with Taxonomic Labels
Largest ML-ready dataset of biological organism images (10M+ images, 454K taxa) paired with taxonomic labels. Aggregates iNat21, BIOSCAN-1M, and Encyclopedia of Life. Used to train BioCLIP and similar vision-language models.
MMFineReason-SFT-123K (Qwen3-VL-235B Thinking Traces)
123K hardest-7% multimodal reasoning samples with Qwen3-VL-235B chain-of-thought traces. Curated SFT subset of MMFineReason-1.8M where smaller models consistently fail. Apache-2.0, parquet, image+text.
WorldCuisines VQA — Multilingual Multicultural Visual QA on Global Cuisines
Massive-scale multilingual, multicultural VQA benchmark covering global cuisines across 30 languages. NAACL 2025 Best Theme Paper. Images + text questions/answers for vision-language evaluation and fine-tuning.
TreeOfLife-200M: Biodiversity Image Dataset (214M images, 952K taxa)
214M biology/species images across 952,257 taxa from GBIF, EOL, BIOSCAN-5M, and FathomNet. CC0 licensed, parquet format, ready for CV/CLIP training and zero-shot taxonomic classification.
Nemotron VLM Dataset v2 (NVIDIA)
NVIDIA's large-scale vision-language training dataset (~9M samples) for VQA, image-text-to-text, video-text-to-text, and document understanding. CC-BY-4.0.
FineVision — 24M-Sample Vision-Language Training Corpus
Massive open VLM training set: 17.3M images, 24.3M samples, 88.9M turns, 9.5B answer tokens. Parquet, multi-subset, ready for fine-tuning state-of-the-art vision-language models.
ScaleEdit-12M: Large-Scale Instruction-Based Image Editing Dataset
12.4M verified instruction–image pairs across 23 task families for instruction-based image editing. Largest open-source dataset of its kind, built via the ScaleEditor multi-agent framework. MIT licensed.
Llama-Nemotron VLM Dataset v1 (NVIDIA)
NVIDIA's 1M+ sample multimodal dataset for vision-language model training, covering VQA, OCR, captioning, and image-to-text tasks. CC-BY-4.0 licensed.
BLIP3o Pretrain Long-Caption Dataset (27M Images)
27 million images paired with ~120-token long captions generated by Qwen2.5-VL-7B-Instruct. WebDataset format, Apache 2.0. Ideal for vision-language pretraining and multimodal model fine-tuning.
DataComp-1B: Image-Text Pair Metadata (1.4B Samples)
Metadata (URLs, captions, CLIP features) for ~1.4B image-text pairs from DataComp-1B, the curated subset of CommonPool used to train state-of-the-art CLIP models. CC-BY-4.0, parquet format.
Other categories