imagesLucasFang/FLUX-Reason-6Mtext-to-imagemultimodalreasoningfluxbilingualchineseenglishsynthetic-imagesfine-tuningapache-2.0

FLUX-Reason-6M Bilingual Text-to-Image Dataset

Free

Open dataset

Sample structure: 90.6 / 100
1 download links issued
Seller: DataBazaar
Sign up to download

Already have an account? Log in

Agent? Connect your account →

Category
Images
Records
5,890,279 rows
Format
PARQUET
Update Frequency
Not documented
Collection Method
auto_imported_huggingface_federated
PII
No flagged field names; not a privacy audit
File Size
~841500.11 MB
Download links issued
1

Source, license and coverage

Supplier documentation. These claims are separate from the automated sample score. A listing edit date is not a data freshness date.

License
apache-2.0
Source / creator
LucasFang/FLUX-Reason-6M
Collection method
Images were synthesized using the state-of-the-art FLUX.1-dev text-to-image model from carefully constructed reasoning-focused prompts. Each image is paired with multiple bilingual descriptions (English and Chinese) totaling ~20M captions across the 6M images. The dataset targets compositional, multi-step, and reasoning-heavy generation scenarios that open-source T2I models historically struggle with.
Coverage start
Not documented
Coverage end
Not documented
Data last updated
Not documented
Update schedule
Not documented

Source documentation ↗

License terms ↗

Embedded binary payloads are explicitly replaced with byte-length descriptors; the preview preserves accompanying text and metadata. All images are synthetic outputs of FLUX.1-dev and inherit any biases, artifacts, or failure modes of that model — they are not photographs or human-curated art. Bilingual coverage is limited to English and Chinese. The dataset emphasizes reasoning prompts which may not represent the full distribution of real-world T2I use cases. Source does not exhaustively document gaps; buyers should validate empirically for their target use case.

Sample structure score: 90.6 / 100

This automated check describes the inspected sample, not factual accuracy, legal rights, representativeness, or the quality of the entire dataset. It is not a customer rating.

Assessed 10 sample records (JSON) on 2026-10-09. All records in the provided sample were checked.

CheckPointsEvidence
Populated cells40.6 / 50284 of 350 top-level cells contain a value. Null, absent and blank values count as missing; zero and false count as populated.
Consistent value types30 / 30284 of 284 populated cells match their column's most common observed type. Types are inferred, not checked against real-world truth.
Consistent record shape20 / 2010 of 10 records have the expected fields. CSV/TSV use the header width; JSON uses the union of observed keys.
Field-level findings and improvements

Check missing cells and mixed types below. Document intentional missing values or mixed types in your field descriptions. Do not fill legitimate unknowns with invented values just to increase this score.

FieldMissing cellsMost common typeOther populated types
id0 / 10string0 / 10
image0 / 10object0 / 10
caption_composition5 / 10string0 / 5
caption_composition_cn5 / 10string0 / 5
bool_caption_composition0 / 10boolean0 / 10
score_composition0 / 10number0 / 10
caption_entity3 / 10string0 / 7
caption_entity_cn3 / 10string0 / 7
bool_caption_entity0 / 10boolean0 / 10
score_entity0 / 10number0 / 10
caption_text1 / 10string0 / 9
caption_text_cn1 / 10string0 / 9
bool_caption_text0 / 10boolean0 / 10
score_text0 / 10number0 / 10
caption_imaginative8 / 10string0 / 2
caption_imaginative_cn8 / 10string0 / 2
bool_caption_imaginative0 / 10boolean0 / 10
score_imaginative0 / 10number0 / 10
caption_style8 / 10string0 / 2
caption_style_cn8 / 10string0 / 2
bool_caption_style0 / 10boolean0 / 10
score_style0 / 10number0 / 10
caption_abstract8 / 10string0 / 2
caption_abstract_cn8 / 10string0 / 2
bool_caption_abstract0 / 10boolean0 / 10
score_abstract0 / 10number0 / 10
caption_original0 / 10string0 / 10
caption_original_cn0 / 10string0 / 10
bool_caption_original0 / 10boolean0 / 10
score_original0 / 10number0 / 10
caption_detail0 / 10string0 / 10
caption_detail_cn0 / 10string0 / 10
bool_caption_detail0 / 10boolean0 / 10
score_image_clarity0 / 10number0 / 10
score_image_structure0 / 10number0 / 10
How the score is calculated, its limitations, and how to correct an assessment →

About this data

Images synthesized by FLUX.1-dev paired with English and Chinese descriptions, designed to develop complex reasoning capabilities in text-to-image generative models.

Retrieve with your agent or Python

Create an account and configure DATABAZAAR_API_KEY. This example retrieves free or already purchased data; it never makes a purchase. For a multi-file dataset, choose a file index from the manifest.

Download the Python example
python3 retrieve-dataset.py ac2d04af-7b9e-46a6-8515-1d4dfc0905b8 --output dataset.bin
Full supplier documentation
## Overview FLUX-Reason-6M is a large-scale text-to-image dataset containing 6 million reasoning-focused images synthesized by the FLUX.1-dev model, paired with 20 million bilingual (English and Chinese) descriptions. The dataset is distributed as Parquet shards with image and text modalities, designed to bridge the performance gap between open-source and leading closed-source text-to-image systems by emphasizing complex compositional and reasoning prompts. ## Schema - image — binary/image — synthesized image generated by FLUX.1-dev - prompt_en — string — English prompt/description - prompt_zh — string — Chinese prompt/description - additional caption/metadata fields — string — extended descriptions and reasoning annotations - (See HF dataset page for full per-shard schema) ## Sources - HuggingFace: https://huggingface.co/datasets/LucasFang/FLUX-Reason-6M — Apache-2.0 - Associated paper: arXiv:2509.09680 ## Methodology Images were synthesized using the state-of-the-art FLUX.1-dev text-to-image model from carefully constructed reasoning-focused prompts. Each image is paired with multiple bilingual descriptions (English and Chinese) totaling ~20M captions across the 6M images. The dataset targets compositional, multi-step, and reasoning-heavy generation scenarios that open-source T2I models historically struggle with. ## Known gaps & limitations All images are synthetic outputs of FLUX.1-dev and inherit any biases, artifacts, or failure modes of that model — they are not photographs or human-curated art. Bilingual coverage is limited to English and Chinese. The dataset emphasizes reasoning prompts which may not represent the full distribution of real-world T2I use cases. Source does not exhaustively document gaps; buyers should validate empirically for their target use case. ## Intended use & out-of-scope - IS for: training and fine-tuning text-to-image generative models with reasoning capabilities, distillation from FLUX.1-dev, bilingual T2I research, prompt-image alignment studies. - NOT for: photorealism benchmarks against real-world images, training models intended for non-synthetic image evaluation, or use cases requiring languages beyond English/Chinese. _Federated dataset: 1,180 parquet shards, 821.78 GB total. Queries and downloads stream through the DataBazaar API._ Original supplier listing: FLUX-Reason-6M: 6M Reasoning-Focused Text-to-Image Dataset 6 million high-quality images synthesized by FLUX.1-dev with 20M bilingual (English/Chinese) descriptions, engineered to instill complex reasoning capabilities in text-to-image generative models.

Schema

NameTypeDescription
idVARCHARUnique identifier for the image record
imageSTRUCT(bytes BLOB, path VARCHAR)JPEG image binary data and file path generated by FLUX.1-dev
caption_compositionVARCHAREnglish description focusing on spatial arrangement and composition
caption_composition_cnVARCHARChinese description focusing on spatial arrangement and composition
bool_caption_compositionBOOLEANWhether composition caption is valid or present
score_compositionINTEGERQuality score for compositional accuracy (0-100 scale)
caption_entityVARCHAREnglish description of objects, entities, and their attributes
caption_entity_cnVARCHARChinese description of objects, entities, and their attributes
bool_caption_entityBOOLEANWhether entity caption is valid or present
score_entityINTEGERQuality score for entity identification accuracy (0-100 scale)
caption_textVARCHAREnglish description of any visible text or written content
caption_text_cnVARCHARChinese description of any visible text or written content
bool_caption_textBOOLEANWhether text caption is valid or present
score_textINTEGERQuality score for text recognition accuracy (0-100 scale)
caption_imaginativeVARCHAREnglish creative or imaginative interpretation of the image
caption_imaginative_cnVARCHARChinese creative or imaginative interpretation of the image
bool_caption_imaginativeBOOLEANWhether imaginative caption is valid or present
score_imaginativeINTEGERQuality score for creative description relevance (0-100 scale)
caption_styleVARCHAREnglish description of artistic style, medium, and visual aesthetics
caption_style_cnVARCHARChinese description of artistic style, medium, and visual aesthetics
bool_caption_styleBOOLEANWhether style caption is valid or present
score_styleINTEGERQuality score for style description accuracy (0-100 scale)
caption_abstractVARCHAREnglish abstract or conceptual interpretation of the image
caption_abstract_cnVARCHARChinese abstract or conceptual interpretation of the image
bool_caption_abstractBOOLEANWhether abstract caption is valid or present
score_abstractINTEGERQuality score for abstract reasoning relevance (0-100 scale)
caption_originalVARCHAREnglish original prompt used to generate the image
caption_original_cnVARCHARChinese original prompt used to generate the image
bool_caption_originalBOOLEANWhether original prompt is valid or present
score_originalINTEGERQuality score for prompt-image alignment (0-100 scale)
caption_detailVARCHAREnglish detailed description with fine-grained visual elements
caption_detail_cnVARCHARChinese detailed description with fine-grained visual elements
bool_caption_detailBOOLEANWhether detail caption is valid or present
score_image_clarityINTEGERImage clarity and definition quality score (0-100 scale)
score_image_structureINTEGERImage structural integrity and composition coherence score (0-100 scale)

Sample Data

Preview a sample of the data before downloading.

Public sample only. Sign in to retrieve the full dataset, including free datasets.

For AI Agents

Via MCP Server
# 1. Add to your agent's MCP config (claude_desktop_config.json or similar):
{
  "mcpServers": {
    "databazaar": { "command": "npx", "args": ["databazaar-mcp"] }
  }
}

# 2. Your agent can then call:
search_datasets({ query: "FLUX-Reason-6M Bilingual Text-" })
// Found: ac2d04af-7b9e-46a6-8515-1d4dfc0905b8
get_download_url({ dataset_id: "ac2d04af-7b9e-46a6-8515-1d4dfc0905b8" })  // free — sign in with MCP OAuth first
Via REST API
# Free dataset — sign in or use your account API key:
curl https://api.databazaar.io/datasets/ac2d04af-7b9e-46a6-8515-1d4dfc0905b8/download-url -H "Authorization: Bearer $DATABAZAAR_API_KEY"