FLUX-Reason-6M Bilingual Text-to-Image Dataset
Source, license and coverage
Supplier documentation. These claims are separate from the automated sample score. A listing edit date is not a data freshness date.
- License
- apache-2.0
- Source / creator
- LucasFang/FLUX-Reason-6M
- Collection method
- Images were synthesized using the state-of-the-art FLUX.1-dev text-to-image model from carefully constructed reasoning-focused prompts. Each image is paired with multiple bilingual descriptions (English and Chinese) totaling ~20M captions across the 6M images. The dataset targets compositional, multi-step, and reasoning-heavy generation scenarios that open-source T2I models historically struggle with.
- Coverage start
- Not documented
- Coverage end
- Not documented
- Data last updated
- Not documented
- Update schedule
- Not documented
Embedded binary payloads are explicitly replaced with byte-length descriptors; the preview preserves accompanying text and metadata. All images are synthetic outputs of FLUX.1-dev and inherit any biases, artifacts, or failure modes of that model — they are not photographs or human-curated art. Bilingual coverage is limited to English and Chinese. The dataset emphasizes reasoning prompts which may not represent the full distribution of real-world T2I use cases. Source does not exhaustively document gaps; buyers should validate empirically for their target use case.
Sample structure score: 90.6 / 100
This automated check describes the inspected sample, not factual accuracy, legal rights, representativeness, or the quality of the entire dataset. It is not a customer rating.
Assessed 10 sample records (JSON) on 2026-10-09. All records in the provided sample were checked.
| Check | Points | Evidence |
|---|---|---|
| Populated cells | 40.6 / 50 | 284 of 350 top-level cells contain a value. Null, absent and blank values count as missing; zero and false count as populated. |
| Consistent value types | 30 / 30 | 284 of 284 populated cells match their column's most common observed type. Types are inferred, not checked against real-world truth. |
| Consistent record shape | 20 / 20 | 10 of 10 records have the expected fields. CSV/TSV use the header width; JSON uses the union of observed keys. |
Field-level findings and improvements
Check missing cells and mixed types below. Document intentional missing values or mixed types in your field descriptions. Do not fill legitimate unknowns with invented values just to increase this score.
| Field | Missing cells | Most common type | Other populated types |
|---|---|---|---|
| id | 0 / 10 | string | 0 / 10 |
| image | 0 / 10 | object | 0 / 10 |
| caption_composition | 5 / 10 | string | 0 / 5 |
| caption_composition_cn | 5 / 10 | string | 0 / 5 |
| bool_caption_composition | 0 / 10 | boolean | 0 / 10 |
| score_composition | 0 / 10 | number | 0 / 10 |
| caption_entity | 3 / 10 | string | 0 / 7 |
| caption_entity_cn | 3 / 10 | string | 0 / 7 |
| bool_caption_entity | 0 / 10 | boolean | 0 / 10 |
| score_entity | 0 / 10 | number | 0 / 10 |
| caption_text | 1 / 10 | string | 0 / 9 |
| caption_text_cn | 1 / 10 | string | 0 / 9 |
| bool_caption_text | 0 / 10 | boolean | 0 / 10 |
| score_text | 0 / 10 | number | 0 / 10 |
| caption_imaginative | 8 / 10 | string | 0 / 2 |
| caption_imaginative_cn | 8 / 10 | string | 0 / 2 |
| bool_caption_imaginative | 0 / 10 | boolean | 0 / 10 |
| score_imaginative | 0 / 10 | number | 0 / 10 |
| caption_style | 8 / 10 | string | 0 / 2 |
| caption_style_cn | 8 / 10 | string | 0 / 2 |
| bool_caption_style | 0 / 10 | boolean | 0 / 10 |
| score_style | 0 / 10 | number | 0 / 10 |
| caption_abstract | 8 / 10 | string | 0 / 2 |
| caption_abstract_cn | 8 / 10 | string | 0 / 2 |
| bool_caption_abstract | 0 / 10 | boolean | 0 / 10 |
| score_abstract | 0 / 10 | number | 0 / 10 |
| caption_original | 0 / 10 | string | 0 / 10 |
| caption_original_cn | 0 / 10 | string | 0 / 10 |
| bool_caption_original | 0 / 10 | boolean | 0 / 10 |
| score_original | 0 / 10 | number | 0 / 10 |
| caption_detail | 0 / 10 | string | 0 / 10 |
| caption_detail_cn | 0 / 10 | string | 0 / 10 |
| bool_caption_detail | 0 / 10 | boolean | 0 / 10 |
| score_image_clarity | 0 / 10 | number | 0 / 10 |
| score_image_structure | 0 / 10 | number | 0 / 10 |
About this data
Images synthesized by FLUX.1-dev paired with English and Chinese descriptions, designed to develop complex reasoning capabilities in text-to-image generative models.
Retrieve with your agent or Python
Create an account and configure DATABAZAAR_API_KEY. This example retrieves free or already purchased data; it never makes a purchase. For a multi-file dataset, choose a file index from the manifest.
Download the Python examplepython3 retrieve-dataset.py ac2d04af-7b9e-46a6-8515-1d4dfc0905b8 --output dataset.bin
Full supplier documentation
Schema
| Name | Type | Description |
|---|---|---|
| id | VARCHAR | Unique identifier for the image record |
| image | STRUCT(bytes BLOB, path VARCHAR) | JPEG image binary data and file path generated by FLUX.1-dev |
| caption_composition | VARCHAR | English description focusing on spatial arrangement and composition |
| caption_composition_cn | VARCHAR | Chinese description focusing on spatial arrangement and composition |
| bool_caption_composition | BOOLEAN | Whether composition caption is valid or present |
| score_composition | INTEGER | Quality score for compositional accuracy (0-100 scale) |
| caption_entity | VARCHAR | English description of objects, entities, and their attributes |
| caption_entity_cn | VARCHAR | Chinese description of objects, entities, and their attributes |
| bool_caption_entity | BOOLEAN | Whether entity caption is valid or present |
| score_entity | INTEGER | Quality score for entity identification accuracy (0-100 scale) |
| caption_text | VARCHAR | English description of any visible text or written content |
| caption_text_cn | VARCHAR | Chinese description of any visible text or written content |
| bool_caption_text | BOOLEAN | Whether text caption is valid or present |
| score_text | INTEGER | Quality score for text recognition accuracy (0-100 scale) |
| caption_imaginative | VARCHAR | English creative or imaginative interpretation of the image |
| caption_imaginative_cn | VARCHAR | Chinese creative or imaginative interpretation of the image |
| bool_caption_imaginative | BOOLEAN | Whether imaginative caption is valid or present |
| score_imaginative | INTEGER | Quality score for creative description relevance (0-100 scale) |
| caption_style | VARCHAR | English description of artistic style, medium, and visual aesthetics |
| caption_style_cn | VARCHAR | Chinese description of artistic style, medium, and visual aesthetics |
| bool_caption_style | BOOLEAN | Whether style caption is valid or present |
| score_style | INTEGER | Quality score for style description accuracy (0-100 scale) |
| caption_abstract | VARCHAR | English abstract or conceptual interpretation of the image |
| caption_abstract_cn | VARCHAR | Chinese abstract or conceptual interpretation of the image |
| bool_caption_abstract | BOOLEAN | Whether abstract caption is valid or present |
| score_abstract | INTEGER | Quality score for abstract reasoning relevance (0-100 scale) |
| caption_original | VARCHAR | English original prompt used to generate the image |
| caption_original_cn | VARCHAR | Chinese original prompt used to generate the image |
| bool_caption_original | BOOLEAN | Whether original prompt is valid or present |
| score_original | INTEGER | Quality score for prompt-image alignment (0-100 scale) |
| caption_detail | VARCHAR | English detailed description with fine-grained visual elements |
| caption_detail_cn | VARCHAR | Chinese detailed description with fine-grained visual elements |
| bool_caption_detail | BOOLEAN | Whether detail caption is valid or present |
| score_image_clarity | INTEGER | Image clarity and definition quality score (0-100 scale) |
| score_image_structure | INTEGER | Image structural integrity and composition coherence score (0-100 scale) |
Sample Data
Preview a sample of the data before downloading.
Public sample only. Sign in to retrieve the full dataset, including free datasets.
For AI Agents
# 1. Add to your agent's MCP config (claude_desktop_config.json or similar):
{
"mcpServers": {
"databazaar": { "command": "npx", "args": ["databazaar-mcp"] }
}
}
# 2. Your agent can then call:
search_datasets({ query: "FLUX-Reason-6M Bilingual Text-" })
// Found: ac2d04af-7b9e-46a6-8515-1d4dfc0905b8
get_download_url({ dataset_id: "ac2d04af-7b9e-46a6-8515-1d4dfc0905b8" }) // free — sign in with MCP OAuth first# Free dataset — sign in or use your account API key: curl https://api.databazaar.io/datasets/ac2d04af-7b9e-46a6-8515-1d4dfc0905b8/download-url -H "Authorization: Bearer $DATABAZAAR_API_KEY"