textMizzenAI/HPDv3human-preferencetext-to-imagerlhfreward-modeldiffusioniccv-2025preference-learningt2imit-license

HPDv3 Human Preference Dataset

Free

Open dataset

Sample structure: 85.7 / 100
3 download links issued
Seller: DataBazaar
Sign up to download

Already have an account? Log in

Agent? Connect your account →

Category
Text
Records
1,154,324 rows
Format
PARQUET
Update Frequency
Not documented
Collection Method
auto_imported_huggingface_federated
PII
No flagged field names; not a privacy audit
File Size
~248.36 MB
Download links issued
3

Source, license and coverage

Supplier documentation. These claims are separate from the automated sample score. A listing edit date is not a data freshness date.

License
mit
Source / creator
MizzenAI/HPDv3
Collection method
The dataset aggregates text-image pairs from multiple text-to-image generators across a wide spectrum of prompt domains and quality levels. Pairwise preference annotations were collected from human annotators comparing two generations for the same or related prompts, with the goal of training reward/preference models (HPSv3) that generalize across styles, qualities, and prompt types. The authors emphasize 'wide-spectrum' coverage to reduce bias toward any single generator family.
Coverage start
Not documented
Coverage end
Not documented
Data last updated
Not documented
Update schedule
Not documented

Source documentation ↗

License terms ↗

Language coverage is English-only. Human preference data inherits annotator demographic and cultural biases that the source does not fully document. Image generator coverage is a snapshot in time and will become stale as new models are released. Source does not document detailed annotator agreement statistics; buyers training reward models should validate calibration empirically and check for leakage against common T2I eval benchmarks.

Sample structure score: 85.7 / 100

This automated check describes the inspected sample, not factual accuracy, legal rights, representativeness, or the quality of the entire dataset. It is not a customer rating.

Assessed 10 sample records (JSON) on 2026-10-09. All records in the provided sample were checked.

CheckPointsEvidence
Populated cells35.7 / 5050 of 70 top-level cells contain a value. Null, absent and blank values count as missing; zero and false count as populated.
Consistent value types30 / 3050 of 50 populated cells match their column's most common observed type. Types are inferred, not checked against real-world truth.
Consistent record shape20 / 2010 of 10 records have the expected fields. CSV/TSV use the header width; JSON uses the union of observed keys.
Field-level findings and improvements

Check missing cells and mixed types below. Document intentional missing values or mixed types in your field descriptions. Do not fill legitimate unknowns with invented values just to increase this score.

FieldMissing cellsMost common typeOther populated types
prompt0 / 10string0 / 10
choice_dist10 / 10unknown0 / 0
confidence10 / 10unknown0 / 0
path10 / 10string0 / 10
path20 / 10string0 / 10
model10 / 10string0 / 10
model20 / 10string0 / 10
How the score is calculated, its limitations, and how to correct an assessment →

About this data

Text-to-image preference dataset comprising 1.08M text-image pairs and 1.17M pairwise human annotations. Used to train HPSv3 reward models for evaluating generated images.

Retrieve with your agent or Python

Create an account and configure DATABAZAAR_API_KEY. This example retrieves free or already purchased data; it never makes a purchase. For a multi-file dataset, choose a file index from the manifest.

Download the Python example
python3 retrieve-dataset.py 24552bae-e4d9-4097-9f81-f7ad484e8b78 --output dataset.bin
Full supplier documentation
## Overview Human Preference Dataset v3 (HPDv3) is a large-scale dataset for training and evaluating human preference models over text-to-image generations. It contains 1.08M text-image pairs and 1.17M annotated pairwise preference comparisons, distributed in JSON format. Released alongside the HPSv3 paper (ICCV 2025) by Mizzen AI and CUHK MMLab, it targets wide-spectrum preference modeling across image quality, prompt alignment, and aesthetics. ## Schema - prompt — string — text prompt used to generate the image - image — string/path — reference to the generated image - model — string — generator model that produced the image - pair_id — string — identifier linking the two images in a pairwise comparison - preferred — int — which image in the pair was preferred by annotators - annotator_id — string — anonymized annotator identifier - score — float — preference score / rating (where available) - +additional metadata columns for source, split, and annotation provenance ## Sources - HuggingFace: https://huggingface.co/datasets/MizzenAI/HPDv3 — MIT license - Paper: arXiv:2508.03789 (HPSv3: Towards Wide-Spectrum Human Preference Score, ICCV 2025) ## Methodology The dataset aggregates text-image pairs from multiple text-to-image generators across a wide spectrum of prompt domains and quality levels. Pairwise preference annotations were collected from human annotators comparing two generations for the same or related prompts, with the goal of training reward/preference models (HPSv3) that generalize across styles, qualities, and prompt types. The authors emphasize 'wide-spectrum' coverage to reduce bias toward any single generator family. ## Known gaps & limitations Language coverage is English-only. Human preference data inherits annotator demographic and cultural biases that the source does not fully document. Image generator coverage is a snapshot in time and will become stale as new models are released. Source does not document detailed annotator agreement statistics; buyers training reward models should validate calibration empirically and check for leakage against common T2I eval benchmarks. ## Intended use & out-of-scope - IS for: training human preference / reward models for text-to-image generation, RLHF/DPO fine-tuning of diffusion models, building T2I evaluation benchmarks, and research on preference modeling. - NOT for: deployment as a sole arbiter of image quality without bias auditing; not deduplicated against external T2I eval suites — leakage risk if used for benchmark training. _Federated dataset: 2 parquet shards, 248.4 MB total. Queries and downloads stream through the DataBazaar API._ Original supplier listing: HPDv3 — Human Preference Dataset v3 (1.08M text-image pairs) Wide-spectrum human preference dataset for text-to-image generation: 1.08M text-image pairs and 1.17M pairwise human preference annotations. MIT-licensed, used to train HPSv3 (ICCV 2025) reward models.

Schema

NameTypeDescription
promptVARCHARText description used to generate the paired images.
choice_distBIGINT[]Array of annotator preference counts for each image in the pair (null if unavailable).
confidenceDOUBLEAnnotator confidence score for the preference judgment (0–1 scale, null if unavailable).
path1VARCHARFile path to the first generated image in the comparison pair.
path2VARCHARFile path to the second generated image in the comparison pair.
model1VARCHARName of the text-to-image model that generated the first image.
model2VARCHARName of the text-to-image model that generated the second image.

Sample Data

Preview a sample of the data before downloading.

Public sample only. Sign in to retrieve the full dataset, including free datasets.

For AI Agents

Via MCP Server
# 1. Add to your agent's MCP config (claude_desktop_config.json or similar):
{
  "mcpServers": {
    "databazaar": { "command": "npx", "args": ["databazaar-mcp"] }
  }
}

# 2. Your agent can then call:
search_datasets({ query: "HPDv3 Human Preference Dataset" })
// Found: 24552bae-e4d9-4097-9f81-f7ad484e8b78
get_download_url({ dataset_id: "24552bae-e4d9-4097-9f81-f7ad484e8b78" })  // free — sign in with MCP OAuth first
Via REST API
# Free dataset — sign in or use your account API key:
curl https://api.databazaar.io/datasets/24552bae-e4d9-4097-9f81-f7ad484e8b78/download-url -H "Authorization: Bearer $DATABAZAAR_API_KEY"