textOpenAssistant/oasst1rlhfinstruction-tuningalignmentmultilingualhuman-feedbackconversationsfine-tuningopen-assistantparquet

OpenAssistant Conversations (OASST1) Multilingual

Free

Open dataset

Sample structure: 95.8 / 100
2 download links issued
Seller: DataBazaar
Sign up to download

Already have an account? Log in

Agent? Connect your account →

Category
Text
Records
88,838 rows
Format
PARQUET
Update Frequency
Not documented
Collection Method
auto_imported_huggingface_federated
PII
No flagged field names; not a privacy audit
File Size
~39.67 MB
Download links issued
2

Source, license and coverage

Supplier documentation. These claims are separate from the automated sample score. A listing edit date is not a data freshness date.

License
apache-2.0
Source / creator
OpenAssistant/oasst1
Collection method
Collected via a worldwide crowd-sourcing effort on the open-assistant.io platform. Volunteers wrote prompts, wrote assistant replies, and rated/ranked other contributors' messages along multiple labeled dimensions. Conversation trees were grown by branching: multiple assistant replies per prompt, ranked by reviewers. Messages flagged by reviewers or detoxify thresholds were filtered, and the released split contains only trees in a 'ready_for_export' state. Train/validation splits are provided by the publisher.
Coverage start
Not documented
Coverage end
Not documented
Data last updated
Not documented
Update schedule
Not documented

Source documentation ↗

License terms ↗

Language distribution is heavily skewed toward English, Spanish, Russian, and German; many of the 35 languages have only a small number of messages. Crowd-sourced contributors are self-selected and not demographically representative. Quality annotations are subjective and reviewer counts per message vary, so per-label confidence is uneven. The dataset captures the state of contributions as of April 2023 and is not updated; OASST2 supersedes it for newer work. Some safety-related content remains in the corpus by design for alignment research — buyers should filter for their use case.

Sample structure score: 95.8 / 100

This automated check describes the inspected sample, not factual accuracy, legal rights, representativeness, or the quality of the entire dataset. It is not a customer rating.

Assessed 10 sample records (JSON) on 2026-10-09. All records in the provided sample were checked.

CheckPointsEvidence
Populated cells45.8 / 50165 of 180 top-level cells contain a value. Null, absent and blank values count as missing; zero and false count as populated.
Consistent value types30 / 30165 of 165 populated cells match their column's most common observed type. Types are inferred, not checked against real-world truth.
Consistent record shape20 / 2010 of 10 records have the expected fields. CSV/TSV use the header width; JSON uses the union of observed keys.
Field-level findings and improvements

Check missing cells and mixed types below. Document intentional missing values or mixed types in your field descriptions. Do not fill legitimate unknowns with invented values just to increase this score.

FieldMissing cellsMost common typeOther populated types
message_id0 / 10string0 / 10
parent_id2 / 10string0 / 8
user_id0 / 10string0 / 10
created_date0 / 10string0 / 10
text0 / 10string0 / 10
role0 / 10string0 / 10
lang0 / 10string0 / 10
review_count0 / 10number0 / 10
review_result0 / 10boolean0 / 10
deleted0 / 10boolean0 / 10
rank3 / 10number0 / 7
synthetic0 / 10boolean0 / 10
model_name10 / 10unknown0 / 0
detoxify0 / 10object0 / 10
message_tree_id0 / 10string0 / 10
tree_state0 / 10string0 / 10
emojis0 / 10object0 / 10
labels0 / 10object0 / 10
How the score is calculated, its limitations, and how to correct an assessment →

About this data

Human-generated assistant conversation messages across 35 languages with quality ratings. Designed for alignment, RLHF, and instruction-tuning research.

Retrieve with your agent or Python

Create an account and configure DATABAZAAR_API_KEY. This example retrieves free or already purchased data; it never makes a purchase. For a multi-file dataset, choose a file index from the manifest.

Download the Python example
python3 retrieve-dataset.py 4bde98a7-6d52-401b-92b8-1f528c2b2da7 --output dataset.bin
Full supplier documentation
## Overview OpenAssistant Conversations (OASST1) is a human-generated, human-annotated assistant-style conversation corpus containing 161,443 messages organized into over 10,000 fully annotated conversation trees, spanning 35 languages. Each message carries quality ratings (461,292 total annotations) covering attributes like helpfulness, creativity, and toxicity. Released April 2023 as parquet, the dataset is widely used for instruction tuning, reward modeling, and RLHF research. ## Schema - message_id — string — unique message identifier - parent_id — string — parent message in the conversation tree (null for root prompts) - text — string — the message content - role — string — 'prompter' or 'assistant' - lang — string — ISO language code - review_count — int — number of human reviews - review_result — bool — whether message passed review - deleted — bool — moderation flag - rank — int — sibling ranking among assistant replies - labels — struct — quality/safety annotations (quality, toxicity, humor, creativity, etc.) with values and counts - tree_state — string — state of the conversation tree (ready_for_export, etc.) - +several metadata columns (emojis, synthetic, model_name, detoxify scores) ## Sources - OpenAssistant/oasst1 on Hugging Face — https://huggingface.co/datasets/OpenAssistant/oasst1 — Apache-2.0 - Paper: Köpf et al., "OpenAssistant Conversations" (arXiv:2304.07327) ## Methodology Collected via a worldwide crowd-sourcing effort on the open-assistant.io platform. Volunteers wrote prompts, wrote assistant replies, and rated/ranked other contributors' messages along multiple labeled dimensions. Conversation trees were grown by branching: multiple assistant replies per prompt, ranked by reviewers. Messages flagged by reviewers or detoxify thresholds were filtered, and the released split contains only trees in a 'ready_for_export' state. Train/validation splits are provided by the publisher. ## Known gaps & limitations Language distribution is heavily skewed toward English, Spanish, Russian, and German; many of the 35 languages have only a small number of messages. Crowd-sourced contributors are self-selected and not demographically representative. Quality annotations are subjective and reviewer counts per message vary, so per-label confidence is uneven. The dataset captures the state of contributions as of April 2023 and is not updated; OASST2 supersedes it for newer work. Some safety-related content remains in the corpus by design for alignment research — buyers should filter for their use case. ## Intended use & out-of-scope - For: instruction tuning, RLHF/DPO reward model training, multilingual assistant fine-tuning, alignment research, conversation-tree evaluation. - Not for: production safety guarantees without additional filtering; benchmarks where leakage into popular open-source instruct models (Llama-2-Chat derivatives, many community models) is a concern — assume contamination across most current open chat LLMs. _Federated dataset: 2 parquet shards, 39.7 MB total. Queries and downloads stream through the DataBazaar API._ Original supplier listing: OpenAssistant Conversations (OASST1) 161,443 human-generated assistant conversation messages across 35 languages with 461,292 quality ratings — a foundational dataset for alignment, RLHF, and instruction-tuning research.

Schema

NameTypeDescription
message_idVARCHAR
parent_idVARCHAR
user_idVARCHAR
created_dateVARCHAR
textVARCHAR
roleVARCHAR
langVARCHAR
review_countINTEGER
review_resultBOOLEAN
deletedBOOLEAN
rankINTEGER
syntheticBOOLEAN
model_nameVARCHAR
detoxifySTRUCT(toxicity DOUBLE, severe_toxicity DOUBLE, obscene DOUBLE, identity_attack DOUBLE, insult DOUBLE, threat DOUBLE, sexual_explicit DOUBLE)
message_tree_idVARCHAR
tree_stateVARCHAR
emojisSTRUCT("name" VARCHAR[], count INTEGER[])
labelsSTRUCT("name" VARCHAR[], "value" DOUBLE[], count INTEGER[])

Sample Data

Preview a sample of the data before downloading.

Public sample only. Sign in to retrieve the full dataset, including free datasets.

For AI Agents

Via MCP Server
# 1. Add to your agent's MCP config (claude_desktop_config.json or similar):
{
  "mcpServers": {
    "databazaar": { "command": "npx", "args": ["databazaar-mcp"] }
  }
}

# 2. Your agent can then call:
search_datasets({ query: "OpenAssistant Conversations (O" })
// Found: 4bde98a7-6d52-401b-92b8-1f528c2b2da7
get_download_url({ dataset_id: "4bde98a7-6d52-401b-92b8-1f528c2b2da7" })  // free — sign in with MCP OAuth first
Via REST API
# Free dataset — sign in or use your account API key:
curl https://api.databazaar.io/datasets/4bde98a7-6d52-401b-92b8-1f528c2b2da7/download-url -H "Authorization: Bearer $DATABAZAAR_API_KEY"