How we assess dataset quality
Our sample structure score measures three observable properties of a provided sample. It is an automated technical check, not a customer rating or a certification of the full dataset. Version: sample-structure-v1.
The exact formula
Score = 50 × populated-cell fraction + 30 × consistent-type fraction + 20 × consistent-record fraction. Each fraction is between 0 and 1. The final score is rounded to one decimal place out of 100; component rounding can produce a 0.1-point display difference.
| Check | Weight | Rule |
|---|---|---|
| Populated cells | 50 | Non-empty cells divided by rows × columns. Null, absent and whitespace-only cells are missing. Zero and false are populated. |
| Consistent value types | 30 | Populated cells matching the most common observed type in their column, divided by all populated cells. With no populated cells, this component is zero. |
| Consistent record shape | 20 | Records with the expected fields divided by inspected records. CSV/TSV use the header width; JSON/JSONL use the union of observed keys. |
The weights are a published product rubric, not an empirically validated prediction of usefulness. Completeness receives the largest weight because missing cells often require extra preparation. Type consistency and record shape measure two other common parsing concerns. None can establish real-world accuracy.
Example: if all records have consistent fields and types but 10% of cells are missing, the score is 45 + 30 + 20 = 95. A complete ten-row sample can score as highly as a complete thousand-row sample. Size, price, downloads, seller identity and popularity never add points.
What we inspect
We inspect the supplied sample, up to 2 MiB, 1,000 records and 256 columns. Samples are not guaranteed random or representative. The assessment records its time, format, inspected row count, whether the row cap was reached, and a SHA-256 fingerprint of the UTF-8 decoded sample text. For an extensionless sample with no declared format, we recognize JSON, TSV or CSV from its content before parsing. The sample format is detected separately from the full dataset format; a Parquet dataset may have a JSON sample.
CSV and TSV support quoted delimiters, escaped quotes and multiline fields. JSON arrays and JSONL records are supported. For JSON wrapper objects, we use a recognized records array (data, records, items, rows, results, entries, listings), otherwise the largest array of objects, otherwise the object itself. Scalar entries become a value column. Checks apply to top-level cells; nested objects and arrays are not recursively validated.
Types are inferred as number, boolean, string, array or object. Numeric-looking strings and true/false strings use their corresponding types; leading-zero identifiers stay strings. Tied types use alphabetical order for reproducibility. We do not verify declared schema types, validate dates, or penalize duplicate records.
What a low score means
A low score means the sample contains missing cells, mixed observed types, or differing record shapes under this rubric. Those properties may be intentional. Sparse scientific measurements, heterogeneous event logs and optional JSON fields can be useful despite a lower score. Read field descriptions and inspect the sample before comparing datasets.
A score of 100 does not prove authenticity, factual correctness, suitability for training, absence of personal information, freshness, representativeness, or permission to reuse the data. License and provenance documentation are shown separately as supplier declarations, not certified by the score.
Unsupported formats, oversized samples, missing files, malformed samples and empty samples show Not assessed with a reason, rather than an invented zero. Media and archive datasets are not treated as low quality because this structural rubric is inapplicable.
Improve or challenge your assessment
Open your listing's field-level findings to see the exact deductions. Correct accidental empty cells, malformed records and unintended type changes. Explain intentional missing values and mixed types in field descriptions. Never invent values or remove valid records just to improve a score.
Use your listing editor to update source, license, coverage and collection documentation. Use its Reassess sample button to retry a failed check or evaluate your current sample. A reassessment uses the same versioned rubric and does not change publication or moderation status.
If the parser or rubric mischaracterizes your data, send feedback with the listing URL, assessment version and the field or record pattern involved. Do not include private dataset contents. We can correct the implementation and reassess; the score cannot be purchased or manually boosted.
Version history
Version 1 removes the former row-count bonus and replaces the ambiguous five-point display with this sample-based 100-point rubric. Old scores are not comparable and are withheld until a current assessment is available. Future formula changes will receive a new version and a catalog reassessment.