Write a Datasheet for a Public Dataset
A practical datasheet template for public tables: document scope, sources, transformations, missing values, reuse terms and release versions.

Before sharing a spreadsheet, write a short datasheet that explains what its rows represent, where the numbers came from and what readers should not infer. Publish that note alongside the interactive table so someone receiving only the link can interpret the data without asking you for context.
For a public dataset, the useful result is a table plus a description of its scope, history and limits—not just a list of column names.
What “datasheet” means here
A technical product datasheet describes a component’s specifications. A dataset datasheet documents the data itself. The research proposal Datasheets for Datasets, developed for the machine-learning community, applies that analogy to dataset motivation, composition, collection and recommended uses, with the aim of improving communication between creators and consumers. (Original paper)
For a public CSV or spreadsheet, adapt that idea to the questions your readers need answered. The template below is a lightweight publishing aid, not the paper’s full questionnaire or a certification.
A data dictionary defines individual fields; a datasheet explains the dataset as a whole, especially whether it is suitable for a reader’s intended use. Keep detailed field definitions in a dataset reference sheet if they would make the datasheet difficult to scan.
A reusable datasheet template
Use these seven sections in a plain-text document or on the web page that introduces your table.
| Section | What to record |
|---|---|
| Identity and purpose | Dataset title, publisher, release identifier and the question it is intended to help answer. |
| Scope and composition | What one row represents; time period, geography and population covered; inclusion and exclusion rules. |
| Sources and collection | Original producer, source links, retrieval dates and collection method. Distinguish the source producer from yourself as publisher. |
| Preparation | Filters, joins, deduplication, calculations, unit conversions and corrections made before publication. |
| Interpretation and quality | Units, missing-value meanings, known gaps, uncertainty, appropriate uses and comparisons the data cannot support. |
| Access and reuse | Link to the public table and source files where available; applicable license or reuse terms; requested attribution. |
| Maintenance | Responsible contact, update schedule or “no updates planned,” and changes since the previous release. |
These sections reflect established web-publishing guidance: W3C recommends providing license information, provenance, quality information, a version indicator and version history. Provenance means the data’s origins and the changes made to it. (W3C Data on the Web Best Practices)
Write specific answers rather than assurances such as “clean data” or “updated regularly.” If the collection method is unknown, say so. If reuse terms are unclear, resolve them before republishing; any license you apply must respect the source’s requirements. (W3C licensing guidance)
Worked example: a small library-visits table
This is an invented demonstration dataset, not a report about real libraries:
library_id,month,visits
LIB-A,2026-04,120
LIB-B,2026-04,0
LIB-C,2026-04,
A concise accompanying datasheet could read:
Dataset: Library visits demonstration, release 1.0.
Purpose and scope: Illustrate publication of monthly count data. Each row represents one fictional library in April 2026. The table contains three libraries, not a complete library system.
Source and preparation: Values were created for this example. No external records, joins or calculations were used.
Fields and missing values:
library_idis a fictional identifier;monthusesYYYY-MM;visitsis a whole-number count. Zero means zero visits. An empty value means the count is unavailable, not zero.Appropriate use: Demonstrating table structure and missing-value handling. Not suitable for evaluating library performance or estimating real attendance.
Maintenance: Static example; no updates planned.
For a real release, add the actual source links, retrieval dates, reuse terms and contact. Notice how the note prevents two misleading readings: treating the blank as zero and treating three rows as complete coverage.
Publish the data and its explanation together
Keep the observation table separate from the datasheet. Do not place prose paragraphs above the CSV header or mix documentation rows into the records.
As of October 5, 2026, TablePage advertises CSV, TSV, XLSX and XLS uploads that generate public dataset pages with filterable tables. Upload only a reviewed, public-safe file—not sensitive information or an unreviewed workbook. Check the datasheet for sensitive information too, including internal notes and contact details not cleared for publication.
Place the datasheet on the page introducing or embedding the table, or in an accessible companion document linked beside the table link. Then check:
- Coverage: Does the note describe exactly the rows in the published file?
- Interpretation: Can a reader distinguish zero, missing and not applicable?
- Traceability: Do source links and transformation notes explain how the published values were produced?
- Release match: Do the file and datasheet carry the same release identifier?
When values or definitions change, update the datasheet with the data and record the change. A shareable table should not retain an explanation of an earlier release.