Glossary

Data dictionary

A data dictionary says what every table and column of a dataset means. This entry lists what a useful one states and how to test it against a sample.

Updated 5 October 20262 min read

A data dictionary describes every table and column of a dataset: its name, meaning, type, allowed values and unit. It lets an engineer or analyst load and read the data correctly without asking the producer what a field means.

What a useful one states

  • The grain of each table: what one row stands for, such as one company on one closed day.
  • The key: the columns that identify a row, and the identifiers that join it to other tables.
  • For each column, its meaning, type, unit and allowed values.
  • What an empty value means: not applicable, not found or not confident.
  • The time convention: the time zone, and whether a date is when something happened or when it was recorded.
  • A version and a date for the document, and how changes are announced.

Reading one before you load

Start with the grain, because it gives you the key to test. In the Fokals dictionary, Hiring Activity is stated per company and closed UTC day, Sales Team Metrics per company and closed week, and Market Series has one row per metric, dimension, window and as-of date. Load a sample and confirm the statement holds:

select company_id, day, count(*)
from company_hiring_daily
group by company_id, day
having count(*) > 1;

An empty result means no company and day appears twice, which is what the stated grain requires. Then read the conventions that change how you load: times are UTC, lists and objects such as the breakdown of postings by job function arrive as JSON in a single cell of a CSV file, and an empty value has a stated meaning.

Dictionary, methodology and manifest

Three documents answer different questions. The dictionary says what each field is, the methodology says how its values are produced, and a data manifest says what one delivery contained. A dictionary does not prove that the files you receive follow it, so test a sample for key uniqueness, allowed values and the rate of empty cells before you build on it.

In Fokals data

The Fokals data dictionary covers the tables of all five datasets, the conventions, the identifiers on every file and the lists of allowed values, such as the job function labels. Label versions are frozen by name, and a breaking change ships as a new version with at least 90 days' notice, so a change that would break your load is announced well before it takes effect.

  • Data manifest: describes one delivery rather than the structure of a dataset.
  • Label version: the frozen name under which labels are produced.
  • Data provenance: the source and observation time behind each record.
  • Data freshness: how recent the data is, which a dictionary states as a refresh rate.

Frequently asked questions

What is the difference between a data dictionary and a schema?

A schema is the structure a database enforces: table names, column names, types and constraints. A data dictionary includes that and adds meaning for people: what a column holds, its unit, its allowed values, what an empty cell means and how often it is refreshed. You can load data with a schema alone, but you cannot interpret it correctly without the dictionary.

What should I check in a vendor's data dictionary before buying?

Check that every table states its grain, that each column has a meaning, a type and a unit, and that fields with a fixed set of values list them. Look for the rule on empty values, the time zone, a version number and a statement of how changes are announced. Then compare the dictionary with a sample file: keys should be unique and values should fall inside the lists.

Is a data dictionary the same as a data catalogue?

No. A data dictionary describes the fields of one dataset. A data catalogue is an inventory of the datasets across an organisation, with owners, lineage and search, and it often links to dictionaries. A vendor supplies a dictionary; a buyer's catalogue records which datasets it has licensed and where they are loaded.

The queries and code on this page are examples to adapt. Test them in your own environment before you rely on them.