From Integrated to Analysis-Ready

Technical integration makes data accessible. Analytical validation determines whether the data accurately represent the intended population, activity, time period, and outcome, and whether they are fit for a specific analysis.

IMPORTANT DISTINCTION: Validation does not question the value of technical integration. It completes the work required for credible analysis.

The Validation Workflow

A dataset becomes analysis-ready through a sequence of questions, tests, and decisions. Use the workflow below to ask the right questions, test what matters, identify potential problems, and document whether the data are fit for the intended use.

What are we trying to answer, and what would the data need to represent?

Document the question, decision, population, unit of analysis, time window, outcomes, and decisions the evidence may inform.

Validate:

  • Purpose and intended decision are documented.
  • Population and unit of analysis are explicit.
  • Time window and outcomes are defined.

Watch for: A validated dataset is reused for a new purpose. Fitness is question-specific. Reassess population, fields, time logic, and limitations whenever the intended use changes.

Do we understand where the data came from and what they mean?

Identify system purpose, source owner, collection process, authoritative fields, known changes, and refresh timing.

Validate:

  • Ownership is confirmed.
  • Definitions and collection practices are understood.
  • Historical changes are documented.

Watch for: Participation is recorded differently across units. Counts may reflect collection practice rather than student behavior. Define minimum collection standards and flag sources that do not meet them.

Does the dataset behave the way we think it does?

Confirm identifiers, grain, uniqueness, data types, null patterns, ranges, duplicates, and record counts.

Validate:

  • The grain is explicit.
  • Duplicate rules are understood.
  • Values, ranges, codes, and dates are plausible.
  • Missingness is examined.

Watch for: Confusing records with students. One student may appear multiple times when data capture activities, transactions, or events. Confirm what each record represents before counting, calculating rates, or joining data.

Do the fields and rules mean what we think they mean?

Verify operational definitions, inclusion and exclusion rules, status codes, dates, and business logic with stewards and program partners.

Validate:

  • Definitions are confirmed with the appropriate owner.
  • Inclusion and exclusion rules are documented.
  • Current and historical definitions are distinguished.

Watch for: Current definitions are applied historically. Programs, statuses, codes, and systems change over time. Version definitions and create period-specific logic where needed.

What happens when we connect this data to something else?

Test joins, unmatched records, one-to-many behavior, temporal alignment, and duplication introduced through connection.

Validate:

  • Join keys are tested.
  • Match rates are documented.
  • Unmatched records are examined.
  • One-to-many relationships are understood.

Watch for: Unmatched records are dropped silently. The analytic population may become systematically different. Report match rates and profile unmatched cases. Also watch for: One-to-many joins inflate counts. Test cardinality and aggregate at the correct stage.

Does what we see align with what we already know?

Compare totals and distributions with trusted reports or source-system counts and investigate material differences.

Validate:

  • Counts reconcile with an authoritative reference.
  • Filters, census dates, and definitions align.
  • Material differences are explained.

Watch for: Totals differ from a known report. Reconcile logic with the authoritative owner rather than simply forcing the numbers to match. Also watch for: Missing data are treated as no participation. Distinguish a true zero from an unknown or unobserved value.

What can this dataset support, and what can it not support?

Record tests, exceptions, limitations, approvals, version, and uses for which the dataset is or is not appropriate.

Validate:

  • Tests and exceptions are documented.
  • Known limitations are stated in plain language.
  • Required reviews are complete.
  • Version, query/code, and refresh information are recorded.
  • Supported and unsupported uses are explicit.

Release rule: Do not label a dataset analysis-ready until the intended use, key tests, unresolved exceptions, and required reviews are documented.

Validation Sign-Off

Validation is most useful when the work can be understood and revisited by someone beyond the person who performed it. Use this record to capture what was validated, any limitations or exceptions, who reviewed the work, and when the dataset should be revisited.

AREAVALIDATION STANDARDCOMPLETE
PurposeQuestion, decision, population, grain, time window, and intended outputs are documented. 
OwnershipSource owner and steward have confirmed definitions and appropriate use. 
CompletenessExpected periods, populations, programs, and required fields are present. 
UniquenessDuplicate rules are understood and tested at the intended grain. 
ValidityValues, ranges, codes, statuses, and date sequences are plausible. 
ConsistencyDefinitions and collection practices are stable—or changes are documented. 
LinkageJoin keys, match rates, unmatched records, and one-to-many relationships are examined. 
ReconciliationCounts and distributions are compared with an authoritative reference. 
MissingnessNulls and systematic missing patterns are evaluated for analytical implications. 
Time logicActivity, enrollment, outcome, and census dates align with the question. 
DerivationsCalculated fields and recodes are reproducible and reviewed. 
LimitationsKnown constraints and excluded uses are written in plain language. 
ApprovalSteward, practitioner, and dissemination reviews are recorded as required. 
ReproducibilityVersion, code/query, refresh date, and documentation location are recorded. 

Analysis-Ready Means Fit for Purpose: A dataset is not analysis-ready simply because it is integrated, accessible, or technically clean. It is analysis-ready when its fitness for the specific intended use has been tested, documented, reviewed, and understood.