Skip to content

Study lifecycle

Use this page when you want to move from a research question to a manuscript-facing, auditable GazeForge analysis without treating the package as a collection of disconnected functions. If you already know the practical task, use Research recipes; if you are preregistering or freezing study metadata, use the Study-design templates.

A complete workflow is not the same thing as validated evidence

Completing every stage below does not automatically validate a tracker, event model, AOI method, population, task, or sampling regime. The lifecycle keeps assumptions and evidence boundaries visible; empirical claims still depend on the relevant validation design. Use the Validation guide and generated Evidence status for current evidence claims.

1 · Plan

Define the observable question

Specify what the gaze data can directly represent, the unit of analysis, acquisition requirements, AOI source, QC rule, and validation plan before fitting models.

Open study templates →

2 · Prepare

Preserve and canonicalise

Keep the source immutable, fingerprint the analysed table, document time/coordinate semantics, compare nominal/native rate with observed timestamp cadence, and create a vendor-neutral canonical derivative.

Run the worked tracker import →

3 · Analyse

QC, events, AOIs, sequences

Add non-destructive QC, start from inspectable event rules, preserve AOI provenance, and retain reviewable semantic sequence outputs. Before inferential modelling, convert reviewed outputs into model-ready tables without losing grouping, exposure, missingness, or censoring.

Choose a research recipe → · Build the statistical handoff →

4 · Validate

Match evidence to the claim

Name the held-out unit, reference labels, acquisition provenance, rate status, calibration, confidence/coverage, and event-level metrics that actually support the intended claim. For learned event models, make participant/split identity and matched held-out rows explicit.

Open the event-model validation clinic →

5 · Freeze

Freeze identity and provenance

Archive the exact software/environment identity, source and output fingerprints, import contract, manifests, certificates, figures, tables, split ledgers, and unresolved evidence boundaries.

Publication readiness →

6 · Report

Write what was actually done

Translate acquisition, preprocessing, QC, model, split, rate, metric, calibration, statistical-handoff, and evidence identity into manuscript language that another team can reconstruct without strengthening the claim in prose.

Open the reporting & interpretation clinic → · Use the validation reporting cookbook →

The eleven-stage research route

Stage Input Action Reviewable output Continue with Do not infer
1. Define the question substantive theory + task define observable gaze construct and unit of analysis analysis/preregistration plan Study templates latent states from gaze alone
2. Record acquisition facts tracker/stimulus setup record hardware, native/nominal rate, geometry, participant/trial identity acquisition/source record Worked tracker import undocumented acquisition facts
3. Preserve source identity original export/table retain immutable source and fingerprint analysed source table source snapshot + checksum/fingerprint Worked tracker import fingerprint = independent validation
4. Canonicalise explicitly source semantics map time, coordinates, identity, optional fields; inspect duplicates/cadence/bounds/row counts import contract + canonical gaze table Adapters & validation successful import = device validity
5. Add QC evidence canonical samples flag anomalies and score trial quality without silent deletion QC columns + trial summaries Research recipes QC flag = invalid observation
6. Build measurement outputs reviewed samples apply event baseline/model; define/review static or dynamic AOIs; derive scanpaths if needed event/AOI/sequence tables Methods overview complex model = superior model
7. Build the statistical handoff reviewed event/AOI outputs + design/coverage preserve participant/trial hierarchy, denominators, missing/zero semantics, and censoring model-ready trial × AOI/event tables + handoff dictionary Analysis handoff aggregation convenience = correct inferential unit
8. Validate the estimand reference labels + split policy evaluate on leakage-safe held-out data with matching metrics split ledger + held-out predictions + sample/event/calibration metrics Event-model validation clinic sample accuracy = temporal event quality
9. Audit rate and provenance acquisition + analysis-rate history distinguish native/nominal, observed cadence, and derived analysis rates rate/sensitivity record Sampling sensitivity derived 60 Hz = native 60 Hz validity; observed cadence = native hardware proof
10. Freeze the evidence bundle final analysis outputs freeze manifests, fingerprints, certificates, code/environment, figures/tables reconstructable archive Publication readiness archive completeness = stronger evidence
11. Report qualified claims frozen bundle write methods/results/captions with explicit evidence boundary manuscript-ready record + reporting derivatives Reporting clinic prose stronger than the frozen evidence

Worked route: tracker export → canonical samples → QC

Before a downstream event/AOI analysis, run the import/QC contract itself:

python examples/07_worked_tracker_import_qc.py \
  --output-dir worked-tracker-import-qc-demo

The deterministic Gazepoint-shaped source uses USER_FILE, MEDIA_ID, TIME, BPOGX, and BPOGY. The script explicitly declares seconds→milliseconds and normalized→pixel conversion, fingerprints and snapshots the source, checks identity, duplicate keys, observed cadence, coordinate bounds, and row-count preservation, then adds non-destructive anomaly flags and trial-quality summaries.

The teaching source deliberately retains a duplicate sample key, off-screen gaze, and a missing coordinate. Those observations are not silently deduplicated, clipped, filled, or deleted. The nominal 60 Hz-shaped source and timestamp-derived observed cadence are reported separately; agreement in the demo is not evidence of native GP3 acquisition.

The output bundle contains five CSVs plus import_contract.json, analysis_plan.json, provenance.json, and workflow_manifest.json. It is classified synthetic_demo_not_empirical_evidence and creates no Gazepoint/GP3, native-60-Hz, event-model, or measurement-validity claim.

Open the worked tracker import → · Use the import clinic →

Worked route: a static advertising/interface study

A concrete worked example is available for a hypothetical static advert/interface with four researcher-defined regions:

brand → claim → product → disclosure

The bundled script uses deterministic synthetic, real-data-shaped gaze and demonstrates source preservation, canonicalisation, QC, transparent I-VT events, fixation centroids, AOI assignment, semantic scanpaths, provenance, and an analysis-plan record.

python examples/04_worked_advertising_study.py \
  --output-dir worked-advertising-study-demo

The example contains no empirical advertising effect and no model-performance claim. It exists to show how a domain study can be structured without turning software output into unsupported evidence.

Open the static worked study →

Worked route: a moving stimulus with dynamic AOIs

For video or moving interfaces, the second worked study makes the geometry-over-time contract explicit:

python examples/05_worked_dynamic_aoi_study.py \
  --output-dir worked-dynamic-aoi-demo

It uses deterministic synthetic fixation rows plus reviewed product, claim, and cta keyframes. Within the observed track range, geometry can be resolved by exact keyframes or bounded interpolation. Deliberate probes before and after the observed range must remain unassigned, so the example mechanically demonstrates no temporal extrapolation.

The output bundle includes the raw fixation table, dynamic keyframes, assignments, semantic scanpaths, interpolation audit, assignment summary, analysis plan, provenance, manifest, and optional figures. It is classified synthetic_demo_not_empirical_evidence: it does not validate a detector/tracker, a device, native 60 Hz acquisition, Gazepoint/GP3, or a substantive psychological effect.

Open the dynamic worked study → · Browse all runnable examples →

Worked route: participant-held-out event-model validation

When the study uses learned event classifiers, run the validation contract separately from downstream substantive analyses:

python examples/06_worked_event_model_validation.py \
  --output-dir worked-event-model-validation-demo \
  --no-figures

The demonstration constructs explicit participant-labelled synthetic events, creates four participant-disjoint folds, verifies zero train/test participant overlap, and evaluates I-VT, Random Forest, and ContextMLP on matched held-out rows. It writes separate sample-level and event-level metric tables, calibration bins, confidence/coverage diagnostics, an illustrative abstention policy, split ledger, predictions, fingerprints, provenance, and a manifest.

The example is deliberately classified synthetic_demo_not_empirical_evidence. A model that leads on one synthetic metric is not thereby superior for another device, task, population, or benchmark, and the demonstration does not establish native 60 Hz or Gazepoint/GP3 validity.

Open the validation clinic → · Run the worked validation study → · Use the reporting cookbook →

Decision points worth freezing before analysis

Before you treat the workflow as confirmatory, freeze the outcome/estimand registry with the Outcome & estimand preregistration clinic, then record at least:

  • the participant/trial/stimulus identifiers that define independent units;
  • the native/nominal acquisition rate, observed timestamp cadence, and any derived analysis rate;
  • the exact import mapping, timestamp unit, coordinate basis, screen geometry, source fingerprint, and duplicate/bounds diagnostics;
  • the source of AOIs and whether AI proposals were human-reviewed;
  • for dynamic AOIs, keyframe timestamps, review state, maximum interpolation gap, overlap rule, and the no-extrapolation policy;
  • the QC review/exclusion rule and whether it was prespecified;
  • the event method, thresholds, model identity, confidence/abstention rule, and training regime;
  • the held-out unit, split ledger, leakage controls, and whether compared models use identical held-out rows;
  • whether model/threshold selection is separated from final confirmatory evaluation;
  • the primary, secondary, and exploratory outcome status, target estimand/contrast, inferential unit, exposure/denominator, missing/zero/censoring policy, and prespecified sensitivity analyses;
  • the primary sample-, event-, and calibration-metric families that match the estimand;
  • the software version or exact commit SHA; and
  • the explicit statement of what the design does not establish.

If those decisions are not yet fixed, label them exploratory rather than backfilling certainty after seeing the outputs. The Study-design templates provide copy-ready records for each of these decisions.

Keep the layers separate

software can run
      ≠
input is scientifically valid
      ≠
model is validated for this setting
      ≠
substantive interpretation is established

That separation is the central reason to keep source data, transformations, predictions, review decisions, validation artifacts, and manuscript claims as distinct records.

Continue with the Outcome & estimand preregistration clinic, Worked tracker import, Research recipes, Study-design templates, the Event-model validation clinic, Validation reporting cookbook, Publication-readiness checklist, Research terminology, and Reproducible reporting.