Skip to content

Methods overview

GazeForge has a broad technical surface, but most studies only need a subset of it. Start from the research task below, then move to the detailed method page that matches the analysis you actually need. For a shorter task-first route with expected artifacts and claim boundaries, use Research recipes.

Methods are not evidence claims

This page organizes software methods and analysis choices. It does not change the generated Evidence status, promote a benchmark, or turn synthetic/demo output into empirical validation. Device adapters likewise do not establish device-specific validity.

Choose by research task

Prepare & quality-check gaze

Use explicit adapters, canonical units, participant/trial identity, and non-destructive quality signals before modelling.

Start: Real-data import clinic
Then: Adapters & validation · Motion-quality gating · Synthetic QC tutorial

Model & validate eye events

Begin with transparent event rules when appropriate, then add learned classifiers only with a leakage-safe held-out design, matched comparisons, calibration, confidence/coverage, and event-level evaluation.

Start: Event-model validation clinic
Run it: Worked validation study
Then: Temporal models · Model comparison · Event-level evaluation · Calibration & dataset holdouts

Build semantic or dynamic AOIs

Represent static regions directly or use reviewable proposals and time-bounded geometry for moving stimuli.

Start: Dynamic AOIs
Run it: Worked dynamic-AOI study
Then: Grounding DINO + SAM 2 · Verified video-frame derivation · Dynamic AOI evaluation

Represent scanpaths & process

Convert reviewed fixation-to-AOI assignments into observable sequence structures for description, similarity, embeddings, and downstream modelling.

Start: Research workflows
Run it: Practical end-to-end workflow

Fit hierarchical distributional models

Use the location-scale family when the scientific question concerns both conditional location and residual scale, including random slopes, covariance structure, calibration, and bootstrap uncertainty.

Start: Hierarchical location-scale models
Then: Correlated location-scale effects · Location random-slope scale model · Full-covariance model

:material-shield-search-outline: Validate, audit & report

Keep split design, sampling-rate handling, calibration, benchmark provenance, frozen evidence, and manuscript-facing software identity visible.

Start: Validation guide
Do event validation: Event-model validation clinic
Plan/report: Study-design templates · Validation reporting cookbook · Reproducible reporting

Method map

Research question Primary method pages Typical reviewable output
How should tracker data enter GazeForge? Real-data import clinic, Adapters & validation canonical gaze table with declared units/rate and source provenance
Which samples or trials need review? Motion-quality gating, Synthetic QC tutorial flags, weights, quality summaries; source rows retained
How should gaze samples become event labels? I-VT tutorial, Temporal models labels/probabilities with model and threshold provenance
How should a learned event model be validated? Event-model validation clinic, Model comparison, Calibration participant/split ledger, matched held-out predictions, probabilities, sample/event metrics, calibration and coverage
How should event performance be evaluated? Event-level evaluation, Stratified performance, Matched-fold differences, Calibration held-out metrics, matched differences, calibration tables with the split unit named
How should moving semantic regions be represented? Dynamic AOIs, Worked dynamic-AOI study, Video-frame derivation, Dynamic AOI evaluation reviewed keyframes, bounded interpolation, no-extrapolation audit, fixation assignments
How can AI propose visual regions without becoming the empirical record? Grounding DINO + SAM 2 backend, Dynamic AOIs proposals plus confidence and review decisions
How should sequence/process structure be represented? Research workflows, Practical workflow semantic scanpaths and provenance-bound exports
What does a gaze-derived measure support me saying? Measurement & interpretation clinic, Research terminology observable/construct bridge + threats + sensitivity/reporting limits
How should a study be preregistered and archived? Study-design templates, Study lifecycle, Publication readiness explicit acquisition/QC/AOI/split/rate/archive records
How can conditional variability be modelled? Hierarchical location-scale, Correlated location-scale location/scale effects with explicit model assumptions
How can random slopes and covariance be represented? Location random-slope scale, Correlated random-slope scale, Full-covariance random-slope scale random-effect/covariance estimates and diagnostics
How should location-scale uncertainty and calibration be checked? Residual calibration, Conditional refit calibration, Hierarchical bootstrap, Bootstrap Monte Carlo precision calibration diagnostics and uncertainty summaries

A useful order of operations

source / tracker export
        ↓
explicit adapter + canonical schema
        ↓
non-destructive QC / reliability evidence
        ↓
transparent or learned event model
        ↓
leakage-safe evaluation + calibration
        ↓
reviewed static/dynamic AOIs / fixation assignments / scanpaths
        ↓
statistics or hierarchical models
        ↓
provenance + evidence boundary + reproducible report

The order is deliberately review-first. AI-generated anomaly flags, event probabilities, or AOI boxes remain inspectable analytic data; they do not silently overwrite the source observations.

Event modelling: baseline before complexity

For a new dataset, a transparent baseline can make model behaviour easier to inspect before a learned classifier is introduced. The I-VT tutorial shows a deterministic pixel-velocity baseline. When expert-labelled events are available, use the Event-model validation clinic to preserve participant identity, leakage checks, matched held-out rows, probabilities, calibration, confidence/coverage, and separate sample/event estimands. Then use model comparison, matched-fold differences, event-level evaluation, stratified performance, and calibration for the required detail.

A method being available in the package does not establish that it is superior for a new population, device, sampling regime, or task. Those are empirical questions that require an appropriate validation design. Likewise, a confidence threshold chosen for one validation design is not a universal abstention cutoff.

Measurement interpretation: observable before construct

When a result moves from “where/when/how long” to a statement about trust, interest, comprehension, persuasion, memory, emotion, or another latent construct, use the Measurement & interpretation clinic. It records the observable, construct bridge, external evidence, validity threats, sensitivity checks, and reporting boundary before the substantive interpretation is frozen.

AOIs and scanpaths: observable structure, not latent state

Semantic and dynamic AOIs describe where reviewed regions are located. Scanpaths describe observable fixation order across those regions. Neither surface, by itself, establishes emotion, persuasion, comprehension, intent, diagnosis, or another latent psychological state.

For video or moving interfaces, use Dynamic AOIs with bounded interpolation and explicit review. The worked dynamic-AOI study demonstrates exact keyframes, interpolation within a declared maximum gap, fixation assignment, semantic sequences, and explicit no extrapolation before/after the observed track. For end-to-end static composition from events to AOI assignments and scanpaths, use the practical workflow.

Distributional modelling

The location-scale family is intentionally separated into pages because the models answer different structural questions. Start with Hierarchical location-scale models, then add correlation or random-slope structure only when the design and estimand require it:

Keep method choice separate from evidence strength

For empirical claims, use the Validation guide and generated Evidence status rather than inferring validity from method availability.

In particular:

  • derived lower-rate evidence remains derived and does not establish native 60 Hz or Gazepoint GP3 validity;
  • source-token-disjoint Hollywood2EM evidence is not participant-disjoint;
  • current Gaze-in-the-Wild participant-disjoint evidence remains task-agnostic while complete authoritative numeric task mapping is unresolved; and
  • current VISUS evidence remains bounded partial public-derivative evidence rather than a full-dataset or native-GP3 validation claim.

Run instead of browse

If you are starting from a real tracker or processed export, use the Real-data import clinic first. If you are validating a learned event classifier, continue with the Event-model validation clinic and run examples/06_worked_event_model_validation.py before adapting the structure to real labelled data. If you know the task but not the API, open Research recipes. If you want executable examples, open the Runnable examples gallery: static studies can start from the worked advertising/interface study, moving stimuli can start from the worked dynamic-AOI study, and learned event validation can start from the worked validation study. Use the Study-design templates to freeze the corresponding preregistration, acquisition, QC, AOI, split, native/derived, and archive records, and the Validation reporting cookbook to keep manuscript wording proportional to the design.