Learning paths¶
GazeForge spans data preparation, quality control, human review, eye-event modelling, semantic AOIs, scanpaths, benchmark validation, and scientific provenance. You do not need to learn every layer before starting.
If you are not yet sure what the package does, begin with the GazeForge Tour. It is the canonical orientation route: one gaze table goes through the main layers and produces ordinary CSV/JSON artifacts you can inspect. If you understand the package but need to turn a research question into a complete study workflow, continue with the First study blueprint. If the task is already clear but the method is not, use the Method chooser; if the files are unfamiliar, use the Artifact & output dictionary.
Choose the path that matches your immediate research question. If you are planning an entire study rather than learning one method, use the Study lifecycle as the orchestration layer from acquisition through publication. If you already know the task you need to perform, the Research recipes page gives the shortest defensible route and the artifacts to retain.
I am new — what does GazeForge do?¶
Run the smallest package-wide walkthrough before choosing a specialist method. See canonicalisation, QC, transparent events, AOIs, scanpaths, provenance, and the resulting output bundle in one place.
Next: GazeForge Tour
Run it: python examples/00_gazeforge_tour.py --output-dir gazeforge-tour-demo
I have a real tracker export¶
Start with the executable import/QC route before event modelling. Preserve the source, declare identity/time/coordinate semantics, compare nominal rate with observed timestamp cadence, inspect duplicate keys and bounds, then add non-destructive QC.
Next: Worked tracker import and QC
Deep guide: Real-data import clinic
QC found problems; what do I exclude?¶
Do not map anomaly flags directly to deletion. Keep pre-review QC immutable, record criteria and denominators, distinguish prespecified from exploratory rules, and create a separate reviewed analysis derivative.
Next: QC review and exclusion ledger
Run it: python examples/08_worked_qc_review_ledger.py --output-dir worked-qc-review-ledger-demo
I want a first result¶
Start with a deterministic synthetic dataset, canonicalise it, add QC flags, and inspect trial-level quality.
Time: about 15 minutes
Next: Synthetic QC tutorial
I need eye-event labels¶
Begin with the transparent I-VT baseline before fitting a learned classifier. This gives you a reference whose decision rule is directly inspectable.
Time: about 20 minutes
Next: I-VT baseline tutorial
I need to validate a learned event model¶
Keep participant identity, held-out folds, probabilities, sample/event metrics, calibration, confidence/coverage, and model-selection boundaries explicit.
Next: Event-model validation clinic
Run it: Worked validation study
I have a video or moving interface¶
Use timestamped dynamic AOI keyframes, bounded interpolation, explicit review, and fixation assignment without extrapolating geometry outside the observed track.
Next: Worked dynamic-AOI study
I need a reviewable manuscript/archive bundle¶
Keep source identity, pre-review QC, reviewed decisions, primary-analysis rows, downstream derivatives, provenance, fingerprints, and reporting metadata as distinct evidence layers.
Next: Research evidence bundle
Run it: python examples/09_worked_research_evidence_bundle.py --output-dir worked-research-evidence-bundle
I need to freeze outcomes and estimands before modelling¶
Register primary/secondary/exploratory outcomes, contrasts, exposure/denominator, missing/zero/censoring semantics, sensitivity checks, and a deviation ledger before results exist.
Next: Outcome & estimand preregistration clinic
Run it: python examples/13_worked_estimand_preregistration.py --output-dir worked-estimand-preregistration
I have a gaze metric and need to know what it supports¶
Separate the observable from the proposed construct, review validity threats, preserve censoring/exposure, and plan justified sensitivity checks before making a substantive interpretation.
Next: Measurement & interpretation clinic
Run it: python examples/12_worked_measurement_interpretation_audit.py --output-dir worked-measurement-interpretation-audit
I need to reconcile denominators, exposure, or censoring¶
Use this route before modelling when a zero may mean several different things, when trial/AOI exposure differs, or when no-fixation latency must remain right-censored.
Start here: Denominator, exposure & censoring clinic
Run it: python examples/16_worked_denominator_exposure_audit.py --output-dir worked-denominator-exposure-audit
Continue to: Analysis handoff
I need to audit model convergence or diagnostics¶
Use this route after a specialist statistical package has fitted the model. Treat non-convergence, singular/boundary states, separation/invalid covariance, and missing diagnostics as stop conditions before interpretation.
Start here: Model diagnostics & convergence clinic
Run it: python examples/17_worked_model_diagnostics_audit.py --output-dir worked-model-diagnostics-audit
Continue to: Sensitivity & robustness clinic · Reporting clinic
:material-chart-error: I need to audit uncertainty or multiplicity before reporting¶
Use this route after the fitted model has passed convergence/diagnostic checks. Keep effect scale, units, interval method/level, multiplicity family, and raw/adjusted inferential fields explicit before manuscript wording is frozen.
Start here: Uncertainty, multiplicity & inferential reporting clinic
Run it: python examples/19_worked_inferential_reporting_audit.py --output-dir worked-inferential-reporting-audit
Continue to: Sensitivity & robustness clinic · Reporting clinic
I need to audit sensitivity or robustness¶
Use this route after the primary estimand and sensitivity plan are frozen and the analysis variants have been executed. Keep prespecified, exploratory, and deviation analyses distinct; preserve failed/unevaluable variants; and compare only same-estimand results directly.
Start here: Sensitivity & robustness clinic
Run it: python examples/15_worked_sensitivity_robustness_audit.py --output-dir worked-sensitivity-robustness-audit
Continue to: Reporting clinic · Publication readiness
I need to write Methods, Results, and captions without overclaiming¶
Translate the frozen evidence identity into claim-safe prose while preserving QC/exclusion, split, sampling, calibration, synthetic/empirical, and observable/latent-state distinctions.
Next: Reporting & interpretation clinic
Run it: python examples/11_worked_manuscript_reporting_bundle.py --output-dir worked-manuscript-reporting-bundle
I am sharing with a reviewer or replicator¶
Use this route after the preregistered estimand records, analysis, interpretation, and reporting artifacts are frozen and the external reader needs to know which files support each statement, what can be rerun, and which inputs require authorized access.
Start here: Reviewer & replication handoff
Run it: python examples/14_worked_reviewer_replication_bundle.py --output-dir worked-reviewer-replication-bundle
Continue to: Publication readiness · API reference
I need defensible empirical evidence¶
Use the benchmark/evidence layer only after the split, labels, sampling condition, and source provenance match the claim you intend to make.
Next: Validation guide
I am planning or preregistering a study¶
Use copy-ready records for acquisition metadata, QC/exclusion rules, AOI provenance, split identity, native/derived sampling, archive manifests, and manuscript Methods.
Next: Study-design templates
:material-shield-search-outline: I need to audit evidence¶
Read the validation matrix, frozen evidence, source-resolution records, and benchmark-specific claim boundaries before interpreting a headline metric.
Next: Validation status
Prefer runnable scripts?¶
Open the Runnable examples gallery for twenty-two deterministic examples/workflows with exact commands, dependencies, expected outputs, and links to the underlying repository files. Start with the GazeForge Tour if you need the package-wide mental model. The worked tracker-import/QC example demonstrates the real-data handoff contract; the QC review/exclusion-ledger clinic demonstrates review and denominator accounting; the worked advertising/interface study demonstrates a static-stimulus design; the worked dynamic-AOI study demonstrates moving regions, bounded interpolation, and explicit no-extrapolation checks; and the worked event-model validation study demonstrates participant-disjoint model comparison with separate sample/event/calibration outputs; and the research evidence bundle demonstrates how to freeze those layers into an archive-facing directory. For a task-first map, start with Research recipes; for deeper technical documentation, use the Methods overview.
A practical progression¶
| Stage | Learn | Produce | Do not claim yet |
|---|---|---|---|
| 0. Import | source identity, units, geometry, nominal rate vs observed cadence | immutable source + import contract + preflight | adapter compatibility = device validity |
| 1. Canonicalise | schema, rate, units, participant/trial boundaries | one vendor-neutral gaze table | comparability across datasets |
| 2. QC | missingness, gaps, off-screen samples, anomaly flags | reviewable QC columns and trial summaries | automatic exclusion validity |
| 3. Review | criteria, scope, decisions, denominators, prespecified vs exploratory status | immutable pre-review QC + review/exclusion ledger + reviewed derivative | reproducible exclusion rule = validated rule |
| 4. Baseline | deterministic I-VT or angular I-VT | inspectable event labels | learned-model superiority |
| 5. Validate | participant-disjoint folds, matched rows, sample/event metrics, calibration/coverage | split ledger + held-out predictions + validation tables | native-device validity from resampled or synthetic data |
| 6. Extend | semantic/dynamic AOIs, scanpaths, hierarchical models | task-specific analytic structures | unsupported psychological inference |
| 7. Freeze | manifests, fingerprints, certificates, source resolution | artifact index + auditable evidence bundle + reviewer README | stronger provenance than the source supports |
For a more complete research route, use the Outcome & estimand preregistration clinic before modelling, then the Grouping, repeated measures & pseudoreplication clinic before specialist fitting when rows repeat within participants/stimuli; the Study lifecycle ties every stage to a reviewable artifact and explicit claim boundary. The Study-design templates make the corresponding records copy-ready.
Which import and QC workflow should I use?¶
Do you have authoritative source metadata?
│
├─ No → stop; recover units/identity/geometry from acquisition or export records
│
└─ Yes
├─ Gazepoint-style fields + known semantics? → adapt_gazepoint_samples()
├─ Other known processed columns? → adapt_processed_table()
├─ Already canonical ms + pixels? → canonicalize_gaze()
└─ Then inspect:
identity · duplicate keys · cadence · bounds · row counts
│
▼
non-destructive QC
│
▼
review + exclusion ledger
Run the worked tracker-import example for the import path, then the QC review/exclusion-ledger clinic before dropping observations. Successful import is a transformation result, not device/model validation; a QC flag is review evidence, not an automatic exclusion.
Which event workflow should I use?¶
Do you already have expert-labelled event data?
│
├─ No
│ ├─ Need a transparent descriptive baseline? → I-VT / angular I-VT
│ └─ Need publishable classifier validation? → acquire or use an audited labelled corpus first
│
└─ Yes
├─ Same participants in train and test? → stop; use participant-disjoint splitting
├─ Only opaque source tokens available? → report source-token-disjoint, not participant-disjoint
├─ Compatible sampling regime? → fit + validate model
├─ Need boundary-sensitive performance? → add event-F1 / temporal IoU / boundary error
├─ Need probability claims? → add Brier / ECE / confidence-coverage
└─ Need lower-rate claims? → preserve native vs derived status + sensitivity
Use the Event-model validation clinic for the full leakage-safe route and the Validation reporting cookbook when converting the final design into manuscript wording.
Which AOI workflow should I use?¶
Static stimulus
└─ define or propose AOIs → human review → freeze AOIs → assign fixations
Video / moving interface
└─ timestamped keyframes → bounded interpolation → review → dynamic fixation assignment
│
└─ no extrapolation outside the observed track
AI-proposed AOIs remain proposals until reviewed. The provenance record should retain model identity, confidence, and any accept/reject/relabel/edit decision. Run the dynamic worked study for a concrete, deterministic example.
Read results visually¶
The results gallery puts the current reviewed benchmark summaries beside their evidence boundaries. It is intended as an orientation layer, not a replacement for frozen reports or validation certificates.
Report the analysis so somebody else can reconstruct it¶
When an analysis becomes manuscript-facing, use the Denominator, exposure & censoring clinic to preserve observation-state mechanics, the Missing-data assumptions & treatment handoff when unavailable measurements require explicit assumptions, and the Measurement & interpretation clinic to audit any substantive gaze claim, then continue with the Sensitivity & robustness clinic, Reporting & interpretation clinic, Reviewer & replication handoff, Publication-readiness checklist, QC review and exclusion-ledger clinic, Validation reporting cookbook, and Reproducible reporting. The reporting surfaces keep import compatibility versus device validity, QC flags versus review/exclusion decisions, demos versus empirical validation, split identity, sample/event metrics, calibration, confidence/coverage, native/derived rate, and model-selection versus confirmatory evaluation explicit. Use Study-design templates to keep the required metadata explicit from preregistration onward.