Method chooser¶
GazeForge contains several ways to import, review, label, summarize, and validate gaze data. The right route depends on what you are trying to learn, what evidence you actually have, and which unit must generalize.
Do not choose a method from a headline metric
A workflow that is useful for one question can be invalid for another. Start from the study unit, reference labels, source identity, timebase, AOI support, and intended generalisation unit. Software availability is not scientific justification.
Quick chooser¶
| Research need | Required input | Generalisation / grouping unit to protect | Start with | Primary outputs | Validation requirement | Do not infer | Continue to |
|---|---|---|---|---|---|---|---|
| Import a tracker export | authoritative source mapping, units, geometry | participant/trial identity must survive import | Worked tracker import + QC | canonical gaze + preflight + import contract | transformation checks | adapter compatibility ≠ device validity | QC review |
| Inspect data quality | canonical gaze + sampling assumptions | participant/trial boundaries | Synthetic QC | anomaly flags + trial quality | review against study policy | flag ≠ invalid observation | QC review |
| Apply exclusions | immutable pre-review QC + criteria | sample/trial/participant scope | QC review & exclusion ledger | criteria + review ledgers + denominator flow | prespecification/review provenance | reproducible rule ≠ validated rule | analysis derivative |
| Create an inspectable event baseline | gaze samples + rate/timebase | trial boundaries | I-VT baseline | event-labelled samples + intervals | threshold sensitivity where relevant | example threshold ≠ universal physiology | Event validation |
| Fit a learned event classifier | expert/reference labels | participant or intended deployment unit | Event-model validation clinic | held-out predictions + sample/event metrics | leakage-safe held-out validation | training fit ≠ validated performance | Validation reporting |
| Compare event models | same held-out rows for every model | same folds / same units | Worked validation study | matched predictions + model summary | paired/matched held-out comparison | one metric ≠ overall superiority | Model comparison |
| Assess probability quality | held-out class probabilities | same held-out units as classifier claim | Calibration | calibration bins + Brier/ECE + confidence/coverage | held-out probabilities only | confidence ≠ correctness | Validation reporting |
| Use static AOIs | reviewed stimulus geometry | stimulus/version identity | First study blueprint | AOI definitions + fixation assignments | AOI construct rationale | AOI membership ≠ psychological state | scanpaths / analysis |
| Use moving AOIs | reviewed timestamped keyframes/tracks | stimulus + timebase | Worked dynamic-AOI study | keyframes + interpolation audit + assignments | support/no-extrapolation checks | detected track ≠ ground truth | scanpaths / dynamic evaluation |
| Build semantic scanpaths | reviewed fixation/AOI assignments | participant/trial sequence identity | Practical workflow | semantic sequence table | assignment provenance | sequence ≠ latent mental state | downstream sequence analysis |
| Freeze outcomes/estimands before modelling | research question + planned observables | declared population/contrast/inferential unit | Outcome & estimand preregistration | outcome/estimand/contrast/sensitivity/deviation registries | primary/secondary/exploratory status frozen before results | preregistration ≠ estimator choice or validity | Analysis handoff |
| Document missing-data assumptions | reconciled observation states + grouping/QC context | participant/trial/stimulus hierarchy | Missing-data assumptions & treatment handoff | source registry + mechanism questions + treatment/sensitivity handoff | assumptions must be study-specific | MCAR/MAR/MNAR and treatment choice are not package outputs | Analysis handoff |
| Audit grouping/repeated measures | model-ready rows + participant/trial/stimulus/AOI identities | declared inferential/generalisation unit | Grouping, repeated measures & pseudoreplication clinic | unit registry + nested/crossed handoff + aggregation/pseudoreplication audit | grouping structure must follow design | repeated rows ≠ independent participants; grouping identity ≠ automatic random-effect syntax | specialist statistical software |
| Build statistical model inputs | reviewed event/AOI outputs + design/coverage | participant/trial hierarchy | Analysis handoff | trial × AOI/event measures + denominators + censoring | preserve missing/zero/exposure semantics | samples/fixations ≠ independent participants | Grouping clinic |
| Interpret a gaze-derived measure | frozen observable + intended substantive claim | declared measurement/inferential unit | Measurement & interpretation clinic | claim registry + threats + sensitivity/reporting boundaries | construct bridge must be explicit | gaze observable ≠ latent construct | Reporting clinic |
| Freeze a study | reviewed analysis derivative + final settings | exact source/software identity | Study lifecycle | manifest + provenance + fingerprints | deterministic reconstruction | reproducibility ≠ external validity | Publication readiness |
| Prepare a paper/archive | reconciled denominators + final results | claim-specific population/unit | Research evidence bundle | artifact index + methods/figures/tables + provenance/manifest | evidence class must match wording | archive completeness ≠ validity | Reporting clinic |
| Write claim-safe Methods/Results/captions | frozen evidence bundle + reporting facts | claim-specific population/unit | Reporting & interpretation clinic | methods/results examples + citation table + evidence boundaries | wording must match evidence identity | prose cannot strengthen evidence | Publication readiness |
| Evaluate benchmark evidence | exact source/provenance + labels | participant/source/dataset identity | Validation guide | evidence status/certificate/report | benchmark-specific | derived/native or token/participant distinctions cannot be collapsed | Evidence status |
Before modelling: outcomes or contrasts are still moving¶
Use the Outcome & estimand preregistration clinic. Freeze primary/secondary/exploratory status, exact observable definitions, inferential unit, exposure/denominator, missing/zero/censoring policy, contrasts, and prespecified sensitivity checks. Record later changes in a deviation ledger rather than rewriting the original registration.
Decision rules that should stop the workflow¶
No expert/reference labels¶
You can run a transparent descriptive event baseline, but do not describe a learned classifier as empirically validated merely because it produces labels or probabilities.
Participant-level generalisation claim¶
Use participant-disjoint splitting. Random row splitting is not a substitute when observations from the same participant are correlated across train and test.
Only opaque source tokens are available¶
Report source-token-disjoint if that is what the evidence supports. Do not rename it participant-disjoint without an authoritative token→participant mapping.
Lower-rate evidence is derived from higher-rate acquisition¶
Call it derived. Derived 60 Hz evidence is not native 60 Hz or native Gazepoint/GP3 validation.
QC anomaly is detected¶
Treat it as review evidence. Do not convert qc_flag=True directly into sample, trial, or participant exclusion unless the study policy explicitly specifies and reviews that decision.
Moving AOI falls outside reviewed temporal support¶
Return unassigned / unsupported according to the declared workflow. Preserve no extrapolation unless a separately justified policy explicitly allows extrapolation.
One model leads on one metric¶
Report the metric-specific result. Sample-level classification, event-boundary fidelity, calibration, confidence/coverage, and downstream utility answer different questions.
A model-ready table contains missing values¶
Stop before replacing them. Use the Denominator, exposure & censoring clinic to preserve observation-state mechanics, then the Missing-data assumptions & treatment handoff when unavailable measurements require explicit assumptions or treatment planning. Preserve participant/trial grouping through the Analysis handoff.
Repeated rows are being treated as independent participants¶
Stop before fitting. Open the Grouping, repeated measures & pseudoreplication clinic. Preserve participant/trial/stimulus/AOI or event identities, distinguish nested from crossed design structure, and keep descriptive aggregation separate from inferential input. GazeForge does not automatically choose random effects, fixed effects, covariance structures, GEE, clustering corrections, or another estimator.
The statistical model fails diagnostics or convergence¶
Do not treat returned coefficients as a valid result. GazeForge does not automatically select or rescue an inferential estimator; resolve the statistical specification and diagnostics in the prespecified specialist analysis environment.
A gaze result is being used as a psychological construct¶
Open the Measurement & interpretation clinic. Separate the observable from the proposed construct, name the external outcome/theory required, review measurement threats, and keep unsupported latent-state language out of the result.
The analysis is frozen but the manuscript wording feels stronger than the evidence¶
Use the Reporting & interpretation clinic. Preserve import/device, QC/exclusion, split identity, calibration/correctness, native/derived rate, synthetic/empirical, and observable/latent-state distinctions in prose and captions.
Which event route should I use?¶
Do you have suitable reference labels?
│
├─ No
│ ├─ Need transparent descriptive segmentation?
│ │ → I-VT / angular I-VT baseline
│ └─ Need a validated learned classifier?
│ → stop; acquire/audit labelled reference data first
│
└─ Yes
├─ Is participant identity available?
│ ├─ Yes → participant-disjoint validation when participants must generalize
│ └─ No → report the weaker identity boundary actually supported
│
├─ Need probability interpretation?
│ → add calibration + confidence/coverage
│
└─ Need temporal boundary performance?
→ add event-F1 / temporal IoU / boundary-sensitive metrics
Which AOI route should I use?¶
Static stimulus?
└─ define/propose AOIs → review → freeze AOIs → assign fixations
Moving stimulus?
└─ reviewed keyframes/tracks
→ verify timebase
→ bounded interpolation
→ no extrapolation outside support
→ assign fixations
→ retain interpolation/assignment audit
AI proposals remain proposals until the study's review policy is satisfied.
Which evidence should I keep?¶
Use the Artifact & output dictionary to identify which CSV/JSON is source evidence, QC/review evidence, an analysis derivative, validation evidence, or reporting/provenance metadata.
Run the Worked research evidence bundle to see those layers assembled into one deterministic archive-facing output directory, then use Publication readiness before freezing manuscript claims.
Method choice is not a ranking¶
This page does not identify a universal “best” event model, AOI method, threshold, or validation metric. It routes methods to questions and evidence conditions. Performance comparisons belong in the corresponding held-out validation context.