Documentation map¶
GazeForge has deep method, validation, benchmark, and workflow documentation. You do not need to read it in navigation order. Start from the task you need to complete, use the shortest runnable route, inspect the artifacts it produces, and only then move into explanation or reference pages.
How this map is organized
The site separates four documentation needs: tutorials for first success, how-to guides for a concrete task, explanations for method/evidence reasoning, and reference for exact API or status information. A page can link across those types, but it should not try to be all four at once.
Software routes are not scientific validation
A successful command, reproducible artifact bundle, or passing software check demonstrates software behavior. It does not by itself establish tracker validity, event-model validity, calibration validity, measurement validity, or general scientific suitability for a study.
Choose by research task¶
| Research task | Prerequisite | Primary start | Run / public surface | Main artifact | Boundary to keep visible | Continue to |
|---|---|---|---|---|---|---|
| Understand GazeForge | none | GazeForge Tour | python examples/00_gazeforge_tour.py --output-dir gazeforge-tour-demo |
source→QC→events→AOIs→scanpaths bundle | synthetic demo ≠ empirical validation | Learning paths |
| Plan a first study | clear research question + acquisition plan | First study blueprint | existing worked static-study route | analysis contract + artifact plan | auditable workflow ≠ construct validity | Outcome/estimand preregistration |
| Freeze outcomes and estimands before modelling | research question + planned measurement definitions | Outcome & estimand preregistration clinic | python examples/13_worked_estimand_preregistration.py --output-dir worked-estimand-preregistration |
outcome + estimand + contrast + sensitivity + deviation registries | preregistration ≠ construct/causal validity; estimator remains unselected | Analysis handoff |
| Choose a method for a known task | explicit research question + available evidence | Method chooser | question/evidence/generalisation decision table | defensible workflow route | software availability ≠ scientific justification | Artifact dictionary |
| Understand output files | one or more GazeForge artifacts | Artifact & output dictionary | artifact role/unit/boundary reference | source/QC/review/analysis/validation/provenance classification | filename order ≠ evidence strength | Research evidence bundle |
| Install or check the environment | supported Python | Getting started | python -m pip install "gazeforge==0.1.0a2" |
importable environment | installation success ≠ measurement validity | Runnable examples |
| Import a tracker export | known source columns/units | Worked tracker import + QC | adapt_gazepoint_samples() / canonical schema |
source, canonical, preflight, QC tables | adapter compatibility ≠ device validity | Real-data import clinic |
| Inspect QC without deleting data | canonical gaze table | Synthetic QC tutorial | ai_flag_anomalies(), score_trial_quality() |
anomaly flags + trial-quality table | flag ≠ invalid sample | QC review & exclusion ledger |
| Review exclusions | preserved pre-review QC table | QC review & exclusion ledger | python examples/08_worked_qc_review_ledger.py --output-dir worked-qc-review-ledger-demo |
sample/trial/participant ledgers + denominator flow | reproducible rule ≠ validated rule | Study lifecycle |
| Label eye events transparently | gaze samples + sampling assumptions | I-VT baseline tutorial | ivt_classify_events() |
event-labelled samples / intervals | example threshold ≠ universal cutoff | Event-model validation clinic |
| Validate learned event models | labelled event data + valid grouping unit | Event-model validation clinic | grouped CV / calibration / event metrics | split ledger + matched held-out predictions + metrics | held-out design must match the intended claim | Validation reporting cookbook |
| Define static or dynamic AOIs | stimulus geometry or reviewed tracks | Research recipes | AOI mapping / dynamic AOI assignment | AOI definitions + assignments + review/audit | AI proposal ≠ ground truth; no silent extrapolation | Worked dynamic-AOI study |
| Build semantic scanpaths | reviewed fixation/AOI assignments | Practical workflow | to_semantic_scanpaths() |
semantic sequence table | sequence representation ≠ latent-state inference | Methods overview |
| Reconcile denominators/exposure/censoring | reviewed trial/AOI/event outputs | Denominator, exposure & censoring clinic | python examples/16_worked_denominator_exposure_audit.py --output-dir worked-denominator-exposure-audit |
status + exposure + rate/proportion + latency-censoring audit | missing/absent/undefined ≠ zero; no-fixation ≠ latency zero | Missing-data assumptions |
| Document missing-data assumptions/treatment handoff | reconciled observation states + grouping/QC context | Missing-data assumptions & treatment handoff | python examples/18_worked_missing_data_assumptions_audit.py --output-dir worked-missing-data-assumptions-audit |
source registry + mechanism questions + non-selecting treatment registry + sensitivity/reporting handoff | MCAR/MAR/MNAR not inferred; treatment not auto-selected | Analysis handoff |
| Build statistical model inputs | reviewed event/AOI outputs + preserved design/coverage | Analysis handoff | python examples/10_worked_analysis_handoff.py --output-dir worked-analysis-handoff-demo |
participant × trial × AOI/event tables + denominators + censoring | missing ≠ zero; samples/fixations are not independent participants | Grouping clinic |
| Audit grouping/repeated measures/pseudoreplication | model-ready rows + participant/trial/stimulus/AOI identities | Grouping, repeated measures & pseudoreplication clinic | python examples/20_worked_grouping_pseudoreplication_audit.py --output-dir worked-grouping-pseudoreplication-audit |
unit registry + grouping structure + independence/aggregation audit + crossed/nested handoff | repeated rows ≠ independent participants; design identity ≠ automatic model term | Model diagnostics |
| Audit what a gaze-derived measure supports | frozen measurement definitions + intended claims | Measurement & interpretation clinic | python examples/12_worked_measurement_interpretation_audit.py --output-dir worked-measurement-interpretation-audit |
claim registry + interpretation matrix + threats + sensitivity/reporting tables | observable ≠ latent construct; audit status ≠ truth label | Sensitivity & robustness clinic |
| Audit model convergence/diagnostics | fitted specialist-model objects + frozen estimand/population identity | Model diagnostics & convergence clinic | python examples/17_worked_model_diagnostics_audit.py --output-dir worked-model-diagnostics-audit |
fit registry + diagnostic state + interpretation gate + replacement linkage | returned coefficients ≠ valid fit; changed estimand ≠ replacement primary | Sensitivity & robustness clinic |
| Audit uncertainty/multiplicity before reporting | diagnostically admissible fitted results + frozen estimand/family identity | Uncertainty, multiplicity & inferential reporting clinic | python examples/19_worked_inferential_reporting_audit.py --output-dir worked-inferential-reporting-audit |
result scale + interval identity + multiplicity family + interpretation gate | converged model ≠ correctly reported inference; raw p ≠ adjusted p | Sensitivity & robustness clinic |
| Audit sensitivity/robustness after analysis | frozen primary estimand + registered/executed variants | Sensitivity & robustness clinic | python examples/15_worked_sensitivity_robustness_audit.py --output-dir worked-sensitivity-robustness-audit |
complete sensitivity registry + execution status + denominator/result comparison + deviations | consistency ≠ validity; changed estimand ≠ direct robustness check | Reporting clinic |
| Reproduce or freeze a study | finalized analysis plan + provenance | Study lifecycle | fingerprints / manifests / deterministic exports | frozen inputs, outputs, provenance | frozen software artifact ≠ external validity | Publication readiness |
| Prepare manuscript/archive evidence | finalized results and denominators | Research evidence bundle | python examples/09_worked_research_evidence_bundle.py --output-dir worked-research-evidence-bundle |
artifact index + source/QC/review/analysis/provenance layers | archive completeness ≠ empirical validity | Reporting clinic |
| Translate frozen evidence into manuscript language | frozen bundle + reconciled denominators | Reporting & interpretation clinic | python examples/11_worked_manuscript_reporting_bundle.py --output-dir worked-manuscript-reporting-bundle |
Methods/results examples + citation table + boundaries + reporting manifest | reporting prose cannot strengthen evidence | Publication readiness |
| Hand a frozen study to a reviewer/replicator | frozen archive + access/licensing status | Reviewer & replication handoff | python examples/14_worked_reviewer_replication_bundle.py --output-dir worked-reviewer-replication-bundle |
claim-artifact map + rerun plan + reproducibility class + limitations + API/hash ledger | inspectability/reruns ≠ scientific validity | Publication readiness |
| Inspect current empirical support | no prerequisite | Evidence status | generated evidence/status pages | Frozen / Reviewed / Bounded / pending status | native/derived and split/identity boundaries remain explicit | Validation status |
Worked routes¶
Route 0 · I am planning outcomes/estimands before analysis¶
- Start with the Outcome & estimand preregistration clinic.
- Register every primary, secondary, and exploratory outcome before model fitting.
- Declare the row/inferential unit, exposure/denominator, time window, missing/zero/censoring semantics, and contrast.
- Prespecify scientifically justified sensitivity checks and a deviation-ledger schema.
- Leave the estimator/model family unselected until the design and outcome support are assessed in specialist statistical software.
- Carry the registry into the Analysis handoff and Measurement & interpretation clinic.
Stop rather than guess: do not silently switch outcomes, convert missing/censored observations to zero, or promote an exploratory result to primary after viewing results.
Route A · I have a Gazepoint-style export and need an analysis table¶
- Read Worked tracker import + QC.
- Freeze the source-column mapping, timestamp unit, coordinate basis, geometry, nominal/native rate, and observed timestamp cadence.
- Use the Real-data import clinic when the export differs from the worked source contract.
- Preserve the pre-review QC table.
- Use the QC review & exclusion ledger before creating the primary-analysis derivative.
- Continue to an event/AOI/scanpath route only after denominators and exclusion decisions reconcile.
- When the analysis is frozen, use the Research evidence bundle pattern to package source identity, decisions, derivatives, and provenance separately.
Stop rather than guess: unknown units, unknown participant/trial identity, unexplained duplicate keys, or an unverified timestamp basis are source-contract problems. Do not repair them by silently coercing the table until it “looks right.”
Route B · I have labelled events and want to compare classifiers¶
- Start with the Event-model validation clinic.
- Define the grouping unit that must remain disjoint across train/test partitions.
- Compare models on matched held-out observations.
- Inspect calibration and event-level temporal behavior, not only sample accuracy.
- Use the Validation reporting cookbook to report split design, denominators, uncertainty, and limitations.
Stop rather than guess: source-token separation is not participant separation, derived 60 Hz is not native 60 Hz, and a synthetic ordering does not establish general model superiority.
Route C · I have a moving stimulus and need semantic AOIs¶
- Use the Worked dynamic-AOI study.
- Preserve reviewed keyframes and the temporal support range.
- Audit interpolation and verify no extrapolation outside reviewed support.
- Keep AI-generated boxes or tracks as proposals until the intended review policy is satisfied.
- Build scanpaths only from the reviewed AOI-assignment derivative.
Stop rather than guess: a detected object, track, or semantic label is not automatically a scientifically valid AOI for the study construct.
Route D · I need a manuscript/archive bundle a reviewer can understand¶
- Start with the Research evidence bundle.
- Open the Artifact & output dictionary when a CSV/JSON role is unclear.
- Preserve source identity and pre-review QC separately from review decisions.
- Build the primary-analysis derivative from the reviewed ledger rather than editing QC in place.
- Freeze an artifact index, analysis plan, provenance, workflow manifest, software identity, and reviewer-facing README.
- Run the Publication readiness checklist before sharing or citing the archive.
Stop rather than guess: a complete archive proves neither measurement validity nor external validity. Archive only the evidence class the study actually supports, and respect participant privacy and source licensing.
Route E · I have reviewed gaze outputs and need statistical model inputs¶
- Start with the Analysis handoff.
- Preserve participant, trial/session, condition, stimulus, and repeated-measures identity.
- Carry observed exposure/denominators into count, rate, proportion, and dwell summaries.
- Keep observed zero, absent-by-design, undefined, and missing states distinct.
- Carry no-fixation latency as explicit censoring when the AOI was observable.
- Generate descriptive participant × condition summaries separately from trial-level inferential inputs.
- Fit the prespecified inferential model in specialist statistical software; do not let the handoff choose an estimator.
Stop rather than guess: never use fillna(0) as a convenience repair, never aggregate away the inferential unit without changing the estimand explicitly, and never interpret a failed-convergence model as a valid result.
Route F · I have gaze-derived measures and need to audit interpretation¶
- Start with the Measurement & interpretation clinic.
- Name the observable separately from the proposed construct.
- Preserve missing, zero, exposure, and censoring semantics.
- Review event/AOI/sampling/QC and generalisation threats.
- Prespecify scientifically justified sensitivity checks rather than selecting favourable variants.
- Record what external outcome or construct-validation evidence is required.
- Route only the supported wording into manuscript reporting.
Stop rather than guess: dwell, fixation count, latency, scanpaths, confidence, and QC flags do not automatically establish trust, persuasion, interest, comprehension, emotion, intent, diagnosis, correctness, or scientific invalidity.
Route G · I have frozen evidence and need manuscript/supplement wording¶
- Start with the Reporting & interpretation clinic.
- Keep import compatibility separate from device validity.
- Report QC evidence separately from review/exclusion decisions.
- Name event/AOI/scanpath identity and its validation boundary.
- Preserve participant/stimulus/source-token/dataset split identity exactly.
- Keep native/nominal, observed cadence, and derived analysis rates distinct.
- Put critical evidence qualifiers in figure/table captions when the visual could be overread.
- Run Publication readiness before submission or archive release.
Stop rather than guess: prose cannot upgrade a synthetic demo into empirical evidence, derived 60 Hz into native 60 Hz, or an observable gaze pattern into trust, persuasion, comprehension, emotion, diagnosis, preference, or intent.
Route H · I need to hand the frozen study to a reviewer or replicator¶
- Start with the Reviewer & replication handoff.
- Classify the archive as
fully_rerunnable_demo,rerunnable_with_private_input, orinspectable_only. - Map each material statement to an exact artifact, access requirement, API/guide route, and interpretation boundary.
- Record exact software/version/full-commit identity plus material environment/configuration information.
- State whether public, private, or restricted inputs are required and never imply that unavailable/restricted source files are bundled.
- Reconcile denominators, exclusions, missing-versus-zero states, censoring, and preregistered outcome/estimand identity before sharing.
- Carry the limitations register and evidence class with the archive.
- Run Publication readiness before release.
Stop rather than guess: matching hashes, deterministic reruns, or reviewer inspectability do not establish device, model, measurement/construct, causal, external, or latent-state validity.
Route I · I have executed sensitivity variants and need a complete robustness audit¶
- Start with the Sensitivity & robustness clinic.
- Keep the registered primary specification as the reference.
- Separate prespecified sensitivity, exploratory sensitivity, and post-registration deviations.
- Verify whether each variant retains the same estimand before comparing it directly with the primary result.
- Report denominator/exposure changes alongside result changes.
- Keep
not_evaluableandnon_convergedconditions in the audit record. - Report the full registered set rather than selecting favourable variants.
- Continue to the Reporting clinic only after the complete sensitivity record is frozen.
Stop rather than guess: do not replace the primary result with a favourable variant, convert a changed-estimand deviation into a direct robustness claim, drop failed variants, or infer scientific validity from specification consistency.
Documentation types¶
Tutorial · first success¶
Use tutorials when you are learning the package and want a controlled path that works end to end. Start with the GazeForge Tour, Synthetic QC tutorial, or I-VT baseline tutorial.
How-to guide · complete a task¶
Use how-to guides when you already know the outcome you need: import a real export, review exclusions, validate a model, audit dynamic AOIs, freeze a study, build an evidence bundle, or prepare a manuscript.
Explanation · understand a method or boundary¶
Use method and governance pages when you need the reasoning behind a choice: Methods overview, Scientific governance, event-level evaluation, calibration, sampling sensitivity, or benchmark-specific evidence pages.
Reference · exact API/status¶
Use API reference, Evidence status, Validation status, and generated benchmark/status pages when you need exact current interfaces or evidence classifications rather than a teaching sequence.
When you are stuck¶
Use the same help order on task pages. The persistent help row at the top of this page links to the Tour, this task map, troubleshooting, runnable examples, and the issue tracker.
The Troubleshooting & Diagnostics guide is for failures, surprising output, uncertain source contracts, and minimal reproducible issue reports. If the problem is a scientific-evidence question rather than a software failure, use Evidence status or the relevant validation page instead.