Frozen benchmark evidence¶
Verified report suites¶
A suite appears here only when its completion manifest and every referenced child report verify successfully. VISUS suites additionally require a verified raw-execution provenance bundle.
| suite | status | report_count | target_sampling_rate_hz | model | reference_stream_id | human_human_agreement_included | source_manifest_fingerprint_sha256 | suite_fingerprint_sha256 |
|---|---|---|---|---|---|---|---|---|
| lund2013-event-validation-v1 | complete | 5 | 60 | 45db53af11a2 | 5dc6d6336b50 |
Frozen reports¶
Only reports whose deterministic fingerprint recomputes successfully are listed.
| benchmark | version | annotation_origin | sampling_origin | reference_strength | sampling_rates_hz | models | report_fingerprint_sha256 |
|---|---|---|---|---|---|---|---|
| Hollywood2EM | 870fa6d6209c9085260918d61433a0a2c70fd497 | human-assisted | resampled | derived-human-reference | 500, 60 | I-VT, RandomForest, ContextMLP | e1f1c030f843 |
| Lund2013 | Andersson-et-al-2017-public-repository | expert-manual | resampled | derived-human-reference | 500, 60 | I-VT, RandomForest, ContextMLP | 0c81c902e1f8 |
| Lund2013 | Andersson-et-al-2017-public-repository | expert-manual | resampled | derived-human-reference | 500, 60 | I-VT, RandomForest, ContextMLP | 4642fc76aaaf |
| Lund2013-human-agreement | Andersson-et-al-2017-public-repository | expert-manual | resampled | derived-human-reference | 60 | 0524d212e59b | |
| Lund2013-human-agreement | Andersson-et-al-2017-public-repository | expert-manual | native | expert-human-reference | 500 | 52b070a91fcc | |
| Lund2013-sampling-sensitivity | Andersson-et-al-2017-public-repository | expert-manual | resampled | derived-human-reference | 500, 120, 90, 60, 30 | I-VT, RandomForest, ContextMLP | b01005c0df0a |
Validated report details¶
The tables below are generated directly from the same fingerprint-validated JSON reports listed above; no performance values are transcribed manually.
Hollywood2EM¶
Sampling: resampled · Reference: derived-human-reference · Report: e1f1c030f843
Overall held-out model performance¶
| Model | Folds | Accuracy | Balanced acc. | Macro-F1 | Event F1 | Event IoU | Brier | ECE |
|---|---|---|---|---|---|---|---|---|
| I-VT | 4 | 0.701 | 0.539 | 0.560 | 0.629 | 0.855 | — | — |
| RandomForest | 4 | 0.753 | 0.728 | 0.740 | 0.440 | 0.860 | 0.342 | 0.023 |
| ContextMLP | 4 | 0.817 | 0.794 | 0.812 | 0.602 | 0.883 | 0.263 | 0.014 |
Lund2013¶
Sampling: resampled · Reference: derived-human-reference · Report: 0c81c902e1f8
Overall held-out model performance¶
| Model | Folds | Accuracy | Balanced acc. | Macro-F1 | Event F1 | Event IoU | Brier | ECE |
|---|---|---|---|---|---|---|---|---|
| I-VT | 5 | 0.682 | 0.396 | 0.301 | 0.624 | 0.922 | — | — |
| RandomForest | 5 | 0.699 | 0.641 | 0.574 | 0.471 | 0.890 | 0.411 | 0.086 |
| ContextMLP | 5 | 0.732 | 0.688 | 0.629 | 0.582 | 0.894 | 0.410 | 0.149 |
Matched-fold model differences¶
Positive Mean improvement A always favours model A; raw Mean A−B keeps the original metric direction. These are descriptive matched-fold differences, not cross-validation significance tests.
| Model A | Model B | Metric | Paired folds | Mean A−B | Mean improvement A | A wins | Ties | B wins |
|---|---|---|---|---|---|---|---|---|
| I-VT | RandomForest | accuracy | 5 | -0.017 | -0.017 | 1 | 0 | 4 |
| I-VT | RandomForest | macro_f1 | 5 | -0.274 | -0.274 | 0 | 0 | 5 |
| I-VT | RandomForest | multiclass_brier_score | 0 | — | — | 0 | 0 | 0 |
| I-VT | RandomForest | expected_calibration_error | 0 | — | — | 0 | 0 | 0 |
| I-VT | RandomForest | event_f1 | 5 | 0.153 | 0.153 | 5 | 0 | 0 |
| I-VT | RandomForest | event_mean_matched_iou | 5 | 0.032 | 0.032 | 5 | 0 | 0 |
| I-VT | ContextMLP | accuracy | 5 | -0.049 | -0.049 | 1 | 0 | 4 |
| I-VT | ContextMLP | macro_f1 | 5 | -0.329 | -0.329 | 0 | 0 | 5 |
| I-VT | ContextMLP | multiclass_brier_score | 0 | — | — | 0 | 0 | 0 |
| I-VT | ContextMLP | expected_calibration_error | 0 | — | — | 0 | 0 | 0 |
| I-VT | ContextMLP | event_f1 | 5 | 0.042 | 0.042 | 3 | 0 | 2 |
| I-VT | ContextMLP | event_mean_matched_iou | 5 | 0.028 | 0.028 | 5 | 0 | 0 |
| RandomForest | ContextMLP | accuracy | 5 | -0.033 | -0.033 | 2 | 0 | 3 |
| RandomForest | ContextMLP | macro_f1 | 5 | -0.055 | -0.055 | 0 | 0 | 5 |
| RandomForest | ContextMLP | multiclass_brier_score | 5 | 0.001 | -0.001 | 3 | 0 | 2 |
| RandomForest | ContextMLP | expected_calibration_error | 5 | -0.062 | 0.062 | 5 | 0 | 0 |
| RandomForest | ContextMLP | event_f1 | 5 | -0.111 | -0.111 | 0 | 0 | 5 |
| RandomForest | ContextMLP | event_mean_matched_iou | 5 | -0.004 | -0.004 | 2 | 0 | 3 |
Performance by stimulus family¶
These are post-hoc summaries of the same held-out predictions used above; models were not refitted by stimulus family.
| Stimulus family | Model | Folds | Held-out rows | Participants | Accuracy | Macro-F1 | Event F1 | Event IoU |
|---|---|---|---|---|---|---|---|---|
| image | ContextMLP | 5 | 7179 | 13 | 0.839 | 0.546 | 0.647 | 0.903 |
| moving_dot | ContextMLP | 5 | 1247 | 10 | 0.641 | 0.493 | 0.359 | 0.871 |
| video | ContextMLP | 5 | 3337 | 9 | 0.572 | 0.604 | 0.462 | 0.857 |
| image | I-VT | 5 | 7179 | 13 | 0.889 | 0.431 | 0.710 | 0.924 |
| moving_dot | I-VT | 5 | 1247 | 10 | 0.124 | 0.165 | 0.242 | 0.950 |
| video | I-VT | 5 | 3337 | 9 | 0.442 | 0.235 | 0.493 | 0.937 |
| image | RandomForest | 5 | 7179 | 13 | 0.831 | 0.522 | 0.569 | 0.891 |
| moving_dot | RandomForest | 5 | 1247 | 10 | 0.419 | 0.339 | 0.210 | 0.892 |
| video | RandomForest | 5 | 3337 | 9 | 0.547 | 0.546 | 0.343 | 0.858 |
Lund2013-human-agreement¶
Sampling: resampled · Reference: derived-human-reference · Report: 0524d212e59b
Human–human annotation agreement¶
| Scope | Aligned samples | Exact agreement | Cohen κ |
|---|---|---|---|
| overall | 12481 | 0.880 | 0.799 |
| image | 7668 | 0.917 | 0.799 |
| moving_dot | 1325 | 0.875 | 0.694 |
| video | 3488 | 0.802 | 0.672 |
Lund2013-human-agreement¶
Sampling: native · Reference: expert-human-reference · Report: 52b070a91fcc
Human–human annotation agreement¶
| Scope | Aligned samples | Exact agreement | Cohen κ |
|---|---|---|---|
| overall | 103878 | 0.893 | 0.815 |
| image | 63849 | 0.932 | 0.822 |
| moving_dot | 10997 | 0.885 | 0.702 |
| video | 29032 | 0.812 | 0.679 |
Lund2013¶
Sampling: resampled · Reference: derived-human-reference · Report: 4642fc76aaaf
Overall held-out model performance¶
| Model | Folds | Accuracy | Balanced acc. | Macro-F1 | Event F1 | Event IoU | Brier | ECE |
|---|---|---|---|---|---|---|---|---|
| I-VT | 5 | 0.637 | 0.388 | 0.287 | 0.626 | 0.921 | — | — |
| RandomForest | 5 | 0.676 | 0.670 | 0.595 | 0.440 | 0.892 | 0.441 | 0.076 |
| ContextMLP | 5 | 0.694 | 0.679 | 0.649 | 0.535 | 0.900 | 0.455 | 0.160 |
Matched-fold model differences¶
Positive Mean improvement A always favours model A; raw Mean A−B keeps the original metric direction. These are descriptive matched-fold differences, not cross-validation significance tests.
| Model A | Model B | Metric | Paired folds | Mean A−B | Mean improvement A | A wins | Ties | B wins |
|---|---|---|---|---|---|---|---|---|
| I-VT | RandomForest | accuracy | 5 | -0.039 | -0.039 | 2 | 0 | 3 |
| I-VT | RandomForest | macro_f1 | 5 | -0.308 | -0.308 | 0 | 0 | 5 |
| I-VT | RandomForest | multiclass_brier_score | 0 | — | — | 0 | 0 | 0 |
| I-VT | RandomForest | expected_calibration_error | 0 | — | — | 0 | 0 | 0 |
| I-VT | RandomForest | event_f1 | 5 | 0.186 | 0.186 | 5 | 0 | 0 |
| I-VT | RandomForest | event_mean_matched_iou | 5 | 0.029 | 0.029 | 5 | 0 | 0 |
| I-VT | ContextMLP | accuracy | 5 | -0.057 | -0.057 | 1 | 0 | 4 |
| I-VT | ContextMLP | macro_f1 | 5 | -0.362 | -0.362 | 0 | 0 | 5 |
| I-VT | ContextMLP | multiclass_brier_score | 0 | — | — | 0 | 0 | 0 |
| I-VT | ContextMLP | expected_calibration_error | 0 | — | — | 0 | 0 | 0 |
| I-VT | ContextMLP | event_f1 | 5 | 0.091 | 0.091 | 5 | 0 | 0 |
| I-VT | ContextMLP | event_mean_matched_iou | 5 | 0.021 | 0.021 | 4 | 0 | 1 |
| RandomForest | ContextMLP | accuracy | 5 | -0.018 | -0.018 | 1 | 0 | 4 |
| RandomForest | ContextMLP | macro_f1 | 5 | -0.054 | -0.054 | 1 | 0 | 4 |
| RandomForest | ContextMLP | multiclass_brier_score | 5 | -0.014 | 0.014 | 2 | 0 | 3 |
| RandomForest | ContextMLP | expected_calibration_error | 5 | -0.084 | 0.084 | 5 | 0 | 0 |
| RandomForest | ContextMLP | event_f1 | 5 | -0.095 | -0.095 | 0 | 0 | 5 |
| RandomForest | ContextMLP | event_mean_matched_iou | 5 | -0.008 | -0.008 | 1 | 0 | 4 |
Performance by stimulus family¶
These are post-hoc summaries of the same held-out predictions used above; models were not refitted by stimulus family.
| Stimulus family | Model | Folds | Held-out rows | Participants | Accuracy | Macro-F1 | Event F1 | Event IoU |
|---|---|---|---|---|---|---|---|---|
| image | ContextMLP | 5 | 7199 | 13 | 0.767 | 0.565 | 0.594 | 0.899 |
| moving_dot | ContextMLP | 5 | 1251 | 10 | 0.699 | 0.596 | 0.426 | 0.944 |
| video | ContextMLP | 5 | 3328 | 9 | 0.623 | 0.593 | 0.472 | 0.878 |
| image | I-VT | 5 | 7199 | 13 | 0.861 | 0.368 | 0.701 | 0.918 |
| moving_dot | I-VT | 5 | 1251 | 10 | 0.153 | 0.232 | 0.335 | 0.916 |
| video | I-VT | 5 | 3328 | 9 | 0.382 | 0.247 | 0.477 | 0.932 |
| image | RandomForest | 5 | 7199 | 13 | 0.766 | 0.527 | 0.525 | 0.898 |
| moving_dot | RandomForest | 5 | 1251 | 10 | 0.421 | 0.392 | 0.222 | 0.877 |
| video | RandomForest | 5 | 3328 | 9 | 0.630 | 0.558 | 0.332 | 0.871 |
Lund2013-sampling-sensitivity¶
Sampling: resampled · Reference: derived-human-reference · Report: b01005c0df0a
Sampling × label-purity settings¶
| Rate Hz | Min purity | Status | Ambiguous | Retained | Participants |
|---|---|---|---|---|---|
| 120.0 | 0.600 | ok | 0.013 | 0.985 | 20 |
| 120.0 | 0.750 | ok | 0.019 | 0.980 | 20 |
| 120.0 | 0.900 | ok | 0.049 | 0.950 | 20 |
| 90.000 | 0.600 | ok | 0.009 | 0.989 | 20 |
| 90.000 | 0.750 | ok | 0.039 | 0.960 | 20 |
| 90.000 | 0.900 | ok | 0.069 | 0.929 | 20 |
| 60.000 | 0.600 | ok | 0.024 | 0.974 | 20 |
| 60.000 | 0.750 | ok | 0.055 | 0.944 | 20 |
| 60.000 | 0.900 | ok | 0.110 | 0.889 | 20 |
| 30.000 | 0.600 | ok | 0.071 | 0.928 | 20 |
| 30.000 | 0.750 | ok | 0.121 | 0.878 | 20 |
| 30.000 | 0.900 | ok | 0.177 | 0.822 | 20 |
Model sensitivity surface¶
| Rate Hz | Min purity | Model | Macro-F1 | Event F1 | Event IoU | Ambiguous | Retained |
|---|---|---|---|---|---|---|---|
| 120.0 | 0.600 | ContextMLP | 0.671 | 0.479 | 0.840 | 0.013 | 0.985 |
| 120.0 | 0.600 | I-VT | 0.244 | 0.528 | 0.874 | 0.013 | 0.985 |
| 120.0 | 0.600 | RandomForest | 0.626 | 0.362 | 0.830 | 0.013 | 0.985 |
| 120.0 | 0.750 | ContextMLP | 0.667 | 0.499 | 0.855 | 0.019 | 0.980 |
| 120.0 | 0.750 | I-VT | 0.247 | 0.535 | 0.884 | 0.019 | 0.980 |
| 120.0 | 0.750 | RandomForest | 0.623 | 0.379 | 0.841 | 0.019 | 0.980 |
| 120.0 | 0.900 | ContextMLP | 0.664 | 0.492 | 0.896 | 0.049 | 0.950 |
| 120.0 | 0.900 | I-VT | 0.280 | 0.531 | 0.936 | 0.049 | 0.950 |
| 120.0 | 0.900 | RandomForest | 0.604 | 0.388 | 0.885 | 0.049 | 0.950 |
| 90.000 | 0.600 | ContextMLP | 0.679 | 0.511 | 0.826 | 0.009 | 0.989 |
| 90.000 | 0.600 | I-VT | 0.244 | 0.568 | 0.851 | 0.009 | 0.989 |
| 90.000 | 0.600 | RandomForest | 0.619 | 0.422 | 0.828 | 0.009 | 0.989 |
| 90.000 | 0.750 | ContextMLP | 0.648 | 0.527 | 0.878 | 0.039 | 0.960 |
| 90.000 | 0.750 | I-VT | 0.278 | 0.584 | 0.907 | 0.039 | 0.960 |
| 90.000 | 0.750 | RandomForest | 0.610 | 0.432 | 0.873 | 0.039 | 0.960 |
| 90.000 | 0.900 | ContextMLP | 0.642 | 0.497 | 0.916 | 0.069 | 0.929 |
| 90.000 | 0.900 | I-VT | 0.278 | 0.575 | 0.956 | 0.069 | 0.929 |
| 90.000 | 0.900 | RandomForest | 0.599 | 0.406 | 0.900 | 0.069 | 0.929 |
| 60.000 | 0.600 | ContextMLP | 0.635 | 0.540 | 0.845 | 0.024 | 0.974 |
| 60.000 | 0.600 | I-VT | 0.253 | 0.597 | 0.889 | 0.024 | 0.974 |
| 60.000 | 0.600 | RandomForest | 0.611 | 0.441 | 0.849 | 0.024 | 0.974 |
| 60.000 | 0.750 | ContextMLP | 0.649 | 0.535 | 0.900 | 0.055 | 0.944 |
| 60.000 | 0.750 | I-VT | 0.287 | 0.626 | 0.921 | 0.055 | 0.944 |
| 60.000 | 0.750 | RandomForest | 0.595 | 0.440 | 0.892 | 0.055 | 0.944 |
| 60.000 | 0.900 | ContextMLP | 0.661 | 0.544 | 0.940 | 0.110 | 0.889 |
| 60.000 | 0.900 | I-VT | 0.243 | 0.632 | 0.967 | 0.110 | 0.889 |
| 60.000 | 0.900 | RandomForest | 0.609 | 0.414 | 0.918 | 0.110 | 0.889 |
| 30.000 | 0.600 | ContextMLP | 0.588 | 0.546 | 0.892 | 0.071 | 0.928 |
| 30.000 | 0.600 | I-VT | 0.259 | 0.635 | 0.913 | 0.071 | 0.928 |
| 30.000 | 0.600 | RandomForest | 0.521 | 0.435 | 0.889 | 0.071 | 0.928 |
| 30.000 | 0.750 | ContextMLP | 0.608 | 0.582 | 0.924 | 0.121 | 0.878 |
| 30.000 | 0.750 | I-VT | 0.250 | 0.659 | 0.930 | 0.121 | 0.878 |
| 30.000 | 0.750 | RandomForest | 0.514 | 0.436 | 0.909 | 0.121 | 0.878 |
| 30.000 | 0.900 | ContextMLP | 0.578 | 0.557 | 0.930 | 0.177 | 0.822 |
| 30.000 | 0.900 | I-VT | 0.228 | 0.646 | 0.948 | 0.177 | 0.822 |
| 30.000 | 0.900 | RandomForest | 0.524 | 0.439 | 0.930 | 0.177 | 0.822 |
What appears on this page¶
A JSON file is listed here only when it follows the GazeForge frozen benchmark-report schema and its deterministic SHA-256 fingerprint recomputes successfully from the benchmark metadata, model metadata, protocol, and metrics. Candidate protocols and configuration manifests are not treated as performance evidence.
Evidence interpretation¶
The table surfaces annotation origin, sampling origin, reference strength, model family, and sampling rate so evidence strength remains visible alongside any future performance result. Derived lower-rate evidence is therefore distinguishable from native-rate recordings, and algorithmic/vendor labels cannot silently appear as human ground truth.
Detailed performance tables are generated only from reports that passed the same integrity check. Unknown future report schemas remain visible in the frozen-report index without GazeForge guessing which nested values should be presented as headline performance metrics.
Current scientific rule¶
The absence of a row is meaningful: implemented benchmark infrastructure, adapters, candidate datasets, and synthetic smoke tests do not become empirical validation merely because they exist in the repository. See the validation status and benchmark evidence pages for work that is implemented but not yet frozen as empirical evidence.