Skip to content

Frozen benchmark evidence

Verified report suites

A suite appears here only when its completion manifest and every referenced child report verify successfully. VISUS suites additionally require a verified raw-execution provenance bundle.

suite status report_count target_sampling_rate_hz model reference_stream_id human_human_agreement_included source_manifest_fingerprint_sha256 suite_fingerprint_sha256
lund2013-event-validation-v1 complete 5 60 45db53af11a2 5dc6d6336b50

Frozen reports

Only reports whose deterministic fingerprint recomputes successfully are listed.

benchmark version annotation_origin sampling_origin reference_strength sampling_rates_hz models report_fingerprint_sha256
Hollywood2EM 870fa6d6209c9085260918d61433a0a2c70fd497 human-assisted resampled derived-human-reference 500, 60 I-VT, RandomForest, ContextMLP e1f1c030f843
Lund2013 Andersson-et-al-2017-public-repository expert-manual resampled derived-human-reference 500, 60 I-VT, RandomForest, ContextMLP 0c81c902e1f8
Lund2013 Andersson-et-al-2017-public-repository expert-manual resampled derived-human-reference 500, 60 I-VT, RandomForest, ContextMLP 4642fc76aaaf
Lund2013-human-agreement Andersson-et-al-2017-public-repository expert-manual resampled derived-human-reference 60 0524d212e59b
Lund2013-human-agreement Andersson-et-al-2017-public-repository expert-manual native expert-human-reference 500 52b070a91fcc
Lund2013-sampling-sensitivity Andersson-et-al-2017-public-repository expert-manual resampled derived-human-reference 500, 120, 90, 60, 30 I-VT, RandomForest, ContextMLP b01005c0df0a

Validated report details

The tables below are generated directly from the same fingerprint-validated JSON reports listed above; no performance values are transcribed manually.

Hollywood2EM

Sampling: resampled · Reference: derived-human-reference · Report: e1f1c030f843

Overall held-out model performance

Model Folds Accuracy Balanced acc. Macro-F1 Event F1 Event IoU Brier ECE
I-VT 4 0.701 0.539 0.560 0.629 0.855
RandomForest 4 0.753 0.728 0.740 0.440 0.860 0.342 0.023
ContextMLP 4 0.817 0.794 0.812 0.602 0.883 0.263 0.014

Lund2013

Sampling: resampled · Reference: derived-human-reference · Report: 0c81c902e1f8

Overall held-out model performance

Model Folds Accuracy Balanced acc. Macro-F1 Event F1 Event IoU Brier ECE
I-VT 5 0.682 0.396 0.301 0.624 0.922
RandomForest 5 0.699 0.641 0.574 0.471 0.890 0.411 0.086
ContextMLP 5 0.732 0.688 0.629 0.582 0.894 0.410 0.149

Matched-fold model differences

Positive Mean improvement A always favours model A; raw Mean A−B keeps the original metric direction. These are descriptive matched-fold differences, not cross-validation significance tests.

Model A Model B Metric Paired folds Mean A−B Mean improvement A A wins Ties B wins
I-VT RandomForest accuracy 5 -0.017 -0.017 1 0 4
I-VT RandomForest macro_f1 5 -0.274 -0.274 0 0 5
I-VT RandomForest multiclass_brier_score 0 0 0 0
I-VT RandomForest expected_calibration_error 0 0 0 0
I-VT RandomForest event_f1 5 0.153 0.153 5 0 0
I-VT RandomForest event_mean_matched_iou 5 0.032 0.032 5 0 0
I-VT ContextMLP accuracy 5 -0.049 -0.049 1 0 4
I-VT ContextMLP macro_f1 5 -0.329 -0.329 0 0 5
I-VT ContextMLP multiclass_brier_score 0 0 0 0
I-VT ContextMLP expected_calibration_error 0 0 0 0
I-VT ContextMLP event_f1 5 0.042 0.042 3 0 2
I-VT ContextMLP event_mean_matched_iou 5 0.028 0.028 5 0 0
RandomForest ContextMLP accuracy 5 -0.033 -0.033 2 0 3
RandomForest ContextMLP macro_f1 5 -0.055 -0.055 0 0 5
RandomForest ContextMLP multiclass_brier_score 5 0.001 -0.001 3 0 2
RandomForest ContextMLP expected_calibration_error 5 -0.062 0.062 5 0 0
RandomForest ContextMLP event_f1 5 -0.111 -0.111 0 0 5
RandomForest ContextMLP event_mean_matched_iou 5 -0.004 -0.004 2 0 3

Performance by stimulus family

These are post-hoc summaries of the same held-out predictions used above; models were not refitted by stimulus family.

Stimulus family Model Folds Held-out rows Participants Accuracy Macro-F1 Event F1 Event IoU
image ContextMLP 5 7179 13 0.839 0.546 0.647 0.903
moving_dot ContextMLP 5 1247 10 0.641 0.493 0.359 0.871
video ContextMLP 5 3337 9 0.572 0.604 0.462 0.857
image I-VT 5 7179 13 0.889 0.431 0.710 0.924
moving_dot I-VT 5 1247 10 0.124 0.165 0.242 0.950
video I-VT 5 3337 9 0.442 0.235 0.493 0.937
image RandomForest 5 7179 13 0.831 0.522 0.569 0.891
moving_dot RandomForest 5 1247 10 0.419 0.339 0.210 0.892
video RandomForest 5 3337 9 0.547 0.546 0.343 0.858

Lund2013-human-agreement

Sampling: resampled · Reference: derived-human-reference · Report: 0524d212e59b

Human–human annotation agreement

Scope Aligned samples Exact agreement Cohen κ
overall 12481 0.880 0.799
image 7668 0.917 0.799
moving_dot 1325 0.875 0.694
video 3488 0.802 0.672

Lund2013-human-agreement

Sampling: native · Reference: expert-human-reference · Report: 52b070a91fcc

Human–human annotation agreement

Scope Aligned samples Exact agreement Cohen κ
overall 103878 0.893 0.815
image 63849 0.932 0.822
moving_dot 10997 0.885 0.702
video 29032 0.812 0.679

Lund2013

Sampling: resampled · Reference: derived-human-reference · Report: 4642fc76aaaf

Overall held-out model performance

Model Folds Accuracy Balanced acc. Macro-F1 Event F1 Event IoU Brier ECE
I-VT 5 0.637 0.388 0.287 0.626 0.921
RandomForest 5 0.676 0.670 0.595 0.440 0.892 0.441 0.076
ContextMLP 5 0.694 0.679 0.649 0.535 0.900 0.455 0.160

Matched-fold model differences

Positive Mean improvement A always favours model A; raw Mean A−B keeps the original metric direction. These are descriptive matched-fold differences, not cross-validation significance tests.

Model A Model B Metric Paired folds Mean A−B Mean improvement A A wins Ties B wins
I-VT RandomForest accuracy 5 -0.039 -0.039 2 0 3
I-VT RandomForest macro_f1 5 -0.308 -0.308 0 0 5
I-VT RandomForest multiclass_brier_score 0 0 0 0
I-VT RandomForest expected_calibration_error 0 0 0 0
I-VT RandomForest event_f1 5 0.186 0.186 5 0 0
I-VT RandomForest event_mean_matched_iou 5 0.029 0.029 5 0 0
I-VT ContextMLP accuracy 5 -0.057 -0.057 1 0 4
I-VT ContextMLP macro_f1 5 -0.362 -0.362 0 0 5
I-VT ContextMLP multiclass_brier_score 0 0 0 0
I-VT ContextMLP expected_calibration_error 0 0 0 0
I-VT ContextMLP event_f1 5 0.091 0.091 5 0 0
I-VT ContextMLP event_mean_matched_iou 5 0.021 0.021 4 0 1
RandomForest ContextMLP accuracy 5 -0.018 -0.018 1 0 4
RandomForest ContextMLP macro_f1 5 -0.054 -0.054 1 0 4
RandomForest ContextMLP multiclass_brier_score 5 -0.014 0.014 2 0 3
RandomForest ContextMLP expected_calibration_error 5 -0.084 0.084 5 0 0
RandomForest ContextMLP event_f1 5 -0.095 -0.095 0 0 5
RandomForest ContextMLP event_mean_matched_iou 5 -0.008 -0.008 1 0 4

Performance by stimulus family

These are post-hoc summaries of the same held-out predictions used above; models were not refitted by stimulus family.

Stimulus family Model Folds Held-out rows Participants Accuracy Macro-F1 Event F1 Event IoU
image ContextMLP 5 7199 13 0.767 0.565 0.594 0.899
moving_dot ContextMLP 5 1251 10 0.699 0.596 0.426 0.944
video ContextMLP 5 3328 9 0.623 0.593 0.472 0.878
image I-VT 5 7199 13 0.861 0.368 0.701 0.918
moving_dot I-VT 5 1251 10 0.153 0.232 0.335 0.916
video I-VT 5 3328 9 0.382 0.247 0.477 0.932
image RandomForest 5 7199 13 0.766 0.527 0.525 0.898
moving_dot RandomForest 5 1251 10 0.421 0.392 0.222 0.877
video RandomForest 5 3328 9 0.630 0.558 0.332 0.871

Lund2013-sampling-sensitivity

Sampling: resampled · Reference: derived-human-reference · Report: b01005c0df0a

Sampling × label-purity settings

Rate Hz Min purity Status Ambiguous Retained Participants
120.0 0.600 ok 0.013 0.985 20
120.0 0.750 ok 0.019 0.980 20
120.0 0.900 ok 0.049 0.950 20
90.000 0.600 ok 0.009 0.989 20
90.000 0.750 ok 0.039 0.960 20
90.000 0.900 ok 0.069 0.929 20
60.000 0.600 ok 0.024 0.974 20
60.000 0.750 ok 0.055 0.944 20
60.000 0.900 ok 0.110 0.889 20
30.000 0.600 ok 0.071 0.928 20
30.000 0.750 ok 0.121 0.878 20
30.000 0.900 ok 0.177 0.822 20

Model sensitivity surface

Rate Hz Min purity Model Macro-F1 Event F1 Event IoU Ambiguous Retained
120.0 0.600 ContextMLP 0.671 0.479 0.840 0.013 0.985
120.0 0.600 I-VT 0.244 0.528 0.874 0.013 0.985
120.0 0.600 RandomForest 0.626 0.362 0.830 0.013 0.985
120.0 0.750 ContextMLP 0.667 0.499 0.855 0.019 0.980
120.0 0.750 I-VT 0.247 0.535 0.884 0.019 0.980
120.0 0.750 RandomForest 0.623 0.379 0.841 0.019 0.980
120.0 0.900 ContextMLP 0.664 0.492 0.896 0.049 0.950
120.0 0.900 I-VT 0.280 0.531 0.936 0.049 0.950
120.0 0.900 RandomForest 0.604 0.388 0.885 0.049 0.950
90.000 0.600 ContextMLP 0.679 0.511 0.826 0.009 0.989
90.000 0.600 I-VT 0.244 0.568 0.851 0.009 0.989
90.000 0.600 RandomForest 0.619 0.422 0.828 0.009 0.989
90.000 0.750 ContextMLP 0.648 0.527 0.878 0.039 0.960
90.000 0.750 I-VT 0.278 0.584 0.907 0.039 0.960
90.000 0.750 RandomForest 0.610 0.432 0.873 0.039 0.960
90.000 0.900 ContextMLP 0.642 0.497 0.916 0.069 0.929
90.000 0.900 I-VT 0.278 0.575 0.956 0.069 0.929
90.000 0.900 RandomForest 0.599 0.406 0.900 0.069 0.929
60.000 0.600 ContextMLP 0.635 0.540 0.845 0.024 0.974
60.000 0.600 I-VT 0.253 0.597 0.889 0.024 0.974
60.000 0.600 RandomForest 0.611 0.441 0.849 0.024 0.974
60.000 0.750 ContextMLP 0.649 0.535 0.900 0.055 0.944
60.000 0.750 I-VT 0.287 0.626 0.921 0.055 0.944
60.000 0.750 RandomForest 0.595 0.440 0.892 0.055 0.944
60.000 0.900 ContextMLP 0.661 0.544 0.940 0.110 0.889
60.000 0.900 I-VT 0.243 0.632 0.967 0.110 0.889
60.000 0.900 RandomForest 0.609 0.414 0.918 0.110 0.889
30.000 0.600 ContextMLP 0.588 0.546 0.892 0.071 0.928
30.000 0.600 I-VT 0.259 0.635 0.913 0.071 0.928
30.000 0.600 RandomForest 0.521 0.435 0.889 0.071 0.928
30.000 0.750 ContextMLP 0.608 0.582 0.924 0.121 0.878
30.000 0.750 I-VT 0.250 0.659 0.930 0.121 0.878
30.000 0.750 RandomForest 0.514 0.436 0.909 0.121 0.878
30.000 0.900 ContextMLP 0.578 0.557 0.930 0.177 0.822
30.000 0.900 I-VT 0.228 0.646 0.948 0.177 0.822
30.000 0.900 RandomForest 0.524 0.439 0.930 0.177 0.822

What appears on this page

A JSON file is listed here only when it follows the GazeForge frozen benchmark-report schema and its deterministic SHA-256 fingerprint recomputes successfully from the benchmark metadata, model metadata, protocol, and metrics. Candidate protocols and configuration manifests are not treated as performance evidence.

Evidence interpretation

The table surfaces annotation origin, sampling origin, reference strength, model family, and sampling rate so evidence strength remains visible alongside any future performance result. Derived lower-rate evidence is therefore distinguishable from native-rate recordings, and algorithmic/vendor labels cannot silently appear as human ground truth.

Detailed performance tables are generated only from reports that passed the same integrity check. Unknown future report schemas remain visible in the frozen-report index without GazeForge guessing which nested values should be presented as headline performance metrics.

Current scientific rule

The absence of a row is meaningful: implemented benchmark infrastructure, adapters, candidate datasets, and synthetic smoke tests do not become empirical validation merely because they exist in the repository. See the validation status and benchmark evidence pages for work that is implemented but not yet frozen as empirical evidence.