Inspect the stored artifact, then independently label whether the frozen task
should pass. The default policy balances domains and evaluator outcomes without
revealing those outcomes during labeling.
Labels are stored locally with full artifact provenance and are
automatically included by benchmark and calibration commands using this corpus.
Queue sampling policy v1 caps
each source run at 4 items.
Actor-comparison readiness: BLOCKED
0 qualifying case(s) ·
0 excluded · policy v1
need at least 24 qualifying cases; found 0
landing-page: need 8 cases; found 0
landing-page: need 3 expected-pass cases; found 0
landing-page: need 3 expected-fail cases; found 0
landing-page: need 3 distinct source runs; found 0
lead-generation: need 8 cases; found 0
lead-generation: need 3 expected-pass cases; found 0
lead-generation: need 3 expected-fail cases; found 0
lead-generation: need 3 distinct source runs; found 0
run #7 · iter 1 ·
landing-page · landing-primary-desktop
Primary action works with a pointer on desktop
Find the page's primary call to action, activate it, and observe a meaningful response such as navigation, scrolling, a dialog, a new window, or changed visible content.
Evaluator observation and runtime evidence are hidden for independent labeling.
run #7 · iter 2 ·
landing-page · landing-primary-desktop
Primary action works with a pointer on desktop
Find the page's primary call to action, activate it, and observe a meaningful response such as navigation, scrolling, a dialog, a new window, or changed visible content.
Evaluator observation and runtime evidence are hidden for independent labeling.