TASTE: measure research judgment without mistaking agreement for discovery

#topic #rsi #original-evaluation

TASTE: measure research judgment without mistaking agreement for discovery

Primary Baig, Joren and Benton, TASTE paper and authors’ shorter post, 28 August 2026, rechecked 4 October. Related benchmark question and feedback-chain test.

Design: Claude-generated proposals, based on 93 seeds from human-authored safety proposals, are evaluated by ten safety researchers, each with 6 months–4 years of experience; four rate each set of three then discuss disagreements in pairs and re-rate. Of 135 generated proposals, the final benchmark contains 92 preference pairs drawn from just 50 proposals; an anchor rater had strong post-discussion confidence and rated the two proposals at least two points apart on a five-point scale. Some proposals recur in up to ten pairs. The estimated human 77% refers to agreement with other raters’ preferences, not later success of projects; different-prompt comparison sometimes draws one held-out rater per proposal and excludes ties. Without the confidence filter, post-discussion and large-score-gap filter, expert agreement is much weaker; the same people disagree on field priors, tractability and even readings of text.

Result: Among tested models the top performer Fable 5 matches selected expert preferences on 60% of comparisons versus estimated expert agreement of 77%; on distinct-prompt pairs it scores 69% on 74 comparisons. Its low-to-max reasoning increases performance from 48% to 60%, while almost all other models are within two standard deviations of chance. On the 18 within-prompt pairs, hiding the motivating prompt improved measured performance by 17 percentage points; the model over-weights whether a proposal responds to the prompt even though the rubric tells judges not to. This shows a presently measurable gap on one research-triage proxy, not a stable upper bound on ability.

Limitations and use: correlated pair labels, ten geographically and professionally narrow raters, confidence-based selection of clearer cases, model-written proposals and no follow-up on which research projects later worked. Choosing plausible proposals and evaluating real scientific returns or novel hazards are distinct. Strongest alternative interpretation: the best model’s gain with reasoning and performance on distinct-prompt pairs show a learnable skill, and AI may increase laboratory output without matching humans on every judgment. A useful independent replication would blind cross-institution researchers to provenance, score proposals selected after model training, pay for matched projects and compare their validated consequences, calibration and costs. This is an unusually direct test of Dru’s research-taste objection, not proof RSI cannot occur.