Posts Tagged
Measurement
AI occupational exposure scores vary with the model doing the rating
Michelle Yin, Hoa Vu and Claudia Persico’s NBER working paper examines the stability of occupational exposure scores produced by large language models. The April 2026 abstract reports a 3.6-fold difference in mean exposure when three models apply the same rubric to identical tasks, with agreement as low as 57%. The abstract also reports that changing the annotator changes downstream empirical estimates. The study therefore concerns measurement reliability as well
