Home»News»AI occupational exposure scores vary with the model doing the rating

AI occupational exposure scores vary with the model doing the rating

Michelle Yin, Hoa Vu and Claudia Persico’s NBER working paper examines the stability of occupational exposure scores produced by large language models. The April 2026 abstract reports a 3.6-fold difference in mean exposure when three models apply the same rubric to identical tasks, with agreement as low as 57%.

The abstract also reports that changing the annotator changes downstream empirical estimates. The study therefore concerns measurement reliability as well as the interpretation of labour-market analysis. Exposure scores describe assessed task exposure; they do not directly count realised job losses.

This item uses the original abstract circulated in the NBER email and the April paper identified on the author’s website. The full PDF could not be retrieved during review. Its exact original publication day was not established, so the timestamp defaults to the email receipt on 27 April 2026 at 06:05:32 Paris time. The working paper has not been peer reviewed.

Ask how a workforce estimate was made For HR teams, a numerical exposure estimate can appear more settled than the process that generated it. Before using it in workforce planning, this editor recommends identifying the model, version, task description, prompt and rubric behind the score.

Teams should also ask whether the result changes materially when a different method or model is used. Recording these choices allows later reviewers to understand what was measured and why an estimate may have changed. An updated score does not necessarily establish that the underlying occupation has changed to the same degree.

Connect scores with the actual work An occupation is not a complete description of the work performed in a particular organisation. Tasks, responsibilities, systems and constraints differ. An employer should examine these local conditions before drawing conclusions about staffing, training or benefit needs.

A score can inform investigation, but decisions affecting employees need a wider evidential basis. Teams can compare the estimate with observed use, workflow experience and feedback from the people doing the work.

The study’s practical contribution is a reason to make measurement choices visible and test their influence. It supports more careful workforce analysis while leaving the organisation responsible for assessing its own circumstances and the consequences of its decisions.

Sources: NBER paper 35110 · Papier auteur, version avril 2026

Previous post

Perceived insurance nonpayment risk shapes protection and retirement choices

Next post

PRA says protected cell captives will miss the UK regime’s initial launch