The metascience unit sets out its thinking on AI-powered research evaluation
The UK metascience unit – the UKRI/BIST team with a remit to “apply the scientific method to the systems, policies and processes of science itself” – has released an explainer on its work relating to AI-driven research evaluation.
The paper, written by co-head of the unit Ben Steyn, sets out the rationale behind the various investments it is making into projects and lines of inquiry – and explains the intention to develop new, AI-driven indicators that will allow for a “new quantitative vocabulary of research quality.”
Steyn summarises the current system of research evaluation as typically involved either expert peer review or traditional bibliometric indicators such as citations, and then highlights the limitations of both: the subjectivity, pace and cost of peer review, and the plethora of issues with using citations as a proxy for quality.
Referencing earlier metascience unit-supported work on AI-based indicators of scientific novelty, he outlines how the intention is now to move this into other areas, such as rigour and impact:
“The Metascience Unit and partners want to cultivate a pipeline of new indicators moving through these stages of validation and critical reflection, with the end result hopefully being a responsible and validated ‘basket of indicators’ which allow for a new quantitative vocabulary of research quality.”
According to Steyn, a “carefully considered basket of task-validated modular indicators” would be preferable to – and more effective than – the use of off-the-shelf AI tools to review papers:
“Simply feeding a paper to an LLM for review, without detailed instruction, risks abdicating responsibility and accountability. Research virtues are amorphous, multi-faceted, and ambiguous when stated in words. New metrics give us a common vocabulary to ground these nebulous concepts in data.”
He also calls for caution in the rush to use AI, concluding that value judgements about research quality are “perhaps the most important task that ought to remain human-driven.”
The second half of the explainer lists ongoing projects supported and/or funded by the metascience unit, both technical scientometric work and more “normative” studies, such as one involving philosophers of science. This project will consider questions such as who ought to decide how a hypothetical “basket” of metrics is weighted, or whether the use of automated assessment is less justified in the arts and humanities.
The metascience unit’s updated information page now explains that its key area of focus in 2026–28 will be science systems in an AI age. Work to be conducted will both test the feasibility and effectiveness of using AI in assessment processes and seek to develop “efficient and scalable models of human review,” such as distributed peer review.
Steyn’s paper finishes by stressing that the metascience unit does not see it as its role to set policy:
“To be clear, our raison d’etre as the Metascience Unit is to build evidence – not to set policy. This is about new exploratory research to test whether AI-driven indicators can be developed, validated, and shown to be reliable, not a commitment to deploy them within UKRI or elsewhere.”
Meanwhile in Kyoto
UKRI has also released a statement on joint metascience work which it will conduct with the US National Science Foundation (NSF). This follows last weekend’s publication by the White House of the Kyoto Vision for a Golden Age of Science, described by the US government as a “joint declaration” on the future of science, supported by 16 other countries including the UK.
UKRI and NSF commit to building on the metascience principles outlined in that declaration (our earlier coverage can be read here). The two organisations will explore cooperation – which will be formalised in the coming months – in areas including measuring research productivity and impact, and improving research review and funding. The partnership will also support the development of an international network of metascience units.