Reviewer 2.0
To misquote Douglas Hofstadter: the backlash against AI in the research system will be bigger than you’re expecting, even once you take into account that it will be bigger than you’re expecting.
The latest flashpoint has been around a funding call where AI was apparently used to reject the bottom half of proposals in a heroically cack-handed way. If you’ve not been keeping up, in a nutshell what seems to have happened is that proposals were uploaded to various different AI models which scored them across various criteria, and “debated” between themselves about the merits and weaknesses.
This led to an incredibly blunt score out of five – to add insult to injury, applicants were told that this rating was shared with them “to provide transparency” – and no further feedback. The impression from those sharing rejections on social media was that applicants were given little to no indication that this approach was going to be used.
This review process was not UKRI-managed, we should stress, but it was apparently a UKRI-funded project. Some of the heat in the situation has undoubtedly been generated by its close proximity to the announcement earlier this month that UKRI is “modernising grant assessment for the age of AI,” including through the use of AI.
There are a few possible angles we might pick up here. The need for a response to snowballing grant application numbers is well-established, as is the pressure on teams, institutions and funders to look more efficient even as they shed experienced staff. For one thing, this is clearly a cautionary tale for institutions thinking they can use AI to do all this extra demand management coming their way: get it wrong, and people are going to be genuinely furious.
On UKRI’s own in-house use of AI-assisted assessment – which is still being trialled, and will look very different from this fiasco when it rolls out further – there just isn’t a world where government departments pursue AI efficiency while UKRI and the research councils do not. Political pressure, whether implicit or explicit, is worth bearing in mind.
If we are going to take something away from this – other than quite how much anger the botched use of artificial intelligence can (rightly) create when used in research assessment – then can we propose that each occasion of massive blowback, of which there will be many more to come, contains some learnings about what not to do?
One: don’t make spurious claims about accuracy and objectivity. Don’t tell applicants “we will soon produce a document detailing our method” (this seems to actually have been said). Everything needs to be clear upfront, and have a heft of evidence behind it.
Two: there are many aspects of this controversy which have infuriated the research community, but work being uploaded to large language models without permission is clearly at the forefront of much distaste. Given that this is currently a bugbear of human peer review – a sense that human reviewers are inevitably uploading your work to American-owned content hoovers for all that funders and journals might ask them not to – then finding ways to rigorously protect researchers’ work could end up being a point in favour of more mechanised processes. Especially when many of these companies are conveniently trying to make headlines in scientific discovery themselves while also handling researchers’ own projects.
Three: we surely can’t jump straight from traditional peer review to “oh, we just dumped everything we were sent into all the LLMs we could think of.” Last week also saw one of the metascience unit projects on AI assessment deliver some initial findings (in pre-print). Among these was that such scoring might help “triangulate” and check for bias – a much more palatable starting point, you would think.
But sadly the rush to create efficiencies, rather than improvements in how reviews work, suggests that this is a row which is only going to escalate.
Metascience unit head Ben Steyn has framed the challenge as one of finding new ways forward in the “new academic world of AI slop, and human slop.” As depressing as that might read, it’s probably a good way of thinking about it. A substantial part of the outrage in this latest set of events is about what appear to be really rather stupid human decisions.
Spotted a politician meddling in the research system? Or a blurry line between academia and government? Email the man himself: haldane [at] researchagenda [dot] news