Artificial intelligence agents are increasingly being used to interpret and even conduct science. But are they able to accurately distinguish science that is solid from science that is subpar?
Valentin Rodionov wanted to find out. To do so, he chose 42 studies that don’t meet the parameters of legitimate research. Some have been retracted from literature; others are considered to be problematic, fraudulent, or pseudoscientific but haven’t formally been retracted or corrected.
Rodionov, an organic chemist at Case Western Reserve University, uploaded sections of his sample papers to 30 AI models, 10 times each. The AI models were mostly well-known large language models (LLMs) created by the tech giants Anthropic, OpenAI, Meta, Google, Mistral AI, and others.