ICPR 2026 · 2026
Prompt injection attacks on llm generated reviews of scientific publications
Why this publication matters
Automating scientific reviews creates a new vulnerability: instructions hidden in the submitted document may influence the reviewer. This study tests that problem and also examines whether language models are too inclined to accept papers. It provides evidence for designing review assistance and evaluations that do not mistake fluent feedback for independent judgment.
Abstract
The ongoing intense discussion on rising LLM usage in the scientific peer-review process has recently been mingled by reports of authors using hidden prompt injections to manipulate review scores. Since the existence of such “attacks” - although seen by some commentators as “self-defense” - would have a great impact on the further debate, this paper investigates the practicability and technical success of the described manipulations. Our systematic evaluation uses 1k reviews of 2024 ICLR papers generated by a wide range of LLMs shows two distinct results: I) very simple prompt injections are indeed highly effective, reaching up to 100% acceptance scores. II) LLM reviews are generally biased toward acceptance (>95% in many models). Both results have great impact on the ongoing discussions on LLM usage in peer-review.
Figures
Cite this paper
@inproceedings{keuper2026promptinjectionattacks86,
title = {{Prompt injection attacks on llm generated reviews of scientific publications}},
author = {Janis Keuper},
booktitle = {International Conference on Pattern Recognition},
year = {2026},
url = {https://arxiv.org/pdf/2509.10248}
}
Figures and abstract are reproduced from the linked research sources. Credit remains with the authors and publishers.