MENU
Skip to main content
All blog posts

Science and research

AgentReview: Exploring Peer Review Dynamics with LLM Agents

AgentReview simulates conference peer review with language-model agents. Changing reviewer bias changed 37.1% of paper decisions in the simulation.

  • Peer reviewed · 2024

Complete abstract from AgentReview: Exploring Peer Review Dynamics with LLM Agents.
Complete abstract from the original paper. Source: ACL Anthology, CC BY 4.0: https://creativecommons.org/licenses/by/4.0/. Yiqiao Jin et al., AgentReview: Exploring Peer Review Dynamics with LLM Agents, EMNLP 2024; reproduced under CC BY 4.0. Source paper, page 1 View full-size excerpt

Peer review shapes what gets published, yet the effect of a reviewer’s temperament is difficult to isolate from expertise, paper quality and the rules of a conference. AgentReview asks whether simulated reviewers can help researchers explore those moving parts without exposing confidential reviews. The answer is promising as a way to study mechanisms, but the experiment is a model of peer review rather than a measurement of actual conference outcomes.

A conference process rendered as agents

The researchers built an agent-based framework around large language models (LLMs). It simulates five stages, from reviewer assessment and author response through area-chair discussion, meta-review and a final decision. Agents receive assigned roles and characteristics, including differences in knowledge, effort, bias and authority. The authors used 500 ICLR submissions from 2020–2023 to generate 53,800 simulated review-process documents. Rather than asking one model to judge a paper, the setup lets the researchers change one feature of the simulated community and observe how the decision process shifts.

That experimental structure matters. Real review outcomes combine many influences that are hard to observe in a private process. In a simulation, the research team can hold most inputs constant while varying a factor such as reviewer bias. This gives a controlled way to formulate hypotheses about review design. It does not establish that the same effect size occurs in a live program committee.

Bias can move a decision

The paper reports that reviewer bias changed 37.1% of paper decisions in the modeled setting. The authors also describe spillover: a reviewer’s behavior can affect colleagues, not only their own rating. For example, simulated conformist reviewers and area chairs reduced the spread of ratings, while an irresponsible reviewer lowered overall reviewer commitment. The simulations also found that revealing author identities for a subset of papers changed some decisions.

The results make peer review look less like a simple average of independent scores. A decision can depend on who reviews, how reviewers interact and what information is visible. The authors report little change when they removed the rebuttal stage, which they interpret as consistent with anchoring: early ratings may persist through later discussion. That finding is specific to this simulation and its prompts, not a general verdict on rebuttals.

What a simulation can and cannot say

AgentReview is useful for exploring scenarios that are ethically or practically difficult to test with confidential reviewer data. It can help researchers ask which review procedures deserve closer measurement. Its privacy advantage comes from generating synthetic interactions rather than analyzing identifiable reviewer records.

The limits are consequential. The agents cannot introduce new experimental evidence during author–reviewer exchanges, and the paper mostly studies factors one at a time even though real reviews involve interacting causes. The authors also do not compare their simulated decisions against a stable human-review baseline. LLM behavior depends on the prompts, model and assumptions used to represent reviewers; it should not be read as direct evidence about real reviewers’ motives or conduct.

Possible implications for practice

For conference organizers, this work offers a sandbox for testing questions about anonymity, discussion and reviewer incentives before attempting changes in a real process. Its strongest contribution is a research method: specify a mechanism, vary it, and see how outcomes respond inside an explicit model. Any proposed reform would still need evidence from real review settings, with safeguards for confidentiality and a careful account of how human judgment differs from the simulated agents.

The paper, “AgentReview: Exploring Peer Review Dynamics with LLM Agents,” appeared in the 2024 Conference on Empirical Methods in Natural Language Processing. This article summarizes the published proceedings version and its stated limitations.

Bibliography & sources

  1. Yiqiao Jin, Qinlin Zhao, Yiyang Wang, Hao Chen, Kaijie Zhu, Yijia Xiao, Jindong Wang. 2024. AgentReview: Exploring Peer Review Dynamics with LLM Agents. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 1208–1226. Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.emnlp-main.70.