MENU
Skip to main content
All blog posts

AI and communication research

AutoPersuade: Evaluating and explaining persuasion with AI

An AI analysis of pro-vegan messages links argument themes to audience ratings from 5,180 paired comparisons.

  • Peer reviewed · 2024
Complete abstract from AutoPersuade: A Framework for Evaluating and Explaining Persuasive Arguments.
Complete abstract from the original paper. Source: Association for Computational Linguistics and Till Raphael Saenger, Musashi Hinck, Justin Grimmer, Brandon M. Stewart. Till Raphael Saenger, Musashi Hinck, Justin Grimmer, Brandon M. Stewart. 2024. “AutoPersuade: A Framework for Evaluating and Explaining Persuasive Arguments.” In *Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing*, 16325–16342. Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.emnlp-main.913. Reproduced under CC BY 4.0; no third-party material is included. Source paper, page 1 View full-size excerpt

Knowing which message won is not the same as knowing why

A survey can reveal which of two messages people prefer, but it may not explain what made one more compelling. AutoPersuade is a framework for connecting reactions to the content of persuasive arguments. It combines audience evaluations with an interpretable topic model, then estimates how the presence of different themes relates to reported persuasiveness.

The researchers demonstrate the workflow with pro-vegan messages. They curated 1,309 arguments: 1,209 derived from 93 existing source arguments and 100 generated by GPT-4. A survey collected 5,180 pairwise comparisons across 1,036 sessions. Each respondent saw five comparisons and selected the more persuasive message. The team used a Bradley–Terry model, a method that summarizes pairwise contests into relative scores, to estimate how arguments performed in the sample (pages 16328–16329).

Interpretable topics connect content and response

AutoPersuade’s SUN model learns topics that describe both argument text and audience reactions. Unlike a conventional topic model that only groups similar documents, SUN is trained to represent features associated with the measured responses as well. Researchers can then estimate the average effect of changing a topic’s prevalence and predict how new arguments might perform. The design aims to balance explanation with prediction: a model that exposes content patterns may be easier to interpret than a black-box score, though it may be less accurate at ranking individual arguments.

The validation results illustrate that distinction. On average, the model predicted responses to new arguments, and an additional study supported estimated average effects across the range of topic prevalence. But selecting arguments predicted to be highly persuasive did not reliably improve the strongest-performing tail. In other words, the framework was more useful for studying broad relationships between content and audience reaction than for automatically finding a winning message (Figure 1 and pages 16326, 16331–16334).

That qualification is important for organizations. Self-reported preference is not the same as behavior change, and a sample evaluating pro-vegan messages cannot establish what works in a different issue or audience. The topic model’s estimates also rely on assumptions that latent message features can be meaningfully separated and that relevant confounding is controlled. A message that resonates in one survey context may be ineffective elsewhere. The authors distinguish their explanatory goal from identifying rhetoric that works universally.

A research aid, not a persuasion recipe

AutoPersuade offers a way to investigate why certain arguments are associated with more favorable responses, especially when a team has many candidate messages. Its contribution is the joined analysis of message content and audience data. The evidence does not support using the model as an automatic generator of maximally persuasive communication. For practitioners, the broader implication is that message analytics can help formulate hypotheses about audience response, while those hypotheses need testing with the intended population and outcomes that matter beyond stated preference.

The public value of such estimates depends on how the argument set, respondent pool and measured reaction are defined before analysis begins. These operational choices shape what a comparison means, so a score should not be treated as a context-free measure of persuasive power or as proof of real-world behavioral change.

Bibliography & sources

  1. Till Raphael Saenger, Musashi Hinck, Justin Grimmer, Brandon M. Stewart. 2024. “AutoPersuade: A Framework for Evaluating and Explaining Persuasive Arguments.” In *Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing*, 16325–16342. Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.emnlp-main.913.