Can AI Debias the News? LLM Interventions Improve Cross-Partisan Receptivity but LLMs Overestimate Their Own Effectiveness
Mirrored from arXiv — NLP / Computation & Language for archival readability. Support the source by reading on the original site.
Computer Science > Computation and Language
Title:Can AI Debias the News? LLM Interventions Improve Cross-Partisan Receptivity but LLMs Overestimate Their Own Effectiveness
Abstract:Partisan news media erode cross-partisan trust, but large language models (LLMs) offer the potential of debiasing such content at scale. Across two pre-registered experiments, we tested whether LLM-generated debiasing of liberal news headlines improves conservative readers' trust-relevant judgments. In Study 1, subtle lexical debiasing (replacing emotive words with moderate synonyms) had no effect on any outcome. Study 2 found that a more substantive reframing intervention significantly increased conservatives' perceived trustworthiness, completeness, and willingness to engage with liberal news headlines, without producing a backfire effect among liberals. In Study 1, the intervention produced robust effects across silicon participants simulated with six different models (o3-mini, o3, GPT-4o mini, GPT-4o, GPT-5 mini, and GPT-5), whereas it had no impact on human readers. In Study 2, the intervention's effects among silicon participants generally aligned directionally with human responses but were significantly larger for some outcomes, and three models (o3, GPT-4o mini, and GPT-4o) incorrectly predicted a liberal backfire effect absent in humans. Moderation analyses revealed that the models' implicit theory of who responds to debiasing diverged from the psychological profile that actually predicted human responsiveness. Most strikingly, in Study 2, participants simulated by each of the six models suggested that debiasing effects would be stronger among participants high in political in-group identification. Yet, no such moderation was observed among human participants. These findings demonstrate that LLM-based debiasing can improve cross-partisan receptivity when targeting ideological framing rather than surface-level language, but that current models lack both the quantitative accuracy and qualitative psychological fidelity to evaluate their own interventions without human oversight.
| Subjects: | Computation and Language (cs.CL); Computers and Society (cs.CY) |
| Cite as: | arXiv:2605.01006 [cs.CL] |
| (or arXiv:2605.01006v3 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2605.01006
arXiv-issued DOI via DataCite
|
Submission history
From: Jonas R. Kunst PhD [view email][v1] Fri, 1 May 2026 18:20:42 UTC (1,574 KB)
[v2] Thu, 7 May 2026 18:16:32 UTC (1,397 KB)
[v3] Fri, 24 Jul 2026 14:29:35 UTC (1,562 KB)
Access Paper:
- View PDF
References & Citations
Bibliographic and Citation Tools
Code, Data and Media Associated with this Article
Demos
Recommenders and Search Tools
arXivLabs: experimental projects with community collaborators
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
More from arXiv — NLP / Computation & Language
-
Unified Hallucination Fuzzing for Multimodal Large Language Models
Aug 11
-
DocAtlas: Long-Document Understanding as Mutable-State Interaction
Aug 11
-
WuYuEval: A Multi-Level Benchmark for Large Language Models in Solid Waste Management
Aug 11
-
Search-G1: Grounded Search Agents via Representation-Based Intrinsic Rewards
Aug 11
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.