跳到正文

Language-model groups overstate consensus when replaying human deliberation on a reasoning task

人工智能 0 次浏览

原文Language-model groups overstate consensus when replaying human deliberation on a reasoning task

作者:Tengfei Shao

来源:arXiv cs.AI(人工智能)

正文

Computer Science > Artificial Intelligence

arXiv:2609.20543v1 (cs)

[Submitted on 17 Sep 2026]

Title:Language-model groups overstate consensus when replaying human deliberation on a reasoning task

Authors:Tengfei Shao

View a PDF of the paper titled Language-model groups overstate consensus when replaying human deliberation on a reasoning task, by Tengfei Shao

View PDF

Abstract:Full-consensus rates are often treated as indicators of collective cognition, yet depend on how participation and final states are operationalized. We replayed 100 held-out human Wason groups with matched large language model (LLM) agent groups, seeding one belief-anchored agent per participant’s pre-discussion answer and scoring agents and people with the same code. Across human scoring definitions, estimates ranged from 24.0% to 57.0%; about one fifth of participants never posted, whereas agents almost always did. Agent groups remained more consensual in two post-unblinding sensitivity analyses: the submit-based comparison (n = 98) yielded gaps of 34.0 and 43.9 percentage points for chat and reasoning modes, and the participation-matched comparison (n = 45) yielded gaps of 34.1 and 44.4 points. These complementary routes reduced different measurement asymmetries yet converged within 0.5 percentage points. The gap persisted without early stopping and under a reparameterization removing the memorizable answer; reasoning-mode groups then agreed nearly unanimously, mostly on incorrect answers. Simulated consensus did not track collective accuracy, and belief-anchored agent groups were biased estimators of the human group-outcome distribution in this setting. These analyses provide a scoring-explicit basis for assessing simulated-group estimates of human deliberative outcomes.

Comments:

37 pages, 4 figures. Preregistration: this https URL . Code and data: this https URL

Subjects:

Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY); Multiagent Systems (cs.MA)

Cite as:

arXiv:2609.20543 [cs.AI]

(or

arXiv:2609.20543v1 [cs.AI] for this version)

https://doi.org/10.48550/arXiv.2609.20543

Focus to learn more

arXiv-issued DOI via DataCite (pending registration)

主题

人工智能 · 自然语言处理 · cs.CY · 多智能体


由「前沿雷达」于 2026-09-20 采集。正文取自原文页面,已保留出处链接。标题与正文版权归原作者所有。

评论

加载中…