<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Cs.CY on 庞玉栋个人博客</title><link>https://pangyd.com/tags/cs.cy/</link><description>Recent content in Cs.CY on 庞玉栋个人博客</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Thu, 17 Sep 2026 23:11:08 +0800</lastBuildDate><atom:link href="https://pangyd.com/tags/cs.cy/index.xml" rel="self" type="application/rss+xml"/><item><title>Language-model groups overstate consensus when replaying human deliberation on a reasoning task</title><link>https://pangyd.com/post/radar-27/</link><pubDate>Thu, 17 Sep 2026 23:11:08 +0800</pubDate><guid>https://pangyd.com/post/radar-27/</guid><description>&lt;blockquote>
&lt;p>&lt;strong>原文&lt;/strong>：&lt;a href="https://arxiv.org/abs/2609.20543v1">Language-model groups overstate consensus when replaying human deliberation on a reasoning task&lt;/a>&lt;/p>
&lt;p>&lt;strong>作者&lt;/strong>：Tengfei Shao&lt;/p>
&lt;p>&lt;strong>来源&lt;/strong>：arXiv cs.AI（人工智能）&lt;/p>
&lt;/blockquote>
&lt;h2 id="正文">正文&lt;/h2>
&lt;p>Computer Science &amp;gt; Artificial Intelligence&lt;/p>
&lt;p>arXiv:2609.20543v1 (cs)&lt;/p>
&lt;p>[Submitted on 17 Sep 2026]&lt;/p>
&lt;p>Title:Language-model groups overstate consensus when replaying human deliberation on a reasoning task&lt;/p>
&lt;p>Authors:Tengfei Shao&lt;/p>
&lt;p>View a PDF of the paper titled Language-model groups overstate consensus when replaying human deliberation on a reasoning task, by Tengfei Shao&lt;/p>
&lt;p>View PDF&lt;/p>
&lt;p>Abstract:Full-consensus rates are often treated as indicators of collective cognition, yet depend on how participation and final states are operationalized. We replayed 100 held-out human Wason groups with matched large language model (LLM) agent groups, seeding one belief-anchored agent per participant&amp;rsquo;s pre-discussion answer and scoring agents and people with the same code. Across human scoring definitions, estimates ranged from 24.0% to 57.0%; about one fifth of participants never posted, whereas agents almost always did. Agent groups remained more consensual in two post-unblinding sensitivity analyses: the submit-based comparison (n = 98) yielded gaps of 34.0 and 43.9 percentage points for chat and reasoning modes, and the participation-matched comparison (n = 45) yielded gaps of 34.1 and 44.4 points. These complementary routes reduced different measurement asymmetries yet converged within 0.5 percentage points. The gap persisted without early stopping and under a reparameterization removing the memorizable answer; reasoning-mode groups then agreed nearly unanimously, mostly on incorrect answers. Simulated consensus did not track collective accuracy, and belief-anchored agent groups were biased estimators of the human group-outcome distribution in this setting. These analyses provide a scoring-explicit basis for assessing simulated-group estimates of human deliberative outcomes.&lt;/p></description></item></channel></rss>