//
自然语言处理
3 篇 · 资讯流Language-model groups overstate consensus when replaying human deliberation on a reasoning task
Full-consensus rates are often treated as indicators of collective cognition, yet depend on how participation and final
原文 ↗KoNeoBench: A Curated Evaluation Dataset for LLM Understanding of Korean Neologisms
Large language models (LLMs) are typically evaluated on static benchmarks, even though natural language constantly evolv
原文 ↗DualSQL: Text-to-SQL with Multi-Agent Reinforcement Learning
State-of-the-art Text-to-SQL systems are typically multi-agent pipelines centered around two fundamental tasks: schema l
原文 ↗