arXiv 自然语言处理· Hongao Zhu (Department of Linguistics, University of California San Diego), Muxiaoqiao Xu (School of Foreign Languages, Shanghai Jiao Tong University), Yikang Liu (School of Computer Science, Shanghai Jiao Tong University), Siyuan Song (School of Foreign Languages, Shanghai Jiao Tong University, Department of Linguistics, University of Texas at Austin), Yuxia Wang (School of Foreign Languages, Shanghai Jiao Tong University), Byung-Doh Oh (Division of Linguistics and Multilingual Studies, Nanyang Technological University), Hai Hu (Department of Language Science and Technology/Division of AI and the Humanities, Hong Kong Polytechnic University)·· 5 小时前AI 评分34
语言模型惊奇度对中文阅读预测能力的系统分析
A Systematic Analysis of the Predictive Power of LM Surprisal in Reading Chinese
AI 导读
研究团队提出 SMS 对齐方案并基于 30B tokens 训练的 Chinese-Pythia 模型(14M-1.4B),证实语言模型 token 级惊奇度(surprisal)对中文阅读时间具有预测能力。
来源:arXiv 自然语言处理 · arxiv.org