arXiv 自然语言处理· Xueping Gao·· 8 小时前AI 评分41
从探测得分到警报策略:大语言模型智能体激活监控器的操作有效性
From Probe Scores to Alarm Policies: Operational Validity of Activation Monitors for Language-Model Agents
AI 导读
研究提出了操作有效性契约,揭示大语言模型智能体的高 AUROC 激活探测得分无法直接转化为低误报预算下的可用警报决策。
来源:arXiv 自然语言处理 · arxiv.org