跳到正文
原文
arXiv 人工智能· Yulin Fu (Beijing University of Posts and Telecommunications), Junren Wang (West China Hospital, Sichuan Provincial Engineering Research Center of Intelligent Diagnosis and Treatment of Breast Diseases), Guangjing Yang (Beijing University of Posts and Telecommunications), Zhangyuan Yu (Beijing University of Posts and Telecommunications), Wanran Sun (Beijing University of Posts and Telecommunications), Jiabao Zhou (Beijing University of Posts and Telecommunications), Jin Yin (West China Hospital, Sichuan Provincial Engineering Research Center of Intelligent Diagnosis and Treatment of Breast Diseases), Qicheng Lao (Beijing University of Posts and Telecommunications)·· 4 小时前AI 评分34

MedBenchAgent:推进医疗 VLM 评测基准构建的系统化自动化

MedBenchAgent: Towards Systematic Automation of Medical VLM Benchmark Construction

AI 导读

研究团队提出多智能体框架 MedBenchAgent,将医疗 VLM 评测基准构建形式化为受限编译,并通过基准中间表示(BIR)分离规范规划与测试项实例化。MedBenchAgent 实现了 90.9% 的 Task-Space F1 分数,抽样的 1,000 个测试项中有 994 项通过人工审核。该框架还成功迁移至专业医学领域并完成了 12 个 VLM 的评测。

来源:arXiv 人工智能 · arxiv.org