跳到正文
原文
Cerebras 官方博客(网页)·· 9 小时前精选AI 评分69

Cerebras 推理平台上线 Gemma 4 31B 多模态模型

Gemma 4 on Cerebras—The Fastest Inference is Now Multimodal June 29, 2026

AI 导读

Cerebras 宣布在其推理平台上推出 Google DeepMind 开源模型 Gemma 4 31B,以超过 1,800 tokens per second(第三方测得 1,851 TPS)的吞吐运行,成为该晶圆级计算平台首个支持图像理解的多模态模型。

推荐理由

原文给出了 1,851 tokens per second 的多模态推理吞吐与首 token 延迟数据,读者可据此评估极速视觉智能体工作流的落地可行性。

来源:Cerebras 官方博客(网页) · cerebras.ai