跳到正文
原文
NVIDIA Developer Blog· Tanya Lenz·· 9 天前精选AI 评分65

NVIDIA Dynamo-Triton 集成 TensorRT 多设备推理以简化多 GPU 模型服务

Simplifying Model Serving Across Multiple GPUs with NVIDIA TensorRT Multi-Device Integration in NVIDIA Dynamo-Triton

AI 导读

NVIDIA 推出 TensorRT 多设备推理能力并在 Dynamo-Triton 中完成集成,旨在简化多 GPU 场景下的模型服务。该能力基于 NCCL 分布式通信集合,允许单个 TensorRT 网络跨多张 GPU 执行,同时保留 TensorRT 的推理优化性能,已自 TensorRT 11.0 起获得完整支持。

推荐理由

原文介绍了基于 NCCL 的多 GPU 分布式推理新特性,读者可据此了解 TensorRT 11.0 的多卡部署优化机制。

来源:NVIDIA Developer Blog · developer.nvidia.com