NVIDIA 发布 Nemotron 3.5 Lightning 与 NeMo Switchyard,加速智能体 AI
摘要
NVIDIA 发布 Nemotron 3.5 Lightning(30B 参数 MoE 开源模型)和 NeMo Switchyard(开源模型路由库),旨在提升智能体 AI 任务的执行效率与成本效益,并增强企业对 AI 部署的控制力。
核心要点
- Nemotron 3.5 Lightning 为 300 亿参数混合专家(MoE)模型,专为多智能体系统中高吞吐的专门任务设计。
- 相比同类模型,输出速度最高提升 4 倍,智能体任务完成时间缩短 30%,且在 PinchBench 基准上保持前沿精度。
- NVIDIA 同步发布 Nemotron-RL-Agentic-Terminal-Pivot 强化学习数据集,用于后训练编码智能体能力。
- NeMo Switchyard 可自动将每个请求路由至最适合的模型,支持开源、专有及 NVIDIA 模型组合,无需重写应用。
- NVIDIA 内部基准显示,NeMo Switchyard 可将任务完成成本降至仅使用 Opus 4.8 时的近三分之一,同时保持前沿精度。
原文佐证
- Together, Nemotron 3.5 Lightning and NeMo Switchyard deliver greater control over how AI is deployed, where it runs and how efficiently it operates — across PCs, workstations, data centers and the cloud.
- The model delivers up to 4x faster output speed, leading to 30% faster agentic task completion compared with other models in its class.
- Internal benchmarks show that NeMo Switchyard maintains frontier-level accuracy while reducing task completion cost to nearly one-third of Opus 4.8 alone.
AI 洞察
NVIDIA 通过开源轻量模型与智能路由的组合,正在抢占智能体 AI 的基础设施制高点。这一策略既能吸引企业采用 NVIDIA 生态,又削弱了对单一超大模型的依赖,可能重塑 AI 应用的成本与部署模式。随着多模型协作成为主流,路由层的价值将上升,NVIDIA 有望成为 Agent 工作流的“操作系统”,也对 Anthropic、OpenAI 等闭源模型厂商形成压力,推动他们提供更灵活的模型选择。