THE APEX TIMES
NVIDIA expands open local AI models and adds routing software to help agents cut inference costs
NVIDIA says its new Nemotron 3.5 Lightning model and NeMo Switchyard routing library are aimed at making it easier to run increasingly capable AI agents on users’ own devices and to lower the cost of multi-step agent workflows.
NVIDIA is pushing harder into “local AI,” the approach of running generative models and agent-like software on nearby hardware such as PCs, workstations, and edge devices rather than relying only on remote cloud services. In a new blog post marking developments in its Local AI community series, the company announced an expanded open-model offering and a separate open source routing library intended to make agent deployments cheaper and more flexible.
The centerpiece is an update to NVIDIA’s Nemotron 3 model family. NVIDIA said it expanded the lineup with Nemotron 3.5 Lightning, describing it as a customizable open 30 billion parameter mixture-of-experts (MoE) model built for “always-on agents.” In MoE models, only a subset of the model’s internal experts is activated for each prompt or step, a design meant to improve efficiency while maintaining strong output quality.
NVIDIA positioned Nemotron 3.5 Lightning around performance and customization. The company said the model can deliver up to 4x faster token generation and 30% faster time to completion compared with other open models in its class. Because the model uses open weights, NVIDIA said developers and AI enthusiasts can fine-tune it using their own examples so the model better matches specific tasks, interests, and workflows.
To help users deploy the model locally, NVIDIA said it worked with deployment and inference ecosystem partners including vLLM, Ollama,, and LM Studio. NVIDIA also said Unsloth provides day-one support with optimized and quantized models through Unsloth Studio. The company said developers can choose model packaging formats including NVFP4 and GGUF, and that those choices matter for balancing speed, memory usage, and compatibility with different local runtimes.
NVIDIA also outlined where Nemotron 3.5 Lightning can run. The post said the model runs locally on NVIDIA RTX PCs, NVIDIA DGX Spark and OEM GB10 systems, and NVIDIA Jetson. It also said it scales up to RTX PRO workstations, NVIDIA DGX Station, GB300 deskside systems, and broader data center and cloud environments. NVIDIA added that systems using “Blackwell,” the company’s latest generation of AI infrastructure, are available through multiple OEMs including Acer, ASUS, Dell Technologies, GIGABYTE, HP, Lenovo, MSI, and Supermicro, with Exxact also named.
Beyond the model itself, NVIDIA introduced what it calls NeMo Switchyard, an open source routing library aimed at controlling the cost of agent workflows. As generative AI adoption grows, NVIDIA said enterprises want ways to keep rising token costs in check without sacrificing access to “frontier” model capability. NeMo Switchyard, according to the blog, automatically directs each step of an agent workflow to the “best-fit model” based on accuracy, speed, and cost.
NVIDIA said the routing approach gives developers flexibility to work across models and providers, using different models for different tasks within the same agent. The company also offered internal benchmark results: it said NeMo Switchyard helped maintain “frontier-level” task completion while reducing benchmark completion cost to roughly one-third of Opus 4.8 alone. The blog did not provide additional methodological detail in the excerpt, including the exact workload mix or how performance and cost were measured, but it framed the result as an outcome of step-by-step routing rather than using a single large model for every part of a multi-step process.
NVIDIA gave examples of the broader agent use case enabled by fine-tuning and tool pairing. The post said that paired fine-tuned models can power more personalized local “agentic” experiences, ranging from an assistant that helps manage email and calendars, to a smart-home agent handling everyday routines, to a coding companion that works alongside developers on a local codebase. It also described NeMo Switchyard as available on GitHub, and Nemotron 3.5 Lightning as starting points for developers through related technical blogs and Jetson AI Lab tutorials.
What is still not fully clear from the announcement is how users should think about tradeoffs between running everything locally versus mixing local and cloud inference, and what deployment paths will be simplest for different developer stacks. The company also emphasized internal benchmark cost reduction, which may not translate directly to every environment. For editorial readers, the practical question is likely to be whether the open-model and open-routing pieces integrate smoothly across real agent frameworks and real hardware constraints, including memory limits and sustained “always-on” workloads.
Why It Matters
- Local AI is increasingly about operational control, where developers want their agents to run closer to the user and integrate with local data and tools.
- Open weights and multiple inference ecosystems can lower friction for experimentation, but they also put pressure on developers to handle deployment complexity themselves.
- Routing step-by-step through different models may become a key lever for enterprises trying to manage token-related costs in multi-step agent workflows.
- If internal benchmark results hold broadly, cost-aware routing could improve the economic case for agents that would otherwise overuse large frontier models.
Key Facts
- NVIDIA expanded its Nemotron 3 model family with Nemotron 3.5 Lightning, an open 30B mixture-of-experts model aimed at always-on local agents.
- NVIDIA said Nemotron 3.5 Lightning can deliver up to 4x faster token generation and 30% faster time to completion versus open models in its class.
- NVIDIA said Nemotron 3.5 Lightning can be deployed locally using partners including vLLM, Ollama,, and LM Studio, with Unsloth support via Unsloth Studio and formats including NVFP4 and GGUF.
- NVIDIA introduced NeMo Switchyard, an open source routing library that selects the best model for each agent workflow step based on accuracy, speed, and cost.
- NVIDIA said NeMo Switchyard reduced benchmark completion cost to roughly one-third of Opus 4.8 alone while maintaining frontier-level task completion, based on internal benchmarks.
- NVIDIA said Nemotron 3.5 Lightning runs on RTX PCs, Jetson, and NVIDIA system offerings including DGX Spark, RTX PRO workstations, DGX Station, and GB300, and that it can also be accessed via OpenRouter, as an NIM microservice, and other ecosystem paths.
Technology Related
AMD says Instinct AI systems are now operating in Saudi Arabia, highlighting a potential ramp tied to additional data-center power
A recent market report frames AMD’s Instinct deployments in Saudi Arabia as a move from plan to production, and points to how incremental data-center capacity, measured in megawatts, could influence investor expectations.
Salesforce says AI-driven revenue momentum is building as Agentforce adoption spreads
In a recent market update circulated by Yahoo Finance, Salesforce management pointed to expanding use of its AI offerings, including agentic workflows and consumption-style pricing, as the company positions its next growth phase.
Salesforce backs HiBob to bolster workforce AI, and adds a new AgentExchange email tool
Salesforce said it is supporting HR-analytics and talent-workforce platform HiBob as part of efforts to connect enterprise data with “powered AI.” The company also announced an AgentExchange email tool aimed at expanding what business agents can do inside everyday workflows.
EverPass Media expands NFL distribution via multi-year Netflix deal for 2026 slate
EverPass Media says it has added Netflix’s five NFL games for the 2026 season to its NFL distribution offering, including the first-ever Thanksgiving Eve game, plus “NFL Honors.”
Broadcom leans harder into VMware AI with a push aimed at enterprise rivals
Broadcom’s VMware AI push is tied to the latest VCF 9.1 release, as the company’s messaging positions it against Nutanix and Microsoft in hybrid cloud and enterprise AI rollouts.
Yahoo Finance points to “buy zones” for Microsoft, Palantir, Shopify and ServiceNow
A market-readout from Yahoo Finance flagged several software and AI-linked names, including Palantir (PLTR), as trading in or near so-called buy zones. The note is framed as technical or timing-oriented, with limited company-specific detail.
Oracle Shares Fall as Investors Focus on Cash Flow Gap and Rising Borrowing Costs
A reported $23.7 billion cash shortfall over Oracle’s last fiscal year and $43 billion in borrowing are drawing attention to the company’s interest-rate exposure, a factor that can quickly change sentiment when Treasury yields are elevated.
Adobe’s next report faces a split view: Citi still expects a beat, but flags lingering risks
After Adobe lowered its annual revenue outlook, one analyst said the company can still deliver a beat-and-raise in fiscal third-quarter results, even as concerns remain.
Palantir’s commercial growth may overtake government revenue sooner than expected, according to a new market model
A widely watched growth-math forecast argues Palantir’s commercial revenue could surpass its government revenue before 2027, driven by a widening gap in the companies’ growth rates.
Netflix shares face another round of debate after new market commentary, but company keeps details scarce
A recent Yahoo Finance-linked article argues Netflix is not finished telling its story, urging investors to stay cautious until more clarity emerges.