NVIDIA Sees Open AI Models Driving a Multi-Model Enterprise Future

NVIDIA (NASDAQ:NVDA) Director of Developer Tech Nader Khalil said open-source AI models are helping broaden access to model customization, accelerate development and support workloads that require local deployment, while closed frontier models will remain part of a “multi-model world.”

Speaking at The Six Five Summit 2026, Khalil described open source as a longstanding catalyst for innovation across software categories, rather than an AI-specific concept. He said the degree to which an AI model is open can vary widely, including whether developers receive access to a model’s weights, architecture or training data.

Open Models and Enterprise Data

Khalil pointed to the evolution of models including RedPajama, Stable LM and Meta’s Llama releases as examples of the industry determining where value resides in AI development. In his view, model weights can be relatively transient, while data is a more durable asset.

“Value was accruing at the data,” Khalil said, arguing that enterprises hold substantial proprietary information that may not be sufficient or economical to use for training a foundation model from scratch. Instead, organizations can post-train or fine-tune an existing model on their own data for specific tasks.

He said a smaller customized model can be faster, less resource-intensive and less expensive for work that falls within a defined domain. However, he said frontier models can remain appropriate for more open-ended or unfamiliar tasks.

Khalil offered his own use of a DGX Station as an example, saying he runs GLM-5.2 locally at roughly 100 tokens per second and uses it alongside Claude and ChatGPT. He said local or open models do not necessarily replace closed models; rather, organizations may use each for different parts of a workflow.

  • Customized smaller models can handle targeted, in-domain tasks.
  • Frontier models can be used for tasks outside a defined domain.
  • Locally deployed models can support workloads where latency or control over data is critical.

Why NVIDIA Releases Open AI Tools

Khalil said NVIDIA releases models, datasets and tooling, including its Nemotron family, because the company sees a bottleneck in enabling enterprises to apply their data to capable, fully open models. He said NVIDIA’s approach includes releasing more than model weights, including datasets that users can adapt or use to train their own models.

He described model selection as a tradeoff between intelligence and speed. Some applications can tolerate slower processing in exchange for greater capability, while others require low latency and may be better served by a smaller, faster model. The appropriate choice depends on the use case, he said.

For workloads at the edge, Khalil said access to model weights is necessary because some applications cannot rely on cloud-hosted frontier intelligence. He cited live sports replay as an example of a latency-sensitive workload where on-premises or edge AI could be necessary.

He also said enterprises that apply their proprietary information to a model can gain an advantage in performance for their particular use cases. “This is a tool for you so that you can customize it,” Khalil said of open-source models.

Agent Harnesses and AI Security

Khalil said AI agents should be treated as systems made up of a model, a harness, tools and skills, and a runtime. He called the harness the place where models are used and where much of the practical value of AI applications is created.

He pointed to features such as chat interfaces, multimodal prompting, memory, web search and coding tools with access to a file system as examples of innovation beyond the underlying model itself. Khalil said the combination of improved models and improved harnesses helped unlock recent progress in agentic AI.

That same layer is central to security, he said. When a failure or unexpected behavior occurs, developers need logs, traces and reasoning traces to investigate it. Khalil said organizations should ensure they own their harness and agent traces, particularly when those systems interact with sensitive intellectual property.

For sensitive workloads, Khalil said companies can run open models that they control and customize. He also said a model router can help identify where requests should be processed and, when appropriate, redact personally identifiable information before sending work to an external model.

Model Routing Across Local, Cloud and Frontier Systems

Khalil said model routers such as NeMo Switchyard could become increasingly important as organizations use multiple models and hardware environments. He described a setup in which a team can use shared API keys to access a model running on a DGX Station, while also drawing on internal compute clusters, cloud resources and frontier-model services.

He said local hardware provides lower latency, direct visibility into where data resides and a lower-cost option than building a data center, while a router and harness can help organizations use different resources based on policy, data sensitivity and performance requirements.

Looking ahead one year, Khalil said he expects users to consume more tokens from open models than from closed frontier models, even as demand for both continues to grow. He said the expansion of open models could also help the industry share security lessons more quickly as companies develop and deploy AI agents.

About NVIDIA (NASDAQ:NVDA)

NVIDIA Corporation, founded in 1993 and headquartered in Santa Clara, California, is a global technology company that designs and develops graphics processing units (GPUs) and system-on-chip (SoC) technologies. Co-founded by Jensen Huang, who serves as president and chief executive officer, along with Chris Malachowsky and Curtis Priem, NVIDIA has grown from a graphics-focused chipmaker into a broad provider of accelerated computing hardware and software for multiple industries.

The company’s product portfolio spans discrete GPUs for gaming and professional visualization (marketed under the GeForce and NVIDIA RTX lines), high-performance data center accelerators used for AI training and inference (including widely adopted platforms such as the A100 and H100 series), and Tegra SoCs for automotive and edge applications.