What NVIDIA is
NVIDIA is a semiconductor and AI systems company whose GPUs have become the dominant compute substrate for training and serving large AI models. Its H200 and Blackwell-generation chips are the hardware most frontier labs run on; its CUDA software stack is the programming model most AI researchers write to. But the company has expanded well beyond silicon: it now publishes open-weights models, runs inference microservices, co-develops architectures with partner labs, and holds financial stakes in the companies that are its largest customers.
Why it matters
Every major AI training run in the events bundle — Anthropic's Claude family, OpenAI's GPT-5.x line, Mistral's 675B Mistral Large 3 (trained on 3,000 NVIDIA H200 GPUs), Apple's AFM 3 — runs on NVIDIA hardware or is co-optimized for it. The company is structurally embedded in the AI supply chain at a depth that makes it difficult to displace quickly, even as customers actively try. OpenAI is pursuing custom silicon with Broadcom (targeting 10 GW by 2029) and AMD (6 GW of Instinct GPUs), yet simultaneously signed a 10 GW datacenter partnership with NVIDIA and accepted $30B of NVIDIA capital in its Series H. Anthropic signed a deal for up to 1 GW of Grace Blackwell and Vera Rubin compute while also expanding Google TPU and Amazon Trainium usage. The pattern is diversification at the margin, not displacement at the core.
The model portfolio: from inference microservices to open-weights frontier
NVIDIA's model strategy has two distinct layers.
NIM and co-developed models — NVIDIA packages third-party models (including Mistral NeMo, released jointly with Mistral AI) as NIM inference microservices, making them available as drop-in, optimized endpoints. This extends NVIDIA's reach into inference without requiring it to own the model weights.
The Nemotron family — NVIDIA's own open-weights model line, released under permissive commercial licenses, now spans the full capability spectrum:
- Nemotron 3 Ultra (550B total / 55B active, hybrid Mamba-transformer MoE, 1M context) ranks as the highest-scoring U.S. open-weights model on the Artificial Analysis Intelligence Index at 47.7–48.2, though it trails leading Chinese models like Kimi K2.6 and DeepSeek V4 Pro.
- Nemotron 3 Super (120B / 12B active, hybrid Mamba-2/transformer/MoE, 1M context) claims 442 tokens/second and leads the PinchBench agentic evaluation, outperforming models with far more total parameters.
- Nemotron 3 Nano 4B targets on-device deployment; Nemotron 3 Nano Omni handles cross-modal document, audio, and video reasoning for agentic workflows.
- Nemotron 3 Embed ranks #1 on the RTEB agentic retrieval benchmark — a practically important result as retrieval quality becomes a bottleneck in production agent systems.
- Audex-30B, from NVIDIA's Nemotron Labs, is a 30B MoE audio-text model trained on 157.4B audio tokens that achieves state-of-the-art audio performance with minimal regression on text reasoning — released with public model checkpoints.
NVIDIA has committed $26 billion over five years to open-weights model development, framing this partly as a strategic response to Chinese labs building capable open-weights models on non-NVIDIA hardware. A healthy open-weights ecosystem drives AI semiconductor adoption regardless of which lab's weights win.
Cosmos — NVIDIA's world-model platform for physical AI. Cosmos 3 was released as the first open omni-model targeting physical AI reasoning and action, with Cosmos 3 Edge following for edge deployment scenarios. These position NVIDIA directly in the robotics and embodied-AI market.
The partnership and investment web
NVIDIA's financial and strategic entanglements with frontier labs are extensive:
- OpenAI: $30B investment in the Series H; separate 10 GW datacenter partnership with first-phase launch in 2026.
- Anthropic: Up to $10B investment; up to 1 GW of Grace Blackwell and Vera Rubin compute; deep technology partnership to co-optimize model performance and future NVIDIA architectures for Anthropic workloads. Anthropic's SpaceX Colossus deal added 220,000+ NVIDIA GPUs to its compute pool.
- Mistral AI: Investor in Mistral's €1.7B Series C; co-developed Mistral NeMo; Mistral is a founding member of the NVIDIA Nemotron Coalition; Mistral Compute (Mistral's sovereign AI infrastructure product) is built on NVIDIA hardware.
- Project Glasswing: NVIDIA is a member of Anthropic's cybersecurity consortium alongside AWS, Apple, Google, Microsoft, and CrowdStrike.
- Apple AFM 3: Apple's new foundation models were built with Google and NVIDIA infrastructure.
AI in NVIDIA's own chip design
NVIDIA uses AI across five stages of its chip design pipeline, as described by chief scientist Bill Dally at GTC 2025. NVCell, a reinforcement learning and genetic algorithm system, redesigns approximately 2,500–3,000 layout cells overnight — work that previously required 10 engineer-months. PrefixRL uses RL to design arithmetic circuits that are 20–30% better than human designs. ChipNeMo and BugNeMo are LLaMA 2-based LLMs fine-tuned on internal GPU documentation for engineering assistance. Dally acknowledged that fully autonomous GPU design from a prompt remains a distant goal, but the internal deployment demonstrates measurable, production-grade gains.
Enterprise software: NemoClaw and the agentic stack
At GTC 2026, NVIDIA unveiled NemoClaw, an enterprise software stack integrating with OpenClaw to add security and governance for agentic deployments, with launch partners including Salesforce, Cisco, and CrowdStrike. NVIDIA's NeMo Retriever topped the ViDoRe v3 leaderboard using a ReACT-based agentic retrieval loop. These moves extend NVIDIA's surface area from hardware and model weights into the enterprise software layer — the same layer where Microsoft, Anthropic, and OpenAI are competing.
The geopolitical fault line
The clearest structural risk in the events bundle is the DeepSeek episode: DeepSeek gave Huawei several weeks of pre-release access to DeepSeek-V4 for hardware optimization while denying the same access to NVIDIA and AMD. This is a direct inversion of the prior norm and signals that Chinese frontier labs are actively aligning their hardware optimization work with domestic chipmakers. A Reuters report (unverified sourcing) claimed DeepSeek-V4 was trained on NVIDIA's most advanced chips despite U.S. export controls — which, if accurate, means NVIDIA hardware is still present in Chinese training runs even as the software and optimization relationships shift away. The trajectory points toward a bifurcated supply chain in which NVIDIA retains dominance in Western markets while losing the optimization partnership that has historically made its chips the default choice.
Where it's heading
The events collectively describe a company that is simultaneously the infrastructure layer everyone depends on and the entity everyone is trying to partially route around. NVIDIA's response has been to deepen its position at every layer — hardware, software, models, enterprise stack, and capital — while betting that open-weights model proliferation (which it funds and enables) will expand the total market for AI compute faster than custom silicon can erode its share. The robotics and physical-AI push via Cosmos represents the next frontier: if embodied AI scales the way language AI has, NVIDIA's early position in that stack could replicate the dynamic it achieved in LLM training.




