NVIDIA Blog https://blogs.nvidia.com/ Fri, 11 Sep 2026 16:54:30 +0000 en-US hourly 1 https://wordpress.org/?v=7.0.4 Skild AI Taps NVIDIA Physical AI to Teach Robots New Tasks From a Single Video https://blogs.nvidia.com/blog/skild-ai-s1-physical-ai/ Thu, 10 Sep 2026 16:30:35 +0000 https://blogs.nvidia.com/?p=98131

Manufacturing floors, warehouses and production lines rarely stay fixed — tasks change, layouts shift and new products arrive, and most robots can’t keep up without significant reprogramming.

Skild AI’s new S1 robot foundation model helps address this, designed to learn previously unseen, long-horizon tasks from a single video demonstration. The model, launched last week, uses video as input to understand and execute the task without updating its weights or undergoing task-specific post-training — a technique called in-context learning.  

Skild built S1 and conducted the research on NVIDIA AI infrastructure, part of a broader collaboration spanning synthetic data generation, model training, simulation and real-world physical AI deployment. The companies are working together to move adaptable robot intelligence from the lab into factories and other dynamic operating environments.

“Learning by experience, and not preprogramming, is the step change that has happened in robotics,” said Deepak Pathak, cofounder and CEO of Skild AI. “NVIDIA Isaac Lab and NVIDIA Cosmos technologies help Skild create the scalable, diverse experience its robots need to learn across many scenarios and embodiments.”

The launch comes as the company reached a $100 million annual revenue run rate 10 months after its first commercial deployment. In that time, Skild has built more than 60 deployment partnerships with work spanning manufacturing, logistics, inspection, security, food preparation and other applications. 

Learning New Work From One Video

Most industrial robots are built for fixed jobs, so each new product, process or layout requires more data, retraining and validation.

S1 takes a different approach: An operator records a video of the desired task and provides it to the model as a prompt. It interprets the demonstrated intent, objects and sequence, then maps them into actions for the robot in front of it — with no retraining — and often for a task not covered by its pretraining dataset. 

S1 can perform unfamiliar tasks lasting up to 10 minutes, including plant potting, pancake making, pour-over coffee brewing and kit assembly. These tasks can span dozens of manipulation steps and require the robot to compose skills in sequences it hasn’t previously performed. 

 

In one plant-potting test, the Skild AI team moved from recording the demonstration to autonomous execution on hardware in just 11 minutes. The model can also adjust when objects move, recover from errors and combine skills in sequences that weren’t explicitly programmed.

In Skild’s tests on new, multistep tasks, its S1 robot succeeded about 66% of the time at each step, compared with 9% for a similar AI system — a more than sevenfold improvement. Skild also estimates that showing the robot one short video example can be as useful as giving it roughly 380 hands-on training examples. A person collecting those examples manually could take 50-100 hours.

From Research to Factory Work

S1 breaks the cycle of needing to constantly retrain robots for new factors by letting operators demonstrate new tasks directly without requiring a new dataset or training run for every change. Where customer agreements permit, experience from Skild’s commercial deployments can inform the broader model and help accelerate future deployments.

That work is already in action on the factory floor. Skild, NVIDIA and Foxconn are deploying the Skild Brain on dual-arm manipulators for high-precision assembly of NVIDIA Blackwell systems. In one demonstrated workflow, a robot installs a busbar and limit block, fastens 16 screws and adapts to disturbances across a multistep task. The work requires precise motion, contact-aware control, sequence tracking and recovery when the scene differs from the plan.

 

NVIDIA Technology Across the Development Cycle

NVIDIA accelerated computing gives Skild the scale to train its shared robot brain using simulation, human video, teleoperation and, where permitted, deployment data. NVIDIA Cosmos open world foundation models help diversify training data and turn video into structured descriptions, while Cosmos Curator helps annotate, filter and organize data at scale.

Skild is extensively using NVIDIA’s open simulation frameworks to train and validate its robot brain before real-world deployment. NVIDIA Omniverse libraries and the NVIDIA Isaac Sim framework provide physically based virtual environments for generating data, testing edge cases and validating behaviors.

Skild further strengthens the skills of its brain through reinforcement learning in Isaac Lab, an open modular robot learning framework. Powered by the Newton physics engine, Isaac Lab helps Skild’s engineers accurately model various physical parameters, such as forces, contact, collision and pressure, and reduce the simulation-to-reality gap.

Skild and NVIDIA are also jointly developing new GPU-accelerated simulation solvers that quickly and accurately model how robots physically touch, grip and manipulate solid objects. They’ll soon be made available to all developers as part of Newton. 

As models move toward production, NVIDIA Nsight tools help engineers find performance bottlenecks during training, and the NVIDIA TensorRT software development kit optimizes inference so robots can respond quickly in the physical world. Together, these technologies connect the data, simulation, training and deployment stages instead of treating them as separate systems.

Read Skild AI’s S1 research and explore the NVIDIA Isaac robotics platform.

]]>
Physical AI Takes the Wheel: How the World’s Robotaxi Leaders Are Building With NVIDIA Technologies https://blogs.nvidia.com/blog/robotaxi-leaders-full-stack-open-platform/ Thu, 10 Sep 2026 16:00:04 +0000 https://blogs.nvidia.com/?p=98141

The global robotaxi market — physical AI’s first commercial breakthrough — is projected to reach $400 billion by 2035, with over 6 million commercial vehicles in operation as driverless fleets are already moving people through some of the world’s busiest and most complex streets.

Deploying a driverless vehicle is one challenge. Scaling a fleet is a next-level computing challenge; it means delivering the same safe, reliable performance across thousands of vehicles. 

Meeting those demands requires enormous amounts of compute across the robotaxi development lifecycle, from preparing and training AI models to simulating and validating driving behavior, as well as real-time processing in the vehicle. 

NVIDIA provides an open platform for AI training, simulation and safety validation, with libraries, software development kits, workflows and models that developers can use alongside their own technology stacks.

Every major robotaxi program operating at commercial scale today is running on NVIDIA’s modular stack, spanning AI training, simulation, in-vehicle computing — or a combination of the three — to develop and deploy fleets at scale. 

What Is a Robotaxi Technology Stack?

A robotaxi technology stack is the end-to-end set of technologies used to develop, validate and deploy autonomous vehicles (AVs) — from data and AI model training to simulation, safety validation and real-time in-vehicle computing. 

NVIDIA’s robotaxi and AV platform brings these capabilities together in a three-computer solution: the model training computer, simulation and validation computer, and in-vehicle computer.

1. Training Computer: NVIDIA DGX

Robotaxi intelligence advances as programs turn growing volumes of fleet data into increasingly capable models. Driving models can be trained on NVIDIA DGX systems. 

The NVIDIA Alpamayo portfolio of open reasoning vision language action (VLA) models, simulation frameworks and physical AI datasets gives developers building blocks they can adapt to their own data, requirements and technology stacks. Its reasoning models help address long-tail AV challenges by breaking complex driving situations into smaller steps, reasoning through each one and selecting the safest trajectory. 

NVIDIA also provides physical AI datasets, reinforcement learning blueprints and recipes for post-training and distillation, helping developers optimize models for their target vehicles.

On a challenging autonomous driving evaluation, adding meta-action and chain-of-thought reasoning data improved a VLA model’s trajectory prediction accuracy, reducing minimum average displacement error — the predicted path’s average deviation from the reference route — by 43%, from 2.08 to 1.18.

2. Simulation and Validation Computer: NVIDIA Omniverse and Cosmos on NVIDIA RTX PRO 

Robotaxi programs can’t rely on physical miles alone to capture rare, long-tail driving scenarios. NVIDIA Omniverse NuRec models reconstruct real-world driving scenarios from sensor data, while NVIDIA Cosmos world foundation models generate physically based variations of them, enabling developers to turn thousands of real-world corner cases into millions of combinations of driving behavior, traffic, weather, lighting and sensor conditions.

From real-world corner cases to thousands of synthetic permutations spanning behavior and content, NVIDIA Cosmos variations expand AV training data and accelerate model deployment.

Running on NVIDIA RTX PRO Servers, NVIDIA Omniverse and Cosmos support closed-loop simulation and validation. The NVIDIA AlpaSim simulation framework extends the workflow for training and evaluating reasoning-based autonomous-driving models, helping developers identify weaknesses before deployment.

3. In-Vehicle Computer and Sensor Architecture: NVIDIA DRIVE Hyperion With DRIVE AGX

NVIDIA DRIVE Hyperion is NVIDIA’s modular in-vehicle compute and sensor reference architecture for level-4-ready robotaxis. DRIVE Hyperion 10 pairs dual NVIDIA DRIVE AGX Thor systems-on-a-chip, built on the NVIDIA Blackwell platform, with 14 high-definition cameras, nine radars, three lidars and 12 ultrasonics for real-time, 360-degree sensor fusion. Its redundant compute and sensing design supports fail-operational driving if a sensor or compute component fails.

The dual DRIVE AGX Thor is designed to run modern AI workloads — including VLA models — for perception, reasoning, path planning and driving actions.

NVIDIA Halos provides a production-ready safety foundation through Halos OS, and a broader validation and certification framework spanning independent inspection, system validation, large-scale simulation and continuous testing from cloud to car.

Robotaxi Leaders Adopting NVIDIA’s Robotaxi Technology Stack

NVIDIA’s robotaxi ecosystem spans every region where commercial robotaxi services are emerging today: Asia, Europe, the Middle East and North America. Across these markets, mobility providers, AV developers and automakers are adopting NVIDIA’s three-computer architecture to train AI models, simulate and validate driving behavior, and deploy autonomous vehicles at scale.

Scaling Robotaxi Services Globally

  • Uber is scaling its fleet of NVIDIA DRIVE Hyperion, with plans to reach 28 cities by 2028. Uber and NVIDIA are also building a robotaxi AI data factory on NVIDIA Cosmos to curate fleet driving data for rare scenarios. Together, Uber and NVIDIA are collaborating with Autobrains, Avride, Lucid, May Mobility, Mercedes-Benz, Momenta, Nissan, Nuro, Pony.ai, Stellantis, Waabi, Wayve, WeRide and Zoox to bring NVIDIA-powered robotaxi services to the Uber platform.
  • May Mobility is planning to operate autonomous ride-hailing services through Uber’s network, while developing its software stack on the NVIDIA DRIVE platform.
  • Bolt uses NVIDIA technologies to develop and scale AVs across Europe.
  • Lyft plans to use NVIDIA DRIVE Hyperion as a reference architecture for future autonomous fleets. May Mobility vehicles are also currently operating on Lyft’s network in Atlanta, powered by NVIDIA DRIVE.
  • Through its partnership with Grab, WeRide plans to bring its DRIVE Hyperion- and DRIVE AGX Thor-based GXR to key markets across Southeast Asia.
  • Waymo partners with NVIDIA to help build its autonomous computing system.

Building Robotaxi Intelligence 

Behind these services, AV developers are using NVIDIA accelerated computing, simulation and in-vehicle platforms to build the intelligence that operators and automakers deploy.

  • Wayve, Nissan and Uber are developing a global robotaxi program using a prototype vehicle that combines Nissan’s vehicle engineering, Wayve’s embodied AI and the NVIDIA DRIVE Hyperion platform.
  • Autobrains is developing robotaxi programs with Uber in Munich and VinFast in Southeast Asia, built on NVIDIA DRIVE Hyperion and enabled by Autobrains’ Agentic AI technology.
  • Zoox uses NVIDIA DRIVE for in-vehicle computing and cloud-based training and simulation. 
  • Momenta is developing its software stack based on NVIDIA DRIVE AGX running on DriveOS. 
  • Pony.ai developed its new-generation autonomous-driving domain controller with NVIDIA DRIVE Hyperion and DRIVE AGX Thor. 
  • Tensor is developing its level 4 Robocar with eight NVIDIA DRIVE AGX Thor systems-on-a-chip in its in-vehicle supercomputer. 
  • Waabi expands into the robotaxi market through a deployment collaboration with Uber; its Waabi Driver platform is built on NVIDIA DRIVE AGX Thor.
  • TIER IV and Isuzu are deploying level 4 autonomous buses built on NVIDIA DRIVE Hyperion and DRIVE AGX Thor.
  • Lenovo is supplying its NVIDIA DRIVE AGX Thor-based AD1 level 4 domain controller for a next-generation robotaxi program with SWM.
  • DeepRoute.ai is developing a new generation of robotaxis built on the NVIDIA DRIVE Hyperion platform with DRIVE AGX Thor.

Bringing Robotaxis Into Production

As these systems move from development into production, automakers are integrating NVIDIA technology into autonomous and robotaxi-ready vehicle programs.

  • Tesla trains its autonomous-driving neural networks on NVIDIA supercomputers. 
  • Mercedes-Benz and NVIDIA are collaborating with Uber to develop a robotaxi ecosystem based on the new S-Class, built on the NVIDIA DRIVE Hyperion architecture, full-stack NVIDIA DRIVE AV L4 software, and NVIDIA Alpamayo open AI models, simulation tools and datasets to support reasoning-based, safety-first autonomy.
  • Stellantis, Wayve and Uber are collaborating to develop and deploy L4 driverless mobility services, leveraging NVIDIA DRIVE Hyperion and AI computing technologies.
  • Lucid, Nuro and Uber are developing a global robotaxi service using NVIDIA DRIVE AGX Thor, part of the DRIVE Hyperion platform.
  • Hyundai Motor and Kia are expanding their collaboration with NVIDIA to develop data-driven autonomous-driving systems built on NVIDIA DRIVE Hyperion. NVIDIA will also explore expanded collaboration with Hyundai Motor Group’s joint venture, Motional, to advance level 4 robotaxi services.
  • Geely, alongside its ecosystem partners, plans to develop and commercialize robotaxis using DRIVE Hyperion.
  • Zeekr, a Geely Auto Group brand, has adopted DRIVE AGX Thor for a centralized domain controller.

From cloud to car, nearly every layer of the robotaxi platform is being developed on NVIDIA accelerated computing. 

Explore NVIDIA’s complete platform for robotaxi development.

]]>
d-Matrix Adopts NVIDIA NVLink Fusion for Rack-Scale XPU Deployment https://blogs.nvidia.com/blog/d-matrix-nvlink-fusion/ Thu, 10 Sep 2026 13:00:21 +0000 https://blogs.nvidia.com/?p=98153

AI inference chipmaker d-Matrix today announced it will use NVLink Fusion to connect its next-generation Raptor XPUs to NVIDIA’s AI infrastructure platform — joining a growing roster of ecosystem partners.

By connecting Raptor to NVIDIA NVLink scale-up and Spectrum-X scale-out networking, the NVIDIA MGX rack architecture and the broader NVIDIA AI platform, NVLink Fusion gives d-Matrix an accelerated, lower-risk path from custom silicon to large-scale deployment. 

“Demand for inference is soaring, but capital, time and energy remain finite,” said Sid Sheth, cofounder and CEO of d-Matrix during a press briefing yesterday. “With NVLink Fusion and MGX, we can integrate our Raptor XPUs into a broadly deployed, liquid-cooled architecture, giving customers a faster, lower-risk path to deploy and scale ultralow-latency inference.”

The NVIDIA AI platform is vertically integrated and horizontally open. NVLink Fusion extends this openness to XPUs and CPUs, allowing silicon companies to focus on their processor innovations while using NVIDIA infrastructure to deploy them at AI factory scale.

From Custom Silicon to Rack-Scale Deployment

Building an XPU is only the first step. Deploying it at AI factory scale requires a complete platform spanning networking, rack architecture, power, cooling, software and a proven supply chain. Each of the steps — sourcing chips, integrating high-speed interfaces, validating a scale-up networking solution, designing and certifying rack architecture — adds time, costs and risks.

NVLink Fusion lets silicon innovators connect directly into NVIDIA’s proven platform. It’s the high-bandwidth, low-latency technology that connects custom XPUs and CPUs to the NVIDIA stack.

By adopting NVLink Fusion, d-Matrix can tap into the NVIDIA MGX ecosystem’s mature, validated rack designs, supply chain, power and cooling infrastructure. By standardizing on a common rack, data centers can be built once and support GPUs, CPUs and XPUs — without requiring a separate rack architecture for each processor type.

NVLink Fusion Integrates d-Matrix XPUs Into AI Factories

Using NVIDIA NVLink, d-Matrix plans to connect its XPUs in a single high-bandwidth, low-latency scale-up domain. Its racks can also work alongside NVIDIA GPU-based systems like NVIDIA Vera Rubin NVL72 for disaggregated inference. 

d-Matrix also plans to integrate NVIDIA Vera CPUs, NVIDIA ConnectX-9 SuperNICs, NVIDIA BlueField-4 DPUs and NVIDIA Spectrum-X Ethernet networking. Together with NVLink and MGX, these technologies give d-Matrix a proven foundation to deploy specialized inference alongside NVIDIA systems within flexible, unified AI factories.

NVLink Fusion Opens NVIDIA AI Factories to Specialized XPU Architectures

The NVIDIA full-stack AI factory platform includes NVIDIA Vera Rubin NVL72, Groq 3 LPX, the Vera CPU rack, Vera BlueField-4 STX storage and Spectrum-6 SPX Ethernet networking. It’s designed to be completely fungible — running every AI workload, model and model architecture — with the best performance per watt and lowest cost per token. 

NVLink Fusion gives customers the flexibility to match the right compute to each workload within a common AI factory platform, opening access to NVIDIA networking, systems, software and global supply chain. Silicon innovators like d-Matrix can use NVLink Fusion to increase performance, accelerate time to market and reduce the risk of deploying semi-custom AI factories.

 

]]>
Boots on the Ground: ‘WARDOGS’ Goes All Out on GeForce NOW at Early-Access Launch https://blogs.nvidia.com/blog/geforce-now-thursday-wardogs/ Thu, 10 Sep 2026 13:00:18 +0000 https://blogs.nvidia.com/?p=98103

Gear up: The latest PC games and major updates are ready to play on GeForce NOW this week. WARDOGS drops onto the cloud at early-access launch, alongside the Valheim 1.0 Deep North update and Bus Simulator 27 — part of nine new titles joining the cloud.

The newest PC releases can demand serious hardware, storage space and upgrade budgets. GeForce NOW puts GeForce RTX-powered performance in the cloud, so members can jump into the newest titles across supported devices without needing expensive PC upgrades to keep up.

From a 100-player tactical firefight to a frozen Viking frontier and a growing bus empire, every adventure is ready when members are — no downloads or storage management required.

Who Let the ‘WARDOGS’ Out?

 

Take on the Control Zone: WARDOGS launches on GeForce NOW today, bringing BULKHEAD’s large-scale tactical first-person shooter to the cloud at early-access launch.

Up to 100 players across three teams battle to control randomized zones in a large-scale military sandbox, where tactical gunplay, combined-arms combat, building and destruction can change the course of a match.

There’s no single route to victory. Outfit a loadout, coordinate a squad, fortify a strategic position or take the fight to the enemy in vehicles. Every team play action earns cash, giving players more ways to invest in gear and vehicles that can turn the tide.

Ultimate members can stream the fight with GeForce RTX 5080-class performance in the cloud for a tactical advantage. Pick up the battle across supported devices without waiting on a huge game install or clearing local storage first.

Ice, Ice in the Deep North, Baby

Sharpen the axes and raise the shields – Valheim has left early access with its 1.0 update, letting members stream the final biome in Iron Gate’s Viking survival saga through Steam and Xbox Store. 

Set sail for the Deep North, a frozen frontier filled with new enemies to face, weapons to craft and strongholds to build before the cold closes in. The 1.0 launch also arrives alongside new platform versions with full crossplay support, making it easier to rally the clan.

Whether returning to a long-running world on a PC or Mac or beginning a fresh voyage on a Steam Deck or Amazon Fire TV Stick, members can head to the Deep North across supported devices without waiting for downloads.

The Wheels on the Bus Go ‘Round in the Cloud

Bus Simulator 27 on GeForce NOW
The route to the top starts here.

Hop into the driver’s seat. Bus Simulator 27 is now streaming from GeForce NOW. 

Drive a fleet of more than 45 officially licensed buses from 13 manufacturers through Felicia Bay, a sunny region with two cities, 20 districts and a full day-and-night cycle. Build routes, manage timetables and grow a successful bus company in Story, Career or Sandbox mode.

And there’s even more to stream from the GeForce NOW library this week:

  • Bus Simulator 27 (New release on Steam, Sept. 8)
  • Halloween: The Game (New release on Steam, Sept. 8)
  • Honeycomb: The World Beyond (New release on Steam, Sept. 8)
  • Wanderburg (New release on Steam, Sept. 8)
  • Shroom and Gloom (New release on Steam, Sept. 10)
  • Train Sim World 7 – Advanced Access (New release on Steam, Sept. 10)
  • WARDOGS (New release on Steam, Sept. 10)
  • Welcome to Elderfield (New release on Steam, Sept. 10)
  • Death Stranding Director’s Cut (Epic Games Store)

Dates listed above reflect when games are released on their respective stores. GeForce NOW availability may vary, as games are onboarded after release and added throughout the week. Keep an eye on GeForce NOW channels and GFN Thursdays for availability updates on announced titles.

One GFN community member recently returned after six years, saying NVIDIA DLSS 5 technology renewed their interest in GeForce NOW cloud gaming. Ready to try it out, too? Start with a day pass to test premium cloud gaming before committing to a membership. For those who decide to keep playing, the cost can be applied toward their first monthly membership.

What are you planning to play this weekend? Let us know on X or in the comments below.

]]>
NVIDIA Brings Real-Time AI to Broadcast, Sports and Global Streaming at IBC https://blogs.nvidia.com/blog/ibc-news-2026/ Wed, 09 Sep 2026 16:00:42 +0000 https://blogs.nvidia.com/?p=98109

At the IBC conference, running Sept. 11-14 in Amsterdam, the creative, technology and business communities are coming together to turn ideas into action and discuss innovations across the media and entertainment industries. More than 44,000 attendees from 170+ countries are gathering to explore 1,300+ exhibitions in 14+ halls and outdoor spaces, with over 600 speakers delivering insights.

Read on to learn more about what NVIDIA’s highlighting at the show.


NVIDIA AI for Media Brings Real-Time Intelligence to Broadcast, Sports and Production Workflows 🔗

Media companies are increasingly integrating AI into live production, sports, news and streaming workflows to unlock richer performance insights, verify video authenticity, and enhance and localize content — all without disrupting trusted broadcast environments.

At IBC 2026 in Amsterdam, NVIDIA announced a major expansion to NVIDIA AI for Media — a collection of GPU-accelerated software development kits (SDKs), NVIDIA NIM microservices, playbooks, and blueprints that enhance audio, video and augmented-reality effects for media and entertainment workflows — to unlock new ways to understand motion, verify and enhance video, localize programming and build AI-powered media applications. 

The NVIDIA Synthetic Video Detector (SVD) NIM microservice, announced earlier this year at SIGGRAPH, helps organizations assess the probability of whether footage is authentic or AI-generated, giving editorial, content-authentication, digital-forensics and media-integrity teams another point of analysis in their review process.

Since its initial release, SVD’s accuracy has reached 99.3% for text-to-video content and 97.7% for image-to-video content, with especially large gains on difficult image-to-video cases. 

Dalet is integrating SVD into a secure, cloud-hosted verification workflow for news organizations. This allows editorial teams to submit footage through SVD, inspect and review the resulting scores and metadata within a Dalet interface. 

TwelveLabs announced the general availability of Compliance by TwelveLabs, its first application built on the company’s video intelligence platform, helping media and broadcast teams rapidly screen content against regional and custom compliance standards. The solution integrates SVD to add frame-level authenticity signals and confidence scores, enabling media teams to identify potentially synthetic media within the same compliance workflow.

Wowza, whose Wowza Streaming Engine media server technology powers more than 35,000 video deployments across over 170 countries, will distribute SVD through the Wowza Video Intelligence Framework. The solution, powered by NVIDIA-accelerated infrastructure, will enable broadcasters, streaming providers and other organizations to analyze live video feeds and extract data around detected objects, scenes and signs of AI generation in real time. It can be deployed and run on premises, at the edge, in the cloud, across hybrid deployments or fully air-gapped, giving organizations greater control over critical media workflows.

NVIDIA 3D Body Pose estimates 2D and 3D human joint locations and angles from video captured by a single camera, helping turn motion into structured data without marker-based capture systems.

For sports organizations, that data can support player and athlete movement tracking, biomechanics and performance analysis, replay enhancement, officiating and adjudication workflows, player-safety applications, and virtual interaction and immersive experiences. 

The technology can also provide structured human-motion data for content-creation workflows. When mapped to a compatible character rig, joint and motion data can serve as input for animation blocking, digital doubles, character retargeting and virtual-production experiences.

Vizrt is using Body Pose technology in live virtual-studio environments, with tracked body movement driving real-time 3D lighting effects such as reflections, shadows and environmental rendering.

Video Frame Generation (VFG) makes video motion appear smoother by using generative AI to create new frames between the original frames of a video. It can increase frame rates by 2x or 4x while preserving visual quality and temporal consistency, enabling more fluid sports, slow-motion replays, live media and other high-motion video experiences. VFG also supports frame-rate conversion and frame boosting for generative AI video workflows.

Ross Video is integrating VFG into its Rio Replay platform to create AI-assisted slow-motion video for sports production.

The work supports 6x slow-motion generation for sports replay. Development is underway toward 8x interpolation, meaning generated intermediate frames can give replay teams smoother motion without requiring every frame to be captured by an ultrahigh-frame-rate source camera. 

NVIDIA Video Super Resolution (VSR) uses AI to upscale video while reducing noise, blur and compression artifacts. New streaming modes let developers choose between real-time performance and higher image quality, while adjustable controls help achieve the desired level of enhancement. VSR also adds 10-bit video support and improves overall performance and quality. VSR is available through the NVIDIA Video Effects SDK and a NIM microservice for use in streaming, broadcast, conferencing, video playback and content-creation applications. 

The technology can support video players, conferencing applications, creator tools, streaming services, transcoders and broadcast systems through a common interface.

NVIDIA TrueHDR converts standard-dynamic-range video into high-dynamic-range output in real time, reaching up to approximately 2,000 nits while preserving local contrast and adapting brightness to the content.

VSR, VFG and TrueHDR can be combined within a single video-effects pipeline — helping media companies enhance existing content libraries for streaming, transcoding, gaming and creator workflows.  

The NVIDIA LipSync and Active Speaker Detection NIM microservices help developers build localization systems for interviews, news, sports, entertainment and other programming where multiple people may appear on screen.

LipSync transforms mouth movement in an input video to match a target audio track while preserving natural head pose, blinking and body movement. The new release improves facial occlusion handling and better preserves teeth, lip and facial textures.

The new Active Speaker Detection NIM microservice no longer requires speaker diarization for multiple audio tracks, adds voice activity detection and expands NIM microservice deployment support through a gRPC interface and broader GPU compatibility.

NDI is using NVIDIA AI for Media, including the NVIDIA LipSync NIM microservice, to enable real-time translation, lip-synced dubbing and regional language adaptation within existing broadcast workflows. By generating multiple language experiences from a common media stream, the approach can help broadcasters reach global audiences while reducing the bandwidth, infrastructure and production complexity traditionally required for multilingual distribution.

Studio Voice includes new Microphone Profiles built on NVIDIA Studio Voice NIM microservices, giving users more control over the tonal character of enhanced speech.

The capability is designed to suppress background noise, reduce room reverberation and improve speech clarity, then shape the enhanced output into a selected microphone profile for more polished live communications, streaming, podcasting and content creation.

Try NVIDIA AI for Media NIM microservices. See the latest NVIDIA and partner workflows at IBC 2026.


NVIDIA Holoscan for Media Provides Open Media Exchange Layer to Build and Connect Live Media Applications 🔗

As broadcasters, streaming services and sports organizations adopt software and AI, the infrastructure behind live content is becoming more flexible, more connected and increasingly built on shared accelerated computing.

The integration of Media Exchange Layer (MXL) with NVIDIA Holoscan for Media accelerates that transition — providing the common exchange layer that helps media applications connect and operate together. 

Holoscan for Media is an open reference architecture and developer toolkit for building AI-powered media functions and applications for software-defined live production. MXL adds an open way for those software-based media functions to exchange live video, audio and data across a distributed environment. 

As production functions move into software, developers can build applications that share accelerated infrastructure, connect dynamically and evolve independently. That can help media companies use infrastructure more efficiently, introduce new capabilities faster and reduce the amount of custom integration required between applications.

The integration also creates a stronger foundation for AI in live media. AI processing, video applications and traditional media functions can increasingly operate on the same accelerated infrastructure and in the same software-defined environment.

For technology vendors, this expands the opportunity to build applications that can work across broader, multi-vendor ecosystems. For media companies, it creates a path toward infrastructure that can adapt as formats, applications and AI capabilities change.

See the demo at IBC in EBU Stand 10.D21. ​Learn more about Holoscan for Media.


NVIDIA Sports Intelligence Playbooks Chart a Path to Multimodal AI for Sports 🔗

Sports is becoming a proving ground for a broader shift in AI: from general-purpose models toward fine-tuned open models built on proprietary data.

NVIDIA Sports Intelligence Playbooks are designed to accelerate that transition. They give leagues, media companies and technology providers structured frameworks to fine-tune NVIDIA open models on their own sports footage and annotations, creating multimodal AI that can understand the rules, players, scoring, strategy and context unique to a sport.

Sports organizations hold large volumes of proprietary video, metadata and performance information that are difficult for competitors to replicate. The playbooks provide a practical blueprint for converting those assets into AI capabilities that can underpin new analytics products, media experiences, automation tools and revenue streams.

The playbooks span the AI lifecycle, including data preparation, fine-tuning, inference, evaluation, optimization and deployment, and bring together NVIDIA technologies including Nemotron, NeMo AutoModel, Megatron Bridge, NIM microservices and NVIDIA accelerated computing. 

By providing an integrated path from model customization to production, Sports Intelligence Playbooks can reduce the cost and complexity of building specialized sports AI while increasing demand across its compute, software and inference stack.

Early testing demonstrates the potential of domain specialization. When evaluated on previously unseen footage using question formats similar to those used in training,  multiple-choice accuracy increased from approximately 53% to 94% and open-ended evaluation from approximately 5.7% to 66%.

Machina Sports is integrating Sports Intelligence Playbooks with its sports-native data, evaluation and agent infrastructure, enabling rights holders to turn proprietary media and expertise into private, deployable intelligence for live production, content and fan experiences.

The opportunity also expands as agentic AI becomes increasingly adopted. With the NVIDIA AI-Q Blueprint, organizations can use their domain-specific sports models as expert intelligence within agents that reason across video, enterprise data and software systems, extending the playbook from sports understanding into decision-making and automation.

Wowza is integrating vision language models, including NVIDIA Cosmos 3 and Nemotron, into the Wowza Video Intelligence Framework, fine-tuned through NVIDIA Sports Intelligence Playbooks to detect sports-specific moments in live streams and reduce time to action. 

Explore NVIDIA Sports Intelligence Playbooks.


NVIDIA Brings Multilingual Content Localization to Live Broadcast 🔗

Reaching global audiences with live programming requires more than translating words. Language nuances, voice, timing, facial movement, captions and onscreen graphics must work together in real time, while preserving the editorial intent and production quality of the original program.

To help broadcasters, sports leagues, rights holders and streaming services bring these elements into a unified, software-defined, real-time localization workflow, NVIDIA is bringing its Content Localization technologies to the NVIDIA Holoscan for Media developer toolkit. Designed for broadcast and streaming developers, the reference workflow enables captions, translated audio, dubbing, synchronized video and localized graphics.

Content Localization with Holoscan for Media provides a reference for how localization technologies can work together in software-defined broadcast applications. Developers can select the capabilities needed for each program, market or distribution channel rather than deploying separate infrastructure for every localized version.

Content Localization with Holoscan for Media incorporates the latest advancements from NVIDIA AI for Media, including improved LipSync when faces are partially obscured and enhanced Active Speaker Detection to help applications identify who’s speaking in multi-person scenes.

Expanding the Reach of Live Programming

Localization can transform the reach and economics of live programming. A shared, composable workflow can help media companies introduce regional coverage faster, serve more audiences and tailor experiences for individual markets — while preserving the timing, visual context and editorial control required for live production.

Technologies from AI-Media, CAMB.AI, Chyron and Panjaya each address a specific part of content localization with Holoscan for Media, from adapting voice and onscreen delivery to creating multilingual captions and translated audio, localizing graphics, and preserving expression and identity across live and on-demand content.

The Content Localization technologies also support file-based, streaming and post-production applications. Developers can use application programming interfaces for on-demand workflows and the Holoscan for Media reference workflow when localization must run as part of a live media environment. Together, they provide a consistent foundation for building multilingual media services across production and distribution.

Learn more about NVIDIA Holoscan for Media and AI for Media.

]]>
Sparks Fly: NVIDIA Accelerates Local AI at IFA 2026 https://blogs.nvidia.com/blog/local-ai-ifa-next-gen-agents-nv-pair-rtx-spark/ Thu, 03 Sep 2026 16:00:59 +0000 https://blogs.nvidia.com/?p=98045

Frontier intelligence is going local. At IFA 2026, NVIDIA, Microsoft and its partners are teaming up to provide faster inference and new tools that make agents easier to set up and run locally on NVIDIA hardware. New compact NVIDIA RTX Spark Windows PCs are also coming in October to give AI enthusiasts, developers and creators more ways to run capable agents locally and securely. 

Today’s announcements include:

  • Simplified local AI support for NVIDIA GPUs is coming in Hermes Agent, OpenClaw and Perplexity Portable Computer.
  • Up to 1.9x faster local inference — new llama.cpp and vLLM optimizations are available now directly and through LM Studio and Ollama. 
  • NVIDIA PAIR — a Personal AI Router tool that intelligently distributes AI inference across the PCs on a user’s local network.
  • NVIDIA RTX Spark arrives in October  — with new Windows PCs from Lenovo and Acer. Electronic Arts, Embark and Ubisoft are among the latest game publishers and developers bringing their blockbuster titles to NVIDIA RTX Spark.

Also, August was a busy month for local AI:

  • Nemotron 3.5 Lightning — which can run on NVIDIA RTX PCs, RTX PRO Workstations, DGX Spark and Jetson — is a 30-billion parameter model that has been launched. Get started with Nemotron 3.5 Lightning today. 
  • Z.ai’s GLM-5.3-Flash is a multimodal mixture-of-experts (MoE) model that’s bringing agentic AI to DGX Station.
  • Qwen has released Qwen3.8-Flash-Next, an open weight multimodal MoE model, which can run locally on DGX Spark and DGX Station, along with Qwen3.8-27B, a 27-billion-parameter open model optimized for local agentic and coding workloads on NVIDIA GPUs. 
  • LTX’s LTX 2.5 is an open-world video generation model optimized for NVIDIA RTX GPUs, DGX Spark and DGX Station, with new NVFP4, FastVideo and ComfyUI enhancements for faster, more memory-efficient local generation.
  • MiniMax-H3 is an open-weight video generation model with synchronized audio that can run locally on NVIDIA GPUs through ComfyUI. FastVideo teamed up with NVIDIA researchers to improve this further by releasing FastH3 — an open-weight, four-step distilled version that improves performance by 7x. Optimized FastVideo recipes for NVIDIA RTX GPUs and DGX Spark are coming soon.
  • Meta’s Muse Glimmer is a 30-billion-parameter open-weight model for coding and agentic workloads that can run locally on GeForce RTX PCs, DGX Spark, DGX Station and Jetson. NVIDIA has also released NVFP4 quantization with DGX Spark support for more memory-efficient local deployment.
  • DeepSeek v4 Flash is a 284-billion-parameter MoE model with 13 billion active parameters that can run locally on 2x DGX Spark cluster and DGX Station.

A Simpler Start for Local Agents

Getting a local agent up and running with local models required some effort — choosing a model, finding a compatible inference server, dialing in quantization settings and keeping everything updated. That friction is disappearing on RTX and DGX systems.

Three of the most widely used agent apps will offer simplified local model setup on Windows, each built on llama.cpp and incorporating NVIDIA’s latest inference optimizations. The new setup experiences are designed to reduce manual configuration and make it easier to get local agents up and running. 

Last month, Perplexity introduced its Portable Computer agent, giving users a simple way to run Perplexity locally on Linux systems like NVIDIA DGX Spark with the models, orchestration and tools packaged into a single app experience.

Perplexity Portable Computer is available on NVIDIA RTX GPUs with at least 24GB VRAM running on Linux, with support on Windows coming soon, bringing that same streamlined setup to a broader group of PC users. Users can run complete workflows locally without consuming credits, while selectively escalating parts of a task to one of 15+ frontier models in the cloud when additional research or reasoning is needed. Portable Computer asks for permission before sending content to the cloud, helping users keep sensitive information on their device. Here’s some example use-cases:

  • Engineering: Review open PRs in a connected GitHub repo and sort them into ready, blocked, stale, and needs review, each tagged with the next step. Docs that fell out of sync with the latest merge get caught and fixed, with a PR opened for the changes.
  • Finance: Point the agent at two years of brokerage summaries, consolidated 1099s and tax returns and have it trace the recurring holdings creating the most avoidable fees and tax drag, with every figure cited to the exact file and page — all without a document ever reaching a chatbot.
  • Startups: Ask why activation went flat, and the agent analyzes the funnel export locally to find where new signups drop off between install and first completed task, then posts the top insights straight to the team’s Slack channel.

Try Portable Computer today.

Hermes Agent — developed by Nous Research — is a general-purpose agent used by millions that excels at reliability and self-improvement. Model- and provider-agnostic, Hermes is built to run all day on local systems, making RTX PCs, RTX PRO workstations and DGX Spark a natural fit.

Configuring a local model in Hermes will provide users with one-click setup across RTX and DGX systems on Windows. The agent will automatically detect the NVIDIA GPU, select an appropriate model and configuration, and run it through integrated llama.cpp with NVIDIA inference optimizations already in place, eliminating manual model downloads and tuning. Support for Linux is coming soon.

Once it is running, Hermes works the way it does anywhere else. It uses tools, maintains context across tasks, remembers information between sessions and creates reusable skills over time, allowing the agent to become more capable with continued use. Running the model locally on a GPU keeps performance fast while keeping data on the system.

One-click local model setup is available now on Windows, with support coming soon to Linux. Learn more about Hermes Agent.

OpenClaw has become one of the defining projects of the open-agent movement — the largest AI project on GitHub, with more than 380K stars and a fast-growing community that’s building tools and skills across research, engineering, project management and everyday productivity.

NVIDIA, Microsoft and OpenClaw have been working together to make that experience easier to set up on Windows PCs. To reduce onboarding friction, the OpenClaw Windows App simplifies the process of setting up an optimized local model on any RTX GPU with at least 24GB of VRAM.

Learn more in the OpenClaw blog.

Faster Inference Gives Local Agents a Boost

Inference performance is critical to keeping local agents responsive. NVIDIA is continuing to collaborate with the open-source llama.cpp and vLLM communities to accelerate agentic workloads across local NVIDIA platforms.

llama.cpp delivers up to 1.9x higher throughput through kernel optimizations on a GeForce RTX 5090, enhanced speculative decoding techniques and faster prefill. ​

vLLM delivers 1.2x on RTX PRO 6000 Blackwell Workstation Edition and up to 1.4x on two DGX Spark clusters. New XQA attention kernels in FlashInfer and backend optimizations help to accelerate inference across both platforms.

These gains are available on the llama.cpp and vLLM inferencing backends. 

Users can also experience these via the LM Studio and Ollama applications.

Tap Idle PCs for More Local AI Compute With NVIDIA PAIR

More than half of U.S. households have two or more PCs, and much of that computing power sits idle throughout the day. NVIDIA Personal AI Router (PAIR) is a free, open source software tool that puts those systems to work together for local AI.

Agentic workflows often break complex tasks into smaller jobs that can run in parallel, but performance can slow when every request is competing for the same GPU. PAIR automatically discovers compatible PCs on a local network and routes independent inference requests to whichever system has capacity. It works with Ollama and LM Studio and can adapt as devices join or leave the network.

For example, a user could ask Hermes to create a “Sunday Reset” plan by sorting through a cluttered inbox and prioritizing what needs attention now, what can wait and what can be skipped. Hermes can split that work across multiple subagents, while PAIR distributes those jobs across available PCs instead of having them all wait on a single GPU.

The result is more compute for local agents, with more tasks running in parallel and the flexibility to move AI workloads to another PC while the main system is being used for gaming, creating or other work.

The NVIDIA PAIR beta is available for Windows, macOS and Linux through both graphical and terminal interfaces, supporting NVIDIA GeForce RTX 20 Series GPUs and newer, NVIDIA RTX PRO workstation GPUs (Turing architecture and newer), NVIDIA DGX Spark and Apple M4 or newer silicon.

Check out the NVIDIA tech blog to get started with NVIDIA PAIR. 

Powerful On Device Photo Editing With Cyberlink PhotoDirector AI PC Mode on RTX Spark

Open image and video models enable artists to experiment with Creative AI models on PCs.  This enables artists to iterate and explore concepts and ideas, without the dreaded token anxiety and keep more of their creative work private and on-device.

CyberLink’s new PhotoDirector AI PC Mode is one of the first applications to integrate these diffusion models directly into a creative software, and turn them into a creative tool at the finger tips of the artists. Coming to PhotoDirector 365 and optimized for NVIDIA RTX Spark when it launches, AI PC Mode users are getting AI-powered editing tools for generative editing, image enhancement, object and distraction removal, background removal and replacement, portrait refinement and the creation of entirely new visuals — with the flexibility to choose between local or cloud processing, depending on the task.

On NVIDIA GPUs, PhotoDirector uses TensorRT-RTX and FP8 to accelerate local AI.

Start using Cyberlink’s PhotoDirector 365 photo editing software and learn more about PhotoDirector AI PC Mode, launching with RTX Spark in October.

NVIDIA RTX Spark Windows PCs Arrive October 2026

NVIDIA RTX Spark is coming this October— and at IFA 2026, partners are showing off their hardware. At IFA, newly announced designs join the existing six OEMs shipping in October. Acer showed its compact desktop RTX Spark concept, and Lenovo announced its Yoga Pro  9n and Yoga 9n 2-in-1.  

RTX Spark is a new beginning for Windows PCs. One PC built for creators, gamers and AI agents. With a powerful 1 Petaflop RTX Blackwell GPU, up to 128GB of unified memory and a highly efficient 20-core Grace CPU, RTX Spark delivers incredible performance and efficiency. This superchip enables high performance thin laptops with all day battery life and compact desktops to power always-on agents. Paired with the new Windows Agent framework, it enables agents that run safely in the background under OS level control.

Last week at Gamescom, Electronic Arts, Embark and Ubisoft were among the latest game publishers and developers bringing their blockbuster titles to NVIDIA RTX Spark Windows PCs. They join the publishers that announced RTX Spark support at COMPUTEX in May, including KRAFTON, NetEase, Riot Games and XBOX. Read more.

Sign up to be notified when RTX Spark laptops and desktops are available.

#ICYMI: More Updates From NVIDIA Local AI 

🎮 NVIDIA Brings New RTX Tech and Games to Gamescom — NVIDIA released DLSS 4.5 Ray Reconstruction, featuring a new second-generation transformer model for improved image quality in ray-traced and path-traced games. Gamescom also brought new RTX announcements for titles including 007 First Light, CONTROL Resonant and Gears of War: E-Day, plus expanded game support for the upcoming NVIDIA RTX Spark.

🐋Introducing DeepSeek Harness — DeepSeek’s new open source harness pairs with DeepSeek-V4-Flash to power local agentic coding workflows on NVIDIA DGX Station and multi-DGX Spark setups.

📊MLPerf Client v2.0 Expands AI PC Benchmarking — MLCommons released MLPerf Client v2.0, developed in collaboration with NVIDIA and other industry leaders. The update adds new benchmarks for agentic AI and image generation, alongside expanded LLM testing for real-world local AI workloads.

Follow NVIDIA RTX Spark on X, Instagram, TikTok and Facebook — and stay informed by subscribing to the NVIDIA Local AI newsletter. Follow NVIDIA Workstation on LinkedIn and X. 

See notice regarding software product information.

]]>
‘NBA 2K27’ With NVIDIA DLSS 5 Leads 28 New Games Coming to GeForce NOW https://blogs.nvidia.com/blog/geforce-now-thursday-september-2026-games-list/ Thu, 03 Sep 2026 13:00:35 +0000 https://blogs.nvidia.com/?p=98028

September is here with 28 more games streaming on GeForce NOW this month, led by a slam dunk: NBA 2K27 with the NVIDIA DLSS 5 3D-Guided Neural Rendering feature.

Through NVIDIA’s close collaboration with Visual Concepts and 2K, DLSS 5 brings a new level of lifelike lighting and material detail to the court — tuned by developers to support the game’s authentic broadcast-style presentation.

Members can also walk the line between human and vampire in Rebel Wolves’ The Blood of Dawnwalker and step into genma-haunted Kyoto in Capcom’s Onimusha: Way of the Sword — part of the 10 games joining the GeForce NOW library this week.

Take the Court With DLSS 5

NVIDIA GeForce NOW NBA2K DLSS 5

GeForce NOW NBA2K DLSS5
Showcase your moves.

NVIDIA, Visual Concepts and 2K have brought DLSS 5 3D-Guided Neural Rendering to NBA 2K27, delivering a new level of sports realism for players streaming from the cloud at launch. 

DLSS 5 enhances and tunes lighting and materials to make the action feel even more lifelike — from how light catches hair and skin to the detail preserved in the facial geometry of real-world athletes — achieving new levels of authenticity. 

GeForce NOW Ultimate members in NVIDIA-operated regions can stream NBA 2K27 across PCs, Macs, handhelds, mobile devices, TVs and more. DLSS 5 is available when streaming from a GeForce RTX 5080-powered rig in the cloud.

Tip off without waiting for downloads or managing storage, and experience NBA 2K27 from the cloud with the visual fidelity of DLSS 5.

Walk the Line Between Day and Night

The Blood of Dawnwalker on GeForce NOW
Human by day. Vampire by night.

Rebel Wolves’ newly launched The Blood of Dawnwalker is streaming from GeForce NOW. The open-world, dark-fantasy action role-playing game transports players to 14th-century Europe, where war, plague and rising vampire powers have reshaped history.

Play as Coen, a young Dawnwalker fighting to save his family while caught between his humanity and cursed vampiric strength. His abilities — and the choices available to him — change between day and night, shaping how players explore, fight and uncover the secrets of the world around them.

Sink teeth into the fresh adventure from nearly any supported device, with Ultimate members tapping into GeForce RTX 5080-class performance in the cloud. Skip the installs and let GeForce NOW handle the heavy lifting, leaving more time to take a bite out of everything Coen’s new world has to offer.

Skip the Downloads and Sharpen the Sword

Onimusha: Way of the Sword on GeForce NOW
Pick up the Oni Gauntlet.

Capcom’s Onimusha: Way of the Sword arrives on GeForce NOW at launch, giving members a way to jump straight into the series’ return instantly from the cloud.

Fight across a twisted vision of Edo-period Kyoto as a lone samurai wielding the mystical Oni Gauntlet. Master precise swordplay, unleash powerful abilities and absorb the souls of fallen enemies while battling relentless Genma beneath mysterious clouds of Malice.

Ultimate members can stream every clash with GeForce RTX 5080-class performance in the cloud across PCs, Macs, handhelds, mobile devices, TVs, the newly supported Firefox browser and more. Cut to the chase by skipping any downloads and storage management and step into battle.

In addition, members can look for the following titles streaming this week:

  • Breathedge 2 (New release on Steam, Aug. 31)
  • Serious Sam: Shatterverse (New release on Steam, Aug. 31)
  • Crimson Moon (New release on Steam, Sept. 1)
  • The Blood of Dawnwalker (New release on Steam, Sept. 2)
  • NBA 2K27 (New release on Steam, Sept. 3)
  • SpeedRunners 2: King of Speed (New release on Steam and Xbox, available on Game Pass, Sept. 3)
  • Onimusha: Way of the Sword (New release on Steam, Sept. 4)
  • How to Fish (Steam)
  • Shelves and Sorcery: Tidy Up the Enchanted Shop (Steam)
  • VHOLUME (Steam)

And look forward to the games coming throughout the month:

  • Bus Simulator 27 (New release on Steam, Sept. 8)
  • Halloween: The Game (New release on Steam, Sept. 8)
  • Honeycomb: The World Beyond (New release on Steam, Sept. 8)
  • Wanderburg (New release on Steam, Sept. 8)
  • Shroom and Gloom (New release on Steam, Sept. 10)
  • WARDOGS (New release on Steam, Sept. 10)
  • Welcome to Elderfield (New release on Steam, Sept. 10)
  • Active Matter (New release on Steam, Sept. 15)
  • Aniimo (New release on Steam, Sept. 15)
  • Train Sim World 7 (New release on Steam, Sept. 15)
  • The Guild 1 Remake: Europa 1410 (New release on Steam, Sept. 17)
  • CONTROL Resonant (New release on Steam, Sept. 24)
  • Nivalis Nights (New release on Steam, Sept. 29)
  • Mixtape (Steam) 
  • Neon White (Steam) 
  • Outer Wilds (Steam) 
  • Stray (Steam)
  • What Remains of Edith Finch (Steam)

All the Extras From August

In addition to the 26 games announced last month, 18 more joined the GeForce NOW library. 

  • Cat Mail Co. (Steam)
  • Monster Hunter Wilds Prologue Demo (Steam)
  • Stars Reach (Steam)
  • DIVE or DIE – Children of Rain (Steam)
  • Escape the Backroom (Xbox, available on Game Pass)
  • Hell Is Us (Xbox, available on Game Pass)
  • Parcel Simulator (Steam)
  • The Thaumaturge (Xbox, available on Game Pass)
  • Tomb Raider IV-VI Remastered (Epic Games Store)
  • STAR WARS Zero Company (EA app and Steam)
  • Baldur’s Gate 3 (GOG)
  • Clair Obscur: Expedition 33 (GOG)
  • Cold Fear (GOG)
  • IRON NEST: Heavy Turret Simulator (Steam)
  • Kingdom Come: Deliverance II (GOG)
  • Mount & Blade: Bannerlord II (GOG)
  • No Man’s Sky (GOG)
  • The Witcher 3: Wild Hunt (GOG)

Dates listed above reflect when games are released on their respective stores. GeForce NOW availability may vary, as games are onboarded after release and added throughout the week. Keep an eye on GeForce NOW channels and GFN Thursdays for availability updates on announced titles.

New to GeForce NOW? Start with a day pass to try premium cloud gaming before committing to a membership. Better yet, the cost of the day pass can be applied toward a first monthly membership — making it easy to try GeForce RTX-powered cloud gaming with the latest releases, then level up for more.

What are you planning to play this weekend? Let us know on X or in the comments below.

]]>
NVIDIA to Acquire Hugging Face https://blogs.nvidia.com/blog/nvidia-to-acquire-hugging-face/ Thu, 03 Sep 2026 11:56:49 +0000 https://blogs.nvidia.com/?p=98049

I’m excited to announce that NVIDIA has agreed to acquire Hugging Face for $12,930,300,000. Together, we will scale Hugging Face’s platform, strengthen its infrastructure and expand access to AI for developers and institutions worldwide.

Over the past decade, Clem, Julien, Thomas and the team at Hugging Face have built something remarkable: a vibrant home for the open model developer community.

More than 18 million developers, researchers and creators use Hugging Face to share more than 3 million models, 500,000 datasets and 1 million applications. More than 200,000 companies use the platform to discover, evaluate, customize and deploy AI.

Hugging Face will remain an open platform for the entire AI ecosystem. Developers will choose the models they want, the frameworks they want, the clouds and inference service providers they want and the computing platforms they want. NVIDIA compute will not be required to build on or deploy through Hugging Face.

Hugging Face will continue to support open source and open weight models from across the ecosystem, from every model builder. It will continue to support multi-cloud and multi-accelerator development and deployment, so builders can use the hardware and infrastructure that best fit their work.

Recently, I coauthored an open letter on the importance of open weights to the AI economy. Joined by leaders from across the industry, we made a simple point: open weights broaden access to AI and help ensure that AI leadership is distributed across companies, institutions and communities.

Open models let startups, businesses, universities and public institutions build on advanced capabilities without training every model from scratch. They enable organizations to match the right model to the right job. That is how AI can advance safely, strengthen cybersecurity and sovereignty, accelerate innovation, and reach factories, hospitals, farms, classrooms and Main Street businesses around the world.

AI advances faster when people can build together.

NVIDIA has been committed to open weight models for years, demonstrated by multiyear investments and major contributions to open source platforms, including Hugging Face. NVIDIA has said that open models, data and tools broaden access to AI, and it has contributed hundreds of open models and datasets to Hugging Face as part of that effort.

  • NVIDIA is the largest contributor of open models and data to Hugging Face, and our contributions continue to grow.
  • NVIDIA has released more than 500 models on Hugging Face and more than 250 open datasets.
  • We build our own models, libraries and tools in the open so developers everywhere can use them, modify them and build on top of them.

As the opportunity for open models accelerates, Hugging Face can serve the global AI community at unprecedented scale. NVIDIA’s infrastructure, engineering and global reach can help improve platform reliability, safety, model evaluation, inference and deployment capabilities, while preserving the open ecosystem that made Hugging Face foundational.

I am honored that Clem came to me as he considered the next chapter of Hugging Face and believed NVIDIA would be a great home for the company, its community and the future of open models. We share this vision, and the Hugging Face team will now bring their passion and expertise to a much larger canvas, with their same iconic 🤗 brand.

To the millions of builders on Hugging Face: thank you for pushing the boundaries of what is possible. We can’t wait to build the future together with you. Together, we will make AI more open, more capable and more accessible to people and institutions around the world.

]]>
NVIDIA and CrowdStrike Strengthen Agentic Cybersecurity Frontier https://blogs.nvidia.com/blog/nvidia-crowdstrike-fal-con-2026/ Tue, 01 Sep 2026 21:19:20 +0000 https://blogs.nvidia.com/?p=98014

“We’re at an inflection point in cybersecurity,” Jensen Huang told a sold-out crowd at CrowdStrike’s Fal.Con 2026 in Las Vegas Tuesday. Attacks are now automated. Defense has to be, too. 

The NVIDIA founder and CEO joined CrowdStrike CEO and founder George Kurtz to announce CrowdStrike SafeMind, its agentic cybersecurity system developed by the CrowdStrike Cyber Superintelligence Lab.

“This is the beginning of a new age of cybersecurity,” Huang told the crowd of 10,000 security professionals. “On the one hand, the adversaries are going to be more armed than ever. On the other hand, all of you are going to be more armed than ever.” 

SafeMind combines CrowdStrike’s purpose-built, beyond frontier-capable models and customized agentic harnesses, with defensive models built on NVIDIA Nemotron, in a continuous coevolution loop where offense and defense repeatedly challenge and improve each other. 

CrowdStrike also announced CrowdStrike Falcon IQ to operationalize Project QuiltWorks through agentic workload automation and expanded its CrowdStrike Guardian AI safety solution.

“We have asymmetric advantages because we have a large community of cybersecurity experts who want to work with each other and keep the world safe,” Huang told the crowd. 

CrowdStrike’s annual conference drew security leaders from financial services, healthcare, the public sector and critical infrastructure.

“The real gap that I saw was that the attackers had frontier AI, and the defenders didn’t,” Kurtz told them. “And that changes now.”

SafeMind

CrowdStrike built SafeMind’s defensive model using NVIDIA Nemotron open models, post-trained with CrowdStrike’s cyber experience and threat data. The SafeMind models are paired with proprietary cybersecurity harnesses optimized to work as an agentic stack.

The result ships natively in the CrowdStrike Falcon platform as SafeMind, CrowdStrike’s agentic cybersecurity system. SafeMind brings offensive and defensive AI together in a continuous coevolution loop, where each side adapts to and strengthens the other. This process continuously hardens the security of the customer environment until attacks are unsuccessful.

“Your decade and a half of security data that we can train on — we can take a frontier model and make it essentially a super AGI that is incredibly good at cybersecurity,” Huang told Kurtz.

“Together with NVIDIA, we built cybersecurity’s first complete agentic system for cybersecurity, including the first frontier models and harness purpose-built for defenders,” Kurtz said. “This isn’t a copilot baked into someone else’s intelligence. It’s not a chatbot with a security skin. It is a frontier-class model built and trained by CrowdStrike on our data in partnership with NVIDIA.”

NVIDIA Nemotron 3 Ultra orchestrates the defensive agent harness. A fine-tuned Nemotron 3 Super powers SafeMind’s rule-generation sub-agent. 

By post-training Nemotron with CrowdStrike data, CrowdStrike internal evaluations showed that the Blue Solano model — based on Nemotron 3 Super — delivered higher accuracy rates than leading frontier models at 99% lower cost. 

While SafeMind can operate as a complete system, the models can be used independently to empower defenders to stay ahead of the adversary. Security experts can also pair their own models with CrowdStrike’s custom harnesses, giving customers the flexibility to use the right models and capabilities for their environment.

“The harness is essentially the exoskeleton of the large language model,” Huang said. “The large language model is the brain. The exoskeleton turns it into an agent — and this exoskeleton doesn’t have to be the same shape and capability for every domain.”

When AI Is the Defense

AI-enabled attacks rose 89% in the past year, and the fastest eCrime breakout time has reached 27 seconds, according to CrowdStrike. Human-speed response isn’t defense. It’s documentation.

“There are many applications in the world where you must have the ability to fine-tune, to post-train — to create an AI that is super good at a particular domain,” Huang said. “Nemotron was created for precisely that. Completely free. Incredibly fast. You have the ability to have an asymmetric advantage against whatever comes your way.”

With open Nemotron as the base, CrowdStrike’s security teams post-trained on their own threat data without sending it to an outside provider, and customized the AI to their environment. 

That’s not possible with a closed frontier model, and in security, the ability to inspect what’s defending matters. 

Red vs. Blue

NVIDIA announced its work testing the CrowdStrike SafeMind models and harnesses in a high-fidelity cyber agent environment running as a simulation of the NVIDIA network. 

The testing runs SafeMind in an offensive-defensive loop for adversarial coevolution. An offensive red-team agent finds the exploit, a blue-team defensive agent closes it and the findings become actionable detections to block attacks. 

The red-agent harness runs Recon, Assault and Compromise sub-agents executing attack paths inside the cyber agent environment. The blue-agent harness monitors via Falcon sensors, generates detection candidates, validates them and promotes them. 

CrowdStrike built the test environment with NVIDIA: a digital twin of NVIDIA’s own accelerated computing infrastructure, validated against NVIDIA’s real threat landscape.

“The basic framework of SafeMind — an adversarial model acting on a digital twin of the environment, with a defender model in a continuous cat-and-mouse loop, eventually learning how to secure itself — this basic framework applies to robotics, edge computing, enterprise computing and just about everything,” Huang said.

CrowdStrike also announced Falcon IQ. NVIDIA Nemotron models help to power the agentic engine at the heart of Charlotte AI AgentWorks, CrowdStrike’s no-code agent development platform where Falcon IQ runs. 

Falcon IQ uses more than 50 agents working together as a unified agentic workforce to automate the most time-intensive workflows in assessment, prioritization and remediation. 

Partners use Falcon IQ to deliver customized findings, recommendations and executive outputs to customers. Charlotte AI AgentWorks enables every Falcon user to build their own agentic security workforce.

The Full Stack

CrowdStrike has thousands of customer organizations generating trillions of daily security events. 

With NVIDIA’s full-stack accelerated computing platform, the collaboration runs from the chips up through the models to the harnesses acting on what those models find. For Kurtz, that’s the point. 

“The crowd in CrowdStrike,” Kurtz added, “is the asymmetry that puts the defenders in a unique position to defeat the adversary.”

]]>
GeForce NOW Gives Gamers More Ways to Play at Gamescom 2026 https://blogs.nvidia.com/blog/geforce-now-thursday-gamescom-2026/ Thu, 27 Aug 2026 13:00:33 +0000 https://blogs.nvidia.com/?p=97971

NVIDIA’s Gamescom announcements are revealing what’s next for GeForce NOW, with new ways to play, more supported devices and platforms, and even more big PC games headed to the cloud.

New NVIDIA DLSS 4.5 technology controls give members more ways to fine-tune gameplay, while expanded support for new Steam devices, GOG single sign-on, Firefox browser support and more Amazon Fire TV devices continue broadening where and how GeForce RTX-powered PC gaming can be enjoyed.

Gamescom is also bringing a blockbuster lineup of PC games to GeForce NOW. CONTROL Resonant, Gears of War: E-Day, Onimusha: Way of the Sword, The Blood of Dawnwalker and STAR WARS Zero Company are all coming to the cloud at launch. 

For a limited time, members can also get CONTROL Resonant with the purchase of a 12-month GeForce NOW Ultimate membership.

The Gamescom extras don’t stop there. To celebrate, GeForce NOW is giving away more than 40 prizes in our Community Giveaway — including gaming hardware and GeForce NOW Ultimate memberships.

There’s even more fun to come. Check out the list of 13 games joining the GeForce NOW library this week.

Next-Level Cloud Gaming

DLSS 4.5 games on GeForce NOW
More ways to choose how to play.

This fall, GeForce NOW is introducing new NVIDIA DLSS 4.5 controls in the cloud, giving Ultimate members more choice over how supported games look and perform, even without the latest hardware. Powered by NVIDIA RTX graphics, each control offers something different: DLSS Super Resolution sharpens image quality, while Dynamic Frame Generation delivers lower latency and stutter for smoother streaming, and Ray Reconstruction enhances ray-traced visuals. Together, these technologies make it easier to tune supported games for image quality, responsiveness or a preferred balance between the two. 

Stay tuned to GFN Thursdays for details about best practices for enabling DLSS 4.5 overrides and how Dynamic Frame Generation can help Ultimate members enable more graphic settings while still getting smooth streaming.

GeForce NOW and Steam hardware
Stream full steam ahead with the Steam ecosystem.

GeForce NOW is also adding community-requested support for Steam Machine with Steam Controller later this fall, and later for Windows and macOS, delivering more ways to play and making it even easier for longtime Steam players to jump into the cloud using familiar hardware. The additions build on existing Steam Deck support and expand compatibility across the Steam ecosystem.

GOG SSO on GeForce NOW
Get your GOG on with one click into the library.

Launching GOG games just got even easier. With GOG single sign-on now available on GeForce NOW, members can securely link their account once, automatically sync supported games and launch them without having to sign in to GOG again — all with a single click.

With fewer steps between opening the app, skipping any downloads and jumping into gameplay, it’s another way GeForce NOW continues streamlining access to PC gaming libraries, making it easier to play the GOG library of games on TV platforms since gamers are automatically logged into their GOG accounts.

Find popular titles, like Cyberpunk 2077 and The Witcher 3: The Wild Hunt, in the app in the dedicated GOG row. Or search and filter for GOG games, then jump into supported titles from almost anywhere — all streamed from the cloud. Check back on GFN Thursdays for the latest games joining the GeForce NOW library, including newly supported games streaming from GOG.

GeForce NOW and Firefox
Outfox nearly any device with GeForce RTX gaming.

Plus, GeForce NOW is now officially supported on Firefox, expanding browser access across all major Windows browsers. Members can jump into supported PC games directly from Firefox without downloading a dedicated app. Ultimate members can enjoy GeForce RTX-powered gaming at up to 1440p and 120 frames per second.

Simply download the latest version of Firefox and visit play.geforcenow.com to get started. The cloud handles downloads, installs and updates, keeping local storage free while games stay ready to play in the cloud. 

GeForce NOW and Fire TV
Bring the fire to living rooms with GeForce NOW.

GeForce NOW is expanding support to additional Amazon Fire TV devices. GeForce NOW first launched on Fire TV devices earlier this year, and will soon expand to the Fire TV Stick 4K Select, a 4K streaming stick that’s ready to game right out of the box. Streaming from the cloud means supported PC games are ready to play using a compatible controller, making it easy to enjoy GeForce NOW on the biggest screen in the house with even more options to play from the comfort of the couch. 

And that’s not all the awesomeness from Amazon. Later this year, Fire TV users will be able to purchase GeForce NOW memberships directly through Amazon, creating a more seamless path from device setup to gameplay.

These new features will arrive alongside some of the biggest games coming to the cloud throughout the year.

Get Ready for Launch

Five newly announced and highly anticipated games are coming to GeForce NOW at launch: CONTROL Resonant, Gears of War: E-Day, Onimusha: Way of the Sword, The Blood of Dawnwalker and STAR WARS Zero Company. 

Better yet, members can lock in one of them early: For a limited time, purchase a 12-month GeForce NOW Ultimate membership and receive CONTROL Resonant at launch at no additional cost, and be first to experience Faden’s next chapter. The offer is available Tuesday, Aug. 25, through Sunday, Sept. 27.

Ultimate members can stream the full lineup with GeForce RTX 5080-class performance in the cloud, including NVIDIA DLSS 4, ray tracing, NVIDIA Reflex and Cinematic Quality Streaming technologies, with up to 5K high dynamic range on supported devices. 

Whether diving into new worlds or returning to iconic franchises, GeForce NOW lets members skip expensive hardware upgrades, massive installs and storage management so they can jump straight into the action the moment each game launches. Here’s a closer look at what’s coming:

Control Resonant bundle on GeForce NOW
Contain the chaos.

Explore a warped Manhattan on the brink of paranatural annihilation in Remedy’s CONTROL Resonant. Harness Dylan Faden’s extraordinary powers to battle the Hiss, the Mold and other reality-bending threats while searching for his sister, Federal Bureau of CONTROL Director Jesse Faden.

Gears of Wars E-day launch to be on GeForce NOW
Emergence begins.

Experience the horror and brutality of Emergence Day in the Coalition’s Gears of War: E-Day. Fourteen years before the original Gears of War, join Marcus Fenix and Dominic Santiago in a gripping origin-story campaign as the Locust Horde first erupt from below, igniting a desperate fight for survival. Rally squads and fight on in multiplayer with a reimagination of Gears’ iconic PvE mode, Horde Siege, or go head to head with a refined versus PvP.

Onimusha Way of the Sword launch on GeForce NOW
Pick up the Oni Gauntlet.

Fight across a twisted vision of Edo-period Kyoto in Capcom’s Onimusha: Way of the Sword. Wield the mystical Oni Gauntlet, master intense swordplay and battle relentless demonic Genma beneath mysterious clouds of Malice.

Blood of the Dawnwalker launch to be on GeForce NOW
Walk the line between day and night.

Step into 14th-century Europe in Bandai Namco’s The Blood of Dawnwalker, where war, plague and rising vampire powers have reshaped history. Play as Coen, a young Dawnwalker torn between humanity and cursed strength, and shape his path while fighting to save his family.

STAR WARS Zero Company launch to be on GeForce NOW
Command the galaxy’s finest.

Lead an elite squad behind enemy lines in STAR WARS Zero Company, a single-player turn-based game set during the twilight of the Clone Wars. Command former Republic officer Hawks and Zero Company through tactical battles and high-stakes missions where strategy, squad composition and player choices shape the fate of the galaxy.

More Ways to Win

Gamescom news for GeForce NOW
Gaming gear and Ultimate memberships are up for grabs.

GeForce NOW is marking Gamescom with a community giveaway featuring more than 40 prizes. Gamers will have a chance to score a MacBook Neo, Steam Deck or Chromebook — each paired with a one-year GeForce NOW Ultimate membership — along with Amazon Fire TV Sticks and GeForce NOW Ultimate memberships ranging from one month to one year.

Head over to the GeForce NOW community on Reddit for all the giveaway details and how to enter.

The Hunt Is On

Squad up and fight for survival in Aliens: Fireteam Elite 2, streaming on GeForce NOW this week at launch. Assemble a four-player fireteam of Colonial Marines and battle through swarms of Xenomorphs, deadly new threats and increasingly desperate encounters. With deeper squad mechanics, new and improved classes, smarter enemies and expanded customization, every mission puts teamwork — and nerves — to the test.

Members can also look for the following games joining the cloud this week:

  • Aliens: Fireteam Elite 2 (New release on Steam, Aug. 25)
  • Resonance: A Plague Tale Legacy (New release on Steam and Xbox, Aug. 27, available on Game Pass)
  • STAR WARS Zero Company (New release on EA app and Steam, Aug. 27)
  • Baldur’s Gate 3 (GOG)
  • Call of Duty: Modern Warfare 4 – Open Beta Weekend (Battle.net, Steam and Xbox available from Aug. 28 10am PDT  to Sep. 1 at 10am PDT. For more information please visit the Call of Duty website).
  • Clair Obscur: Expedition 33 (GOG)
  • Cold Fear (GOG)
  • IRON NEST: Heavy Turret Simulator (Steam)
  • Kingdom Come: Deliverance II (GOG)
  • Mount & Blade: Bannerlord II (GOG)
  • No Man’s Sky (GOG)
  • Paradise Killer (Steam)
  • The Witcher 3: Wild Hunt (GOG)

Dates listed above reflect when games are released on their respective stores. GeForce NOW availability may vary, as games are onboarded after release and added throughout the week. Keep an eye on GeForce NOW channels and GFN Thursdays for availability updates on announced titles.

To experience GeForce NOW, gamers can start with a GeForce NOW day pass and try premium cloud gaming before committing to a membership. Better yet, the cost of a day pass can be applied toward a first membership purchase, making it even easier to level up when ready.

Check back every GFN Thursday for new games, new features and even more ways to play on GeForce NOW. Which Gamescom announcement has you most excited? Let us know on X or in the comments below. 

]]>
Delivering Vera: NVIDIA’s First CPU Built for Agents Is Shipping Now https://blogs.nvidia.com/blog/vera-cpu-delivery/ Thu, 27 Aug 2026 13:00:17 +0000 https://blogs.nvidia.com/?p=92992 ]]> NVIDIA NVLink Fusion Expands With NVHBM Custom High-Bandwidth Memory https://blogs.nvidia.com/blog/nvlink-fusion-nvhbm-custom-high-bandwidth-memory/ Wed, 26 Aug 2026 21:05:30 +0000 https://blogs.nvidia.com/?p=97957

The next wave of AI is placing new demands on infrastructure. 

As AI agents and trillion-parameter workloads become mainstream, the performance of AI infrastructure depends not only on compute, but on how compute, memory, storage, networking and software are designed together as a unified system.

To help hyperscalers and AI innovators build the next generation of semi-custom AI infrastructure, NVIDIA today expanded NVIDIA NVLink Fusion with NVIDIA NVHBM, a next-generation high-bandwidth memory technology that brings higher memory performance and efficiency to XPUs. It will be validated and offered by leading memory partners, extending this advanced memory capability to NVLink Fusion customers.

Traditional HBM architectures place the memory controller on the XPU die, consuming valuable silicon area that could otherwise be dedicated to compute. NVHBM, built on the same technology that NVIDIA will use for future GPUs, integrates NVIDIA’s custom memory controller into the HBM base die. 

By integrating the memory controller into the 3D HBM stack instead of the XPU, NVHBM delivers up to 30% greater memory bandwidth and 15% lower HBM power consumption, and frees up to 25% more area on XPU compute die compared with standard HBM4E.

NVIDIA is establishing a standard NVHBM implementation, available from multiple memory providers. This reduces the engineering effort required to integrate and qualify memory across multiple suppliers — giving NVLink Fusion customers a faster path for bringing custom AI chips to market. 

Amazon’s Annapurna Labs will be the first to work on NVHBM as part of its broader collaboration with NVIDIA around NVLink Fusion.

AWS and NVIDIA Continue NVLink Fusion Collaboration

Amazon’s Annapurna Labs will work with NVIDIA on NVHBM technology and the NVLink scale-up architecture to enhance performance and efficiency for AI workloads.

This builds on AWS’s previously announced support for NVLink Fusion. Annapurna Labs will support NVLink Fusion with its next-generation Trainium chips starting with Trainium4, which would allow Amazon chips and NVIDIA GPUs to work together with common rack-scale architecture.

“NVHBM represents a new architectural approach to advancing high-bandwidth memory performance and efficiency,” said Nafea Bshara, vice president of Annapurna Labs at Amazon. “We look forward to this technology collaboration to benefit future AWS infrastructure designs.” 

Vertically Integrated and Horizontally Open

NVLink Fusion enables partners to connect custom XPUs and CPUs to NVIDIA’s rack-scale platform.

Partners can access NVIDIA NVLink chiplets, NVLink-C2C, NVLink Switches and NVIDIA MGX systems and racks, as well as a broad ecosystem of CPU partners, ASIC designers, system manufacturers and technology providers. 

Offered with each generation of NVIDIA’s rack-scale system architecture, NVLink Fusion allows hyperscalers and AI-native companies to focus engineering resources on XPU innovation while using a proven technology stack for scale-up and scale-out networking, rack-scale systems and software — creating a faster, lower-risk path to deploying semi-custom AI infrastructure.

Learn more about NVLink and NVLink Fusion.

]]>
Leading Publishers Bring Blockbuster PC Games and Technology to NVIDIA RTX Spark https://blogs.nvidia.com/blog/gamescom-rtx-spark-pc-games-technology/ Tue, 25 Aug 2026 15:30:48 +0000 https://blogs.nvidia.com/?p=97918

NVIDIA is bringing the next wave of RTX gaming to the Gamescom conference running this week in Cologne, Germany, with support for new games, anti-cheat technologies and increased visual quality. 

Electronic Arts, Embark and Ubisoft are among the latest game publishers and developers bringing their blockbuster titles to NVIDIA RTX Spark ahead of its launch this fall. They join the publishers that announced RTX Spark support at COMPUTEX in May, including KRAFTON, NetEase, Riot Games and XBOX. 

NVIDIA is working closely with game developers and ecosystem partners to ensure today’s top PC games — including EA’s EA SPORTS F1 25 and Apex Legends, Ubisoft’s Anno 117: Pax Romana and Embark Studios’ ARC Raiders and THE FINALS — run seamlessly on RTX Spark devices.

“We’re excited to bring Anno 117: Pax Romana to a new generation of RTX-powered PCs,” said Stéphane Jankowski, Ubisoft executive producer of Anno 117 : Pax Romana. “Our teams build the game to be seen wide: full districts, long sightlines, thousands of details at once. Seeing RTX Spark hold that level of fidelity from a single chip is genuinely impressive.”

RTX Spark reinvents personal computing, bringing personal AI agents, advanced content creation and high-performance gaming together in a platform engineered from the ground up for the next wave of Windows PCs. For gamers, RTX Spark combines Windows devices with NVIDIA RTX technologies, including DLSS and Reflex. 

Bringing Anti-Cheat to RTX Spark

Supporting modern PC gaming goes beyond running the game itself. Many of today’s biggest online games rely on sophisticated anti-cheat technologies, to enable fair play across the multiple platforms these games run on.  

NVIDIA has been working with EA to bring native EA Javelin Anticheat support to RTX Spark, including for EA’s biggest multiplayer games, such as EA SPORTS F1 25. RTX Spark launches this fall. 

RTX Delivers Next-Generation Gaming

RTX Spark is just one part of NVIDIA’s lineup of gaming announcements at Gamescom, including:

  • DLSS 4.5 Ray Reconstruction is available now, featuring NVIDIA’s latest second-generation transformer model to deliver even higher image quality in ray- and path-traced content. It replaces traditional denoisers with an NVIDIA supercomputer-trained AI network to reconstruct higher-quality pixels in the parts of a ray-traced frame where rays were not sampled.
  • Path tracing comes to CONTROL Resonant and 007 First Light to deliver new levels of visual fidelity. 
  • NVIDIA ACE technologies are coming to Aniimo in early 2027, letting players naturally interact with their creatures.
  • Gears of War: E-Day includes RTX Mega Geometry integration, advancing image quality on GeForce RTX GPUs, alongside ray tracing and DLSS 4.5.
  • RTX Remix keeps remastering PC classics, with a major Painkiller RTX mod update available now and The Elder Scrolls III: Morrowind RTX coming soon.
  • GeForce NOW gives members more ways to play with new DLSS 4.5 controls, expanded support for Steam Machine and other devices and platforms, plus blockbuster PC games including CONTROL Resonant, Gears of War: E-Day and more coming to the cloud at launch.

Find more on NVIDIA’s announcements at Gamescom in this GeForce article. 

]]>
How XPUs Meet a World-Class AI Factory https://blogs.nvidia.com/blog/nvlink-fusion-xpu-ai-factory/ Mon, 24 Aug 2026 15:00:54 +0000 https://blogs.nvidia.com/?p=97869

To generate intelligence at scale, AI factories run continuously, and their economics are defined by delivered output: tokens per second, tokens per watt, cost per token, utilization and uptime. 

That requires AI infrastructure designed and built as a full factory, not a collection of individual accelerators.

Hyperscalers and AI-native companies building custom XPUs must consider not just XPU design, but the design and development of the entire AI platform, including scale-up and scale-out networking, rack-scale architecture, production factory software and a robust supplier ecosystem. 

At AI factory scale, this path is complex and costly, and represents a fundamental obstacle to getting XPUs to market quickly. 

Breaking the constraint means combining custom XPUs with proven, mature infrastructure — allowing builders to focus innovation where it matters most while harnessing established technology for the rest. 

NVLink Fusion delivers on that need, connecting XPUs to NVIDIA’s world-leading AI infrastructure to increase performance, accelerate time to market and mitigate risk for semi-custom AI factories.

Unlock XPU Performance With Fast Scale-Up

For modern workloads such as running trillion-parameter models, mixture-of-experts architectures and agentic AI, if the scale-up fabric cannot keep up, utilization drops and cost per token rises.

A scale-up networking solution must excel on three dimensions: 

  • Delivered performance: End-to-end network performance, in-network compute and  mature software integration.
  • Factory resiliency: Uptime, continuous health monitoring and telemetry, and component-level serviceability while the factory keeps running.
  • Platform maturity: Reduced operational risk by using a mature technology stack with a demonstrated track record of large-scale deployments and realized return on investment. 

As an example, NVLink Fusion brings XPUs into the NVIDIA NVLink scale-up domain. Sixth-generation NVLink provides leading high-bandwidth, low-latency networking across a 72-XPU domain. The end-to-end latency for XPU-to-XPU transfers is 3x lower than alternative solutions based on off-the-shelf Ethernet, and the packet rate is 10x higher.

For end-to-end performance, NVIDIA GB300 NVL72 systems help deliver significantly higher throughput and better interactivity compared with configurations that don’t use NVL72, and future NVLink roadmap configurations include domains of up to 1,152 accelerators and co-packaged optics.

A Pareto chart comparing GB300 NVL72 with B300 inference throughput performance in tokens per second per GPU on DeepSeek-V4-Pro at ISL=1K and OSL=1K sampled at various interactivity points in tokens per second per user. GB300 NVL72 is more than 10x the throughput of B300 in the middle of the Pareto between 70 and 100 tokens per second per user.
The 72-GPU NVLink scale-up domain enables GB300 NVL72 to deliver higher per-GPU throughput and interactivity compared with NVIDIA B300. Results from NVIDIA’s AI Inference Performance Benchmarks page.

NVLink Fusion also includes NVIDIA NVLink-C2C for connecting XPUs to NVIDIA Vera CPUs or other ecosystem CPUs, delivering up to 6x the energy efficiency of a PCIe interface — helping remove barriers between control and compute for agentic systems.

A Proven Stack and Ecosystem for Development and Deployment

Teams developing custom XPUs often underestimate the effort and complexity of turning XPU innovation into data center deployment. This includes:

  • Integrating high-speed CPU and scale-up interfaces
  • Sourcing and validating a scale-up network solution
  • Designing compute and switch trays
  • Designing and validating a rack architecture, including cooling and power
  • Integrating security and storage
  • Managing a complex supplier ecosystem

The ideal platform provides all of this, allowing teams to focus on targeted innovation while using proven solutions for the rest.

NVLink Fusion is supported by an ecosystem designed for rapid development, integration and deployment, spanning ASIC design, CPU, and IP and optical interconnect partners.

“NVLink Fusion gives customers the ability to choose the CPU architecture, the performance level, the software capabilities that best meet their needs for the workloads that they care about,” said Tim Wilson, vice president and general manager of data center silicon engineering at Intel.

NVLink Fusion adopters can also use the NVIDIA MGX rack-scale architecture and the same supply chain used for MGX-based systems such as NVIDIA Vera Rubin NVL72. Manufacturing partners manage design and integration, while MGX suppliers provide the building blocks for rack, cooling, power and emerging 800 VDC designs.

“With Vera Rubin [NVL72], we are looking at almost 100% automation of system builds in the manufacturing line,” said Jack Luoh, head of product and solution at QCT and Quanta Computer. “Most of those investments can be leveraged if the XPU leverages NVLink Fusion.”

The NVIDIA AI infrastructure platform is vertically integrated and horizontally open. NVLink Fusion adopters can optionally incorporate NVIDIA Rubin GPUs, Vera CPUs, co-packaged optics switches, ConnectX SuperNICs, BlueField DPUs, Mission Control software and full-rack solutions including NVIDIA Vera Rubin NVL72, Vera CPU Rack, LPX, STX and SPX.

Managing Risk With Infrastructure Standardization

AI factory planning doesn’t wait for silicon. Power procurement, facility design, cooling, rack layout and network architecture begin long before the final accelerator mix is available. A data center locked to one chip can become a schedule risk.

Different workloads may favor different accelerators, including XPUs, GPUs, CPUs and LPUs. GPU systems may work alongside semi-custom systems for training, post-training, reasoning, retrieval and serving.

“The value of the NVLink Fusion program is … [customers] can deploy their rack-level solution with the NVIDIA GPU, and then they can decouple the development of their XPU and put it at a different pace,” said Vince Hu, corporate senior vice president and general manager of the data center and computing business group at MediaTek.

NVLink Fusion addresses these challenges  through a unified architecture. XPU- and GPU-based systems such as Vera Rubin NVL72 can share rack footprints, networking, cooling, power delivery and management systems. Operators can move forward with buildout while deferring the precise silicon mix, then reprovision capacity as workload demand, silicon supply and business priorities change. 

“NVLink Fusion allows the hyperscalers or the custom ASIC designers to integrate their own custom CPU or XPU and bridges the NVIDIA technology with a third-party process to create a unified rack-scale architecture,” said Lie-Szu Juang, chair and chief strategy officer at GUC.

Designed, Validated and Operated as a Factory

Factory buildout is expensive, and mistakes can require costly rework. Infrastructure must be validated before construction begins. NVLink Fusion aligns with the NVIDIA DSX reference architecture for AI factories: codesigning buildings, power, cooling, compute and networking. The NVIDIA Omniverse DSX AI Factory Blueprint provides a digital twin and open reference design for gigawatt-scale AI factories, enabling partners to model facilities and technology together before deployment.

At the rack level, serviceability is part of performance. Reference compute trays feature 100% liquid cooling with no fans, cables or hoses, and allow trays to be removed while the rest of the rack remains operational. NVLink Switch trays are also liquid cooled and support continued operation during service.

“With NVLink Fusion we can use proven NVL72 rack design to have time-to-market, and we can have access to multiple suppliers to help us to deliver more into the hands of our customers,” said CC Lee, senior hardware development manager at Annapurna Labs, an Amazon company.

Software completes the factory. NVIDIA NCCL for distributed workloads, NVIDIA Dynamo and NIXL for disaggregation and NVIDIA Mission Control for cluster management, telemetry and debugging help operators run mixed AI infrastructure as a coordinated system.

With NVLink Fusion, XPUs can now meet a world-class AI platform, enabling hyperscalers and AI-native companies to build unified, semi-custom AI factories that combine the strengths of many builders into infrastructure no one company could build alone.

Learn more about NVLink Fusion.

]]>
With Groq 3 LPX in Full Production, NVIDIA Extends Vera Rubin Inference for Agents https://blogs.nvidia.com/blog/vera-rubin-lpx-spectrum-x-nvlink-fusion/ Mon, 24 Aug 2026 15:00:41 +0000 https://blogs.nvidia.com/?p=97768

The next era of AI inference won’t be defined by a single breakthrough chip, network or system. It’ll be defined by how every layer of the AI factory works together. That’s why NVIDIA is extending Vera Rubin NVL72 with fast token generation for agentic systems.

Announced today, the NVIDIA Vera Rubin rack-scale system NVIDIA Groq 3 LPX is in full production. In an Artificial Analysis benchmark running Gemma 4 31B, an open source agentic model, it delivered 3,400 output tokens per second for 100,000-token long-context use cases critical to agentic systems, 4x faster than the nearest alternative platform. 

Industry partners worldwide are adopting Vera Rubin platform solutions. SpaceXAI announced that NVIDIA Vera CPUs will power its next generation of agentic AI. CoreWeave has deployed into production Spectrum-X Multiplane, which connects NVIDIA Vera Rubin racks using multiple parallel switches to provide high-bandwidth, flat and lossless AI networks. Nebius is the first AI cloud to adopt NVIDIA Groq 3 LPX.

As AI shifts from training to reasoning and agentic, inference has become the new frontier. Agentic AI systems are generating more tokens, processing dramatically larger context windows and increasingly collaborating with other AI systems to solve complex problems. 

These workloads demand a new class of infrastructure optimized not just for performance but for throughput, responsiveness and economics at unprecedented scale. 

At the Hot Chips conference this week in Palo Alto, California, NVIDIA is showcasing how extreme codesign is reshaping the AI factory from end to end. By architecting compute, networking and inference acceleration as a unified system, NVIDIA is helping customers build infrastructure purpose-built for the emerging demands of long-context inference and multi-agent systems.

Extreme Codesign Optimizes for Performance

Extreme codesign is the guiding principle behind NVIDIA platforms. Vera Rubin is engineered to accelerate inference as agents reason over increasingly long sequences. 

NVIDIA Spectrum-X Ethernet moves those massive data flows efficiently across AI factories, and NVIDIA Groq 3 LPX is built to generate tokens at ultrafast speeds. Together, they show how NVIDIA is optimizing every stage of the AI pipeline, from context and communication to generation, as part of a single, integrated AI factory architecture.

NVIDIA Groq 3 LPX brings a new low-latency inference architecture designed to work alongside Vera Rubin NLV72, the most versatile AI factory platform, helping enterprises and cloud providers deliver the low latency, extreme throughput and scalable economics required for agentic applications.

Breakthrough performance comes not from optimizing individual components in isolation, but from codesigning every layer of the stack. From networking and context processing to large-scale inference, NVIDIA’s full-stack platform turns AI factories into integrated engines for intelligence, built to turn ever-growing volumes of tokens into revenue.


Tuesday, Aug. 24, 8:00 a.m. PT 🔗

NVIDIA Partners Adopt Vera Rubin for Lowest Token Costs

Nebius, a leading AI cloud, is first to adopt NVIDIA Groq 3 LPX, giving developers access to leading token generation speeds for highly responsive agentic AI applications. 

Adding NVIDIA Groq 3 LPX to NVIDIA Vera Rubin NVL72 in  Nebius Token Factory will boost inference performance so developers can build highly interactive agents, coding systems and other real-time AI experiences at scale.

Connecting NVIDIA Vera Rubin racks, CoreWeave is deploying Spectrum-X Multiplane in production, unlocking advances for its AI cloud infrastructure.


Tuesday, Aug. 24, 8:00 a.m. PT 🔗

SpaceXAI Adopts NVIDIA Vera CPUs for Agentic AI

SpaceXAI plans to build and scale its future AI architecture around NVIDIA Vera Rubin, from data centers on Earth to orbital satellites. The company plans to deploy NVIDIA Vera CPUs to accelerate the CPU-intensive work behind agentic AI, including orchestration, tool use, code execution, data processing and simulation. 

The SpaceXAI partnership extends NVIDIA’s full-stack AI platform to SpaceXAI, bringing together Vera CPUs, NVIDIA accelerated computing, networking and software to advance AI at unprecedented scale.

Designed for the agentic era, Vera Rubin provides leading per-core performance, exceptional memory bandwidth and predictable performance under load, helping agents complete tasks faster and keeping valuable GPU infrastructure fully utilized.


Tuesday, Aug. 24, 8:00 a.m. PT 🔗

NVIDIA Groq 3 LPX: The Interactive AI Inference Accelerator

Codesigned with the Vera Rubin NVL72 platform, NVIDIA Groq 3 LPX is helping AI factories deliver tokens at the lowest latency for agentic workloads.

Agentic AI is creating a new performance challenge: decode latency. As AI agents reason, use tools and interact with other systems, they generate responses one token at a time, causing even tiny delays to multiply across complex chains of work. To keep agents operating at the pace users expect, NVIDIA Groq 3 LPX extends the Vera Rubin NVL72 platform with specialized acceleration for token generation. 

NVIDIA Rubin GPUs handle large-scale context processing while LPX accelerates latency-sensitive decode workloads. The result is faster, more predictable token generation that helps AI factories deliver responsive reasoning, smoother agent interactions and greater infrastructure efficiency. 

Together, Rubin GPUs and LPUs are designed to eliminate the traditional tradeoff between speed and throughput, helping AI providers deliver responsive, large-scale inference for the next generation of agentic AI applications.

Building the Token Factory

As the industry shifts from model training to serving intelligence at scale, infrastructure must evolve into what NVIDIA describes as a “token factory” capable of delivering performance, throughput, intelligence integrity and economic efficiency simultaneously. Agentic AI systems increasingly communicate with other AI systems, access multiple data sources and maintain large amounts of context, creating unprecedented demand for fast inference.

NVIDIA Groq 3 LPX was designed for exactly these workloads. As an extension of the Vera Rubin NVL72, it enables ultrafast responsiveness even across massive context windows while helping service providers maximize throughput and infrastructure utilization. 

Extreme Codesign for Inference

Unlike standalone accelerators, NVIDIA Groq 3 LPX combines the strengths of GPUs and LPUs through extreme codesign. Rubin GPUs and LPUs jointly compute every layer of an AI model, enabling new levels of inference performance for agentic workloads. 

At scale, fleets of LPUs operate as a giant processor optimized for deterministic inference. A rack-scale NVIDIA Groq 3 LPX deployment can include 256 LP30 accelerators connected through direct chip-to-chip links, creating a highly efficient inference engine built for modern AI factories.

Designed for the Agentic AI Era

As reasoning models grow and agentic workflows generate ever more tokens, the infrastructure required to serve them must evolve. NVIDIA Groq 3 LPX extends the Vera Rubin NVL72 platform with a purpose-built inference architecture designed to maximize responsiveness, throughput and efficiency, helping power the next generation of AI factories.

And this is only the beginning, more optimizations, more models, more performance when paired with Vera Rubin NVL72 — new levels of throughput and interactivity are coming. Stay tuned. 


Tuesday, Aug. 24, 8:00 a.m. PT 🔗

NVIDIA Spectrum-X Multiplane Enables Massive AI Factory Scale on a Flatter, More Resilient Network​

As AI factories grow massive, the network has become a critical engine of performance. At Hot Chips, NVIDIA is spotlighting Spectrum-X Multiplane — the latest in the hardware-accelerated Spectrum-X Ethernet architecture that lets Ethernet scale to unprecedented size while avoiding the latency, jitter and cost of adding another network tier.

NVIDIA Spectrum-X Ethernet is designed as an end-to-end, AI-optimized Ethernet platform, combining NVIDIA Spectrum-X Ethernet switches, SuperNICs and software to improve the performance and efficiency of Ethernet-based AI infrastructure for AI factories and clouds. The platform is designed to deliver 1.6x better AI networking performance compared with off-the-shelf Ethernet, while providing consistent, predictable performance in multi-tenant environments.

Multiplane Unlocks Scale Without the Tradeoffs of a New Tier

Scaling an AI factory beyond today’s largest clusters traditionally means adding a third network tier, which adds latency, slows things down unpredictably and drives up the cost of cabling, optics and power. Spectrum-X Multiplane takes a simpler approach: It splits each server’s network connection into several independent paths, or “planes,” each running its own lightweight two-tier network. The result is a flat, simple network that scales to 512,000 GPUs, without the added cost and complexity of a third tier.

This all happens automatically. A dedicated hardware engine inside the NVIDIA ConnectX SuperNIC manages traffic across the planes and instantly reroutes around any failure, so applications and software simply see one fast, reliable connection. In an eight-plane topology, if one plane fails, the network still maintains about 90% of its total bandwidth, with hardware recovery that’s 11x faster than software-based multiplane load balancing. This translates to 1.6x higher AI factory output.

Built Through Extreme Codesign

That reliability comes from extreme codesign of Vera Rubin NVL72, spanning switch silicon, SuperNICs and software. Spectrum-X SN6000 series switches, based on the 102.4Tb/s Spectrum-6 Ethernet ASIC and ConnectX-9 SuperNICs, supporting up to 1,600Gb/s per GPU, are purpose-built for Vera Rubin NVL72 AI factories. Spectrum-XGS Ethernet extends that same codesign across data centers, letting multiple facilities function as a single AI super-factory and accelerating multi-site NCCL collectives by 1.9x.


Tuesday, Aug. 24, 8:00 a.m. PT 🔗

NVIDIA Introduces Scale-In Infrastructure for Agentic AI Factories, Powered by BlueField-4, DOCA

NVIDIA is introducing NVIDIA Scale-In, the fifth pillar of NVIDIA AI networking and a new class of accelerated network infrastructure for agentic AI factories. Scale-In extends purpose-built acceleration to the infrastructure services that secure, manage and operate the AI factory.

Powered by the NVIDIA BlueField-4 processor and NVIDIA DOCA software platform and connected over NVIDIA Spectrum-X Ethernet, NVIDIA Scale-In transforms the traditional north-south access network into a unified, accelerated infrastructure domain.

Cloud computing brought software-defined networking, composability and elasticity to the data center, enabling users, applications, data and services to scale dynamically. 

Agentic AI represents the next platform shift. AI factories bring together massive accelerated compute with growing numbers of users, applications and autonomous agents, all continuously interacting with data, storage and services. This transforms the demands on the infrastructure that brings AI to life. Networking, storage, cybersecurity and operations must now be accelerated alongside AI compute, combining software-defined flexibility with purpose-built hardware acceleration and full-stack codesign. 

NVIDIA Scale-In delivers multi-tenant networking, high-performance storage access, in-silicon security, elastic provisioning and real-time observability, while keeping infrastructure processing independent of host compute resources. By accelerating and codesigning these services as part of the AI factory, Scale-In helps security, data access and operations scale alongside AI compute. The result is secure, efficient and manageable shared infrastructure for deploying and operating agentic AI at massive scale.


Tuesday, Aug. 24, 8:00 a.m. PT 🔗

NVIDIA NVLink Fusion Connects XPUs to NVIDIA’s Leading AI Platform

NVIDIA NVLink Fusion brings custom silicon into NVIDIA’s world-leading AI infrastructure platform, enabling hyperscalers and AI-native companies to build semi-custom AI factories with greater performance, flexibility and speed.

As AI models grow in size and complexity, raw compute alone is not enough. AI factories require high-bandwidth, low-latency scale-up networking, proven rack-scale architectures and a full ecosystem spanning power, cooling, management software and supply chain. NVLink Fusion addresses these challenges by connecting custom XPUs and CPUs to NVIDIA’s scale-up and scale-out technology stack.

The platform includes sixth-generation NVIDIA NVLink and NVLink Switch purpose-built scale-up networking, as well as NVLink-C2C for energy-efficient connectivity between XPUs and CPUs. Through the NVIDIA MGX ecosystem, adopters can also use production-proven rack designs, components, manufacturing partner solutions and open, extensible software for distributed computing, disaggregated workloads and cluster management.

By standardizing GPU- and XPU-based systems on a unified architecture, NVLink Fusion helps decouple data center buildout from silicon readiness. Operators can share rack footprints, networking, cooling, power delivery and management systems, then adjust the mix of GPUs and XPUs as supply and workload requirements evolve.

NVLink Fusion extends the NVIDIA AI platform’s vertically integrated, horizontally open approach to custom silicon. It gives partners the freedom to innovate where they differentiate while drawing on NVIDIA technologies across compute, networking, infrastructure and software — creating a single, flexible AI factory that no one company could build alone.

]]>
Up to 30x More Work Per Watt: NVIDIA Vera Rubin NVL72 Sets a New Efficiency Standard for AI Agents https://blogs.nvidia.com/blog/vera-rubin-nvl72-efficiency-ai-agents/ Mon, 24 Aug 2026 15:00:19 +0000 https://blogs.nvidia.com/?p=97826

According to OpenRouter data, agentic AI workloads consume 15x more tokens than a simple chat request. Why? 

Consider what happens when an AI agent researches a company for an investment decision. The agent queries financial databases, searches news and filings, invokes a sub-agent to run peer comparisons and model valuations, then synthesizes everything into a recommendation. Agents and sub-agents keep reasoning until the task is done, driving increased token demand. With every step, the accumulated tokens become the input to the next, making long-context handling central to agentic AI performance.

The same pattern plays out across every agentic use case, from software development to customer service to deep research. 

As agentic AI moves into production across industries, the infrastructure running it needs to meet that token demand efficiently. 

New measured performance data shows NVIDIA Vera Rubin NVL72 systems deliver up to 30x higher throughput per megawatt than NVIDIA GB300 NVL72 on agentic workloads. NVIDIA measured this inference throughput data using the SemiAnalysis AgentX workload, consisting of recorded real-world agentic coding sessions, with actual context growth, tool calls and sub-agent spawning preserved. For power-constrained AI factories, that translates directly into 30x more agentic work for the same energy footprint.

These early results for Vera Rubin NVL72 demonstrate NVIDIA’s accelerated pace of innovation. With continuous software optimizations, performance across both Vera Rubin NVL72 and GB300 NVL72 will continue to improve.

Vera Rubin NVL72: 30x Higher Throughput per Megawatt and 35x Lower Token Cost

Agentic workloads look fundamentally different from chat or document summarization, where input and output sequences typically range from 1K to 8K tokens. In agentic sessions, context accumulates across steps and can reach hundreds of thousands of input tokens, with wide variability in both input and output lengths across requests. Performance measurement must evolve to capture the full agent workflow rather than a single inference request.

The results below reflect performance measured on real-world agentic coding trajectories.

In SemiAnalysis AgentX, the NVIDIA Blackwell platform delivers leading performance across multiple agentic models including Kimi K3, MiniMax M3, GLM5.3, Qwen3.5 and DeepSeek V4 Pro. 

For example, GB300 NVL72 delivers up to 15x better throughput per megawatt than the NVIDIA Hopper architecture on the DeepSeek V4 Pro model, giving customers a high-performance foundation to run agentic workloads. This leap reflects the advantage of a larger scale-up GPU domain and codesigned software in delivering significantly better inference efficiency.

Vera Rubin extends that advantage, lifting the performance across the entire Pareto curve, to deliver as much as 30x higher throughput per megawatt than GB300 NVL72 on the DeepSeek V4 Pro model. These early results, measured using the SemiAnalysis AgentX workload and currently pending SemiAnalysis review, don’t yet reflect Vera CPU performance for tool calling. 

NVIDIA DSX MaxLPS technologies manages power across the GPU, rack and workload levels to provision up to 40% more GPUs within the same megawatt budget, pushing throughput per megawatt further at AI factory scale. 

Throughput per megawatt also directly impacts the cost of every token produced. At up to 35x lower cost per million tokens than GB300 NVL72, Vera Rubin NVL72 can run agents continuously, at scale, across the full breadth of customers’ workloads. 

For power-constrained AI factories, throughput per megawatt determines AI factory revenue and cost per million tokens determines the profit margin on that revenue.

Extreme Codesign for Agentic Scale

Modern inference optimization spans a range of techniques that are especially critical for agentic AI. NVIDIA Vera Rubin NVL72 enables all of these and more through extreme codesign across every layer of the platform to deliver multifold performance gains.

  • Disaggregated serving separates context processing (prefill) from response generation (decode) so each scales independently.
  • Rate matching synchronizes the speeds at which prefill GPUs and decode GPUs produce tokens to maximize efficiency.
  • Large-scale expert parallelism distributes expert sub-networks in mixture-of-experts models across the scale-up GPU domain. 
  • Distributed KV-caching extends memory across the scale-up GPU domain, while KV-cache offloading tiers less-active context to host and storage, keeping previously processed context accessible without recomputation.
  • KV-aware routing directs incoming requests to the GPUs that already hold the relevant cached context, reducing redundant computation across long sessions. 
  • Fused CUDA kernels like MegaMoE combine many computation and inter-GPU communication operations into a single execution pass, keeping GPUs active rather than waiting for data.

NVIDIA Rubin GPUs’ enhanced fifth-generation Tensor Cores and the third-generation Transformer Engine accelerate both prefill and decode stages of inference. NVFP4 quantization compresses model weights to 4-bit precision, reducing memory footprint and increasing throughput without sacrificing output quality. 

The NVL72 scale-up domain, a defining architecture across Vera Rubin and Grace Blackwell, enables the high-bandwidth and low-latency inter-GPU communication essential for techniques such as large-scale expert parallelism and distributed KV-caching. Purpose-built to power this scale-up domain, NVIDIA NVLink interconnect technology and NVLink Switches, now in their sixth generation, deliver 10x higher packet rates and 3x lower latency than off-the-shelf Ethernet alternatives.

Spanning optimized CUDA kernels, inference runtimes like NVIDIA TensorRT LLM and serving frameworks like NVIDIA Dynamo, NVIDIA’s software stack is codesigned with the hardware to enable inference optimizations.

While the results above reflect current Vera Rubin NVL72 performance, the full platform is a seven-chip architecture that also includes the NVIDIA Vera CPU, Groq 3 LPU, NVLink 6 Switch, BlueField-4 DPU, Spectrum-6 SPX and ConnectX-9 SuperNIC, all purpose-built for AI factories deploying agents at scale.

Extreme codesign also extends to NVIDIA’s co-engineering with its partners. Vera Rubin is in full production and is scaling across the ecosystem. 

Learn more about the NVIDIA Vera Rubin platform.

]]>
Bring the Fire: Play Games on GeForce NOW With New Firefox Browser Support https://blogs.nvidia.com/blog/geforce-now-thursday-firefox/ Thu, 20 Aug 2026 13:00:24 +0000 https://blogs.nvidia.com/?p=97798

It’s a new way into the cloud. 

GeForce NOW welcomes Firefox support to the cloud, opening up another way to jump into high-performance PC gaming straight from the browser, starting today. Whether on a school laptop or everyday PC, it’s now even easier to play supported PC games without downloading a dedicated app.

Plus, discover 12 new games streaming from the GeForce NOW library this week, led by Gallipoli. 

No Downloads. Just Firefox.

GeForce NOW on Firefox
This cloud is on fire with new ways to play on Firefox.

GeForce NOW is now supported on Firefox, expanding access across even more devices. 

Members can outfox big downloads and jump into high-performance PC gaming directly from the browser — no waiting on lengthy game installs or updates. The experience delivers GeForce RTX-powered performance at up to 1440p and 120 frames per second for GeForce NOW Ultimate members. 

Simply download the latest version of Firefox, head to play.geforcenow.com and start playing. The cloud handles the heavy lifting, so games are ready to play without taking up valuable storage space on the device. Firefox joins Chrome, Edge and Opera in supporting GeForce NOW on Windows browsers.

Deploy the Games

Land on the beaches of the Gallipoli Peninsula in BlackMill Games’ Gallipoli, the latest entry in the WW1 Game Series. Fight as soldiers of the British and Ottoman Empires in intense 25v25 objective-based battles across coastal shores, deserts and urban battlefields, with 10 historical classes and over 50 authentic weapons and equipment options.

Teamwork is key in its grounded, one-shot-one-kill combat. Take command as an Officer, provide cover as a Heavy Machine Gunner or Sniper, keep the squad going as a Stretcher Bearer and more while navigating challenging campaigns.

Head to the frontlines from the cloud with no lengthy game downloads or installs taking up local storage space. Play directly in the newly supported Firefox browser, or take the action across PCs, Macs, Chromebooks, handhelds and more with GeForce NOW cloud saves.

In addition, members can look for the following games streaming throughout the week:

  • Stars Reach (New release on Steam, Aug. 18)
  • The Sinking City 2 (New release on Steam, Aug. 18)
  • Gallipoli (New release on Steam, Aug. 20)
  • DIVE or DIE – Children of Rain (Steam)
  • Escape the Backroom (Xbox, available on Game Pass)
  • Hell Is Us (Xbox, available on Game Pass)
  • High on Life (Steam)
  • High on Life 2 (Steam and Xbox, available on Game Pass)
  • Parcel Simulator (Steam)
  • Starvester (Steam)
  • The Thaumaturge (Xbox, available on Game Pass)
  • Tomb Raider IV-VI Remastered (Epic Games Store)

Dates listed above reflect when games are released on their respective stores. GeForce NOW availability may vary, as games are onboarded after release and added throughout the week. Keep an eye on GeForce NOW channels and GFN Thursdays for availability updates on announced titles.

To try GeForce NOW in Firefox, start with a GeForce NOW day pass and experience premium cloud gaming directly from the browser. If members later upgrade to a Performance or Ultimate membership, the cost of the day pass can be applied toward the first membership purchase.

What are you planning to play this weekend? Let us know on X or in the comments below.

]]>
Securing the Infrastructure of Intelligence https://blogs.nvidia.com/blog/securing-the-infrastructure-of-intelligence/ Mon, 17 Aug 2026 12:34:51 +0000 https://blogs.nvidia.com/?p=97752

AI factories are the defining infrastructure of the AI era — where compute transforms energy and data into intelligence that powers every business, industry and country.

In the AI economy, compute is revenue.

AI factories require a full stack of critical resources: advanced chips, packaging, memory and networking — as well as land, power and shell.

Just as NVIDIA has used its scale, long-term visibility and supply-chain partnerships to secure critical semiconductor resources, we are now applying that same discipline to secure LPS capacity exclusively for NVIDIA AI factories.

Today, we are partnering with SB Energy to secure LPS capacity at the exceptional PORTS-Pike Technology Campus in Portsmouth, Ohio, to host NVIDIA compute. OpenAI will be the tenant.

LPS: The Next Strategic Resource

For the vast majority of NVIDIA customers, securing LPS has long been a part of their infrastructure strategy.

The world’s largest cloud service providers and investment-grade enterprises have balance sheets, infrastructure expertise and long-term contracts to secure LPS independently. They build and operate AI factories using NVIDIA accelerated computing, networking, systems and software.

This model will continue to represent most of NVIDIA’s business.

But frontier AI labs are different.

Frontier AI labs have extraordinary demand for training and inference compute, but many are growing faster than their balance sheets and long-term credit profiles can support. They may have strong customer demand and rapidly growing revenue yet still lack the decades-long infrastructure contracts and investment-grade financing capacity needed to secure the AI factory infrastructure independently.

Their growth is increasingly constrained not by algorithms or customer demand, but by the availability of compute.

For these companies, more compute means more intelligence, more products, more users and more revenue. NVIDIA is helping provide the infrastructure that powers this flywheel.

PORTS-Pike: A Site for Generations of NVIDIA Compute

OpenAI will build and operate a world-class AI factory at PORTS-Pike. The AI factory will use NVIDIA’s full-stack DSX AI factory platform, including GPUs, CPUs, networking and infrastructure software.  

The initial deployment is expected to provide 4.25 gigawatts of AI factory capacity. Each generation of NVIDIA AI factory systems deployed at PORTS-Pike could represent approximately 1.5 million NVIDIA GPUs, or approximately $150 billion to $200 billion in NVIDIA revenue. Over 20 years, the site can support multiple upgrade cycles. 

This is the essential economic point: the LPS commitment secures a long-lived AI factory site, while the NVIDIA compute inside can be upgraded repeatedly. Each new generation can deliver greater production, more intelligence and better economics. 

NVIDIA may also choose to extend the arrangement at PORTS-Pike beyond the initial 4.25 gigawatts to secure the remaining capacity of 3.75 gigawatts. 

OpenAI and NVIDIA Expanding Compute Opportunity

More broadly, OpenAI has committed to substantial deployments of NVIDIA AI infrastructure through 2030. OpenAI’s existing and planned commitments represent approximately 12 gigawatts of NVIDIA compute, with an opportunity to expand to approximately 16 gigawatts if NVIDIA extends the PORTS-Pike arrangement beyond the initial 4.25 gigawatts. 

At these levels, the opportunity represents roughly $600 billion of NVIDIA compute through 2030. 

The Important Questions

What is NVIDIA guaranteeing, and for how long?

NVIDIA is supporting the LPS infrastructure at PORTS-Pike for approximately 4 gigawatts over a 20-year term, securing a site on which NVIDIA compute will be exclusively deployed.  

Our support is limited to defined portions of lease and power payments, along with a specified residual-value commitment — not the full cost of the site or all of the tenant’s obligations.   

The guarantee will become effective in phases as data centers are placed in service between 2028 and 2030. As OpenAI makes lease payments and capacity comes online, NVIDIA’s remaining exposure declines.

Why is NVIDIA guaranteeing PORTS-Pike?

LPS has become a critical constraint on AI factory deployment. NVIDIA is selectively securing exceptional sites where we can host multiple generations of NVIDIA compute and serve durable customer demand. 

The productive life of the site extends through multiple generations of NVIDIA systems, each capable of producing more intelligence and more revenue than the generation before.

Is this circular financing?

No. OpenAI will pay the lease.

NVIDIA uses its scale and long-term visibility to secure PORTS-Pike to host NVIDIA compute. This is the same discipline we apply to supply-chain management: we secure critical inputs when we have visibility into customer demand and when doing so enables long-term productive capacity.

What happens to PORTS-Pike if OpenAI does not use the site in the future?

NVIDIA compute is versatile, fungible and broadly adopted. The capacity can be resold to another qualified tenant across NVIDIA’s global ecosystem of cloud service providers, enterprises, AI labs and startups. 

CUDA makes NVIDIA compute more than hardware. It gives developers and NVIDIA engineers a common platform to continually improve installed systems. 

CUDA makes NVIDIA compute versatile. Versatility makes it fungible. Fungibility drives utilization and durability — making NVIDIA compute a productive asset: rentable and financeable. 

The value of an exceptional site, like PORTS-Pike, is not limited to one customer or one generation of compute. NVIDIA’s standardized platform, broad developer ecosystem and large market of potential users support the ability to redeploy productive capacity over time.

How much LPS will NVIDIA secure?

It will be strategic and disciplined. 

Most NVIDIA customers will continue to secure their own LPS.  The vast majority of LPS hosting NVIDIA compute will continue to be secured directly by CSPs, enterprises, sovereign AI builders and other customers. 

NVIDIA will focus selectively on exceptional sites where visible, durable demand can support multiple generations of NVIDIA compute.

The Infrastructure of Intelligence

PORTS-Pike represents the next step in NVIDIA’s journey. 

We began by building accelerated computing chips. We then expanded to systems, networking, CUDA and full-stack AI factories. Today, we are helping secure the critical infrastructure required to build these factories. 

NVIDIA is the full-stack AI infrastructure platform. 

We are investing in the long-lived foundations of AI factories so our customers can deploy the most productive compute platform in the world, generation after generation. 

By securing the critical resources needed to host NVIDIA compute, we can help the world’s most innovative companies build the AI factories that will power the age of intelligence.

]]>