HTML 77.2%
TypeScript 10.5%
Python 9.6%
JavaScript 2.5%
1<?xml version="1.0" encoding="UTF-8"?><feed2 xmlns="http://www.w3.org/2005/Atom"3 xmlns:thr="http://purl.org/syndication/thread/1.0"4 xml:lang=""5 xml:base="https://developer.nvidia.com/blog/wp-atom.php"6 >7 <title type="text">NVIDIA Technical Blog</title>8 <subtitle type="text">News and tutorials for developers, data scientists, and IT admins</subtitle>910 <updated>2026-09-10T18:22:39Z</updated>1112 <link rel="alternate" type="text/html" href="https://developer.nvidia.com/blog" />13 <id>https://developer.nvidia.com/blog/feed/</id>14 <link rel="self" type="application/atom+xml" href="https://developer.nvidia.com/blog/feed/" />1516 17 <entry>18 <author>19 <name>Elizabeth Goodman</name>20 </author>21 <title type="html"><![CDATA[How Full-Stack NIM Optimizations Deliver 2.5x More Users on Nemotron 3 Ultra]]></title>22 <link rel="alternate" type="text/html" href="https://developer.nvidia.com/blog/how-full-stack-nim-optimizations-deliver-2-5x-more-users-on-nemotron-3-ultra/" />23 <id>https://developer.nvidia.com/blog/?p=122372</id>24 <updated>2026-09-10T16:55:39Z</updated>25 <published>2026-09-10T16:55:32Z</published>26 <category scheme="https://developer.nvidia.com/blog" term="Agentic AI / Generative AI" /><category scheme="https://developer.nvidia.com/blog" term="Developer Tools & Techniques" /><category scheme="https://developer.nvidia.com/blog" term="Build AI Agents" /><category scheme="https://developer.nvidia.com/blog" term="Inference Performance" /><category scheme="https://developer.nvidia.com/blog" term="NIM" /> <summary type="html"><![CDATA[<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4_1480x833-768x432.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" fetchpriority="high" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4_1480x833-768x432.jpg 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4_1480x833-179x101.jpg 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4_1480x833-300x169.jpg 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4_1480x833-625x352.jpg 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4_1480x833-645x363.jpg 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4_1480x833-660x370.jpg 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4_1480x833-500x281.jpg 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4_1480x833-160x90.jpg 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4_1480x833-362x204.jpg 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4_1480x833-195x110.jpg 195w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4_1480x833-1024x576.jpg 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4_1480x833-960x540.jpg 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4_1480x833.webp 1480w" sizes="(max-width: 768px) 100vw, 768px" title="image4_1480x833" />Deploying a large language model is only the first step toward production-ready serving. Production teams also need to serve as many concurrent users as...]]></summary>27 <content type="html" xml:base="https://developer.nvidia.com/blog/how-full-stack-nim-optimizations-deliver-2-5x-more-users-on-nemotron-3-ultra/"><![CDATA[<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4_1480x833-768x432.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4_1480x833-768x432.jpg 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4_1480x833-179x101.jpg 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4_1480x833-300x169.jpg 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4_1480x833-625x352.jpg 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4_1480x833-645x363.jpg 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4_1480x833-660x370.jpg 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4_1480x833-500x281.jpg 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4_1480x833-160x90.jpg 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4_1480x833-362x204.jpg 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4_1480x833-195x110.jpg 195w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4_1480x833-1024x576.jpg 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4_1480x833-960x540.jpg 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4_1480x833.webp 1480w" sizes="(max-width: 768px) 100vw, 768px" title="image4_1480x833" />Deploying a large language model is only the first step toward production-ready serving. Production teams also need to serve as many concurrent users as...<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4_1480x833-768x432.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4_1480x833-768x432.jpg 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4_1480x833-179x101.jpg 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4_1480x833-300x169.jpg 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4_1480x833-625x352.jpg 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4_1480x833-645x363.jpg 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4_1480x833-660x370.jpg 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4_1480x833-500x281.jpg 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4_1480x833-160x90.jpg 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4_1480x833-362x204.jpg 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4_1480x833-195x110.jpg 195w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4_1480x833-1024x576.jpg 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4_1480x833-960x540.jpg 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4_1480x833.webp 1480w" sizes="auto, (max-width: 768px) 100vw, 768px" title="image4_1480x833" /><p>Deploying a large language model is only the first step toward production-ready serving. Production teams also need to serve as many concurrent users as possible on available GPU infrastructure while preserving the interactivity that keeps applications responsive. That tradeoff matters even more for agentic AI workloads, where prompts can be long, context can be reused across steps…</p>28<p><a href="https://developer.nvidia.com/blog/how-full-stack-nim-optimizations-deliver-2-5x-more-users-on-nemotron-3-ultra/" rel="nofollow" data-wpel-link="internal" target="_self">Source</a></p>]]></content>29 <link rel="replies" type="text/html" href="https://developer.nvidia.com/blog/how-full-stack-nim-optimizations-deliver-2-5x-more-users-on-nemotron-3-ultra/#comments" thr:count="0"/>30 <link rel="replies" type="application/atom+xml" href="https://developer.nvidia.com/blog/how-full-stack-nim-optimizations-deliver-2-5x-more-users-on-nemotron-3-ultra/feed/" thr:count="0"/>31 <thr:total>0</thr:total>32 </entry>33 <entry>34 <author>35 <name>Elizabeth Goodman</name>36 </author>37 <title type="html"><![CDATA[High-Throughput Structure Prediction with BioNeMo Inference Runtime]]></title>38 <link rel="alternate" type="text/html" href="https://developer.nvidia.com/blog/high-throughput-structure-prediction-with-bionemo-inference-runtime/" />39 <id>https://developer.nvidia.com/blog/?p=122343</id>40 <updated>2026-09-09T23:45:42Z</updated>41 <published>2026-09-10T15:00:00Z</published>42 <category scheme="https://developer.nvidia.com/blog" term="Data Center / Cloud" /><category scheme="https://developer.nvidia.com/blog" term="Data Science" /><category scheme="https://developer.nvidia.com/blog" term="Simulation / Modeling / Design" /><category scheme="https://developer.nvidia.com/blog" term="BioNeMo" /><category scheme="https://developer.nvidia.com/blog" term="CUDA" /><category scheme="https://developer.nvidia.com/blog" term="Drug Discovery" /><category scheme="https://developer.nvidia.com/blog" term="Healthcare & Life Sciences" /><category scheme="https://developer.nvidia.com/blog" term="HPC / Scientific Computing" /> <summary type="html"><![CDATA[<img width="600" height="338" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4-1.webp" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4-1.webp 600w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4-1-179x101.png 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4-1-300x169.png 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4-1-500x282.png 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4-1-160x90.png 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4-1-362x204.png 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4-1-195x110.png 195w" sizes="auto, (max-width: 600px) 100vw, 600px" title="image4" />Biomolecular structure prediction is now often run at proteome scale, where the goal is to move an entire worklist through the pipeline efficiently. NVIDIA...]]></summary>43 <content type="html" xml:base="https://developer.nvidia.com/blog/high-throughput-structure-prediction-with-bionemo-inference-runtime/"><![CDATA[<img width="600" height="338" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4-1.webp" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4-1.webp 600w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4-1-179x101.png 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4-1-300x169.png 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4-1-500x282.png 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4-1-160x90.png 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4-1-362x204.png 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4-1-195x110.png 195w" sizes="auto, (max-width: 600px) 100vw, 600px" title="image4" />Biomolecular structure prediction is now often run at proteome scale, where the goal is to move an entire worklist through the pipeline efficiently. NVIDIA...<img width="600" height="338" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4-1.webp" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4-1.webp 600w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4-1-179x101.png 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4-1-300x169.png 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4-1-500x282.png 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4-1-160x90.png 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4-1-362x204.png 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4-1-195x110.png 195w" sizes="auto, (max-width: 600px) 100vw, 600px" title="image4" /><p>Biomolecular structure prediction is now often run at proteome scale, where the goal is to move an entire worklist through the pipeline efficiently. NVIDIA BioNeMo Inference Runtime (BioIR) helps accelerate supported biomolecular structure-prediction models on NVIDIA GPUs while keeping the familiar PyTorch workflow. It uses optimized kernels and, where applicable, CUDA Graphs to speed model…</p>44<p><a href="https://developer.nvidia.com/blog/high-throughput-structure-prediction-with-bionemo-inference-runtime/" rel="nofollow" data-wpel-link="internal" target="_self">Source</a></p>]]></content>45 <link rel="replies" type="text/html" href="https://developer.nvidia.com/blog/high-throughput-structure-prediction-with-bionemo-inference-runtime/#comments" thr:count="0"/>46 <link rel="replies" type="application/atom+xml" href="https://developer.nvidia.com/blog/high-throughput-structure-prediction-with-bionemo-inference-runtime/feed/" thr:count="0"/>47 <thr:total>0</thr:total>48 </entry>49 <entry>50 <author>51 <name>Elizabeth Goodman</name>52 </author>53 <title type="html"><![CDATA[From Wafer-Out to First Token: Codifying Supply Chain Expertise with Nemotron and Palantir Foundry]]></title>54 <link rel="alternate" type="text/html" href="https://developer.nvidia.com/blog/from-wafer-out-to-first-token-codifying-supply-chain-expertise-with-nemotron-and-palantir-foundry/" />55 <id>https://developer.nvidia.com/blog/?p=122401</id>56 <updated>2026-09-10T05:31:11Z</updated>57 <published>2026-09-10T09:00:00Z</published>58 <category scheme="https://developer.nvidia.com/blog" term="Agentic AI / Generative AI" /><category scheme="https://developer.nvidia.com/blog" term="AI Agent" /><category scheme="https://developer.nvidia.com/blog" term="Blackwell" /><category scheme="https://developer.nvidia.com/blog" term="cuOpt" /><category scheme="https://developer.nvidia.com/blog" term="GB200" /><category scheme="https://developer.nvidia.com/blog" term="LLMs" /><category scheme="https://developer.nvidia.com/blog" term="NeMo" /><category scheme="https://developer.nvidia.com/blog" term="Nemotron" /><category scheme="https://developer.nvidia.com/blog" term="Vera Rubin" /> <summary type="html"><![CDATA[<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-768x432.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-768x432.jpg 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-179x101.jpg 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-300x169.jpg 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-625x352.jpg 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-1536x864.jpg 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-645x363.jpg 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-660x370.jpg 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-500x281.jpg 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-160x90.jpg 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-362x204.jpg 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-196x110.jpg 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-1024x576.jpg 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-960x540.jpg 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4.webp 1920w" sizes="auto, (max-width: 768px) 100vw, 768px" title="image1" />NVIDIA has one of the largest and most complex supply chains in the world, and its performance is measured from wafer-out to first token. The interval is in two...]]></summary>59 <content type="html" xml:base="https://developer.nvidia.com/blog/from-wafer-out-to-first-token-codifying-supply-chain-expertise-with-nemotron-and-palantir-foundry/"><![CDATA[<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-768x432.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-768x432.jpg 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-179x101.jpg 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-300x169.jpg 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-625x352.jpg 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-1536x864.jpg 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-645x363.jpg 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-660x370.jpg 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-500x281.jpg 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-160x90.jpg 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-362x204.jpg 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-196x110.jpg 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-1024x576.jpg 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-960x540.jpg 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4.webp 1920w" sizes="auto, (max-width: 768px) 100vw, 768px" title="image1" />NVIDIA has one of the largest and most complex supply chains in the world, and its performance is measured from wafer-out to first token. The interval is in two...<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-768x432.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-768x432.jpg 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-179x101.jpg 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-300x169.jpg 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-625x352.jpg 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-1536x864.jpg 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-645x363.jpg 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-660x370.jpg 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-500x281.jpg 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-160x90.jpg 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-362x204.jpg 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-196x110.jpg 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-1024x576.jpg 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-960x540.jpg 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4.webp 1920w" sizes="auto, (max-width: 768px) 100vw, 768px" title="image1" /><p>NVIDIA has one of the largest and most complex supply chains in the world, and its performance is measured from wafer-out to first token. The interval is in two parts. Time-to-rack runs from silicon leaving the fab to an assembled system arriving on a data center floor. Time-to-token covers everything thereafter: power, cooling, networking, and the software stack that makes the infrastructure…</p>60<p><a href="https://developer.nvidia.com/blog/from-wafer-out-to-first-token-codifying-supply-chain-expertise-with-nemotron-and-palantir-foundry/" rel="nofollow" data-wpel-link="internal" target="_self">Source</a></p>]]></content>61 <link rel="replies" type="text/html" href="https://developer.nvidia.com/blog/from-wafer-out-to-first-token-codifying-supply-chain-expertise-with-nemotron-and-palantir-foundry/#comments" thr:count="0"/>62 <link rel="replies" type="application/atom+xml" href="https://developer.nvidia.com/blog/from-wafer-out-to-first-token-codifying-supply-chain-expertise-with-nemotron-and-palantir-foundry/feed/" thr:count="0"/>63 <thr:total>0</thr:total>64 </entry>65 <entry>66 <author>67 <name>Tanya Lenz</name>68 </author>69 <title type="html"><![CDATA[When to Use Encode-Prefill-Decode Disaggregation to Accelerate Multimodal Model Serving]]></title>70 <link rel="alternate" type="text/html" href="https://developer.nvidia.com/blog/when-to-use-encode-prefill-decode-disaggregation-to-accelerate-multimodal-model-serving/" />71 <id>https://developer.nvidia.com/blog/?p=122311</id>72 <updated>2026-09-09T20:31:11Z</updated>73 <published>2026-09-09T20:31:04Z</published>74 <category scheme="https://developer.nvidia.com/blog" term="Agentic AI / Generative AI" /><category scheme="https://developer.nvidia.com/blog" term="Computer Vision / Video Analytics" /><category scheme="https://developer.nvidia.com/blog" term="Developer Tools & Techniques" /><category scheme="https://developer.nvidia.com/blog" term="AI Agent" /><category scheme="https://developer.nvidia.com/blog" term="AI Inference" /><category scheme="https://developer.nvidia.com/blog" term="Low-Latency Inference" /><category scheme="https://developer.nvidia.com/blog" term="Mixture of Experts (MoE)" /><category scheme="https://developer.nvidia.com/blog" term="NVFP4" /> <summary type="html"><![CDATA[<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-768x432.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-768x432.jpg 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-179x101.jpg 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-300x169.jpg 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-625x352.jpg 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-1536x864.jpg 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-645x363.jpg 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-660x370.jpg 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-500x281.jpg 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-160x90.jpg 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-362x204.jpg 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-195x110.jpg 195w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-1024x576.jpg 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-960x540.jpg 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo.webp 1999w" sizes="auto, (max-width: 768px) 100vw, 768px" title="multimodal-dynamo" />Encode-prefill-decode (EPD) disaggregation is an inference optimization technique for multimodal models that separates the vision encoder stage from the prefill...]]></summary>75 <content type="html" xml:base="https://developer.nvidia.com/blog/when-to-use-encode-prefill-decode-disaggregation-to-accelerate-multimodal-model-serving/"><![CDATA[<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-768x432.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-768x432.jpg 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-179x101.jpg 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-300x169.jpg 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-625x352.jpg 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-1536x864.jpg 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-645x363.jpg 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-660x370.jpg 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-500x281.jpg 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-160x90.jpg 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-362x204.jpg 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-195x110.jpg 195w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-1024x576.jpg 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-960x540.jpg 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo.webp 1999w" sizes="auto, (max-width: 768px) 100vw, 768px" title="multimodal-dynamo" />Encode-prefill-decode (EPD) disaggregation is an inference optimization technique for multimodal models that separates the vision encoder stage from the prefill...<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-768x432.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-768x432.jpg 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-179x101.jpg 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-300x169.jpg 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-625x352.jpg 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-1536x864.jpg 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-645x363.jpg 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-660x370.jpg 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-500x281.jpg 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-160x90.jpg 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-362x204.jpg 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-195x110.jpg 195w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-1024x576.jpg 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-960x540.jpg 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo.webp 1999w" sizes="auto, (max-width: 768px) 100vw, 768px" title="multimodal-dynamo" /><p>Encode-prefill-decode (EPD) disaggregation is an inference optimization technique for multimodal models that separates the vision encoder stage from the prefill and decode stages. It is most effective for image-heavy prompts, short-to-medium outputs, and quantized mixture-of-experts (MoE) models. This post shows when and how to use EPD disaggregation with NVIDIA Dynamo to achieve up to 5x…</p>76<p><a href="https://developer.nvidia.com/blog/when-to-use-encode-prefill-decode-disaggregation-to-accelerate-multimodal-model-serving/" rel="nofollow" data-wpel-link="internal" target="_self">Source</a></p>]]></content>77 <link rel="replies" type="text/html" href="https://developer.nvidia.com/blog/when-to-use-encode-prefill-decode-disaggregation-to-accelerate-multimodal-model-serving/#comments" thr:count="0"/>78 <link rel="replies" type="application/atom+xml" href="https://developer.nvidia.com/blog/when-to-use-encode-prefill-decode-disaggregation-to-accelerate-multimodal-model-serving/feed/" thr:count="0"/>79 <thr:total>0</thr:total>80 </entry>81 <entry>82 <author>83 <name>Jonathan Bentz</name>84 </author>85 <title type="html"><![CDATA[CUDA Toolkit 13.4 Adds Windows on Arm Support and Greater Control over Shared GPUs]]></title>86 <link rel="alternate" type="text/html" href="https://developer.nvidia.com/blog/cuda-toolkit-13-4-adds-windows-on-arm-support-and-greater-control-over-shared-gpus/" />87 <id>https://developer.nvidia.com/blog/?p=121255</id>88 <updated>2026-09-09T20:24:19Z</updated>89 <published>2026-09-09T20:24:12Z</published>90 <category scheme="https://developer.nvidia.com/blog" term="Developer Tools & Techniques" /><category scheme="https://developer.nvidia.com/blog" term="Simulation / Modeling / Design" /><category scheme="https://developer.nvidia.com/blog" term="Top Stories" /><category scheme="https://developer.nvidia.com/blog" term="CUDA" /><category scheme="https://developer.nvidia.com/blog" term="CUDA Tile" /><category scheme="https://developer.nvidia.com/blog" term="Nsight Tools - Compute" /> <summary type="html"><![CDATA[<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/06/cuda-python-768x432.png" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/06/cuda-python-768x432.png 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/06/cuda-python-179x101.png 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/06/cuda-python-300x169.png 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/06/cuda-python-625x352.png 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/06/cuda-python-1536x864.png 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/06/cuda-python-645x363.png 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/06/cuda-python-660x370.png 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/06/cuda-python-500x281.png 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/06/cuda-python-160x90.png 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/06/cuda-python-362x204.png 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/06/cuda-python-196x110.png 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/06/cuda-python-1024x576.png 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/06/cuda-python-960x540.png 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/06/cuda-python.webp 1920w" sizes="auto, (max-width: 768px) 100vw, 768px" title="cuda-python" />Every NVIDIA CUDA Toolkit release adds functionality and performance improvements that help developers get more from NVIDIA GPUs and the broader NVIDIA software...]]></summary>91 <content type="html" xml:base="https://developer.nvidia.com/blog/cuda-toolkit-13-4-adds-windows-on-arm-support-and-greater-control-over-shared-gpus/"><![CDATA[<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/06/cuda-python-768x432.png" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/06/cuda-python-768x432.png 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/06/cuda-python-179x101.png 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/06/cuda-python-300x169.png 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/06/cuda-python-625x352.png 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/06/cuda-python-1536x864.png 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/06/cuda-python-645x363.png 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/06/cuda-python-660x370.png 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/06/cuda-python-500x281.png 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/06/cuda-python-160x90.png 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/06/cuda-python-362x204.png 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/06/cuda-python-196x110.png 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/06/cuda-python-1024x576.png 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/06/cuda-python-960x540.png 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/06/cuda-python.webp 1920w" sizes="auto, (max-width: 768px) 100vw, 768px" title="cuda-python" />Every NVIDIA CUDA Toolkit release adds functionality and performance improvements that help developers get more from NVIDIA GPUs and the broader NVIDIA software...<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/06/cuda-python-768x432.png" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/06/cuda-python-768x432.png 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/06/cuda-python-179x101.png 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/06/cuda-python-300x169.png 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/06/cuda-python-625x352.png 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/06/cuda-python-1536x864.png 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/06/cuda-python-645x363.png 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/06/cuda-python-660x370.png 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/06/cuda-python-500x281.png 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/06/cuda-python-160x90.png 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/06/cuda-python-362x204.png 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/06/cuda-python-196x110.png 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/06/cuda-python-1024x576.png 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/06/cuda-python-960x540.png 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/06/cuda-python.webp 1920w" sizes="auto, (max-width: 768px) 100vw, 768px" title="cuda-python" /><p>Every NVIDIA CUDA Toolkit release adds functionality and performance improvements that help developers get more from NVIDIA GPUs and the broader NVIDIA software platform. CUDA Toolkit 13.4 adds support for Windows on Arm. CUDA applications have long been supported on Arm platforms through Linux; this release extends that capability to the Windows on Arm platform.</p>92<p><a href="https://developer.nvidia.com/blog/cuda-toolkit-13-4-adds-windows-on-arm-support-and-greater-control-over-shared-gpus/" rel="nofollow" data-wpel-link="internal" target="_self">Source</a></p>]]></content>93 <link rel="replies" type="text/html" href="https://developer.nvidia.com/blog/cuda-toolkit-13-4-adds-windows-on-arm-support-and-greater-control-over-shared-gpus/#comments" thr:count="0"/>94 <link rel="replies" type="application/atom+xml" href="https://developer.nvidia.com/blog/cuda-toolkit-13-4-adds-windows-on-arm-support-and-greater-control-over-shared-gpus/feed/" thr:count="0"/>95 <thr:total>0</thr:total>96 </entry>97 <entry>98 <author>99 <name>Elizabeth Goodman</name>100 </author>101 <title type="html"><![CDATA[Introducing CUDA Rust: Two Tracks for Writing GPU Kernels]]></title>102 <link rel="alternate" type="text/html" href="https://developer.nvidia.com/blog/introducing-cuda-rust-two-tracks-for-writing-gpu-kernels/" />103 <id>https://developer.nvidia.com/blog/?p=122285</id>104 <updated>2026-09-04T23:07:14Z</updated>105 <published>2026-09-08T12:00:00Z</published>106 <category scheme="https://developer.nvidia.com/blog" term="Data Science" /><category scheme="https://developer.nvidia.com/blog" term="Developer Tools & Techniques" /><category scheme="https://developer.nvidia.com/blog" term="CUDA" /><category scheme="https://developer.nvidia.com/blog" term="CUDA Tile" /><category scheme="https://developer.nvidia.com/blog" term="Dynamo" /><category scheme="https://developer.nvidia.com/blog" term="NeMo Retriever" /><category scheme="https://developer.nvidia.com/blog" term="Open Source" /><category scheme="https://developer.nvidia.com/blog" term="Programming Languages / Compilers" /> <summary type="html"><![CDATA[<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-1-768x432.png" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-1-768x432.png 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-1-179x101.png 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-1-300x169.png 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-1-625x352.png 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-1-1536x864.png 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-1-645x363.png 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-1-660x370.png 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-1-500x281.png 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-1-160x90.png 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-1-362x204.png 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-1-196x110.png 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-1-1024x576.png 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-1-960x540.png 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-1.webp 1920w" sizes="auto, (max-width: 768px) 100vw, 768px" title="image1" />In September 2026, NVIDIA announced it is leaning into native GPU programming in Rust. CUDA C++ and CUDA Python are mature, enterprise-grade toolchains, and...]]></summary>107 <content type="html" xml:base="https://developer.nvidia.com/blog/introducing-cuda-rust-two-tracks-for-writing-gpu-kernels/"><![CDATA[<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-1-768x432.png" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-1-768x432.png 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-1-179x101.png 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-1-300x169.png 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-1-625x352.png 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-1-1536x864.png 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-1-645x363.png 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-1-660x370.png 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-1-500x281.png 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-1-160x90.png 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-1-362x204.png 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-1-196x110.png 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-1-1024x576.png 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-1-960x540.png 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-1.webp 1920w" sizes="auto, (max-width: 768px) 100vw, 768px" title="image1" />In September 2026, NVIDIA announced it is leaning into native GPU programming in Rust. CUDA C++ and CUDA Python are mature, enterprise-grade toolchains, and...<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-1-768x432.png" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-1-768x432.png 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-1-179x101.png 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-1-300x169.png 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-1-625x352.png 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-1-1536x864.png 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-1-645x363.png 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-1-660x370.png 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-1-500x281.png 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-1-160x90.png 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-1-362x204.png 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-1-196x110.png 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-1-1024x576.png 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-1-960x540.png 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-1.webp 1920w" sizes="auto, (max-width: 768px) 100vw, 768px" title="image1" /><p>In September 2026, NVIDIA announced it is leaning into native GPU programming in Rust. CUDA C++ and CUDA Python are mature, enterprise-grade toolchains, and NVIDIA will be growing and maturing CUDA Rust into 2027 and beyond The systems layer of AI spans inference engines, serving infrastructure, drivers, and agent runtimes, and it churns constantly as models and techniques change.</p>108<p><a href="https://developer.nvidia.com/blog/introducing-cuda-rust-two-tracks-for-writing-gpu-kernels/" rel="nofollow" data-wpel-link="internal" target="_self">Source</a></p>]]></content>109 <link rel="replies" type="text/html" href="https://developer.nvidia.com/blog/introducing-cuda-rust-two-tracks-for-writing-gpu-kernels/#comments" thr:count="0"/>110 <link rel="replies" type="application/atom+xml" href="https://developer.nvidia.com/blog/introducing-cuda-rust-two-tracks-for-writing-gpu-kernels/feed/" thr:count="0"/>111 <thr:total>0</thr:total>112 </entry>113 <entry>114 <author>115 <name>Tanya Lenz</name>116 </author>117 <title type="html"><![CDATA[Building a Memory-Driven Agent with NVIDIA NemoClaw]]></title>118 <link rel="alternate" type="text/html" href="https://developer.nvidia.com/blog/building-a-memory-driven-agent-with-nvidia-nemoclaw/" />119 <id>https://developer.nvidia.com/blog/?p=122286</id>120 <updated>2026-09-04T18:05:02Z</updated>121 <published>2026-09-04T18:04:55Z</published>122 <category scheme="https://developer.nvidia.com/blog" term="Agentic AI / Generative AI" /><category scheme="https://developer.nvidia.com/blog" term="Developer Tools & Techniques" /><category scheme="https://developer.nvidia.com/blog" term="AI Agent" /><category scheme="https://developer.nvidia.com/blog" term="Build AI Agents" /><category scheme="https://developer.nvidia.com/blog" term="LLMs" /><category scheme="https://developer.nvidia.com/blog" term="NemoClaw" /><category scheme="https://developer.nvidia.com/blog" term="OpenShell" /><category scheme="https://developer.nvidia.com/blog" term="Retrieval Augmented Generation (RAG)" /> <summary type="html"><![CDATA[<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/ai-agent-skills-768x432.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/ai-agent-skills-768x432.jpg 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/ai-agent-skills-179x101.jpg 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/ai-agent-skills-300x169.jpg 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/ai-agent-skills-625x352.jpg 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/ai-agent-skills-1536x864.jpg 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/ai-agent-skills-645x363.jpg 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/ai-agent-skills-660x370.jpg 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/ai-agent-skills-500x281.jpg 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/ai-agent-skills-160x90.jpg 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/ai-agent-skills-362x204.jpg 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/ai-agent-skills-196x110.jpg 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/ai-agent-skills-1024x576.jpg 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/ai-agent-skills-960x540.jpg 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/ai-agent-skills.webp 1920w" sizes="auto, (max-width: 768px) 100vw, 768px" title="ai-agent-skills" />Enterprise work spans messages, decisions, projects, and obligations that change over time. An AI agent that starts without this context must reconstruct it...]]></summary>123 <content type="html" xml:base="https://developer.nvidia.com/blog/building-a-memory-driven-agent-with-nvidia-nemoclaw/"><![CDATA[<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/ai-agent-skills-768x432.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/ai-agent-skills-768x432.jpg 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/ai-agent-skills-179x101.jpg 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/ai-agent-skills-300x169.jpg 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/ai-agent-skills-625x352.jpg 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/ai-agent-skills-1536x864.jpg 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/ai-agent-skills-645x363.jpg 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/ai-agent-skills-660x370.jpg 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/ai-agent-skills-500x281.jpg 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/ai-agent-skills-160x90.jpg 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/ai-agent-skills-362x204.jpg 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/ai-agent-skills-196x110.jpg 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/ai-agent-skills-1024x576.jpg 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/ai-agent-skills-960x540.jpg 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/ai-agent-skills.webp 1920w" sizes="auto, (max-width: 768px) 100vw, 768px" title="ai-agent-skills" />Enterprise work spans messages, decisions, projects, and obligations that change over time. An AI agent that starts without this context must reconstruct it...<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/ai-agent-skills-768x432.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/ai-agent-skills-768x432.jpg 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/ai-agent-skills-179x101.jpg 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/ai-agent-skills-300x169.jpg 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/ai-agent-skills-625x352.jpg 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/ai-agent-skills-1536x864.jpg 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/ai-agent-skills-645x363.jpg 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/ai-agent-skills-660x370.jpg 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/ai-agent-skills-500x281.jpg 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/ai-agent-skills-160x90.jpg 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/ai-agent-skills-362x204.jpg 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/ai-agent-skills-196x110.jpg 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/ai-agent-skills-1024x576.jpg 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/ai-agent-skills-960x540.jpg 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/ai-agent-skills.webp 1920w" sizes="auto, (max-width: 768px) 100vw, 768px" title="ai-agent-skills" /><p>Enterprise work spans messages, decisions, projects, and obligations that change over time. An AI agent that starts without this context must reconstruct it before contributing. To provide agents with this necessary context, our team used NVIDIA NemoClaw to build a memory-driven Chief of Staff. It maintains a human-readable knowledge layer called the self model: an agent memory of relevant…</p>124<p><a href="https://developer.nvidia.com/blog/building-a-memory-driven-agent-with-nvidia-nemoclaw/" rel="nofollow" data-wpel-link="internal" target="_self">Source</a></p>]]></content>125 <link rel="replies" type="text/html" href="https://developer.nvidia.com/blog/building-a-memory-driven-agent-with-nvidia-nemoclaw/#comments" thr:count="0"/>126 <link rel="replies" type="application/atom+xml" href="https://developer.nvidia.com/blog/building-a-memory-driven-agent-with-nvidia-nemoclaw/feed/" thr:count="0"/>127 <thr:total>0</thr:total>128 </entry>129 <entry>130 <author>131 <name>Elizabeth Goodman</name>132 </author>133 <title type="html"><![CDATA[Frontier Reasoning Reaches the Edge: How to Deploy and Optimize Models on NVIDIA Jetson]]></title>134 <link rel="alternate" type="text/html" href="https://developer.nvidia.com/blog/frontier-reasoning-reaches-the-edge-how-to-deploy-and-optimize-models-on-nvidia-jetson/" />135 <id>https://developer.nvidia.com/blog/?p=122184</id>136 <updated>2026-09-04T16:21:13Z</updated>137 <published>2026-09-04T16:21:04Z</published>138 <category scheme="https://developer.nvidia.com/blog" term="Agentic AI / Generative AI" /><category scheme="https://developer.nvidia.com/blog" term="Edge Computing" /><category scheme="https://developer.nvidia.com/blog" term="JetPack" /><category scheme="https://developer.nvidia.com/blog" term="Jetson" /><category scheme="https://developer.nvidia.com/blog" term="Jetson Orin" /><category scheme="https://developer.nvidia.com/blog" term="Physical AI" /><category scheme="https://developer.nvidia.com/blog" term="Thor" /><category scheme="https://developer.nvidia.com/blog" term="Tutorial" /> <summary type="html"><![CDATA[<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image2-3-768x432.png" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image2-3-768x432.png 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image2-3-179x101.png 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image2-3-300x169.png 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image2-3-625x352.png 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image2-3-1536x864.png 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image2-3-645x363.png 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image2-3-660x370.png 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image2-3-500x281.png 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image2-3-160x90.png 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image2-3-362x204.png 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image2-3-196x110.png 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image2-3-1024x576.png 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image2-3-960x540.png 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image2-3.webp 1920w" sizes="auto, (max-width: 768px) 100vw, 768px" title="image2" />Running reasoning and agentic AI at the edge has been harder than it needs to be. Until recently, models capable of multi-step reasoning were too large to run...]]></summary>139 <content type="html" xml:base="https://developer.nvidia.com/blog/frontier-reasoning-reaches-the-edge-how-to-deploy-and-optimize-models-on-nvidia-jetson/"><![CDATA[<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image2-3-768x432.png" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image2-3-768x432.png 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image2-3-179x101.png 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image2-3-300x169.png 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image2-3-625x352.png 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image2-3-1536x864.png 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image2-3-645x363.png 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image2-3-660x370.png 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image2-3-500x281.png 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image2-3-160x90.png 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image2-3-362x204.png 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image2-3-196x110.png 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image2-3-1024x576.png 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image2-3-960x540.png 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image2-3.webp 1920w" sizes="auto, (max-width: 768px) 100vw, 768px" title="image2" />Running reasoning and agentic AI at the edge has been harder than it needs to be. Until recently, models capable of multi-step reasoning were too large to run...<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image2-3-768x432.png" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image2-3-768x432.png 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image2-3-179x101.png 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image2-3-300x169.png 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image2-3-625x352.png 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image2-3-1536x864.png 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image2-3-645x363.png 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image2-3-660x370.png 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image2-3-500x281.png 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image2-3-160x90.png 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image2-3-362x204.png 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image2-3-196x110.png 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image2-3-1024x576.png 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image2-3-960x540.png 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image2-3.webp 1920w" sizes="auto, (max-width: 768px) 100vw, 768px" title="image2" /><p>Running reasoning and agentic AI at the edge has been harder than it needs to be. Until recently, models capable of multi-step reasoning were too large to run locally on edge hardware. Developers building agents have had to route inference through a data center, adding network dependency, increasing costs, and exposing data that may need to stay on device. That constraint is lifting.</p>140<p><a href="https://developer.nvidia.com/blog/frontier-reasoning-reaches-the-edge-how-to-deploy-and-optimize-models-on-nvidia-jetson/" rel="nofollow" data-wpel-link="internal" target="_self">Source</a></p>]]></content>141 <link rel="replies" type="text/html" href="https://developer.nvidia.com/blog/frontier-reasoning-reaches-the-edge-how-to-deploy-and-optimize-models-on-nvidia-jetson/#comments" thr:count="0"/>142 <link rel="replies" type="application/atom+xml" href="https://developer.nvidia.com/blog/frontier-reasoning-reaches-the-edge-how-to-deploy-and-optimize-models-on-nvidia-jetson/feed/" thr:count="0"/>143 <thr:total>0</thr:total>144 </entry>145 <entry>146 <author>147 <name>Elizabeth Goodman</name>148 </author>149 <title type="html"><![CDATA[How to Carry User Identity Across Federated Kubernetes and AI Platforms]]></title>150 <link rel="alternate" type="text/html" href="https://developer.nvidia.com/blog/how-to-carry-user-identity-across-federated-kubernetes-and-ai-platforms/" />151 <id>https://developer.nvidia.com/blog/?p=122235</id>152 <updated>2026-09-03T22:36:46Z</updated>153 <published>2026-09-03T22:36:02Z</published>154 <category scheme="https://developer.nvidia.com/blog" term="AI Platforms/Deployment" /><category scheme="https://developer.nvidia.com/blog" term="Data Center / Cloud" /><category scheme="https://developer.nvidia.com/blog" term="Developer Tools & Techniques" /><category scheme="https://developer.nvidia.com/blog" term="AI Platform" /><category scheme="https://developer.nvidia.com/blog" term="Cloud Services" /><category scheme="https://developer.nvidia.com/blog" term="Infrastructure" /><category scheme="https://developer.nvidia.com/blog" term="Kubernetes" /><category scheme="https://developer.nvidia.com/blog" term="Software-Defined Data Center" /> <summary type="html"><![CDATA[<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/kai-scheduler-representation-768x432.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/kai-scheduler-representation-768x432.jpg 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/kai-scheduler-representation-179x101.jpg 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/kai-scheduler-representation-300x169.jpg 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/kai-scheduler-representation-625x352.jpg 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/kai-scheduler-representation-1536x864.jpg 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/kai-scheduler-representation-645x363.jpg 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/kai-scheduler-representation-660x370.jpg 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/kai-scheduler-representation-500x281.jpg 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/kai-scheduler-representation-160x90.jpg 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/kai-scheduler-representation-362x204.jpg 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/kai-scheduler-representation-196x110.jpg 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/kai-scheduler-representation-1024x576.jpg 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/kai-scheduler-representation-960x540.jpg 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/kai-scheduler-representation.webp 1920w" sizes="auto, (max-width: 768px) 100vw, 768px" title="kai-scheduler-representation" />Modern AI platforms are no longer a single application behind one login screen. A user may start in a central portal, open a governed dataset, launch a notebook...]]></summary>155 <content type="html" xml:base="https://developer.nvidia.com/blog/how-to-carry-user-identity-across-federated-kubernetes-and-ai-platforms/"><![CDATA[<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/kai-scheduler-representation-768x432.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/kai-scheduler-representation-768x432.jpg 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/kai-scheduler-representation-179x101.jpg 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/kai-scheduler-representation-300x169.jpg 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/kai-scheduler-representation-625x352.jpg 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/kai-scheduler-representation-1536x864.jpg 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/kai-scheduler-representation-645x363.jpg 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/kai-scheduler-representation-660x370.jpg 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/kai-scheduler-representation-500x281.jpg 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/kai-scheduler-representation-160x90.jpg 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/kai-scheduler-representation-362x204.jpg 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/kai-scheduler-representation-196x110.jpg 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/kai-scheduler-representation-1024x576.jpg 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/kai-scheduler-representation-960x540.jpg 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/kai-scheduler-representation.webp 1920w" sizes="auto, (max-width: 768px) 100vw, 768px" title="kai-scheduler-representation" />Modern AI platforms are no longer a single application behind one login screen. A user may start in a central portal, open a governed dataset, launch a notebook...<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/kai-scheduler-representation-768x432.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/kai-scheduler-representation-768x432.jpg 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/kai-scheduler-representation-179x101.jpg 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/kai-scheduler-representation-300x169.jpg 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/kai-scheduler-representation-625x352.jpg 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/kai-scheduler-representation-1536x864.jpg 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/kai-scheduler-representation-645x363.jpg 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/kai-scheduler-representation-660x370.jpg 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/kai-scheduler-representation-500x281.jpg 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/kai-scheduler-representation-160x90.jpg 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/kai-scheduler-representation-362x204.jpg 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/kai-scheduler-representation-196x110.jpg 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/kai-scheduler-representation-1024x576.jpg 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/kai-scheduler-representation-960x540.jpg 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/kai-scheduler-representation.webp 1920w" sizes="auto, (max-width: 768px) 100vw, 768px" title="kai-scheduler-representation" /><p>Modern AI platforms are no longer a single application behind one login screen. A user may start in a central portal, open a governed dataset, launch a notebook where that data resides, and invoke an assistant that calls services in another cluster. The workflow feels unified, but identity crosses control-plane and data-plane boundaries at every step. That is where conventional single sign-on…</p>156<p><a href="https://developer.nvidia.com/blog/how-to-carry-user-identity-across-federated-kubernetes-and-ai-platforms/" rel="nofollow" data-wpel-link="internal" target="_self">Source</a></p>]]></content>157 <link rel="replies" type="text/html" href="https://developer.nvidia.com/blog/how-to-carry-user-identity-across-federated-kubernetes-and-ai-platforms/#comments" thr:count="0"/>158 <link rel="replies" type="application/atom+xml" href="https://developer.nvidia.com/blog/how-to-carry-user-identity-across-federated-kubernetes-and-ai-platforms/feed/" thr:count="0"/>159 <thr:total>0</thr:total>160 </entry>161 <entry>162 <author>163 <name>Tanya Lenz</name>164 </author>165 <title type="html"><![CDATA[NVIDIA PAIR Virtual Inference Router Expands Available Compute on Your Local Network]]></title>166 <link rel="alternate" type="text/html" href="https://developer.nvidia.com/blog/nvidia-pair-virtual-inference-router-expands-available-compute-on-your-local-network/" />167 <id>https://developer.nvidia.com/blog/?p=121926</id>168 <updated>2026-09-02T21:33:39Z</updated>169 <published>2026-09-03T16:00:00Z</published>170 <category scheme="https://developer.nvidia.com/blog" term="Content Creation / Rendering" /><category scheme="https://developer.nvidia.com/blog" term="Edge Computing" /><category scheme="https://developer.nvidia.com/blog" term="AI Agent" /><category scheme="https://developer.nvidia.com/blog" term="AI Inference" /><category scheme="https://developer.nvidia.com/blog" term="DGX Spark" /><category scheme="https://developer.nvidia.com/blog" term="Gaming" /><category scheme="https://developer.nvidia.com/blog" term="GeForce" /><category scheme="https://developer.nvidia.com/blog" term="Open Source" /> <summary type="html"><![CDATA[<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/nvidia-pair-virtual-inference-router-768x432.png" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/nvidia-pair-virtual-inference-router-768x432.png 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/nvidia-pair-virtual-inference-router-179x101.png 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/nvidia-pair-virtual-inference-router-300x169.png 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/nvidia-pair-virtual-inference-router-625x352.png 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/nvidia-pair-virtual-inference-router-1536x864.png 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/nvidia-pair-virtual-inference-router-645x363.png 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/nvidia-pair-virtual-inference-router-660x370.png 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/nvidia-pair-virtual-inference-router-500x281.png 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/nvidia-pair-virtual-inference-router-160x90.png 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/nvidia-pair-virtual-inference-router-362x204.png 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/nvidia-pair-virtual-inference-router-196x110.png 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/nvidia-pair-virtual-inference-router-1024x576.png 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/nvidia-pair-virtual-inference-router-960x540.png 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/nvidia-pair-virtual-inference-router.webp 1920w" sizes="auto, (max-width: 768px) 100vw, 768px" title="nvidia-pair-virtual-inference-router" />AI agents are learning to do more by working together. A lead agent can break a complex task into smaller jobs and assign those jobs to specialized subagents....]]></summary>171 <content type="html" xml:base="https://developer.nvidia.com/blog/nvidia-pair-virtual-inference-router-expands-available-compute-on-your-local-network/"><![CDATA[<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/nvidia-pair-virtual-inference-router-768x432.png" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/nvidia-pair-virtual-inference-router-768x432.png 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/nvidia-pair-virtual-inference-router-179x101.png 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/nvidia-pair-virtual-inference-router-300x169.png 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/nvidia-pair-virtual-inference-router-625x352.png 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/nvidia-pair-virtual-inference-router-1536x864.png 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/nvidia-pair-virtual-inference-router-645x363.png 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/nvidia-pair-virtual-inference-router-660x370.png 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/nvidia-pair-virtual-inference-router-500x281.png 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/nvidia-pair-virtual-inference-router-160x90.png 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/nvidia-pair-virtual-inference-router-362x204.png 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/nvidia-pair-virtual-inference-router-196x110.png 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/nvidia-pair-virtual-inference-router-1024x576.png 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/nvidia-pair-virtual-inference-router-960x540.png 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/nvidia-pair-virtual-inference-router.webp 1920w" sizes="auto, (max-width: 768px) 100vw, 768px" title="nvidia-pair-virtual-inference-router" />AI agents are learning to do more by working together. A lead agent can break a complex task into smaller jobs and assign those jobs to specialized subagents....<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/nvidia-pair-virtual-inference-router-768x432.png" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/nvidia-pair-virtual-inference-router-768x432.png 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/nvidia-pair-virtual-inference-router-179x101.png 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/nvidia-pair-virtual-inference-router-300x169.png 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/nvidia-pair-virtual-inference-router-625x352.png 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/nvidia-pair-virtual-inference-router-1536x864.png 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/nvidia-pair-virtual-inference-router-645x363.png 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/nvidia-pair-virtual-inference-router-660x370.png 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/nvidia-pair-virtual-inference-router-500x281.png 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/nvidia-pair-virtual-inference-router-160x90.png 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/nvidia-pair-virtual-inference-router-362x204.png 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/nvidia-pair-virtual-inference-router-196x110.png 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/nvidia-pair-virtual-inference-router-1024x576.png 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/nvidia-pair-virtual-inference-router-960x540.png 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/nvidia-pair-virtual-inference-router.webp 1920w" sizes="auto, (max-width: 768px) 100vw, 768px" title="nvidia-pair-virtual-inference-router" /><p>AI agents are learning to do more by working together. A lead agent can break a complex task into smaller jobs and assign those jobs to specialized subagents. Additionally, users are starting to run multiple agent sessions at the same time. Multi-agent workflows for accomplishing complex tasks are also becoming more common. This breadth-first approach can improve the speed of task completion…</p>172<p><a href="https://developer.nvidia.com/blog/nvidia-pair-virtual-inference-router-expands-available-compute-on-your-local-network/" rel="nofollow" data-wpel-link="internal" target="_self">Source</a></p>]]></content>173 <link rel="replies" type="text/html" href="https://developer.nvidia.com/blog/nvidia-pair-virtual-inference-router-expands-available-compute-on-your-local-network/#comments" thr:count="0"/>174 <link rel="replies" type="application/atom+xml" href="https://developer.nvidia.com/blog/nvidia-pair-virtual-inference-router-expands-available-compute-on-your-local-network/feed/" thr:count="0"/>175 <thr:total>0</thr:total>176 </entry>177 <entry>178 <author>179 <name>Elizabeth Goodman</name>180 </author>181 <title type="html"><![CDATA[The Modern CUDA Toolbox in Practice: A Step-by-Step Optimization Walkthrough]]></title>182 <link rel="alternate" type="text/html" href="https://developer.nvidia.com/blog/the-modern-cuda-toolbox-in-practice-a-step-by-step-optimization-walkthrough/" />183 <id>https://developer.nvidia.com/blog/?p=122046</id>184 <updated>2026-09-02T17:16:06Z</updated>185 <published>2026-09-02T17:15:57Z</published>186 <category scheme="https://developer.nvidia.com/blog" term="Developer Tools & Techniques" /><category scheme="https://developer.nvidia.com/blog" term="CUDA" /> <summary type="html"><![CDATA[<img width="768" height="431" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/keyboard_16x9-768x431.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/keyboard_16x9-768x431.jpg 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/keyboard_16x9-179x100.jpg 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/keyboard_16x9-300x168.jpg 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/keyboard_16x9-625x351.jpg 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/keyboard_16x9-645x362.jpg 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/keyboard_16x9-660x370.jpg 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/keyboard_16x9-500x280.jpg 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/keyboard_16x9-160x90.jpg 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/keyboard_16x9-362x203.jpg 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/keyboard_16x9-196x110.jpg 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/keyboard_16x9-1024x574.jpg 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/keyboard_16x9-960x538.jpg 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/keyboard_16x9.webp 1480w" sizes="auto, (max-width: 768px) 100vw, 768px" title="keyboard_16x9" />NVIDIA CUDA remains the foundation of GPU-accelerated computing, powering everything from scientific simulations to large-scale AI training. But writing...]]></summary>187 <content type="html" xml:base="https://developer.nvidia.com/blog/the-modern-cuda-toolbox-in-practice-a-step-by-step-optimization-walkthrough/"><![CDATA[<img width="768" height="431" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/keyboard_16x9-768x431.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/keyboard_16x9-768x431.jpg 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/keyboard_16x9-179x100.jpg 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/keyboard_16x9-300x168.jpg 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/keyboard_16x9-625x351.jpg 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/keyboard_16x9-645x362.jpg 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/keyboard_16x9-660x370.jpg 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/keyboard_16x9-500x280.jpg 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/keyboard_16x9-160x90.jpg 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/keyboard_16x9-362x203.jpg 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/keyboard_16x9-196x110.jpg 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/keyboard_16x9-1024x574.jpg 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/keyboard_16x9-960x538.jpg 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/keyboard_16x9.webp 1480w" sizes="auto, (max-width: 768px) 100vw, 768px" title="keyboard_16x9" />NVIDIA CUDA remains the foundation of GPU-accelerated computing, powering everything from scientific simulations to large-scale AI training. But writing...<img width="768" height="431" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/keyboard_16x9-768x431.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/keyboard_16x9-768x431.jpg 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/keyboard_16x9-179x100.jpg 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/keyboard_16x9-300x168.jpg 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/keyboard_16x9-625x351.jpg 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/keyboard_16x9-645x362.jpg 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/keyboard_16x9-660x370.jpg 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/keyboard_16x9-500x280.jpg 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/keyboard_16x9-160x90.jpg 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/keyboard_16x9-362x203.jpg 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/keyboard_16x9-196x110.jpg 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/keyboard_16x9-1024x574.jpg 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/keyboard_16x9-960x538.jpg 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/keyboard_16x9.webp 1480w" sizes="auto, (max-width: 768px) 100vw, 768px" title="keyboard_16x9" /><p></p>188<p><a href="https://developer.nvidia.com/blog/the-modern-cuda-toolbox-in-practice-a-step-by-step-optimization-walkthrough/" rel="nofollow" data-wpel-link="internal" target="_self">Source</a></p>]]></content>189 <link rel="replies" type="text/html" href="https://developer.nvidia.com/blog/the-modern-cuda-toolbox-in-practice-a-step-by-step-optimization-walkthrough/#comments" thr:count="0"/>190 <link rel="replies" type="application/atom+xml" href="https://developer.nvidia.com/blog/the-modern-cuda-toolbox-in-practice-a-step-by-step-optimization-walkthrough/feed/" thr:count="0"/>191 <thr:total>0</thr:total>192 </entry>193 <entry>194 <author>195 <name>Tanya Lenz</name>196 </author>197 <title type="html"><![CDATA[Co-Designing AI Models Using Speculative Decoding for Faster LLM Inference]]></title>198 <link rel="alternate" type="text/html" href="https://developer.nvidia.com/blog/co-designing-ai-models-using-speculative-decoding-for-faster-llm-inference/" />199 <id>https://developer.nvidia.com/blog/?p=122024</id>200 <updated>2026-09-02T23:06:37Z</updated>201 <published>2026-09-02T16:04:19Z</published>202 <category scheme="https://developer.nvidia.com/blog" term="Agentic AI / Generative AI" /><category scheme="https://developer.nvidia.com/blog" term="Data Center / Cloud" /><category scheme="https://developer.nvidia.com/blog" term="Developer Tools & Techniques" /><category scheme="https://developer.nvidia.com/blog" term="AI Inference" /><category scheme="https://developer.nvidia.com/blog" term="Inference Performance" /><category scheme="https://developer.nvidia.com/blog" term="LLMs" /><category scheme="https://developer.nvidia.com/blog" term="NVFP4" /><category scheme="https://developer.nvidia.com/blog" term="Training AI Models" /> <summary type="html"><![CDATA[<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/07/llm-optimize-deploy-768x432.png" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/07/llm-optimize-deploy-768x432.png 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/07/llm-optimize-deploy-179x101.png 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/07/llm-optimize-deploy-300x169.png 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/07/llm-optimize-deploy-625x352.png 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/07/llm-optimize-deploy-1536x864.png 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/07/llm-optimize-deploy-645x363.png 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/07/llm-optimize-deploy-660x370.png 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/07/llm-optimize-deploy-500x281.png 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/07/llm-optimize-deploy-160x90.png 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/07/llm-optimize-deploy-362x204.png 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/07/llm-optimize-deploy-196x110.png 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/07/llm-optimize-deploy-1024x576.png 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/07/llm-optimize-deploy-960x540.png 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/07/llm-optimize-deploy.webp 1877w" sizes="auto, (max-width: 768px) 100vw, 768px" title="llm-optimize-deploy" />This post is the third in a series on AI model co-design. It explores how to accelerate LLM inference while maintaining accuracy using speculative decoding and...]]></summary>203 <content type="html" xml:base="https://developer.nvidia.com/blog/co-designing-ai-models-using-speculative-decoding-for-faster-llm-inference/"><![CDATA[<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/07/llm-optimize-deploy-768x432.png" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/07/llm-optimize-deploy-768x432.png 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/07/llm-optimize-deploy-179x101.png 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/07/llm-optimize-deploy-300x169.png 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/07/llm-optimize-deploy-625x352.png 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/07/llm-optimize-deploy-1536x864.png 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/07/llm-optimize-deploy-645x363.png 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/07/llm-optimize-deploy-660x370.png 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/07/llm-optimize-deploy-500x281.png 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/07/llm-optimize-deploy-160x90.png 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/07/llm-optimize-deploy-362x204.png 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/07/llm-optimize-deploy-196x110.png 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/07/llm-optimize-deploy-1024x576.png 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/07/llm-optimize-deploy-960x540.png 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/07/llm-optimize-deploy.webp 1877w" sizes="auto, (max-width: 768px) 100vw, 768px" title="llm-optimize-deploy" />This post is the third in a series on AI model co-design. It explores how to accelerate LLM inference while maintaining accuracy using speculative decoding and...<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/07/llm-optimize-deploy-768x432.png" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/07/llm-optimize-deploy-768x432.png 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/07/llm-optimize-deploy-179x101.png 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/07/llm-optimize-deploy-300x169.png 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/07/llm-optimize-deploy-625x352.png 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/07/llm-optimize-deploy-1536x864.png 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/07/llm-optimize-deploy-645x363.png 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/07/llm-optimize-deploy-660x370.png 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/07/llm-optimize-deploy-500x281.png 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/07/llm-optimize-deploy-160x90.png 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/07/llm-optimize-deploy-362x204.png 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/07/llm-optimize-deploy-196x110.png 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/07/llm-optimize-deploy-1024x576.png 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/07/llm-optimize-deploy-960x540.png 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/07/llm-optimize-deploy.webp 1877w" sizes="auto, (max-width: 768px) 100vw, 768px" title="llm-optimize-deploy" /><p>This post is the third in a series on AI model co-design. It explores how to accelerate LLM inference while maintaining accuracy using speculative decoding and offers five guidelines for selecting draft length and draft mechanism across the Pareto frontier. For a discussion of how model design choices impact both throughput and interactivity without sacrificing accuracy, see AI Model Co…</p>204<p><a href="https://developer.nvidia.com/blog/co-designing-ai-models-using-speculative-decoding-for-faster-llm-inference/" rel="nofollow" data-wpel-link="internal" target="_self">Source</a></p>]]></content>205 <link rel="replies" type="text/html" href="https://developer.nvidia.com/blog/co-designing-ai-models-using-speculative-decoding-for-faster-llm-inference/#comments" thr:count="0"/>206 <link rel="replies" type="application/atom+xml" href="https://developer.nvidia.com/blog/co-designing-ai-models-using-speculative-decoding-for-faster-llm-inference/feed/" thr:count="0"/>207 <thr:total>0</thr:total>208 </entry>209 <entry>210 <author>211 <name>Michelle Horton</name>212 </author>213 <title type="html"><![CDATA[Building an Adaptive Agentic Cybersecurity System with NVIDIA Nemotron]]></title>214 <link rel="alternate" type="text/html" href="https://developer.nvidia.com/blog/building-an-adaptive-agentic-cybersecurity-system-with-nvidia-nemotron/" />215 <id>https://developer.nvidia.com/blog/?p=121995</id>216 <updated>2026-09-01T19:00:37Z</updated>217 <published>2026-09-01T17:00:04Z</published>218 <category scheme="https://developer.nvidia.com/blog" term="Agentic AI / Generative AI" /><category scheme="https://developer.nvidia.com/blog" term="Top Stories" /><category scheme="https://developer.nvidia.com/blog" term="Trustworthy AI / Cybersecurity" /><category scheme="https://developer.nvidia.com/blog" term="NeMo" /><category scheme="https://developer.nvidia.com/blog" term="Security for AI" /> <summary type="html"><![CDATA[<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-Adaptive-Security-e1788211755260-768x432.webp" class="webfeedsFeaturedVisual wp-post-image" alt="An illustration showing agentic security." style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-Adaptive-Security-e1788211755260-768x432.webp 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-Adaptive-Security-e1788211755260-179x101.webp 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-Adaptive-Security-e1788211755260-300x169.webp 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-Adaptive-Security-e1788211755260-625x352.webp 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-Adaptive-Security-e1788211755260-1536x865.webp 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-Adaptive-Security-e1788211755260-2048x1153.webp 2048w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-Adaptive-Security-e1788211755260-645x363.webp 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-Adaptive-Security-e1788211755260-657x370.webp 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-Adaptive-Security-e1788211755260-500x282.webp 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-Adaptive-Security-e1788211755260-160x90.webp 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-Adaptive-Security-e1788211755260-362x204.webp 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-Adaptive-Security-e1788211755260-195x110.webp 195w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-Adaptive-Security-e1788211755260-1024x577.webp 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-Adaptive-Security-e1788211755260-960x540.webp 960w" sizes="auto, (max-width: 768px) 100vw, 768px" title="Agentic-Adaptive-Security" />AI is changing the pace of cybersecurity. Agentic systems can coordinate work and pursue complex objectives over long horizons. Security teams are beginning to...]]></summary>219 <content type="html" xml:base="https://developer.nvidia.com/blog/building-an-adaptive-agentic-cybersecurity-system-with-nvidia-nemotron/"><![CDATA[<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-Adaptive-Security-e1788211755260-768x432.webp" class="webfeedsFeaturedVisual wp-post-image" alt="An illustration showing agentic security." style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-Adaptive-Security-e1788211755260-768x432.webp 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-Adaptive-Security-e1788211755260-179x101.webp 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-Adaptive-Security-e1788211755260-300x169.webp 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-Adaptive-Security-e1788211755260-625x352.webp 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-Adaptive-Security-e1788211755260-1536x865.webp 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-Adaptive-Security-e1788211755260-2048x1153.webp 2048w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-Adaptive-Security-e1788211755260-645x363.webp 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-Adaptive-Security-e1788211755260-657x370.webp 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-Adaptive-Security-e1788211755260-500x282.webp 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-Adaptive-Security-e1788211755260-160x90.webp 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-Adaptive-Security-e1788211755260-362x204.webp 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-Adaptive-Security-e1788211755260-195x110.webp 195w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-Adaptive-Security-e1788211755260-1024x577.webp 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-Adaptive-Security-e1788211755260-960x540.webp 960w" sizes="auto, (max-width: 768px) 100vw, 768px" title="Agentic-Adaptive-Security" />AI is changing the pace of cybersecurity. Agentic systems can coordinate work and pursue complex objectives over long horizons. Security teams are beginning to...<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-Adaptive-Security-e1788211755260-768x432.webp" class="webfeedsFeaturedVisual wp-post-image" alt="An illustration showing agentic security." style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-Adaptive-Security-e1788211755260-768x432.webp 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-Adaptive-Security-e1788211755260-179x101.webp 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-Adaptive-Security-e1788211755260-300x169.webp 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-Adaptive-Security-e1788211755260-625x352.webp 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-Adaptive-Security-e1788211755260-1536x865.webp 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-Adaptive-Security-e1788211755260-2048x1153.webp 2048w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-Adaptive-Security-e1788211755260-645x363.webp 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-Adaptive-Security-e1788211755260-657x370.webp 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-Adaptive-Security-e1788211755260-500x282.webp 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-Adaptive-Security-e1788211755260-160x90.webp 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-Adaptive-Security-e1788211755260-362x204.webp 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-Adaptive-Security-e1788211755260-195x110.webp 195w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-Adaptive-Security-e1788211755260-1024x577.webp 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-Adaptive-Security-e1788211755260-960x540.webp 960w" sizes="auto, (max-width: 768px) 100vw, 768px" title="Agentic-Adaptive-Security" /><p>AI is changing the pace of cybersecurity. Agentic systems can coordinate work and pursue complex objectives over long horizons. Security teams are beginning to apply agents across security operations, but many implementations remain anchored to existing alerts, predefined workflows, and known attack behaviors. The harder problem is identifying what defenses miss and turning those gaps into…</p>220<p><a href="https://developer.nvidia.com/blog/building-an-adaptive-agentic-cybersecurity-system-with-nvidia-nemotron/" rel="nofollow" data-wpel-link="internal" target="_self">Source</a></p>]]></content>221 <link rel="replies" type="text/html" href="https://developer.nvidia.com/blog/building-an-adaptive-agentic-cybersecurity-system-with-nvidia-nemotron/#comments" thr:count="0"/>222 <link rel="replies" type="application/atom+xml" href="https://developer.nvidia.com/blog/building-an-adaptive-agentic-cybersecurity-system-with-nvidia-nemotron/feed/" thr:count="0"/>223 <thr:total>0</thr:total>224 </entry>225 <entry>226 <author>227 <name>Elizabeth Goodman</name>228 </author>229 <title type="html"><![CDATA[How to Size GPUs for AI Inference and TCO Without Overspending]]></title>230 <link rel="alternate" type="text/html" href="https://developer.nvidia.com/blog/how-to-size-gpus-for-ai-inference-and-tco-without-overspending/" />231 <id>https://developer.nvidia.com/blog/?p=121993</id>232 <updated>2026-08-31T19:28:23Z</updated>233 <published>2026-09-01T15:00:00Z</published>234 <category scheme="https://developer.nvidia.com/blog" term="Agentic AI / Generative AI" /><category scheme="https://developer.nvidia.com/blog" term="Data Center / Cloud" /><category scheme="https://developer.nvidia.com/blog" term="MLOps" /><category scheme="https://developer.nvidia.com/blog" term="AI Factory" /><category scheme="https://developer.nvidia.com/blog" term="AI Inference" /><category scheme="https://developer.nvidia.com/blog" term="Inference Performance" /> <summary type="html"><![CDATA[<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image4-5-768x432.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image4-5-768x432.jpg 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image4-5-179x101.jpg 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image4-5-300x169.jpg 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image4-5-625x352.jpg 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image4-5-1536x864.jpg 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image4-5-645x363.jpg 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image4-5-660x370.jpg 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image4-5-500x281.jpg 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image4-5-160x90.jpg 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image4-5-362x204.jpg 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image4-5-195x110.jpg 195w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image4-5-1024x576.jpg 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image4-5-960x540.jpg 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image4-5.webp 1999w" sizes="auto, (max-width: 768px) 100vw, 768px" title="image4" />The surge in AI adoption is transforming everything from chatbots to content generation. Still, a common pain point remains: How can organizations confidently...]]></summary>235 <content type="html" xml:base="https://developer.nvidia.com/blog/how-to-size-gpus-for-ai-inference-and-tco-without-overspending/"><![CDATA[<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image4-5-768x432.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image4-5-768x432.jpg 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image4-5-179x101.jpg 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image4-5-300x169.jpg 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image4-5-625x352.jpg 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image4-5-1536x864.jpg 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image4-5-645x363.jpg 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image4-5-660x370.jpg 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image4-5-500x281.jpg 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image4-5-160x90.jpg 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image4-5-362x204.jpg 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image4-5-195x110.jpg 195w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image4-5-1024x576.jpg 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image4-5-960x540.jpg 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image4-5.webp 1999w" sizes="auto, (max-width: 768px) 100vw, 768px" title="image4" />The surge in AI adoption is transforming everything from chatbots to content generation. Still, a common pain point remains: How can organizations confidently...<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image4-5-768x432.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image4-5-768x432.jpg 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image4-5-179x101.jpg 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image4-5-300x169.jpg 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image4-5-625x352.jpg 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image4-5-1536x864.jpg 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image4-5-645x363.jpg 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image4-5-660x370.jpg 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image4-5-500x281.jpg 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image4-5-160x90.jpg 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image4-5-362x204.jpg 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image4-5-195x110.jpg 195w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image4-5-1024x576.jpg 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image4-5-960x540.jpg 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image4-5.webp 1999w" sizes="auto, (max-width: 768px) 100vw, 768px" title="image4" /><p>The surge in AI adoption is transforming everything from chatbots to content generation. Still, a common pain point remains: How can organizations confidently size GPU resources for inference workloads and optimize Total Cost of Ownership (TCO)? With a dizzying mix of latency targets, model choices, quirky traffic patterns, and budget constraints, it’s easy to feel lost in the weeds…</p>236<p><a href="https://developer.nvidia.com/blog/how-to-size-gpus-for-ai-inference-and-tco-without-overspending/" rel="nofollow" data-wpel-link="internal" target="_self">Source</a></p>]]></content>237 <link rel="replies" type="text/html" href="https://developer.nvidia.com/blog/how-to-size-gpus-for-ai-inference-and-tco-without-overspending/#comments" thr:count="0"/>238 <link rel="replies" type="application/atom+xml" href="https://developer.nvidia.com/blog/how-to-size-gpus-for-ai-inference-and-tco-without-overspending/feed/" thr:count="0"/>239 <thr:total>0</thr:total>240 </entry>241 <entry>242 <author>243 <name>Michelle Horton</name>244 </author>245 <title type="html"><![CDATA[Run NVIDIA BioNeMo NIM Microservices for Protein Structure Prediction in Claude Science]]></title>246 <link rel="alternate" type="text/html" href="https://developer.nvidia.com/blog/run-nvidia-bionemo-nim-microservices-for-protein-structure-prediction-in-claude-science/" />247 <id>https://developer.nvidia.com/blog/?p=121350</id>248 <updated>2026-08-31T16:23:58Z</updated>249 <published>2026-08-31T16:30:00Z</published>250 <category scheme="https://developer.nvidia.com/blog" term="Agentic AI / Generative AI" /><category scheme="https://developer.nvidia.com/blog" term="Developer Tools & Techniques" /><category scheme="https://developer.nvidia.com/blog" term="Simulation / Modeling / Design" /><category scheme="https://developer.nvidia.com/blog" term="Drug Discovery" /> <summary type="html"><![CDATA[<img width="768" height="431" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/AdobeStock_842772489-e1787092703768-768x431.webp" class="webfeedsFeaturedVisual wp-post-image" alt="A picture of a protein molecule." style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/AdobeStock_842772489-e1787092703768-768x431.webp 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/AdobeStock_842772489-e1787092703768-179x100.webp 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/AdobeStock_842772489-e1787092703768-300x168.webp 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/AdobeStock_842772489-e1787092703768-625x351.webp 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/AdobeStock_842772489-e1787092703768-1536x862.webp 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/AdobeStock_842772489-e1787092703768-645x362.webp 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/AdobeStock_842772489-e1787092703768-660x370.webp 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/AdobeStock_842772489-e1787092703768-500x281.webp 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/AdobeStock_842772489-e1787092703768-160x90.webp 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/AdobeStock_842772489-e1787092703768-362x203.webp 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/AdobeStock_842772489-e1787092703768-196x110.webp 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/AdobeStock_842772489-e1787092703768-1024x575.webp 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/AdobeStock_842772489-e1787092703768-960x540.webp 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/AdobeStock_842772489-e1787092703768.webp 1843w" sizes="auto, (max-width: 768px) 100vw, 768px" title="" />Agentic AI is changing how research is done. AI scientists can read papers, propose hypotheses, call models, and determine which experiments to prioritize next....]]></summary>251 <content type="html" xml:base="https://developer.nvidia.com/blog/run-nvidia-bionemo-nim-microservices-for-protein-structure-prediction-in-claude-science/"><![CDATA[<img width="768" height="431" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/AdobeStock_842772489-e1787092703768-768x431.webp" class="webfeedsFeaturedVisual wp-post-image" alt="A picture of a protein molecule." style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/AdobeStock_842772489-e1787092703768-768x431.webp 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/AdobeStock_842772489-e1787092703768-179x100.webp 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/AdobeStock_842772489-e1787092703768-300x168.webp 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/AdobeStock_842772489-e1787092703768-625x351.webp 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/AdobeStock_842772489-e1787092703768-1536x862.webp 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/AdobeStock_842772489-e1787092703768-645x362.webp 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/AdobeStock_842772489-e1787092703768-660x370.webp 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/AdobeStock_842772489-e1787092703768-500x281.webp 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/AdobeStock_842772489-e1787092703768-160x90.webp 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/AdobeStock_842772489-e1787092703768-362x203.webp 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/AdobeStock_842772489-e1787092703768-196x110.webp 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/AdobeStock_842772489-e1787092703768-1024x575.webp 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/AdobeStock_842772489-e1787092703768-960x540.webp 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/AdobeStock_842772489-e1787092703768.webp 1843w" sizes="auto, (max-width: 768px) 100vw, 768px" title="" />Agentic AI is changing how research is done. AI scientists can read papers, propose hypotheses, call models, and determine which experiments to prioritize next....<img width="768" height="431" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/AdobeStock_842772489-e1787092703768-768x431.webp" class="webfeedsFeaturedVisual wp-post-image" alt="A picture of a protein molecule." style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/AdobeStock_842772489-e1787092703768-768x431.webp 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/AdobeStock_842772489-e1787092703768-179x100.webp 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/AdobeStock_842772489-e1787092703768-300x168.webp 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/AdobeStock_842772489-e1787092703768-625x351.webp 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/AdobeStock_842772489-e1787092703768-1536x862.webp 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/AdobeStock_842772489-e1787092703768-645x362.webp 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/AdobeStock_842772489-e1787092703768-660x370.webp 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/AdobeStock_842772489-e1787092703768-500x281.webp 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/AdobeStock_842772489-e1787092703768-160x90.webp 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/AdobeStock_842772489-e1787092703768-362x203.webp 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/AdobeStock_842772489-e1787092703768-196x110.webp 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/AdobeStock_842772489-e1787092703768-1024x575.webp 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/AdobeStock_842772489-e1787092703768-960x540.webp 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/AdobeStock_842772489-e1787092703768.webp 1843w" sizes="auto, (max-width: 768px) 100vw, 768px" title="" /><p>Agentic AI is changing how research is done. AI scientists can read papers, propose hypotheses, call models, and determine which experiments to prioritize next. First proving their value in software engineering, coding agents now write, test, and ship production code. Scientific research can be more demanding and iterative. Researchers continually evaluate evidence, refine hypotheses…</p>252<p><a href="https://developer.nvidia.com/blog/run-nvidia-bionemo-nim-microservices-for-protein-structure-prediction-in-claude-science/" rel="nofollow" data-wpel-link="internal" target="_self">Source</a></p>]]></content>253 <link rel="replies" type="text/html" href="https://developer.nvidia.com/blog/run-nvidia-bionemo-nim-microservices-for-protein-structure-prediction-in-claude-science/#comments" thr:count="0"/>254 <link rel="replies" type="application/atom+xml" href="https://developer.nvidia.com/blog/run-nvidia-bionemo-nim-microservices-for-protein-structure-prediction-in-claude-science/feed/" thr:count="0"/>255 <thr:total>0</thr:total>256 </entry>257 <entry>258 <author>259 <name>Michelle Horton</name>260 </author>261 <title type="html"><![CDATA[Scale AV Perception Across Vehicle Platforms with NVIDIA Omniverse NuRec]]></title>262 <link rel="alternate" type="text/html" href="https://developer.nvidia.com/blog/scale-av-perception-across-vehicle-platforms-with-nvidia-omniverse-nurec/" />263 <id>https://developer.nvidia.com/blog/?p=121899</id>264 <updated>2026-08-28T19:07:31Z</updated>265 <published>2026-08-31T16:00:00Z</published>266 <category scheme="https://developer.nvidia.com/blog" term="Developer Tools & Techniques" /><category scheme="https://developer.nvidia.com/blog" term="Robotics" /><category scheme="https://developer.nvidia.com/blog" term="Simulation / Modeling / Design" /><category scheme="https://developer.nvidia.com/blog" term="autonomous vehicles" /><category scheme="https://developer.nvidia.com/blog" term="Physical AI" /> <summary type="html"><![CDATA[<img width="600" height="338" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/NuRec-AV.gif" class="webfeedsFeaturedVisual wp-post-image" alt="Figure showing the original view and the new view after using NuRec to re-render a video." style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" title="NuRec-AV" />A perception stack is shaped by the vehicle that carries it. Move the same software to a new carline—for example, from an SUV to a sedan or another vehicle...]]></summary>267 <content type="html" xml:base="https://developer.nvidia.com/blog/scale-av-perception-across-vehicle-platforms-with-nvidia-omniverse-nurec/"><![CDATA[<img width="600" height="338" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/NuRec-AV.gif" class="webfeedsFeaturedVisual wp-post-image" alt="Figure showing the original view and the new view after using NuRec to re-render a video." style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" title="NuRec-AV" />A perception stack is shaped by the vehicle that carries it. Move the same software to a new carline—for example, from an SUV to a sedan or another vehicle...<img width="600" height="338" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/NuRec-AV.gif" class="webfeedsFeaturedVisual wp-post-image" alt="Figure showing the original view and the new view after using NuRec to re-render a video." style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" title="NuRec-AV" /><p>A perception stack is shaped by the vehicle that carries it. Move the same software to a new carline—for example, from an SUV to a sedan or another vehicle variant in the portfolio—and its perception of the world changes. The sensor placement, calibration, fields of view, occlusions, body geometry, timing, and coverage all shift. A traffic light may appear in a different part of the frame.</p>268<p><a href="https://developer.nvidia.com/blog/scale-av-perception-across-vehicle-platforms-with-nvidia-omniverse-nurec/" rel="nofollow" data-wpel-link="internal" target="_self">Source</a></p>]]></content>269 <link rel="replies" type="text/html" href="https://developer.nvidia.com/blog/scale-av-perception-across-vehicle-platforms-with-nvidia-omniverse-nurec/#comments" thr:count="0"/>270 <link rel="replies" type="application/atom+xml" href="https://developer.nvidia.com/blog/scale-av-perception-across-vehicle-platforms-with-nvidia-omniverse-nurec/feed/" thr:count="0"/>271 <thr:total>0</thr:total>272 </entry>273 <entry>274 <author>275 <name>Tanya Lenz</name>276 </author>277 <title type="html"><![CDATA[Deploy an Open Model from Checkpoint to Inference in Two Commands with NVIDIA TensorRT Model Connect]]></title>278 <link rel="alternate" type="text/html" href="https://developer.nvidia.com/blog/deploy-an-open-model-from-checkpoint-to-inference-in-two-commands-with-nvidia-tensorrt-model-connect/" />279 <id>https://developer.nvidia.com/blog/?p=121956</id>280 <updated>2026-08-28T17:06:37Z</updated>281 <published>2026-08-28T17:06:28Z</published>282 <category scheme="https://developer.nvidia.com/blog" term="Data Center / Cloud" /><category scheme="https://developer.nvidia.com/blog" term="Developer Tools & Techniques" /><category scheme="https://developer.nvidia.com/blog" term="Edge Computing" /><category scheme="https://developer.nvidia.com/blog" term="AI Inference" /><category scheme="https://developer.nvidia.com/blog" term="C++" /><category scheme="https://developer.nvidia.com/blog" term="Open Source" /><category scheme="https://developer.nvidia.com/blog" term="PyTorch" /> <summary type="html"><![CDATA[<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/ai-use-cases-1-768x432.png" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/ai-use-cases-1-768x432.png 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/ai-use-cases-1-179x101.png 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/ai-use-cases-1-300x169.png 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/ai-use-cases-1-625x352.png 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/ai-use-cases-1-645x363.png 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/ai-use-cases-1-660x370.png 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/ai-use-cases-1-500x281.png 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/ai-use-cases-1-160x90.png 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/ai-use-cases-1-362x204.png 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/ai-use-cases-1-196x110.png 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/ai-use-cases-1-1024x576.png 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/ai-use-cases-1-960x540.png 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/ai-use-cases-1.webp 1209w" sizes="auto, (max-width: 768px) 100vw, 768px" title="ai-use-cases" />Open AI models are evolving faster than ever, but bringing them into native applications can still require model-specific conversion, preprocessing,...]]></summary>283 <content type="html" xml:base="https://developer.nvidia.com/blog/deploy-an-open-model-from-checkpoint-to-inference-in-two-commands-with-nvidia-tensorrt-model-connect/"><![CDATA[<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/ai-use-cases-1-768x432.png" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/ai-use-cases-1-768x432.png 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/ai-use-cases-1-179x101.png 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/ai-use-cases-1-300x169.png 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/ai-use-cases-1-625x352.png 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/ai-use-cases-1-645x363.png 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/ai-use-cases-1-660x370.png 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/ai-use-cases-1-500x281.png 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/ai-use-cases-1-160x90.png 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/ai-use-cases-1-362x204.png 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/ai-use-cases-1-196x110.png 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/ai-use-cases-1-1024x576.png 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/ai-use-cases-1-960x540.png 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/ai-use-cases-1.webp 1209w" sizes="auto, (max-width: 768px) 100vw, 768px" title="ai-use-cases" />Open AI models are evolving faster than ever, but bringing them into native applications can still require model-specific conversion, preprocessing,...<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/ai-use-cases-1-768x432.png" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/ai-use-cases-1-768x432.png 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/ai-use-cases-1-179x101.png 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/ai-use-cases-1-300x169.png 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/ai-use-cases-1-625x352.png 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/ai-use-cases-1-645x363.png 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/ai-use-cases-1-660x370.png 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/ai-use-cases-1-500x281.png 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/ai-use-cases-1-160x90.png 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/ai-use-cases-1-362x204.png 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/ai-use-cases-1-196x110.png 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/ai-use-cases-1-1024x576.png 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/ai-use-cases-1-960x540.png 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/ai-use-cases-1.webp 1209w" sizes="auto, (max-width: 768px) 100vw, 768px" title="ai-use-cases" /><p>Open AI models are evolving faster than ever, but bringing them into native applications can still require model-specific conversion, preprocessing, post-processing, and runtime code. NVIDIA TensorRT Model Connect open collection of reference implementations helps to address this challenge. TensorRT Model Connect shows you how to run supported models with NVIDIA TensorRT in native C++…</p>284<p><a href="https://developer.nvidia.com/blog/deploy-an-open-model-from-checkpoint-to-inference-in-two-commands-with-nvidia-tensorrt-model-connect/" rel="nofollow" data-wpel-link="internal" target="_self">Source</a></p>]]></content>285 <link rel="replies" type="text/html" href="https://developer.nvidia.com/blog/deploy-an-open-model-from-checkpoint-to-inference-in-two-commands-with-nvidia-tensorrt-model-connect/#comments" thr:count="0"/>286 <link rel="replies" type="application/atom+xml" href="https://developer.nvidia.com/blog/deploy-an-open-model-from-checkpoint-to-inference-in-two-commands-with-nvidia-tensorrt-model-connect/feed/" thr:count="0"/>287 <thr:total>0</thr:total>288 </entry>289 <entry>290 <author>291 <name>Farshad Ghodsian</name>292 </author>293 <title type="html"><![CDATA[NVIDIA NVLink Fusion Brings NVHBM to Next-Generation AI Infrastructure]]></title>294 <link rel="alternate" type="text/html" href="https://developer.nvidia.com/blog/nvidia-nvlink-fusion-brings-nvhbm-to-next-generation-ai-infrastructure/" />295 <id>https://developer.nvidia.com/blog/?p=120848</id>296 <updated>2026-08-26T21:08:40Z</updated>297 <published>2026-08-26T21:06:58Z</published>298 <category scheme="https://developer.nvidia.com/blog" term="Data Center / Cloud" /><category scheme="https://developer.nvidia.com/blog" term="Networking / Communications" /><category scheme="https://developer.nvidia.com/blog" term="Top Stories" /><category scheme="https://developer.nvidia.com/blog" term="AI Factory" /> <summary type="html"><![CDATA[<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/NVHBM-768x432.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="Decorative image." style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/NVHBM-768x432.jpg 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/NVHBM-179x101.jpg 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/NVHBM-300x169.jpg 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/NVHBM-625x352.jpg 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/NVHBM-1536x864.jpg 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/NVHBM-645x363.jpg 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/NVHBM-660x370.jpg 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/NVHBM-500x281.jpg 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/NVHBM-160x90.jpg 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/NVHBM-362x204.jpg 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/NVHBM-196x110.jpg 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/NVHBM-1024x576.jpg 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/NVHBM-960x540.jpg 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/NVHBM.webp 1920w" sizes="auto, (max-width: 768px) 100vw, 768px" title="NVHBM" />AI factories must support increasingly large models and more complex reasoning workloads. To keep up with the insatiable compute demands of AI workloads,...]]></summary>299 <content type="html" xml:base="https://developer.nvidia.com/blog/nvidia-nvlink-fusion-brings-nvhbm-to-next-generation-ai-infrastructure/"><![CDATA[<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/NVHBM-768x432.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="Decorative image." style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/NVHBM-768x432.jpg 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/NVHBM-179x101.jpg 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/NVHBM-300x169.jpg 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/NVHBM-625x352.jpg 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/NVHBM-1536x864.jpg 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/NVHBM-645x363.jpg 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/NVHBM-660x370.jpg 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/NVHBM-500x281.jpg 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/NVHBM-160x90.jpg 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/NVHBM-362x204.jpg 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/NVHBM-196x110.jpg 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/NVHBM-1024x576.jpg 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/NVHBM-960x540.jpg 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/NVHBM.webp 1920w" sizes="auto, (max-width: 768px) 100vw, 768px" title="NVHBM" />AI factories must support increasingly large models and more complex reasoning workloads. To keep up with the insatiable compute demands of AI workloads,...<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/NVHBM-768x432.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="Decorative image." style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/NVHBM-768x432.jpg 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/NVHBM-179x101.jpg 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/NVHBM-300x169.jpg 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/NVHBM-625x352.jpg 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/NVHBM-1536x864.jpg 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/NVHBM-645x363.jpg 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/NVHBM-660x370.jpg 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/NVHBM-500x281.jpg 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/NVHBM-160x90.jpg 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/NVHBM-362x204.jpg 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/NVHBM-196x110.jpg 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/NVHBM-1024x576.jpg 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/NVHBM-960x540.jpg 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/NVHBM.webp 1920w" sizes="auto, (max-width: 768px) 100vw, 768px" title="NVHBM" /><p>AI factories must support increasingly large models and more complex reasoning workloads. To keep up with the insatiable compute demands of AI workloads, hyperscalers and AI-native companies are developing custom AI accelerators, or XPUs. Deploying these accelerators at scale requires high-bandwidth memory (HBM) to keep compute fed, sufficient package and silicon area for more compute…</p>300<p><a href="https://developer.nvidia.com/blog/nvidia-nvlink-fusion-brings-nvhbm-to-next-generation-ai-infrastructure/" rel="nofollow" data-wpel-link="internal" target="_self">Source</a></p>]]></content>301 <link rel="replies" type="text/html" href="https://developer.nvidia.com/blog/nvidia-nvlink-fusion-brings-nvhbm-to-next-generation-ai-infrastructure/#comments" thr:count="0"/>302 <link rel="replies" type="application/atom+xml" href="https://developer.nvidia.com/blog/nvidia-nvlink-fusion-brings-nvhbm-to-next-generation-ai-infrastructure/feed/" thr:count="0"/>303 <thr:total>0</thr:total>304 </entry>305 <entry>306 <author>307 <name>Tanya Lenz</name>308 </author>309 <title type="html"><![CDATA[How to Train a Cross-Embodiment Robot Navigation Policy with AI Agents]]></title>310 <link rel="alternate" type="text/html" href="https://developer.nvidia.com/blog/how-to-train-a-cross-embodiment-robot-navigation-policy-with-ai-agents/" />311 <id>https://developer.nvidia.com/blog/?p=121922</id>312 <updated>2026-08-26T22:34:18Z</updated>313 <published>2026-08-26T20:05:06Z</published>314 <category scheme="https://developer.nvidia.com/blog" term="Agentic AI / Generative AI" /><category scheme="https://developer.nvidia.com/blog" term="Robotics" /><category scheme="https://developer.nvidia.com/blog" term="Simulation / Modeling / Design" /><category scheme="https://developer.nvidia.com/blog" term="AI Agent" /><category scheme="https://developer.nvidia.com/blog" term="Physical AI" /><category scheme="https://developer.nvidia.com/blog" term="Reinforcement Learning" /> <summary type="html"><![CDATA[<img width="600" height="338" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/robot-quad-composite-1-1.gif" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" title="robot-quad-composite (1)" />Navigation enables a robot to turn perception and motion into purposeful autonomy. Unlike locomotion, which produces stable movement, navigation must be used to...]]></summary>315 <content type="html" xml:base="https://developer.nvidia.com/blog/how-to-train-a-cross-embodiment-robot-navigation-policy-with-ai-agents/"><![CDATA[<img width="600" height="338" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/robot-quad-composite-1-1.gif" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" title="robot-quad-composite (1)" />Navigation enables a robot to turn perception and motion into purposeful autonomy. Unlike locomotion, which produces stable movement, navigation must be used to...<img width="600" height="338" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/robot-quad-composite-1-1.gif" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" title="robot-quad-composite (1)" /><p>Navigation enables a robot to turn perception and motion into purposeful autonomy. Unlike locomotion, which produces stable movement, navigation must be used to continuously localize the robot, interpret changing surroundings, select a route, and avoid obstacles to reach a goal safely. Moving this capability to a new robot or scene can require new data, simulation assets, robot interfaces…</p>316<p><a href="https://developer.nvidia.com/blog/how-to-train-a-cross-embodiment-robot-navigation-policy-with-ai-agents/" rel="nofollow" data-wpel-link="internal" target="_self">Source</a></p>]]></content>317 <link rel="replies" type="text/html" href="https://developer.nvidia.com/blog/how-to-train-a-cross-embodiment-robot-navigation-policy-with-ai-agents/#comments" thr:count="0"/>318 <link rel="replies" type="application/atom+xml" href="https://developer.nvidia.com/blog/how-to-train-a-cross-embodiment-robot-navigation-policy-with-ai-agents/feed/" thr:count="0"/>319 <thr:total>0</thr:total>320 </entry>321 <entry>322 <author>323 <name>Michelle Horton</name>324 </author>325 <title type="html"><![CDATA[Experiment with Qwen3.8-Flash-Next on NVIDIA GB300 NVL72 for Agentic Coding]]></title>326 <link rel="alternate" type="text/html" href="https://developer.nvidia.com/blog/experiment-with-qwen3-8-flash-next-on-nvidia-gb300-nvl72-for-agentic-coding/" />327 <id>https://developer.nvidia.com/blog/?p=121855</id>328 <updated>2026-09-09T20:48:39Z</updated>329 <published>2026-08-26T17:07:12Z</published>330 <category scheme="https://developer.nvidia.com/blog" term="Agentic AI / Generative AI" /><category scheme="https://developer.nvidia.com/blog" term="Developer Tools & Techniques" /><category scheme="https://developer.nvidia.com/blog" term="Top Stories" /><category scheme="https://developer.nvidia.com/blog" term="GB300 NVL72" /><category scheme="https://developer.nvidia.com/blog" term="Mixture of Experts (MoE)" /> <summary type="html"><![CDATA[<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-AI-Qwen-768x432.png" class="webfeedsFeaturedVisual wp-post-image" alt="Decorative image." style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-AI-Qwen-768x432.png 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-AI-Qwen-179x101.png 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-AI-Qwen-300x169.png 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-AI-Qwen-625x352.png 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-AI-Qwen-1536x864.png 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-AI-Qwen-645x363.png 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-AI-Qwen-660x370.png 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-AI-Qwen-500x281.png 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-AI-Qwen-160x90.png 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-AI-Qwen-362x204.png 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-AI-Qwen-196x110.png 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-AI-Qwen-1024x576.png 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-AI-Qwen-960x540.png 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-AI-Qwen.webp 1920w" sizes="auto, (max-width: 768px) 100vw, 768px" title="Agentic-AI-Qwen" />Alibaba released the model weights for Qwen3.8-Flash-Next as a preview of the upcoming Qwen4 architecture for developers to experiment with and evaluate. It’s...]]></summary>331 <content type="html" xml:base="https://developer.nvidia.com/blog/experiment-with-qwen3-8-flash-next-on-nvidia-gb300-nvl72-for-agentic-coding/"><![CDATA[<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-AI-Qwen-768x432.png" class="webfeedsFeaturedVisual wp-post-image" alt="Decorative image." style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-AI-Qwen-768x432.png 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-AI-Qwen-179x101.png 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-AI-Qwen-300x169.png 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-AI-Qwen-625x352.png 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-AI-Qwen-1536x864.png 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-AI-Qwen-645x363.png 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-AI-Qwen-660x370.png 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-AI-Qwen-500x281.png 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-AI-Qwen-160x90.png 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-AI-Qwen-362x204.png 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-AI-Qwen-196x110.png 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-AI-Qwen-1024x576.png 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-AI-Qwen-960x540.png 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-AI-Qwen.webp 1920w" sizes="auto, (max-width: 768px) 100vw, 768px" title="Agentic-AI-Qwen" />Alibaba released the model weights for Qwen3.8-Flash-Next as a preview of the upcoming Qwen4 architecture for developers to experiment with and evaluate. It’s...<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-AI-Qwen-768x432.png" class="webfeedsFeaturedVisual wp-post-image" alt="Decorative image." style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-AI-Qwen-768x432.png 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-AI-Qwen-179x101.png 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-AI-Qwen-300x169.png 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-AI-Qwen-625x352.png 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-AI-Qwen-1536x864.png 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-AI-Qwen-645x363.png 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-AI-Qwen-660x370.png 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-AI-Qwen-500x281.png 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-AI-Qwen-160x90.png 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-AI-Qwen-362x204.png 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-AI-Qwen-196x110.png 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-AI-Qwen-1024x576.png 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-AI-Qwen-960x540.png 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/Agentic-AI-Qwen.webp 1920w" sizes="auto, (max-width: 768px) 100vw, 768px" title="Agentic-AI-Qwen" /><p>Alibaba released the model weights for Qwen3.8-Flash-Next as a preview of the upcoming Qwen4 architecture for developers to experiment with and evaluate. It’s a multimodal mixture-of-experts (MoE) model with a 125B-parameter main model supplemented by an additional 51B N-gram embeddings, with 6B parameters activated per token. It has a native 262,144-token context window, extensible to 1M tokens…</p>332<p><a href="https://developer.nvidia.com/blog/experiment-with-qwen3-8-flash-next-on-nvidia-gb300-nvl72-for-agentic-coding/" rel="nofollow" data-wpel-link="internal" target="_self">Source</a></p>]]></content>333<link href="https://developer.download.nvidia.com/video/devblog/Qwen3.8.mp4" rel="enclosure" length="1597613" type="video/mp4" />334 <link rel="replies" type="text/html" href="https://developer.nvidia.com/blog/experiment-with-qwen3-8-flash-next-on-nvidia-gb300-nvl72-for-agentic-coding/#comments" thr:count="0"/>335 <link rel="replies" type="application/atom+xml" href="https://developer.nvidia.com/blog/experiment-with-qwen3-8-flash-next-on-nvidia-gb300-nvl72-for-agentic-coding/feed/" thr:count="0"/>336 <thr:total>0</thr:total>337 </entry>338 <entry>339 <author>340 <name>Michelle Horton</name>341 </author>342 <title type="html"><![CDATA[Restore LLM Inference Capacity in Seconds with Shadow Engine Recovery in NVIDIA Dynamo]]></title>343 <link rel="alternate" type="text/html" href="https://developer.nvidia.com/blog/restore-llm-inference-capacity-in-seconds-with-shadow-engine-recovery-in-nvidia-dynamo/" />344 <id>https://developer.nvidia.com/blog/?p=121821</id>345 <updated>2026-08-25T22:03:16Z</updated>346 <published>2026-08-25T20:57:54Z</published>347 <category scheme="https://developer.nvidia.com/blog" term="Agentic AI / Generative AI" /><category scheme="https://developer.nvidia.com/blog" term="Data Science" /><category scheme="https://developer.nvidia.com/blog" term="Developer Tools & Techniques" /><category scheme="https://developer.nvidia.com/blog" term="CUDA-X" /><category scheme="https://developer.nvidia.com/blog" term="NCCL" /> <summary type="html"><![CDATA[<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/unnamed-17-768x432.png" class="webfeedsFeaturedVisual wp-post-image" alt="Decorative image." style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/unnamed-17-768x432.png 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/unnamed-17-179x101.png 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/unnamed-17-300x169.png 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/unnamed-17-625x352.png 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/unnamed-17-1536x864.png 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/unnamed-17-645x363.png 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/unnamed-17-660x370.png 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/unnamed-17-500x281.png 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/unnamed-17-160x90.png 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/unnamed-17-362x204.png 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/unnamed-17-196x110.png 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/unnamed-17-1024x576.png 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/unnamed-17-960x540.png 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/unnamed-17.webp 2048w" sizes="auto, (max-width: 768px) 100vw, 768px" title="unnamed-17" />When an LLM engine process fails, the standard recovery path involves a cold restart. This requires loading weights into HBM from storage, compiling kernels,...]]></summary>348 <content type="html" xml:base="https://developer.nvidia.com/blog/restore-llm-inference-capacity-in-seconds-with-shadow-engine-recovery-in-nvidia-dynamo/"><![CDATA[<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/unnamed-17-768x432.png" class="webfeedsFeaturedVisual wp-post-image" alt="Decorative image." style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/unnamed-17-768x432.png 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/unnamed-17-179x101.png 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/unnamed-17-300x169.png 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/unnamed-17-625x352.png 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/unnamed-17-1536x864.png 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/unnamed-17-645x363.png 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/unnamed-17-660x370.png 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/unnamed-17-500x281.png 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/unnamed-17-160x90.png 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/unnamed-17-362x204.png 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/unnamed-17-196x110.png 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/unnamed-17-1024x576.png 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/unnamed-17-960x540.png 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/unnamed-17.webp 2048w" sizes="auto, (max-width: 768px) 100vw, 768px" title="unnamed-17" />When an LLM engine process fails, the standard recovery path involves a cold restart. This requires loading weights into HBM from storage, compiling kernels,...<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/unnamed-17-768x432.png" class="webfeedsFeaturedVisual wp-post-image" alt="Decorative image." style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/unnamed-17-768x432.png 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/unnamed-17-179x101.png 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/unnamed-17-300x169.png 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/unnamed-17-625x352.png 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/unnamed-17-1536x864.png 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/unnamed-17-645x363.png 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/unnamed-17-660x370.png 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/unnamed-17-500x281.png 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/unnamed-17-160x90.png 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/unnamed-17-362x204.png 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/unnamed-17-196x110.png 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/unnamed-17-1024x576.png 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/unnamed-17-960x540.png 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/unnamed-17.webp 2048w" sizes="auto, (max-width: 768px) 100vw, 768px" title="unnamed-17" /><p>When an LLM engine process fails, the standard recovery path involves a cold restart. This requires loading weights into HBM from storage, compiling kernels, and capturing NVIDIA CUDA graphs. For large models, initialization can take several minutes, during which surviving workers must absorb the displaced traffic. Shadow engine recovery, available as a preview feature in NVIDIA Dynamo…</p>349<p><a href="https://developer.nvidia.com/blog/restore-llm-inference-capacity-in-seconds-with-shadow-engine-recovery-in-nvidia-dynamo/" rel="nofollow" data-wpel-link="internal" target="_self">Source</a></p>]]></content>350 <link rel="replies" type="text/html" href="https://developer.nvidia.com/blog/restore-llm-inference-capacity-in-seconds-with-shadow-engine-recovery-in-nvidia-dynamo/#comments" thr:count="0"/>351 <link rel="replies" type="application/atom+xml" href="https://developer.nvidia.com/blog/restore-llm-inference-capacity-in-seconds-with-shadow-engine-recovery-in-nvidia-dynamo/feed/" thr:count="0"/>352 <thr:total>0</thr:total>353 </entry>354 <entry>355 <author>356 <name>Elizabeth Goodman</name>357 </author>358 <title type="html"><![CDATA[CUDA Python 1.0: Stable APIs, One Foundation, Full Platform Access]]></title>359 <link rel="alternate" type="text/html" href="https://developer.nvidia.com/blog/cuda-python-1-0-stable-apis-one-foundation-full-platform-access/" />360 <id>https://developer.nvidia.com/blog/?p=121791</id>361 <updated>2026-08-24T17:52:50Z</updated>362 <published>2026-08-25T15:00:00Z</published>363 <category scheme="https://developer.nvidia.com/blog" term="Data Science" /><category scheme="https://developer.nvidia.com/blog" term="Developer Tools & Techniques" /><category scheme="https://developer.nvidia.com/blog" term="CUDA" /><category scheme="https://developer.nvidia.com/blog" term="Numba" /><category scheme="https://developer.nvidia.com/blog" term="Python" /><category scheme="https://developer.nvidia.com/blog" term="RAPIDS" /> <summary type="html"><![CDATA[<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-5-768x432.png" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-5-768x432.png 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-5-179x101.png 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-5-300x169.png 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-5-625x352.png 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-5-1536x864.png 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-5-645x363.png 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-5-660x370.png 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-5-500x281.png 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-5-160x90.png 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-5-362x204.png 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-5-195x110.png 195w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-5-1024x576.png 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-5-960x540.png 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-5.webp 1999w" sizes="auto, (max-width: 768px) 100vw, 768px" title="image3" />For years, a Python developer who needed a GPU had two realistic choices: Learn NVIDIA CUDA C++ well enough to write an extension, set up a build toolchain, and...]]></summary>364 <content type="html" xml:base="https://developer.nvidia.com/blog/cuda-python-1-0-stable-apis-one-foundation-full-platform-access/"><![CDATA[<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-5-768x432.png" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-5-768x432.png 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-5-179x101.png 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-5-300x169.png 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-5-625x352.png 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-5-1536x864.png 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-5-645x363.png 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-5-660x370.png 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-5-500x281.png 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-5-160x90.png 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-5-362x204.png 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-5-195x110.png 195w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-5-1024x576.png 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-5-960x540.png 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-5.webp 1999w" sizes="auto, (max-width: 768px) 100vw, 768px" title="image3" />For years, a Python developer who needed a GPU had two realistic choices: Learn NVIDIA CUDA C++ well enough to write an extension, set up a build toolchain, and...<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-5-768x432.png" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-5-768x432.png 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-5-179x101.png 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-5-300x169.png 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-5-625x352.png 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-5-1536x864.png 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-5-645x363.png 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-5-660x370.png 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-5-500x281.png 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-5-160x90.png 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-5-362x204.png 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-5-195x110.png 195w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-5-1024x576.png 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-5-960x540.png 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-5.webp 1999w" sizes="auto, (max-width: 768px) 100vw, 768px" title="image3" /><p>For years, a Python developer who needed a GPU had two realistic choices: Learn NVIDIA CUDA C++ well enough to write an extension, set up a build toolchain, and maintain bindings back to Python, which most people never did; or move up the stack and let someone else’s library do it, namely PyTorch, CuPy, or RAPIDS. The second option is why the Python GPU ecosystem thrives. But it has limits.</p>365<p><a href="https://developer.nvidia.com/blog/cuda-python-1-0-stable-apis-one-foundation-full-platform-access/" rel="nofollow" data-wpel-link="internal" target="_self">Source</a></p>]]></content>366 <link rel="replies" type="text/html" href="https://developer.nvidia.com/blog/cuda-python-1-0-stable-apis-one-foundation-full-platform-access/#comments" thr:count="0"/>367 <link rel="replies" type="application/atom+xml" href="https://developer.nvidia.com/blog/cuda-python-1-0-stable-apis-one-foundation-full-platform-access/feed/" thr:count="0"/>368 <thr:total>0</thr:total>369 </entry>370 <entry>371 <author>372 <name>Elizabeth Goodman</name>373 </author>374 <title type="html"><![CDATA[Giga-Scale AI and the Ethernet Evolution: How Spectrum-X Ethernet Rewrites the Rules]]></title>375 <link rel="alternate" type="text/html" href="https://developer.nvidia.com/blog/giga-scale-ai-ethernet-evolution-spectrum-x-ethernet-rewrites-rules/" />376 <id>https://developer.nvidia.com/blog/?p=121595</id>377 <updated>2026-08-27T17:48:53Z</updated>378 <published>2026-08-24T15:08:39Z</published>379 <category scheme="https://developer.nvidia.com/blog" term="Data Center / Cloud" /><category scheme="https://developer.nvidia.com/blog" term="Networking / Communications" /><category scheme="https://developer.nvidia.com/blog" term="AI Agent" /><category scheme="https://developer.nvidia.com/blog" term="AI Factory" /><category scheme="https://developer.nvidia.com/blog" term="AI Networking" /><category scheme="https://developer.nvidia.com/blog" term="Infrastructure" /><category scheme="https://developer.nvidia.com/blog" term="Internet/Communications" /><category scheme="https://developer.nvidia.com/blog" term="Spectrum-X" /> <summary type="html"><![CDATA[<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/featured_image-768x432.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/featured_image-768x432.jpg 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/featured_image-179x101.jpg 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/featured_image-300x169.jpg 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/featured_image-625x352.jpg 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/featured_image-1536x864.jpg 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/featured_image-645x363.jpg 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/featured_image-660x370.jpg 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/featured_image-500x281.jpg 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/featured_image-160x90.jpg 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/featured_image-362x204.jpg 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/featured_image-196x110.jpg 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/featured_image-1024x576.jpg 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/featured_image-960x540.jpg 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/featured_image.webp 1920w" sizes="auto, (max-width: 768px) 100vw, 768px" title="featured_image" />The massive growth of generative AI has fundamentally altered data center design. As distributed model training scales to span hundreds of thousands of GPUs,...]]></summary>380 <content type="html" xml:base="https://developer.nvidia.com/blog/giga-scale-ai-ethernet-evolution-spectrum-x-ethernet-rewrites-rules/"><![CDATA[<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/featured_image-768x432.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/featured_image-768x432.jpg 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/featured_image-179x101.jpg 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/featured_image-300x169.jpg 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/featured_image-625x352.jpg 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/featured_image-1536x864.jpg 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/featured_image-645x363.jpg 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/featured_image-660x370.jpg 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/featured_image-500x281.jpg 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/featured_image-160x90.jpg 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/featured_image-362x204.jpg 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/featured_image-196x110.jpg 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/featured_image-1024x576.jpg 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/featured_image-960x540.jpg 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/featured_image.webp 1920w" sizes="auto, (max-width: 768px) 100vw, 768px" title="featured_image" />The massive growth of generative AI has fundamentally altered data center design. As distributed model training scales to span hundreds of thousands of GPUs,...<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/featured_image-768x432.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/featured_image-768x432.jpg 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/featured_image-179x101.jpg 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/featured_image-300x169.jpg 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/featured_image-625x352.jpg 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/featured_image-1536x864.jpg 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/featured_image-645x363.jpg 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/featured_image-660x370.jpg 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/featured_image-500x281.jpg 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/featured_image-160x90.jpg 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/featured_image-362x204.jpg 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/featured_image-196x110.jpg 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/featured_image-1024x576.jpg 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/featured_image-960x540.jpg 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/featured_image.webp 1920w" sizes="auto, (max-width: 768px) 100vw, 768px" title="featured_image" /><p>The massive growth of generative AI has fundamentally altered data center design. As distributed model training scales to span hundreds of thousands of GPUs, the scale-out network connecting these nodes has emerged as a first-order performance bottleneck. For decades, traditional off-the-shelf Ethernet has been the undisputed king of enterprise and cloud networking. It is cheap, standardized…</p>381<p><a href="https://developer.nvidia.com/blog/giga-scale-ai-ethernet-evolution-spectrum-x-ethernet-rewrites-rules/" rel="nofollow" data-wpel-link="internal" target="_self">Source</a></p>]]></content>382 <link rel="replies" type="text/html" href="https://developer.nvidia.com/blog/giga-scale-ai-ethernet-evolution-spectrum-x-ethernet-rewrites-rules/#comments" thr:count="0"/>383 <link rel="replies" type="application/atom+xml" href="https://developer.nvidia.com/blog/giga-scale-ai-ethernet-evolution-spectrum-x-ethernet-rewrites-rules/feed/" thr:count="0"/>384 <thr:total>0</thr:total>385 </entry>386 <entry>387 <author>388 <name>Elizabeth Goodman</name>389 </author>390 <title type="html"><![CDATA[NVIDIA Vera Rubin and Blackwell Set a New Standard for Agentic AI Performance per Watt ]]></title>391 <link rel="alternate" type="text/html" href="https://developer.nvidia.com/blog/nvidia-vera-rubin-and-blackwell-set-a-new-standard-for-agentic-ai-performance-per-watt/" />392 <id>https://developer.nvidia.com/blog/?p=121715</id>393 <updated>2026-08-24T15:23:55Z</updated>394 <published>2026-08-24T15:00:05Z</published>395 <category scheme="https://developer.nvidia.com/blog" term="Agentic AI / Generative AI" /><category scheme="https://developer.nvidia.com/blog" term="Data Center / Cloud" /><category scheme="https://developer.nvidia.com/blog" term="AI Agent" /><category scheme="https://developer.nvidia.com/blog" term="Blackwell" /><category scheme="https://developer.nvidia.com/blog" term="Cloud Networking" /><category scheme="https://developer.nvidia.com/blog" term="Cloud Services" /><category scheme="https://developer.nvidia.com/blog" term="Nemotron" /><category scheme="https://developer.nvidia.com/blog" term="Software-Defined Data Center" /> <summary type="html"><![CDATA[<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-4-768x432.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-4-768x432.jpg 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-4-179x101.jpg 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-4-300x169.jpg 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-4-625x352.jpg 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-4-1536x864.jpg 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-4-645x363.jpg 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-4-660x370.jpg 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-4-500x281.jpg 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-4-160x90.jpg 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-4-362x204.jpg 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-4-196x110.jpg 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-4-1024x576.jpg 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-4-960x540.jpg 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-4.webp 1920w" sizes="auto, (max-width: 768px) 100vw, 768px" title="image3" />AI agents have expanded inference from single-turn interactions into multi-step workflows that reason, invoke tools, coordinate subagents, and carry growing...]]></summary>396 <content type="html" xml:base="https://developer.nvidia.com/blog/nvidia-vera-rubin-and-blackwell-set-a-new-standard-for-agentic-ai-performance-per-watt/"><![CDATA[<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-4-768x432.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-4-768x432.jpg 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-4-179x101.jpg 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-4-300x169.jpg 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-4-625x352.jpg 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-4-1536x864.jpg 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-4-645x363.jpg 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-4-660x370.jpg 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-4-500x281.jpg 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-4-160x90.jpg 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-4-362x204.jpg 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-4-196x110.jpg 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-4-1024x576.jpg 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-4-960x540.jpg 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-4.webp 1920w" sizes="auto, (max-width: 768px) 100vw, 768px" title="image3" />AI agents have expanded inference from single-turn interactions into multi-step workflows that reason, invoke tools, coordinate subagents, and carry growing...<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-4-768x432.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-4-768x432.jpg 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-4-179x101.jpg 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-4-300x169.jpg 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-4-625x352.jpg 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-4-1536x864.jpg 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-4-645x363.jpg 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-4-660x370.jpg 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-4-500x281.jpg 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-4-160x90.jpg 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-4-362x204.jpg 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-4-196x110.jpg 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-4-1024x576.jpg 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-4-960x540.jpg 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image3-4.webp 1920w" sizes="auto, (max-width: 768px) 100vw, 768px" title="image3" /><p>AI agents have expanded inference from single-turn interactions into multi-step workflows that reason, invoke tools, coordinate subagents, and carry growing context from one turn to the next. The scale of this shift is now visible in raw consumption: across 100 trillion tokens of real-world usage, OpenRouter’s State of AI report found that average prompt tokens per request grew roughly fourfold…</p>397<p><a href="https://developer.nvidia.com/blog/nvidia-vera-rubin-and-blackwell-set-a-new-standard-for-agentic-ai-performance-per-watt/" rel="nofollow" data-wpel-link="internal" target="_self">Source</a></p>]]></content>398 <link rel="replies" type="text/html" href="https://developer.nvidia.com/blog/nvidia-vera-rubin-and-blackwell-set-a-new-standard-for-agentic-ai-performance-per-watt/#comments" thr:count="0"/>399 <link rel="replies" type="application/atom+xml" href="https://developer.nvidia.com/blog/nvidia-vera-rubin-and-blackwell-set-a-new-standard-for-agentic-ai-performance-per-watt/feed/" thr:count="0"/>400 <thr:total>0</thr:total>401 </entry>402 <entry>403 <author>404 <name>Michelle Horton</name>405 </author>406 <title type="html"><![CDATA[NVIDIA BlueField-4 Powers New Scale-In Network Infrastructure for Agentic AI Factories]]></title>407 <link rel="alternate" type="text/html" href="https://developer.nvidia.com/blog/nvidia-bluefield-4-powers-new-scale-in-network-infrastructure-for-agentic-ai-factories/" />408 <id>https://developer.nvidia.com/blog/?p=121527</id>409 <updated>2026-08-24T15:01:04Z</updated>410 <published>2026-08-24T15:00:00Z</published>411 <category scheme="https://developer.nvidia.com/blog" term="Data Center / Cloud" /><category scheme="https://developer.nvidia.com/blog" term="Networking / Communications" /><category scheme="https://developer.nvidia.com/blog" term="Top Stories" /><category scheme="https://developer.nvidia.com/blog" term="AI Factory" /><category scheme="https://developer.nvidia.com/blog" term="BlueField DPU" /><category scheme="https://developer.nvidia.com/blog" term="ConnectX" /><category scheme="https://developer.nvidia.com/blog" term="DSX" /><category scheme="https://developer.nvidia.com/blog" term="Grace CPU" /><category scheme="https://developer.nvidia.com/blog" term="Vera CPU" /><category scheme="https://developer.nvidia.com/blog" term="Vera Rubin" /> <summary type="html"><![CDATA[<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/BlueField-4-Scale-In-768x432.png" class="webfeedsFeaturedVisual wp-post-image" alt="BlueField-4 render." style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/BlueField-4-Scale-In-768x432.png 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/BlueField-4-Scale-In-179x101.png 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/BlueField-4-Scale-In-300x169.png 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/BlueField-4-Scale-In-625x352.png 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/BlueField-4-Scale-In-1536x864.png 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/BlueField-4-Scale-In-645x363.png 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/BlueField-4-Scale-In-660x370.png 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/BlueField-4-Scale-In-500x281.png 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/BlueField-4-Scale-In-160x90.png 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/BlueField-4-Scale-In-362x204.png 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/BlueField-4-Scale-In-196x110.png 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/BlueField-4-Scale-In-1024x576.png 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/BlueField-4-Scale-In-960x540.png 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/BlueField-4-Scale-In.webp 1920w" sizes="auto, (max-width: 768px) 100vw, 768px" title="BlueField-4-Scale-In" />Traditional cloud infrastructure was designed for predictable, general-purpose workloads and standard interfaces. Agentic AI factories connect diverse users,...]]></summary>412 <content type="html" xml:base="https://developer.nvidia.com/blog/nvidia-bluefield-4-powers-new-scale-in-network-infrastructure-for-agentic-ai-factories/"><![CDATA[<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/BlueField-4-Scale-In-768x432.png" class="webfeedsFeaturedVisual wp-post-image" alt="BlueField-4 render." style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/BlueField-4-Scale-In-768x432.png 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/BlueField-4-Scale-In-179x101.png 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/BlueField-4-Scale-In-300x169.png 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/BlueField-4-Scale-In-625x352.png 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/BlueField-4-Scale-In-1536x864.png 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/BlueField-4-Scale-In-645x363.png 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/BlueField-4-Scale-In-660x370.png 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/BlueField-4-Scale-In-500x281.png 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/BlueField-4-Scale-In-160x90.png 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/BlueField-4-Scale-In-362x204.png 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/BlueField-4-Scale-In-196x110.png 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/BlueField-4-Scale-In-1024x576.png 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/BlueField-4-Scale-In-960x540.png 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/BlueField-4-Scale-In.webp 1920w" sizes="auto, (max-width: 768px) 100vw, 768px" title="BlueField-4-Scale-In" />Traditional cloud infrastructure was designed for predictable, general-purpose workloads and standard interfaces. Agentic AI factories connect diverse users,...<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/BlueField-4-Scale-In-768x432.png" class="webfeedsFeaturedVisual wp-post-image" alt="BlueField-4 render." style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/BlueField-4-Scale-In-768x432.png 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/BlueField-4-Scale-In-179x101.png 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/BlueField-4-Scale-In-300x169.png 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/BlueField-4-Scale-In-625x352.png 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/BlueField-4-Scale-In-1536x864.png 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/BlueField-4-Scale-In-645x363.png 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/BlueField-4-Scale-In-660x370.png 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/BlueField-4-Scale-In-500x281.png 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/BlueField-4-Scale-In-160x90.png 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/BlueField-4-Scale-In-362x204.png 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/BlueField-4-Scale-In-196x110.png 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/BlueField-4-Scale-In-1024x576.png 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/BlueField-4-Scale-In-960x540.png 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/BlueField-4-Scale-In.webp 1920w" sizes="auto, (max-width: 768px) 100vw, 768px" title="BlueField-4-Scale-In" /><p>Traditional cloud infrastructure was designed for predictable, general-purpose workloads and standard interfaces. Agentic AI factories connect diverse users, agents, applications, data sources, and storage systems to massively accelerated compute at multi-terabit bandwidth per server, making dedicated DPU processing essential for line-rate networking, storage, and security.</p>413<p><a href="https://developer.nvidia.com/blog/nvidia-bluefield-4-powers-new-scale-in-network-infrastructure-for-agentic-ai-factories/" rel="nofollow" data-wpel-link="internal" target="_self">Source</a></p>]]></content>414 <link rel="replies" type="text/html" href="https://developer.nvidia.com/blog/nvidia-bluefield-4-powers-new-scale-in-network-infrastructure-for-agentic-ai-factories/#comments" thr:count="0"/>415 <link rel="replies" type="application/atom+xml" href="https://developer.nvidia.com/blog/nvidia-bluefield-4-powers-new-scale-in-network-infrastructure-for-agentic-ai-factories/feed/" thr:count="0"/>416 <thr:total>0</thr:total>417 </entry>418 <entry>419 <author>420 <name>Michelle Horton</name>421 </author>422 <title type="html"><![CDATA[Solving Agentic AI Fleet Challenges with NVIDIA Vera CPU]]></title>423 <link rel="alternate" type="text/html" href="https://developer.nvidia.com/blog/solving-agentic-ai-fleet-challenges-with-nvidia-vera-cpu/" />424 <id>https://developer.nvidia.com/blog/?p=121644</id>425 <updated>2026-08-21T23:09:15Z</updated>426 <published>2026-08-24T15:00:00Z</published>427 <category scheme="https://developer.nvidia.com/blog" term="Agentic AI / Generative AI" /><category scheme="https://developer.nvidia.com/blog" term="Data Center / Cloud" /><category scheme="https://developer.nvidia.com/blog" term="Top Stories" /><category scheme="https://developer.nvidia.com/blog" term="AI Agent" /><category scheme="https://developer.nvidia.com/blog" term="AI Factory" /><category scheme="https://developer.nvidia.com/blog" term="Vera CPU" /><category scheme="https://developer.nvidia.com/blog" term="Vera Rubin" /> <summary type="html"><![CDATA[<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/nvidia-vera-e1787349682277-768x432.webp" class="webfeedsFeaturedVisual wp-post-image" alt="Vera CPU render." style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/nvidia-vera-e1787349682277-768x432.webp 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/nvidia-vera-e1787349682277-179x101.webp 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/nvidia-vera-e1787349682277-300x169.webp 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/nvidia-vera-e1787349682277-625x351.webp 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/nvidia-vera-e1787349682277-645x363.webp 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/nvidia-vera-e1787349682277-658x370.webp 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/nvidia-vera-e1787349682277-500x281.webp 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/nvidia-vera-e1787349682277-160x90.webp 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/nvidia-vera-e1787349682277-362x204.webp 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/nvidia-vera-e1787349682277-196x110.webp 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/nvidia-vera-e1787349682277-1024x576.webp 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/nvidia-vera-e1787349682277-960x540.webp 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/nvidia-vera-e1787349682277.webp 1195w" sizes="auto, (max-width: 768px) 100vw, 768px" title="nvidia-vera" />AI factories are interconnected systems where fleet economics depend on how efficiently the entire stack converts power and capital into completed agent tasks....]]></summary>428 <content type="html" xml:base="https://developer.nvidia.com/blog/solving-agentic-ai-fleet-challenges-with-nvidia-vera-cpu/"><![CDATA[<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/nvidia-vera-e1787349682277-768x432.webp" class="webfeedsFeaturedVisual wp-post-image" alt="Vera CPU render." style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/nvidia-vera-e1787349682277-768x432.webp 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/nvidia-vera-e1787349682277-179x101.webp 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/nvidia-vera-e1787349682277-300x169.webp 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/nvidia-vera-e1787349682277-625x351.webp 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/nvidia-vera-e1787349682277-645x363.webp 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/nvidia-vera-e1787349682277-658x370.webp 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/nvidia-vera-e1787349682277-500x281.webp 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/nvidia-vera-e1787349682277-160x90.webp 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/nvidia-vera-e1787349682277-362x204.webp 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/nvidia-vera-e1787349682277-196x110.webp 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/nvidia-vera-e1787349682277-1024x576.webp 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/nvidia-vera-e1787349682277-960x540.webp 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/nvidia-vera-e1787349682277.webp 1195w" sizes="auto, (max-width: 768px) 100vw, 768px" title="nvidia-vera" />AI factories are interconnected systems where fleet economics depend on how efficiently the entire stack converts power and capital into completed agent tasks....<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/nvidia-vera-e1787349682277-768x432.webp" class="webfeedsFeaturedVisual wp-post-image" alt="Vera CPU render." style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/nvidia-vera-e1787349682277-768x432.webp 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/nvidia-vera-e1787349682277-179x101.webp 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/nvidia-vera-e1787349682277-300x169.webp 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/nvidia-vera-e1787349682277-625x351.webp 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/nvidia-vera-e1787349682277-645x363.webp 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/nvidia-vera-e1787349682277-658x370.webp 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/nvidia-vera-e1787349682277-500x281.webp 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/nvidia-vera-e1787349682277-160x90.webp 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/nvidia-vera-e1787349682277-362x204.webp 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/nvidia-vera-e1787349682277-196x110.webp 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/nvidia-vera-e1787349682277-1024x576.webp 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/nvidia-vera-e1787349682277-960x540.webp 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/nvidia-vera-e1787349682277.webp 1195w" sizes="auto, (max-width: 768px) 100vw, 768px" title="nvidia-vera" /><p>AI factories are interconnected systems where fleet economics depend on how efficiently the entire stack converts power and capital into completed agent tasks. While GPUs run the models, CPUs handle orchestration, tool execution, and sandboxed computation. Unlike conventional computing with stable runtime profiles, agentic workloads are unpredictable and highly variable. Based on telemetry from…</p>429<p><a href="https://developer.nvidia.com/blog/solving-agentic-ai-fleet-challenges-with-nvidia-vera-cpu/" rel="nofollow" data-wpel-link="internal" target="_self">Source</a></p>]]></content>430 <link rel="replies" type="text/html" href="https://developer.nvidia.com/blog/solving-agentic-ai-fleet-challenges-with-nvidia-vera-cpu/#comments" thr:count="0"/>431 <link rel="replies" type="application/atom+xml" href="https://developer.nvidia.com/blog/solving-agentic-ai-fleet-challenges-with-nvidia-vera-cpu/feed/" thr:count="0"/>432 <thr:total>0</thr:total>433 </entry>434 <entry>435 <author>436 <name>Tanya Lenz</name>437 </author>438 <title type="html"><![CDATA[How NVIDIA Groq 3 LPX Unlocks Ultrafast Interactivity at Long Context on NVIDIA Vera Rubin]]></title>439 <link rel="alternate" type="text/html" href="https://developer.nvidia.com/blog/how-nvidia-groq-3-lpx-unlocks-ultrafast-interactivity-at-long-context-on-nvidia-vera-rubin/" />440 <id>https://developer.nvidia.com/blog/?p=121675</id>441 <updated>2026-08-31T22:07:34Z</updated>442 <published>2026-08-24T15:00:00Z</published>443 <category scheme="https://developer.nvidia.com/blog" term="Agentic AI / Generative AI" /><category scheme="https://developer.nvidia.com/blog" term="Data Center / Cloud" /><category scheme="https://developer.nvidia.com/blog" term="Developer Tools & Techniques" /><category scheme="https://developer.nvidia.com/blog" term="AI Factory" /><category scheme="https://developer.nvidia.com/blog" term="AI Inference" /><category scheme="https://developer.nvidia.com/blog" term="Groq 3 LPX" /><category scheme="https://developer.nvidia.com/blog" term="Inference Performance" /><category scheme="https://developer.nvidia.com/blog" term="Low-Latency Inference" /><category scheme="https://developer.nvidia.com/blog" term="Rubin GPU" /><category scheme="https://developer.nvidia.com/blog" term="Training AI Models" /><category scheme="https://developer.nvidia.com/blog" term="Vera Rubin" /><category scheme="https://developer.nvidia.com/blog" term="Vera Rubin NVL72" /> <summary type="html"><![CDATA[<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/gpu-architecture-groq3-lpx-rack-768x432.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/gpu-architecture-groq3-lpx-rack-768x432.jpg 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/gpu-architecture-groq3-lpx-rack-179x101.jpg 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/gpu-architecture-groq3-lpx-rack-300x169.jpg 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/gpu-architecture-groq3-lpx-rack-625x352.jpg 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/gpu-architecture-groq3-lpx-rack-1536x864.jpg 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/gpu-architecture-groq3-lpx-rack-645x363.jpg 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/gpu-architecture-groq3-lpx-rack-660x370.jpg 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/gpu-architecture-groq3-lpx-rack-500x281.jpg 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/gpu-architecture-groq3-lpx-rack-160x90.jpg 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/gpu-architecture-groq3-lpx-rack-362x204.jpg 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/gpu-architecture-groq3-lpx-rack-196x110.jpg 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/gpu-architecture-groq3-lpx-rack-1024x576.jpg 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/gpu-architecture-groq3-lpx-rack-960x540.jpg 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/gpu-architecture-groq3-lpx-rack.webp 1920w" sizes="auto, (max-width: 768px) 100vw, 768px" title="gpu-architecture-groq3-lpx-rack" />NVIDIA Groq 3 LPX is the interactive AI inference accelerator for the NVIDIA Vera Rubin platform. At the core of the platform is NVIDIA Vera Rubin NVL72, the...]]></summary>444 <content type="html" xml:base="https://developer.nvidia.com/blog/how-nvidia-groq-3-lpx-unlocks-ultrafast-interactivity-at-long-context-on-nvidia-vera-rubin/"><![CDATA[<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/gpu-architecture-groq3-lpx-rack-768x432.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/gpu-architecture-groq3-lpx-rack-768x432.jpg 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/gpu-architecture-groq3-lpx-rack-179x101.jpg 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/gpu-architecture-groq3-lpx-rack-300x169.jpg 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/gpu-architecture-groq3-lpx-rack-625x352.jpg 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/gpu-architecture-groq3-lpx-rack-1536x864.jpg 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/gpu-architecture-groq3-lpx-rack-645x363.jpg 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/gpu-architecture-groq3-lpx-rack-660x370.jpg 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/gpu-architecture-groq3-lpx-rack-500x281.jpg 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/gpu-architecture-groq3-lpx-rack-160x90.jpg 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/gpu-architecture-groq3-lpx-rack-362x204.jpg 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/gpu-architecture-groq3-lpx-rack-196x110.jpg 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/gpu-architecture-groq3-lpx-rack-1024x576.jpg 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/gpu-architecture-groq3-lpx-rack-960x540.jpg 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/gpu-architecture-groq3-lpx-rack.webp 1920w" sizes="auto, (max-width: 768px) 100vw, 768px" title="gpu-architecture-groq3-lpx-rack" />NVIDIA Groq 3 LPX is the interactive AI inference accelerator for the NVIDIA Vera Rubin platform. At the core of the platform is NVIDIA Vera Rubin NVL72, the...<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/gpu-architecture-groq3-lpx-rack-768x432.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/gpu-architecture-groq3-lpx-rack-768x432.jpg 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/gpu-architecture-groq3-lpx-rack-179x101.jpg 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/gpu-architecture-groq3-lpx-rack-300x169.jpg 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/gpu-architecture-groq3-lpx-rack-625x352.jpg 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/gpu-architecture-groq3-lpx-rack-1536x864.jpg 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/gpu-architecture-groq3-lpx-rack-645x363.jpg 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/gpu-architecture-groq3-lpx-rack-660x370.jpg 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/gpu-architecture-groq3-lpx-rack-500x281.jpg 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/gpu-architecture-groq3-lpx-rack-160x90.jpg 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/gpu-architecture-groq3-lpx-rack-362x204.jpg 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/gpu-architecture-groq3-lpx-rack-196x110.jpg 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/gpu-architecture-groq3-lpx-rack-1024x576.jpg 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/gpu-architecture-groq3-lpx-rack-960x540.jpg 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/gpu-architecture-groq3-lpx-rack.webp 1920w" sizes="auto, (max-width: 768px) 100vw, 768px" title="gpu-architecture-groq3-lpx-rack" /><p>NVIDIA Groq 3 LPX is the interactive AI inference accelerator for the NVIDIA Vera Rubin platform. At the core of the platform is NVIDIA Vera Rubin NVL72, the most versatile machine ever built, delivering high throughput and interactivity across the widest range of AI workloads—from small to large models, both open and closed. Groq 3 LPX, when paired with Vera Rubin NVL72, extends the platform’s…</p>445<p><a href="https://developer.nvidia.com/blog/how-nvidia-groq-3-lpx-unlocks-ultrafast-interactivity-at-long-context-on-nvidia-vera-rubin/" rel="nofollow" data-wpel-link="internal" target="_self">Source</a></p>]]></content>446 <link rel="replies" type="text/html" href="https://developer.nvidia.com/blog/how-nvidia-groq-3-lpx-unlocks-ultrafast-interactivity-at-long-context-on-nvidia-vera-rubin/#comments" thr:count="0"/>447 <link rel="replies" type="application/atom+xml" href="https://developer.nvidia.com/blog/how-nvidia-groq-3-lpx-unlocks-ultrafast-interactivity-at-long-context-on-nvidia-vera-rubin/feed/" thr:count="0"/>448 <thr:total>0</thr:total>449 </entry>450 <entry>451 <author>452 <name>Tanya Lenz</name>453 </author>454 <title type="html"><![CDATA[Maximizing AI Factory Performance per Watt with NVIDIA DSX MaxLPS]]></title>455 <link rel="alternate" type="text/html" href="https://developer.nvidia.com/blog/maximizing-ai-factory-performance-per-watt-with-nvidia-dsx-maxlps/" />456 <id>https://developer.nvidia.com/blog/?p=121508</id>457 <updated>2026-08-31T20:56:08Z</updated>458 <published>2026-08-24T15:00:00Z</published>459 <category scheme="https://developer.nvidia.com/blog" term="Data Center / Cloud" /><category scheme="https://developer.nvidia.com/blog" term="AI Factory" /><category scheme="https://developer.nvidia.com/blog" term="AI Inference" /><category scheme="https://developer.nvidia.com/blog" term="Vera Rubin NVL72" /> <summary type="html"><![CDATA[<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/03/data-center-768x432.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/03/data-center-768x432.jpg 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/03/data-center-179x101.jpg 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/03/data-center-300x169.jpg 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/03/data-center-625x352.jpg 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/03/data-center-1536x864.jpg 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/03/data-center-2048x1152.jpg 2048w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/03/data-center-645x363.jpg 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/03/data-center-660x370.jpg 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/03/data-center-500x281.jpg 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/03/data-center-160x90.jpg 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/03/data-center-362x204.jpg 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/03/data-center-196x110.jpg 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/03/data-center-1024x576.jpg 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/03/data-center-960x540.jpg 960w" sizes="auto, (max-width: 768px) 100vw, 768px" title="data-center" />AI factories are power-constrained industrial systems. The question is no longer how many GPUs fit in a data center, but how much AI output each available...]]></summary>460 <content type="html" xml:base="https://developer.nvidia.com/blog/maximizing-ai-factory-performance-per-watt-with-nvidia-dsx-maxlps/"><![CDATA[<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/03/data-center-768x432.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/03/data-center-768x432.jpg 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/03/data-center-179x101.jpg 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/03/data-center-300x169.jpg 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/03/data-center-625x352.jpg 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/03/data-center-1536x864.jpg 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/03/data-center-2048x1152.jpg 2048w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/03/data-center-645x363.jpg 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/03/data-center-660x370.jpg 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/03/data-center-500x281.jpg 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/03/data-center-160x90.jpg 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/03/data-center-362x204.jpg 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/03/data-center-196x110.jpg 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/03/data-center-1024x576.jpg 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/03/data-center-960x540.jpg 960w" sizes="auto, (max-width: 768px) 100vw, 768px" title="data-center" />AI factories are power-constrained industrial systems. The question is no longer how many GPUs fit in a data center, but how much AI output each available...<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/03/data-center-768x432.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/03/data-center-768x432.jpg 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/03/data-center-179x101.jpg 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/03/data-center-300x169.jpg 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/03/data-center-625x352.jpg 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/03/data-center-1536x864.jpg 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/03/data-center-2048x1152.jpg 2048w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/03/data-center-645x363.jpg 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/03/data-center-660x370.jpg 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/03/data-center-500x281.jpg 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/03/data-center-160x90.jpg 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/03/data-center-362x204.jpg 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/03/data-center-196x110.jpg 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/03/data-center-1024x576.jpg 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/03/data-center-960x540.jpg 960w" sizes="auto, (max-width: 768px) 100vw, 768px" title="data-center" /><p>AI factories are power-constrained industrial systems. The question is no longer how many GPUs fit in a data center, but how much AI output each available megawatt can deliver. For AI inference workloads, this makes application-level performance per watt the key metric for measuring AI factory efficiency. Not every megawatt translates to revenue-generating compute. Power distribution, cooling…</p>461<p><a href="https://developer.nvidia.com/blog/maximizing-ai-factory-performance-per-watt-with-nvidia-dsx-maxlps/" rel="nofollow" data-wpel-link="internal" target="_self">Source</a></p>]]></content>462 <link rel="replies" type="text/html" href="https://developer.nvidia.com/blog/maximizing-ai-factory-performance-per-watt-with-nvidia-dsx-maxlps/#comments" thr:count="0"/>463 <link rel="replies" type="application/atom+xml" href="https://developer.nvidia.com/blog/maximizing-ai-factory-performance-per-watt-with-nvidia-dsx-maxlps/feed/" thr:count="0"/>464 <thr:total>0</thr:total>465 </entry>466 <entry>467 <author>468 <name>Elizabeth Goodman</name>469 </author>470 <title type="html"><![CDATA[GPU-Accelerated Clustering for Financial Instruments at Scale]]></title>471 <link rel="alternate" type="text/html" href="https://developer.nvidia.com/blog/gpu-accelerated-clustering-for-financial-instruments-at-scale/" />472 <id>https://developer.nvidia.com/blog/?p=121550</id>473 <updated>2026-08-21T16:21:21Z</updated>474 <published>2026-08-21T16:21:04Z</published>475 <category scheme="https://developer.nvidia.com/blog" term="Data Science" /><category scheme="https://developer.nvidia.com/blog" term="Simulation / Modeling / Design" /><category scheme="https://developer.nvidia.com/blog" term="Data Analytics / Processing" /><category scheme="https://developer.nvidia.com/blog" term="Financial Services" /><category scheme="https://developer.nvidia.com/blog" term="NeMo" /><category scheme="https://developer.nvidia.com/blog" term="Nemotron" /> <summary type="html"><![CDATA[<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image10_1920x1080-768x432.jpeg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image10_1920x1080-768x432.jpeg 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image10_1920x1080-179x101.jpeg 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image10_1920x1080-300x169.jpeg 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image10_1920x1080-625x352.jpeg 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image10_1920x1080-1536x864.jpeg 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image10_1920x1080-645x363.jpeg 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image10_1920x1080-660x370.jpeg 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image10_1920x1080-500x281.jpeg 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image10_1920x1080-160x90.jpeg 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image10_1920x1080-362x204.jpeg 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image10_1920x1080-196x110.jpeg 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image10_1920x1080-1024x576.jpeg 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image10_1920x1080-960x540.jpeg 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image10_1920x1080.webp 1920w" sizes="auto, (max-width: 768px) 100vw, 768px" title="image10_1920x1080" />Use AdaptGrow, a GPU-accelerated matrix factorization algorithm, to turn rolling correlation and tail-dependence matrices into hard clusters, soft factor...]]></summary>476 <content type="html" xml:base="https://developer.nvidia.com/blog/gpu-accelerated-clustering-for-financial-instruments-at-scale/"><![CDATA[<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image10_1920x1080-768x432.jpeg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image10_1920x1080-768x432.jpeg 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image10_1920x1080-179x101.jpeg 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image10_1920x1080-300x169.jpeg 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image10_1920x1080-625x352.jpeg 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image10_1920x1080-1536x864.jpeg 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image10_1920x1080-645x363.jpeg 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image10_1920x1080-660x370.jpeg 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image10_1920x1080-500x281.jpeg 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image10_1920x1080-160x90.jpeg 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image10_1920x1080-362x204.jpeg 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image10_1920x1080-196x110.jpeg 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image10_1920x1080-1024x576.jpeg 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image10_1920x1080-960x540.jpeg 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image10_1920x1080.webp 1920w" sizes="auto, (max-width: 768px) 100vw, 768px" title="image10_1920x1080" />Use AdaptGrow, a GPU-accelerated matrix factorization algorithm, to turn rolling correlation and tail-dependence matrices into hard clusters, soft factor...<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image10_1920x1080-768x432.jpeg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image10_1920x1080-768x432.jpeg 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image10_1920x1080-179x101.jpeg 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image10_1920x1080-300x169.jpeg 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image10_1920x1080-625x352.jpeg 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image10_1920x1080-1536x864.jpeg 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image10_1920x1080-645x363.jpeg 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image10_1920x1080-660x370.jpeg 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image10_1920x1080-500x281.jpeg 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image10_1920x1080-160x90.jpeg 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image10_1920x1080-362x204.jpeg 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image10_1920x1080-196x110.jpeg 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image10_1920x1080-1024x576.jpeg 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image10_1920x1080-960x540.jpeg 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/image10_1920x1080.webp 1920w" sizes="auto, (max-width: 768px) 100vw, 768px" title="image10_1920x1080" /><p>Use AdaptGrow, a GPU-accelerated matrix factorization algorithm, to turn rolling correlation and tail-dependence matrices into hard clusters, soft factor loadings, and structural-break signals at single-GPU and multi-node scale Quant strategies routinely group instruments for portfolio construction, risk aggregation, statistical arbitrage, and trade surveillance. Incorrect groupings can make…</p>477<p><a href="https://developer.nvidia.com/blog/gpu-accelerated-clustering-for-financial-instruments-at-scale/" rel="nofollow" data-wpel-link="internal" target="_self">Source</a></p>]]></content>478 <link rel="replies" type="text/html" href="https://developer.nvidia.com/blog/gpu-accelerated-clustering-for-financial-instruments-at-scale/#comments" thr:count="0"/>479 <link rel="replies" type="application/atom+xml" href="https://developer.nvidia.com/blog/gpu-accelerated-clustering-for-financial-instruments-at-scale/feed/" thr:count="0"/>480 <thr:total>0</thr:total>481 </entry>482 <entry>483 <author>484 <name>Tanya Lenz</name>485 </author>486 <title type="html"><![CDATA[NVIDIA AVO Reaches 100% on ARC-AGI-3, Demonstrating a Frontier-Level General-Purpose Architecture for Long-Horizon Autonomous Agents]]></title>487 <link rel="alternate" type="text/html" href="https://developer.nvidia.com/blog/nvidia-avo-reaches-100-on-arc-agi-3-demonstrating-a-frontier-level-general-purpose-architecture-for-long-horizon-autonomous-agents/" />488 <id>https://developer.nvidia.com/blog/?p=121575</id>489 <updated>2026-08-21T21:08:48Z</updated>490 <published>2026-08-21T13:00:00Z</published>491 <category scheme="https://developer.nvidia.com/blog" term="Agentic AI / Generative AI" /><category scheme="https://developer.nvidia.com/blog" term="Top Stories" /><category scheme="https://developer.nvidia.com/blog" term="Trustworthy AI / Cybersecurity" /><category scheme="https://developer.nvidia.com/blog" term="AI Agent" /><category scheme="https://developer.nvidia.com/blog" term="NVIDIA Research" /> <summary type="html"><![CDATA[<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/agentic-ai-visual-cybersecurity-avo-1-768x432.png" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/agentic-ai-visual-cybersecurity-avo-1-768x432.png 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/agentic-ai-visual-cybersecurity-avo-1-179x101.png 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/agentic-ai-visual-cybersecurity-avo-1-300x169.png 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/agentic-ai-visual-cybersecurity-avo-1-625x352.png 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/agentic-ai-visual-cybersecurity-avo-1-1536x864.png 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/agentic-ai-visual-cybersecurity-avo-1-2048x1152.png 2048w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/agentic-ai-visual-cybersecurity-avo-1-645x363.png 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/agentic-ai-visual-cybersecurity-avo-1-660x370.png 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/agentic-ai-visual-cybersecurity-avo-1-500x281.png 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/agentic-ai-visual-cybersecurity-avo-1-160x90.png 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/agentic-ai-visual-cybersecurity-avo-1-362x204.png 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/agentic-ai-visual-cybersecurity-avo-1-196x110.png 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/agentic-ai-visual-cybersecurity-avo-1-1024x576.png 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/agentic-ai-visual-cybersecurity-avo-1-960x540.png 960w" sizes="auto, (max-width: 768px) 100vw, 768px" title="agentic-ai-visual-cybersecurity-avo" />A frontier language model is only one component of an AI agent. The surrounding agent system—often called a harness—determines how the model receives...]]></summary>492 <content type="html" xml:base="https://developer.nvidia.com/blog/nvidia-avo-reaches-100-on-arc-agi-3-demonstrating-a-frontier-level-general-purpose-architecture-for-long-horizon-autonomous-agents/"><![CDATA[<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/agentic-ai-visual-cybersecurity-avo-1-768x432.png" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/agentic-ai-visual-cybersecurity-avo-1-768x432.png 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/agentic-ai-visual-cybersecurity-avo-1-179x101.png 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/agentic-ai-visual-cybersecurity-avo-1-300x169.png 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/agentic-ai-visual-cybersecurity-avo-1-625x352.png 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/agentic-ai-visual-cybersecurity-avo-1-1536x864.png 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/agentic-ai-visual-cybersecurity-avo-1-2048x1152.png 2048w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/agentic-ai-visual-cybersecurity-avo-1-645x363.png 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/agentic-ai-visual-cybersecurity-avo-1-660x370.png 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/agentic-ai-visual-cybersecurity-avo-1-500x281.png 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/agentic-ai-visual-cybersecurity-avo-1-160x90.png 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/agentic-ai-visual-cybersecurity-avo-1-362x204.png 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/agentic-ai-visual-cybersecurity-avo-1-196x110.png 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/agentic-ai-visual-cybersecurity-avo-1-1024x576.png 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/agentic-ai-visual-cybersecurity-avo-1-960x540.png 960w" sizes="auto, (max-width: 768px) 100vw, 768px" title="agentic-ai-visual-cybersecurity-avo" />A frontier language model is only one component of an AI agent. The surrounding agent system—often called a harness—determines how the model receives...<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/agentic-ai-visual-cybersecurity-avo-1-768x432.png" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/agentic-ai-visual-cybersecurity-avo-1-768x432.png 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/agentic-ai-visual-cybersecurity-avo-1-179x101.png 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/agentic-ai-visual-cybersecurity-avo-1-300x169.png 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/agentic-ai-visual-cybersecurity-avo-1-625x352.png 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/agentic-ai-visual-cybersecurity-avo-1-1536x864.png 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/agentic-ai-visual-cybersecurity-avo-1-2048x1152.png 2048w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/agentic-ai-visual-cybersecurity-avo-1-645x363.png 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/agentic-ai-visual-cybersecurity-avo-1-660x370.png 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/agentic-ai-visual-cybersecurity-avo-1-500x281.png 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/agentic-ai-visual-cybersecurity-avo-1-160x90.png 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/agentic-ai-visual-cybersecurity-avo-1-362x204.png 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/agentic-ai-visual-cybersecurity-avo-1-196x110.png 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/agentic-ai-visual-cybersecurity-avo-1-1024x576.png 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/agentic-ai-visual-cybersecurity-avo-1-960x540.png 960w" sizes="auto, (max-width: 768px) 100vw, 768px" title="agentic-ai-visual-cybersecurity-avo" /><p>A frontier language model is only one component of an AI agent. The surrounding agent system—often called a harness—determines how the model receives context, uses tools, maintains state, responds to feedback, recovers from failure, and sustains progress over long-running tasks. The challenge is how to build the agent architecture that makes frontier language models work reliably on extended…</p>493<p><a href="https://developer.nvidia.com/blog/nvidia-avo-reaches-100-on-arc-agi-3-demonstrating-a-frontier-level-general-purpose-architecture-for-long-horizon-autonomous-agents/" rel="nofollow" data-wpel-link="internal" target="_self">Source</a></p>]]></content>494 <link rel="replies" type="text/html" href="https://developer.nvidia.com/blog/nvidia-avo-reaches-100-on-arc-agi-3-demonstrating-a-frontier-level-general-purpose-architecture-for-long-horizon-autonomous-agents/#comments" thr:count="1"/>495 <link rel="replies" type="application/atom+xml" href="https://developer.nvidia.com/blog/nvidia-avo-reaches-100-on-arc-agi-3-demonstrating-a-frontier-level-general-purpose-architecture-for-long-horizon-autonomous-agents/feed/" thr:count="1"/>496 <thr:total>1</thr:total>497 </entry>498 </feed>