HTML 77.2%
TypeScript 10.5%
Python 9.6%
JavaScript 2.5%
1<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"2 xmlns:content="http://purl.org/rss/1.0/modules/content/"3 xmlns:wfw="http://wellformedweb.org/CommentAPI/"4 xmlns:dc="http://purl.org/dc/elements/1.1/"5 xmlns:atom="http://www.w3.org/2005/Atom"6 xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"7 xmlns:slash="http://purl.org/rss/1.0/modules/slash/"8 >910<channel>11 <title>Microsoft Research</title>12 <atom:link href="https://www.microsoft.com/en-us/research/feed/" rel="self" type="application/rss+xml" />13 <link>https://www.microsoft.com/en-us/research/</link>14 <description></description>15 <lastBuildDate>Mon, 31 Aug 2026 13:43:23 +0000</lastBuildDate>16 <language>en-US</language>17 <sy:updatePeriod>18 hourly </sy:updatePeriod>19 <sy:updateFrequency>20 1 </sy:updateFrequency>21 <generator>https://wordpress.org/?v=7.0.4</generator>22 <item>23 <title>GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models</title>24 <link>https://www.microsoft.com/en-us/research/blog/gigapath-flash-and-gigatime-flash-toward-population-scale-discovery-with-efficient-pathology-foundation-models/</link>25 26 <dc:creator><![CDATA[Naoto Usuyama, Jeya Maria Jose Valanarasu, Tristan Naumann]]></dc:creator>27 <pubDate>Mon, 31 Aug 2026 16:00:00 +0000</pubDate>28 <category><![CDATA[Research Blog]]></category>29 <guid isPermaLink="false">https://www.microsoft.com/en-us/research/blog/gigapath-flash-and-gigatime-flash-toward-population-scale-discovery-with-efficient-pathology-foundation-models/</guid>3031 <description><![CDATA[<p>What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration.</p>32<p>The post <a href="https://www.microsoft.com/en-us/research/blog/gigapath-flash-and-gigatime-flash-toward-population-scale-discovery-with-efficient-pathology-foundation-models/">GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models</a> appeared first on <a href="https://www.microsoft.com/en-us/research">Microsoft Research</a>.</p>33]]></description>34 <content:encoded><![CDATA[35<figure class="wp-block-image size-full"><img fetchpriority="high" decoding="async" width="1400" height="788" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/GigaPathFlash-BlogHeroFeature-1400x788_NEW.jpg" alt="Overview of the GigaPath/GigaTIME model family. GigaPath-Flash provides efficient tile and slide encoders at 22M parameters. GigaTIME-Flash predicts spatial proteomics from H&E, replacing the CNN backbone with the distilled ViT-S encoder. " class="wp-image-1184911" srcset="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/GigaPathFlash-BlogHeroFeature-1400x788_NEW.jpg 1400w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/GigaPathFlash-BlogHeroFeature-1400x788_NEW-300x169.jpg 300w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/GigaPathFlash-BlogHeroFeature-1400x788_NEW-1024x576.jpg 1024w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/GigaPathFlash-BlogHeroFeature-1400x788_NEW-768x432.jpg 768w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/GigaPathFlash-BlogHeroFeature-1400x788_NEW-1066x600.jpg 1066w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/GigaPathFlash-BlogHeroFeature-1400x788_NEW-655x368.jpg 655w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/GigaPathFlash-BlogHeroFeature-1400x788_NEW-240x135.jpg 240w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/GigaPathFlash-BlogHeroFeature-1400x788_NEW-640x360.jpg 640w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/GigaPathFlash-BlogHeroFeature-1400x788_NEW-960x540.jpg 960w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/GigaPathFlash-BlogHeroFeature-1400x788_NEW-1280x720.jpg 1280w" sizes="(max-width: 1400px) 100vw, 1400px" /></figure>36373839<div style="padding-bottom:0;padding-top:0" class="wp-block-msr-immersive-section alignfull row">40 41 <div class="container">42 <div class="wp-block-msr-immersive-section__inner wp-block-msr-immersive-section__inner--narrow">43 <div class="wp-block-columns mb-10 pb-1 pr-1 is-layout-flex wp-container-core-columns-is-layout-8f761849 wp-block-columns-is-layout-flex" style="box-shadow:var(--wp--preset--shadow--outlined)">44<div class="wp-block-column is-layout-flow wp-block-column-is-layout-flow">45<h2 id="at-a-glance" class="wp-block-heading h3">At a glance</h2>46474849<ul class="wp-block-list">50<li>The Flash family extends GigaPath and GigaTIME with dramatically improved efficiency, making large-scale pathology research more accessible and practical.</li>51525354<li>A distilled pathology foundation model backbone reduces computational requirements without sacrificing performance, enabling repeated analyses across larger patient cohorts.</li>55565758<li>These open models support population-scale discovery, helping researchers investigate disease biology, biomarkers, and clinical outcomes across diverse cancer datasets.</li>59</ul>60616263<p class="wp-block-paragraph"><a class="msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall" href="https://aka.ms/gigapath" target="_blank" rel="noopener noreferrer"><em>GigaPath</em><span class="sr-only"> (opens in new tab)</span></a><em> and </em><a class="msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall" href="https://aka.ms/gigatime" target="_blank" rel="noopener noreferrer"><em>GigaTIME</em><span class="sr-only"> (opens in new tab)</span></a><em> demonstrated how foundation models can support whole-slide analysis and tumor microenvironment modeling from routinely collected pathology data. GigaPath-Flash and GigaTIME-Flash make these capabilities substantially more efficient, enabling researchers to analyze larger cohorts, run more experiments, and move toward population-scale discovery.</em> <em>GigaPath-Flash and GigaTIME-Flash are research models. They are not intended or validated for clinical use, including diagnosis, prognosis, treatment selection, or other patient-care decisions. Performance may vary across datasets, scanners, institutions, populations, and use cases.</em></p>64</div>65</div> </div>66 </div>6768 </div>69707172<h2 id="the-scale-opportunity-in-computational-pathology" class="wp-block-heading">The scale opportunity in computational pathology</h2>73747576<p class="wp-block-paragraph">Histopathology is among the richest and most widely available sources of information in cancer research. Every tissue biopsy produces a whole-slide image that captures cellular morphology at subcellular resolution — and hospitals generate millions of these slides each year. This data contains information relevant to diagnosis, prognosis, treatment selection, and the biology of the tumor microenvironment.</p>77787980<p class="wp-block-paragraph">Foundation models have begun to unlock this information at scale. But whole-slide images are large — often exceeding a gigapixel — and applying a foundation model to even a single slide requires processing thousands of image tiles. When a research question involves tens of thousands of patients, the computational cost grows quickly. And population-scale discovery is not a single model run: it requires repeated cycles of feature extraction, statistical analysis, hypothesis testing, and validation across patient subgroups, biomarkers, and clinical endpoints. </p>81828384<p class="wp-block-paragraph">Computational cost limits the number of patients, datasets, tasks, and hypotheses that researchers can study. To realize the full potential of pathology foundation models, we need models that can be applied repeatedly and affordably across large patient populations.</p>85868788<h2 id="from-gigapath-and-gigatime-to-the-flash-family" class="wp-block-heading">From GigaPath and GigaTIME to the -Flash family</h2>89909192<p class="wp-block-paragraph"><a class="msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall" href="https://aka.ms/gigapath-paper" type="link" id="https://aka.ms/gigapath-paper" target="_blank" rel="noopener noreferrer">GigaPath (Nature, 2024)<span class="sr-only"> (opens in new tab)</span></a> is a whole-slide foundation model pretrained on large-scale real-world histopathology data from Providence. Unlike models that operate only at the tile level, GigaPath learns contextualized representations of entire slides, capturing both local cellular patterns and global tissue architecture.</p>93949596<p class="wp-block-paragraph"><a class="msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall" href="https://aka.ms/gigatime-paper" type="link" id="https://aka.ms/gigatime-paper" target="_blank" rel="noopener noreferrer">GigaTIME (Cell, 2026)<span class="sr-only"> (opens in new tab)</span></a> extends this line of work to tumor microenvironments. Trained on 40 million cells with paired H&E and multiplex immunofluorescence (mIF) data, GigaTIME translates routine H&E images into virtual spatial proteomics maps across 21 protein channels. Applied to over 14,000 cancer patients, it generated a virtual population that uncovered more than 1,200 statistically significant associations between immune cell states and clinical biomarkers.</p>979899100<p class="wp-block-paragraph">GigaPath and GigaTIME addressed the scale of pathology data and biological discovery. And now, the Flash family of models addresses scale of experimentation.</p>101102103104<figure class="wp-block-image aligncenter size-full"><img decoding="async" width="980" height="417" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/fig1_overview_GPF_GTF.png" alt="Figure 1: Overview of the GigaPath/GigaTIME model family. GigaPath-Flash provides efficient tile and slide encoders at 22M parameters. GigaTIME-Flash predicts spatial proteomics from H&E, replacing the CNN backbone with the distilled ViT-S encoder. " class="wp-image-1184838" srcset="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/fig1_overview_GPF_GTF.png 980w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/fig1_overview_GPF_GTF-300x128.png 300w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/fig1_overview_GPF_GTF-768x327.png 768w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/fig1_overview_GPF_GTF-240x102.png 240w" sizes="(max-width: 980px) 100vw, 980px" /><figcaption class="wp-element-caption">Figure 1: Overview of the GigaPath/GigaTIME model family. GigaPath-Flash provides efficient tile and slide encoders at 22M parameters. GigaTIME-Flash predicts spatial proteomics from H&E, replacing the CNN backbone with the distilled ViT-S encoder.</figcaption></figure>105106107108<h2 id="introducing-gigapath-flash-and-gigatime-flash" class="wp-block-heading">Introducing GigaPath-Flash and GigaTIME-Flash</h2>109110111112<p class="wp-block-paragraph">The Flash family shares a core design goal: preserving useful pathology representations while substantially reducing the computational resources required to generate and use them. Both models are built on a common efficient backbone — a compact ViT-S tile encoder distilled from the original billion-parameter GigaPath encoder — and both are released under the Apache 2.0 license. </p>113114115116<h3 id="gigapath-flash" class="wp-block-heading">GigaPath-Flash</h3>117118119120<p class="wp-block-paragraph">GigaPath-Flash is an efficient foundation model for whole-slide representation learning. It combines a 22M-parameter ViT-S tile encoder with a 21M-parameter LongNet slide encoder. The tile encoder is distilled from the original GigaPath ViT-g teacher, transferring the representational capacity of a billion-parameter model into a backbone that is an order of magnitude smaller. The slide encoder contextualizes all tile embeddings via dilated attention, scaling linearly with the number of tiles.</p>121122123124<p class="wp-block-paragraph">On slide-level classification benchmarks (PANDA prostate grading and EBRAINS brain tumor subtyping), GigaPath-Flash achieves the lowest inference cost among whole-slide pretrained models while retaining competitive performance — scoring within 3% of the original GigaPath at roughly 50 times less compute.</p>125126127128<figure class="wp-block-image aligncenter size-full"><img decoding="async" width="847" height="536" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/fig2_gigapath_slide_benchmark_efficiency_GPF_GTF.png" alt="Figure 2: Efficiency–performance trade-off on whole-slide benchmarks. GigaPath-Flash (red, top-left) achieves competitive performance at substantially lower computational cost than other whole-slide pretrained models." class="wp-image-1184839" srcset="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/fig2_gigapath_slide_benchmark_efficiency_GPF_GTF.png 847w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/fig2_gigapath_slide_benchmark_efficiency_GPF_GTF-300x190.png 300w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/fig2_gigapath_slide_benchmark_efficiency_GPF_GTF-768x486.png 768w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/fig2_gigapath_slide_benchmark_efficiency_GPF_GTF-240x152.png 240w" sizes="(max-width: 847px) 100vw, 847px" /><figcaption class="wp-element-caption">Figure 2: Efficiency–performance trade-off on whole-slide benchmarks. GigaPath-Flash (red, top-left) achieves competitive performance at substantially lower computational cost than other whole-slide pretrained models.</figcaption></figure>129130131132<h3 id="gigatime-flash" class="wp-block-heading">GigaTIME-Flash</h3>133134135136<p class="wp-block-paragraph">GigaTIME-Flash replaces the CNN backbone of the original GigaTIME with the GigaPath-Flash ViT-S encoder, paired with a lightweight convolutional decoder for H&E-to-mIF translation. The model is fine-tuned using LoRA adapters, keeping the pretrained encoder weights largely frozen.</p>137138139140<p class="wp-block-paragraph">On both in-distribution and out-of-distribution cohorts spanning brain, breast, colon, and lung cancers, GigaTIME-Flash matches or improves upon the original GigaTIME in spatial protein prediction quality. The gains are particularly notable on out-of-distribution data, suggesting that the foundation model backbone improves generalization to previously unseen tissue types.</p>141142143144<figure class="wp-block-image aligncenter size-full"><img loading="lazy" decoding="async" width="1203" height="612" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/fig3_gigatime_mean_pearson_bars_GPF_GTF.png" alt="Figure 3: Mean windowed Pearson correlation for GigaTIME and GigaTIME-Flash on the GigaTIME test set and four out-of-distribution Prov-TMA cohorts. GigaTIME-Flash matches or improves upon the original across all cohorts." class="wp-image-1184840" srcset="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/fig3_gigatime_mean_pearson_bars_GPF_GTF.png 1203w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/fig3_gigatime_mean_pearson_bars_GPF_GTF-300x153.png 300w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/fig3_gigatime_mean_pearson_bars_GPF_GTF-1024x521.png 1024w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/fig3_gigatime_mean_pearson_bars_GPF_GTF-768x391.png 768w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/fig3_gigatime_mean_pearson_bars_GPF_GTF-240x122.png 240w" sizes="auto, (max-width: 1203px) 100vw, 1203px" /><figcaption class="wp-element-caption">Figure 3: Mean windowed Pearson correlation for GigaTIME and GigaTIME-Flash on the GigaTIME test set and four out-of-distribution Prov-TMA cohorts. GigaTIME-Flash matches or improves upon the original across all cohorts.</figcaption></figure>145146147148<h2 id="efficiency-without-giving-up-the-foundation" class="wp-block-heading">Efficiency without giving up the foundation</h2>149150151152<p class="wp-block-paragraph">The efficiency gains of the Flash models are substantial:</p>153154155156<figure class="wp-block-table"><table><thead><tr><th>Model</th><th>Type</th><th>Efficiency gain</th></tr></thead><tbody><tr><td>GigaPath-Flash</td><td>Whole-slide Foundation Model</td><td>~50× less compute, 97% of predictive performance compared to GigaPath</td></tr><tr><td>GigaTIME-Flash</td><td>Spatial Proteomics</td><td>~6× faster, ~8× less memory, better predictive performance compared to GigaTIME</td></tr></tbody></table></figure>157158159160<p class="wp-block-paragraph">For a single slide, these differences reduce runtime and hardware requirements. Across tens of thousands of slides, they can determine whether an experiment is practical at all. To illustrate, we estimate the wall-clock time for generating virtual mIF across cohorts of different sizes on a single A100 GPU, assuming approximately 10,000 tiles per slide:</p>161162163164<figure class="wp-block-table"><table class="has-fixed-layout"><thead><tr><th>Model</th><th>1,000 slides</th><th>100K slides</th><th>1M slides</th></tr></thead><tbody><tr><td><strong>GigaTIME-Flash</strong></td><td>~2 GPU-hours</td><td>~7 GPU-days</td><td>~70 GPU-days</td></tr><tr><td>GigaTIME</td><td>~7 GPU-hours</td><td>~30 GPU-days</td><td>~300 GPU-days</td></tr></tbody></table><figcaption class="wp-element-caption">Estimates assume ~10,000 tiles per slide, batch size 128, single NVIDIA A100 GPU. Actual runtime depends on slide size, tiling resolution, and hardware.</figcaption></figure>165166167168<figure class="wp-block-image aligncenter size-full"><img loading="lazy" decoding="async" width="842" height="375" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/fig4_gigatime_efficiency_plot_GPF_GTF.png" alt="Figure 4: GigaTIME efficiency scaling. Left: throughput (tiles/sec) vs. batch size. Right: peak GPU memory (GB) vs. batch size. GigaTIME-Flash scales to over 1,600 tiles/sec while using a fraction of the memory." class="wp-image-1184841" srcset="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/fig4_gigatime_efficiency_plot_GPF_GTF.png 842w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/fig4_gigatime_efficiency_plot_GPF_GTF-300x134.png 300w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/fig4_gigatime_efficiency_plot_GPF_GTF-768x342.png 768w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/fig4_gigatime_efficiency_plot_GPF_GTF-240x107.png 240w" sizes="auto, (max-width: 842px) 100vw, 842px" /><figcaption class="wp-element-caption">Figure 4: GigaTIME efficiency scaling. Left: throughput (tiles/sec) vs. batch size. Right: peak GPU memory (GB) vs. batch size. GigaTIME-Flash scales to over 1,600 tiles/sec while using a fraction of the memory.</figcaption></figure>169170171172<h2 id="an-open-model-release" class="wp-block-heading">An open model release</h2>173174175176<p class="wp-block-paragraph">Both GigaPath-Flash and GigaTIME-Flash are released as open-weight models under the Apache 2.0 license. Model weights and code are available on HuggingFace:</p>177178179180<ul class="wp-block-list">181<li><a class="msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall" href="https://aka.ms/GigaPath-Flash" type="link" id="aka.ms/GigaPath-Flash" target="_blank" rel="noopener noreferrer">GigaPath-Flash<span class="sr-only"> (opens in new tab)</span></a></li>182183184185<li><a class="msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall" href="https://aka.ms/GigaTIME-Flash" type="link" id="aka.ms/GigaTIME-Flash" target="_blank" rel="noopener noreferrer">GigaTIME-Flash<span class="sr-only"> (opens in new tab)</span></a></li>186187188189<li><a href="https://www.microsoft.com/en-us/research/publication/gigapath-flash-and-gigatime-flash-efficient-pathology-foundation-models-for-whole-slide-and-tumor-microenvironment-analysis/" type="msr-research-item" id="1180651">GigaPath-Flash and GigaTIME-Flash: Efficient Pathology Foundation Models for Whole-Slide and Tumor Microenvironment Analysis</a></li>190</ul>191192193194<p class="wp-block-paragraph">This is an early research release. Our current evaluations cover a limited set of benchmarks and cohorts, and broader validation across tasks, scanners, and patient populations is still needed. We expect the most valuable applications of these models to include scientific questions, cohorts, and use cases beyond those in our initial experiments. We welcome community evaluation, and we are equally interested in reports of where these models work well and where they fall short.</p>195196197198<p class="wp-block-paragraph">Downstream clinical applications will require additional multi-institutional and prospective validation.</p>199200201202 <div class="border-bottom border-top border-gray-300 mt-5 mb-5 msr-promo text-center text-md-left alignwide" data-bi-aN="promo" data-bi-id="999693">203 204205 <p class="msr-promo__label text-gray-800 text-center text-uppercase">206 <span class="px-4 bg-white display-inline-block font-weight-semibold small">Spotlight: Event Series</span>207 </p>208 209 <div class="row pt-3 pb-4 align-items-center">210 <div class="msr-promo__media col-12 col-md-5">211 <a class="bg-gray-300 display-block" href="https://www.microsoft.com/en-us/research/event/microsoft-research-forum/past-episodes/?OCID=msr_researchforum_MCR_Blog_Promo" aria-label="Microsoft Research Forum" data-bi-cn="Microsoft Research Forum" target="_blank">212 <img decoding="async" class="w-100 display-block" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2025/05/Research-Forum-hero_1400x788.jpg" alt="Research Forum | abstract background with colorful hexagons" />213 </a>214 </div>215 216 <div class="msr-promo__content p-3 px-5 col-12 col-md">217218 <h2 class="h4">Microsoft Research Forum</h2>219 220 <p id="microsoft-research-forum" class="large">Join us for a continuous exchange of ideas about research in the era of general AI. Watch the latest episodes on demand.</p>221 222 <div class="wp-block-buttons justify-content-center justify-content-md-start">223 <div class="wp-block-button">224 <a href="https://www.microsoft.com/en-us/research/event/microsoft-research-forum/past-episodes/?OCID=msr_researchforum_MCR_Blog_Promo" aria-describedby="microsoft-research-forum" class="btn btn-brand glyph-append glyph-append-chevron-right" data-bi-cn="Microsoft Research Forum" target="_blank">225 Watch on-demand </a>226 </div>227 </div>228 </div><!--/.msr-promo__content-->229 </div><!--/.msr-promo__inner-wrap-->230<span id="label-external-link" class="sr-only" aria-hidden="true">Opens in a new tab</span> </div><!--/.msr-promo-->231 232233234<h2 id="efficiency-as-an-enabler-of-discovery" class="wp-block-heading">Efficiency as an enabler of discovery</h2>235236237238<p class="wp-block-paragraph">GigaPath and GigaTIME demonstrated what pathology foundation models can learn from whole slides and tumor tissue. GigaPath-Flash and GigaTIME-Flash are a step toward making those capabilities usable across larger populations, more experiments, and a broader research community.</p>239240241242<p class="wp-block-paragraph">By making pathology foundation models more efficient, we hope to expand the scale of the scientific questions researchers can ask.</p>243244245246<h2 id="acknowledgements" class="wp-block-heading">Acknowledgements</h2>247248249250<p class="wp-block-paragraph">GigaPath-Flash and GigaTIME-Flash are joint work across Microsoft Research, the University of Washington, and Providence. For technical details, see the <a href="https://www.microsoft.com/en-us/research/publication/gigapath-flash-and-gigatime-flash-efficient-pathology-foundation-models-for-whole-slide-and-tumor-microenvironment-analysis/" type="msr-research-item" id="1180651">paper</a>.</p>251252253254<p class="wp-block-paragraph"><em><strong>Paper co-authors</strong>: Naoto Usuyama, Jeya Maria Jose Valanarasu, Sicong Yao, Hanwen Xu, Jaspreet Bagga, Guanghui Qin, Robert E. Kramer, Cliff Wong, Soohee Lee, Hao Qiu, Theodore Zhengde Zhao, Racheli Ben Shimol, Angela Crabtree, Kevin Matlock, Eduardo Alejandro Lozano Garcia, Naiteek Sangani, Alberto Santamaria-Pang, Maximilian Rokuss, Yashna Hasija, Naisargi Manishkumar Patel, Jason Entenmann, Alexandra Q. Bartlett, Bill J. Wright, Bernard A. Fox, Brian Piening, Sheng Zhang, Sheng Wang, Tristan Naumann, Carlo Bifulco, Hoifung Poon</em></p>255<span id="label-external-link" class="sr-only" aria-hidden="true">Opens in a new tab</span><p>The post <a href="https://www.microsoft.com/en-us/research/blog/gigapath-flash-and-gigatime-flash-toward-population-scale-discovery-with-efficient-pathology-foundation-models/">GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models</a> appeared first on <a href="https://www.microsoft.com/en-us/research">Microsoft Research</a>.</p>256]]></content:encoded>257 258 259 260 </item>261 <item>262 <title>Broadening access to Skala creates a faster path to predictive DFT </title>263 <link>https://www.microsoft.com/en-us/research/blog/broadening-access-to-skala-creates-a-faster-path-to-predictive-dft/</link>264 265 <dc:creator><![CDATA[Sebastian Ehlert, Stefano Battaglia, Thijs Vogels, Jan Hermann, Jens Wehner, Giulia Luise, Klaas Giesbertz, Chin-Wei Huang, Aaron Kaplan, Kate Milton, Stephanie Marisa Lanius, Derk Kooi, P. Bernát Szabó, Gregor Simm, Rianne van den Berg, Paola Gori Giorgi]]></dc:creator>266 <pubDate>Thu, 20 Aug 2026 16:00:00 +0000</pubDate>267 <category><![CDATA[Research Blog]]></category>268 <guid isPermaLink="false">https://www.microsoft.com/en-us/research/?p=1184167</guid>269270 <description><![CDATA[<p>Skala 1.1, the updated deep-learning exchange-correlation functional from Microsoft Research, provides greater accuracy, expanded accessibility across the computational chemistry ecosystem, and a living benchmark to track computational performance. </p>271<p>The post <a href="https://www.microsoft.com/en-us/research/blog/broadening-access-to-skala-creates-a-faster-path-to-predictive-dft/">Broadening access to Skala creates a faster path to predictive DFT </a> appeared first on <a href="https://www.microsoft.com/en-us/research">Microsoft Research</a>.</p>272]]></description>273 <content:encoded><![CDATA[274<figure class="wp-block-image size-full"><img loading="lazy" decoding="async" width="1400" height="788" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Skala-BlogHeroFeature-1400x788-1-1.jpg" alt="Schematic of the Skala architecture, showing how meta-GGA electronic features are transformed through point-wise processing and non-local atomic interactions to predict density functional theory energies." class="wp-image-1184452" srcset="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Skala-BlogHeroFeature-1400x788-1-1.jpg 1400w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Skala-BlogHeroFeature-1400x788-1-1-300x169.jpg 300w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Skala-BlogHeroFeature-1400x788-1-1-1024x576.jpg 1024w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Skala-BlogHeroFeature-1400x788-1-1-768x432.jpg 768w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Skala-BlogHeroFeature-1400x788-1-1-1066x600.jpg 1066w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Skala-BlogHeroFeature-1400x788-1-1-655x368.jpg 655w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Skala-BlogHeroFeature-1400x788-1-1-240x135.jpg 240w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Skala-BlogHeroFeature-1400x788-1-1-640x360.jpg 640w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Skala-BlogHeroFeature-1400x788-1-1-960x540.jpg 960w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Skala-BlogHeroFeature-1400x788-1-1-1280x720.jpg 1280w" sizes="auto, (max-width: 1400px) 100vw, 1400px" /></figure>275276277278<div style="padding-bottom:0;padding-top:0" class="wp-block-msr-immersive-section alignfull row">279 280 <div class="container">281 <div class="wp-block-msr-immersive-section__inner wp-block-msr-immersive-section__inner--narrow">282 <div class="wp-block-columns mb-10 pb-1 pr-1 is-layout-flex wp-container-core-columns-is-layout-8f761849 wp-block-columns-is-layout-flex" style="box-shadow:var(--wp--preset--shadow--outlined)">283<div class="wp-block-column is-layout-flow wp-block-column-is-layout-flow">284<h2 id="at-a-glance" class="wp-block-heading h3">At a glance</h2>285286287288<ul class="wp-block-list">289<li>Skala 1.1 demonstrates the continuously improving nature of Microsoft Research’s deep-learning DFT approach: trained on 2.5× more data than its predecessor, it delivers substantially higher accuracy across key molecular simulation challenges, including thermochemistry, reaction kinetics, and molecular structure prediction.</li>290291292293<li>Skala is now available in <strong>CP2K</strong> and is being integrated into <strong>Psi4</strong>, <strong>FHI-aims</strong>, <strong>ORCA</strong> and <strong>VASP</strong>, bringing next-generation DFT accuracy closer to the communities that rely on these codes every day.</li>294295296297<li>Microsoft Research is also introducing a living benchmark that will track the computational performance of successive, increasingly optimized Skala releases to help the community measure and accelerate progress toward ever greater accuracy and efficiency.</li>298299300301<li>Together, these developments mark another milestone toward a future in which computational chemistry simulations are both predictive and integrated in all relevant scientific and industrial workflows.</li>302</ul>303</div>304</div> </div>305 </div>306307 </div>308309310311<p class="wp-block-paragraph"><strong>Bringing density functional theory (DFT) to predictive accuracy is a journey, not a single breakthrough.</strong> Since <a href="https://www.microsoft.com/en-us/research/blog/breaking-bonds-breaking-ground-advancing-the-accuracy-of-computational-chemistry-with-deep-learning/" target="_blank" rel="noreferrer noopener">introducing Skala</a>, our deep-learning exchange-correlation functional, we have continued to advance along two complementary fronts: improving accuracy and expanding accessibility across the computational chemistry ecosystem.</p>312313314315<figure class="wp-block-image aligncenter size-full"><img loading="lazy" decoding="async" width="1785" height="2560" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/fig1-blog_SKALA1.1-scaled.png" alt="Fig. 1: Table of errors for Skala-1.1 and competing density functionals on the 55 subsets of GMTKN55. Skala-1.1 delivers the lowest error on 32 subsets, indicating broad and consistent accuracy across a wide range of chemical properties and reaction types." class="wp-image-1184173" srcset="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/fig1-blog_SKALA1.1-scaled.png 1785w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/fig1-blog_SKALA1.1-209x300.png 209w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/fig1-blog_SKALA1.1-714x1024.png 714w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/fig1-blog_SKALA1.1-768x1101.png 768w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/fig1-blog_SKALA1.1-1071x1536.png 1071w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/fig1-blog_SKALA1.1-1428x2048.png 1428w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/fig1-blog_SKALA1.1-126x180.png 126w" sizes="auto, (max-width: 1785px) 100vw, 1785px" /><figcaption class="wp-element-caption">Figure 1: Accuracy of Skala-1.1 for thermochemistry, kinetics, and non-covalent interactions. At the computational cost of a meta-GGA functional, Skala 1.1 outperforms the best, most expensive global hybrid functionals, ranking first (earning gold medals) in 32 of the 55 categories of the widely used GMTKN55 benchmark, which spans a broad range of chemical problems.</figcaption></figure>316317318319<p class="wp-block-paragraph">On the accuracy front, the <a class="msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall" href="https://arxiv.org/abs/2506.14665" target="_blank" rel="noopener noreferrer">release of <strong>Skala-1.1</strong><span class="sr-only"> (opens in new tab)</span></a> provides the first demonstration of the continuous-improvement paradigm underlying Skala. Trained on 2.5x more data than the first public version of Skala, the updated model delivers substantially improved performance across key challenges in molecular simulation, including main-group thermochemistry, reaction kinetics, and molecular structure prediction.</p>320321322323<p class="wp-block-paragraph">But accuracy alone is not enough. DFT is the computational engine behind a vast range of scientific and industrial workflows, spanning chemistry, materials science, catalysis, energy technologies, and drug discovery. To have real-world impact, advanced functionals must be accessible where scientists already perform their calculations. That is why we are also expanding the Skala ecosystem through collaborations with leading electronic-structure software developers.</p>324325326327<p class="wp-block-paragraph">Today, we are announcing that Skala is available in <strong>CP2K</strong> and is being integrated into <strong>Psi4</strong>, <strong>FHI-aims</strong>, <strong>ORCA</strong> and <strong>VASP</strong>, bringing next-generation DFT accuracy closer to the communities that rely on these codes every day. Alongside these integration efforts, we are introducing a living benchmark that tracks the computational performance of successive, increasingly optimized Skala releases. By providing a transparent and continuously updated reference for implementations across software packages and hardware platforms, this resource will help the community measure and accelerate progress toward ever greater accuracy and efficiency. </p>328329330331<p class="wp-block-paragraph">Together, these developments mark another milestone toward a future in which computational chemistry simulations are both predictive and accessible across a broader range of relevant scientific and industrial workflows.</p>332333334335<p class="wp-block-paragraph">Want to learn more about Skala and why DFT plays such an important role in in-silico discovery? Read also <a class="msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall" href="https://aka.ms/skaladft/blog" type="link" id="https://aka.ms/skaladft/blog" target="_blank" rel="noopener noreferrer">our first blog post<span class="sr-only"> (opens in new tab)</span></a>.</p>336337338339<h2 id="skala-as-a-continuously-improving-functional" class="wp-block-heading">Skala as a continuously improving functional</h2>340341342343<p class="wp-block-paragraph">Unlike the traditional “functional zoo”, where new functionals accumulate without replacing older ones, Skala follows a different philosophy: each release is designed to supersede the previous one. As new data, model architectures, and training strategies become available, the model improves while maintaining the same practical computational cost.</p>344345346347<p class="wp-block-paragraph"><strong>Skala-1.1</strong> is the latest demonstration of this approach. It achieves a weighted average error of <strong>2.8 kcal/mol on GMTKN55</strong>, a widely used benchmark suite comprising 55 categories of chemistry, including thermochemistry, reaction barriers, and noncovalent interactions. This level of accuracy surpasses today’s leading global (range-separated) hybrid functionals while retaining the efficiency of a semi-local functional. Beyond energies, Skala-1.1 also provides highly accurate electron densities, dipole moments, and molecular geometries.</p>348349350351<p class="wp-block-paragraph">These advances were enabled by major expansions of the <strong><a class="msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall" href="https://www.nature.com/articles/s41597-026-07200-8" type="link" id="https://www.nature.com/articles/s41597-026-07200-8" target="_blank" rel="noopener noreferrer">Microsoft Research Accurate Chemistry Collection<span class="sr-only"> (opens in new tab)</span></a> (MSR-ACC)</strong>, our large-scale collection of high-accuracy quantum-chemistry reference data generated with expensive wavefunction methods. For Skala-1.1, we added new categories, including electron affinities and noncovalent clusters, increasing both the size and, crucially, the diversity of the training data. This data-driven approach allows Skala to improve systematically with each generation, moving us closer to a truly scalable and predictive DFT framework.</p>352353354355<h2 id="available-where-scientists-work" class="wp-block-heading">Available where scientists work</h2>356357358359<p class="wp-block-paragraph">To fully realize the potential of Skala’s continuously evolving approach to DFT, we need dedicated infrastructure that allows new releases to be rapidly and seamlessly integrated into the major software packages used by scientists in industry and academia. In turn, this will establish the fast feedback loop essential for accelerating Skala’s ongoing development.</p>360361362363<p class="wp-block-paragraph">We first made Skala available through our <a class="msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall" href="https://github.com/microsoft/skala" type="link" id="https://github.com/microsoft/skala" target="_blank" rel="noopener noreferrer">open-source community release<span class="sr-only"> (opens in new tab)</span></a>, built on (GPU4)<a class="msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall" href="https://pyscf.org/" type="link" id="https://pyscf.org/" target="_blank" rel="noopener noreferrer">PySCF<span class="sr-only"> (opens in new tab)</span></a> and <a class="msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall" href="https://aka.ms/SkalaDFT/ASE" type="link" id="https://aka.ms/SkalaDFT/ASE" target="_blank" rel="noopener noreferrer">integrated with ASE<span class="sr-only"> (opens in new tab)</span></a>. This enables researchers to evaluate and apply Skala with minimal effort while benefiting from highly optimized CPU and GPU performance.</p>364365366367<p class="wp-block-paragraph">But no single software package can meet the needs of every application or research community. Computational chemistry and materials science rely on a rich ecosystem of electronic-structure codes, each shaped over decades to tackle specific scientific and industrial challenges. Bringing Skala to this broader ecosystem has therefore been a major focus of the past year. We are fortunate to build on the remarkable foundations created by the DFT community and grateful to the many researchers and developers who are helping to make Skala available within the software platforms that scientists use every day.</p>368369370371 <div class="border-bottom border-top border-gray-300 mt-5 mb-5 msr-promo text-center text-md-left alignwide" data-bi-aN="promo" data-bi-id="999693">372 373374 <p class="msr-promo__label text-gray-800 text-center text-uppercase">375 <span class="px-4 bg-white display-inline-block font-weight-semibold small">Spotlight: Event Series</span>376 </p>377 378 <div class="row pt-3 pb-4 align-items-center">379 <div class="msr-promo__media col-12 col-md-5">380 <a class="bg-gray-300 display-block" href="https://www.microsoft.com/en-us/research/event/microsoft-research-forum/past-episodes/?OCID=msr_researchforum_MCR_Blog_Promo" aria-label="Microsoft Research Forum" data-bi-cn="Microsoft Research Forum" target="_blank">381 <img decoding="async" class="w-100 display-block" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2025/05/Research-Forum-hero_1400x788.jpg" alt="Research Forum | abstract background with colorful hexagons" />382 </a>383 </div>384 385 <div class="msr-promo__content p-3 px-5 col-12 col-md">386387 <h2 class="h4">Microsoft Research Forum</h2>388 389 <p id="microsoft-research-forum" class="large">Join us for a continuous exchange of ideas about research in the era of general AI. Watch the latest episodes on demand.</p>390 391 <div class="wp-block-buttons justify-content-center justify-content-md-start">392 <div class="wp-block-button">393 <a href="https://www.microsoft.com/en-us/research/event/microsoft-research-forum/past-episodes/?OCID=msr_researchforum_MCR_Blog_Promo" aria-describedby="microsoft-research-forum" class="btn btn-brand glyph-append glyph-append-chevron-right" data-bi-cn="Microsoft Research Forum" target="_blank">394 Watch on-demand </a>395 </div>396 </div>397 </div><!--/.msr-promo__content-->398 </div><!--/.msr-promo__inner-wrap-->399<span id="label-external-link" class="sr-only" aria-hidden="true">Opens in a new tab</span> </div><!--/.msr-promo-->400 401402403<h2 id="from-community-release-to-native-integrations" class="wp-block-heading">From community release to native integrations</h2>404405406407<p class="wp-block-paragraph">In collaboration with the team of Prof. Thomas D. Kühne at the <a class="msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall" href="https://casus.science" type="link" id="https://casus.science" target="_blank" rel="noopener noreferrer">Center for Advanced Systems Understanding (CASUS)<span class="sr-only"> (opens in new tab)</span></a>, Skala has been successfully integrated into the open-source <a class="msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall" href="https://www.cp2k.org/" type="link" id="https://www.cp2k.org/" target="_blank" rel="noopener noreferrer">CP2K<span class="sr-only"> (opens in new tab)</span></a> package. With more than 25 years of development, CP2K is a powerhouse for DFT simulations, particularly for large-scale systems and long-timescale molecular dynamics, while also providing a rich portfolio of high-accuracy electronic-structure methods. Skala expands the frontiers of what is possible within CP2K, delivering a step change in DFT accuracy while preserving the computational efficiency needed for simulations at scale. We are excited to see how CP2K’s scale and versatility, combined with Skala’s continuously improving accuracy, will enable new scientific applications and discoveries in the years ahead.</p>408409410411<p class="wp-block-paragraph">There is more to come. Together with its vibrant developer’s community , we are actively integrating Skala into the open-source <a class="msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall" href="https://psicode.org/" type="link" id="https://psicode.org/" target="_blank" rel="noopener noreferrer">Psi4<span class="sr-only"> (opens in new tab)</span></a> package, an essential platform for molecular electronic-structure research. Combined with the PySCF-based Skala Community Edition, this will make Skala available in three widely used open-source quantum chemistry packages.</p>412413414415<p class="wp-block-paragraph">Beyond open-source software, we are working closely with leading developers behind <a class="msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall" href="https://www.fhi-aims.org/" target="_blank" rel="noopener noreferrer">FHI-aims<span class="sr-only"> (opens in new tab)</span></a>, <a class="msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall" href="https://www.faccts.de/orca/" target="_blank" rel="noopener noreferrer">ORCA<span class="sr-only"> (opens in new tab)</span></a>, and <a class="msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall" href="https://vasp.at/" target="_blank" rel="noopener noreferrer">VASP<span class="sr-only"> (opens in new tab)</span></a>, with the goal of making Skala broadly accessible across the major software platforms used in computational chemistry and materials science.</p>416417418419<h2 id="validating-accuracy-across-implementations-cp2k-as-case-study" class="wp-block-heading">Validating accuracy across implementations: CP2K as case study</h2>420421422423<p class="wp-block-paragraph">Thorough testing is essential for any new implementation. We want to ensure that Skala delivers consistent accuracy across different codes and computational settings. Together with the CP2K team, we developed a comprehensive suite of integration tests to verify that Skala produces numerically correct and reliable results. We are particularly grateful to the CASUS team, whose deep expertise in the numerical verification of computational methods was instrumental in designing and validating this testing framework.</p>424425426427<figure class="wp-block-image aligncenter size-full is-resized"><img loading="lazy" decoding="async" width="2007" height="2007" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/cp2k_vs_pyscf_error-3_SKALA.png" alt="Fig. 2: Plot comparing signed errors for Skala-1.1 using the CP2K and PySCF implementations on a representative subset of GMTKN55. The two implementations show nearly identical results, agreeing within 0.04 kcal/mol and confirming that the CP2K integration reproduces the accuracy of the community release." class="wp-image-1184807" style="width:556px;height:auto" srcset="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/cp2k_vs_pyscf_error-3_SKALA.png 2007w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/cp2k_vs_pyscf_error-3_SKALA-300x300.png 300w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/cp2k_vs_pyscf_error-3_SKALA-1024x1024.png 1024w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/cp2k_vs_pyscf_error-3_SKALA-150x150.png 150w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/cp2k_vs_pyscf_error-3_SKALA-768x768.png 768w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/cp2k_vs_pyscf_error-3_SKALA-1536x1536.png 1536w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/cp2k_vs_pyscf_error-3_SKALA-180x180.png 180w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/cp2k_vs_pyscf_error-3_SKALA-360x360.png 360w" sizes="auto, (max-width: 2007px) 100vw, 2007px" /><figcaption class="wp-element-caption">Figure 2: Signed errors relative to high-accuracy reference values for a representative subset of GMTKN55, comparing the CP2K and PySCF implementations of Skala-1.1 using as closely matched numerical settings as possible. The two implementations agree to within 0.1 kcal/mol MAD across the entire subset.</figcaption></figure>428429430431<p class="wp-block-paragraph">A detailed discussion of the implementation, validation strategy, and testing infrastructure for Skala in CP2K can be found in our joint paper with the CASUS team: “<a href="https://www.microsoft.com/en-us/research/publication/molecular-implementation-of-the-machine-learned-skalaexchange-correlation-functional-in-cp2k-through-gauxc/">Molecular Implementation of the Machine-Learned Skala Exchange-Correlation Functional in CP2K through GauXC</a>.”</p>432433434435<h2 id="a-living-performance-report-for-skala" class="wp-block-heading">A living performance report for Skala</h2>436437438439<p class="wp-block-paragraph">Accuracy and broad availability only translate into scientific impact if Skala is also fast. Today, Skala can deliver performance comparable to semi-local meta-GGAs on both CPU’s (with an overhead that disappears for molecules with more than 20-30 atoms) and GPUs, and we are committed to preserving that efficiency as it is integrated across the electronic-structure software ecosystem.</p>440441442443<p class="wp-block-paragraph">But performance is not a fixed property. New Skala releases, improvements in libraries such as GauXC, and hardware-specific optimizations continuously improve efficiency and reveal new opportunities for further gains. Capturing this progress requires more than a single benchmark snapshot.</p>444445446447<p class="wp-block-paragraph">To provide a transparent and up-to-date view of Skala’s performance, we are publishing a <a class="msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall" href="https://aka.ms/SkalaDFT/timing" type="link" id="https://aka.ms/SkalaDFT/timing" target="_blank" rel="noopener noreferrer">benchmarking harness together with a living performance report<span class="sr-only"> (opens in new tab)</span></a> that will be updated as new optimizations become available. This report tracks performance across a range of tasks and hardware platforms, while the harness enables package developers to benchmark, validate, and improve their own Skala implementations.</p>448449450451<figure class="wp-block-image aligncenter size-full"><img loading="lazy" decoding="async" width="2518" height="1108" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/fig3_SKALA1.1.png" alt="Fig. 3 (figure attached): Benchmark of the computational cost of Skala-1.1 relative to r2SCAN, B3LYP, and M06-2X on GPUs and CPUs. Skala-1.1 achieves performance comparable to r2SCAN on GPUs and approaches semilocal-functional cost on CPUs for larger systems, while remaining significantly less expensive than hybrid functionals." class="wp-image-1184172" srcset="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/fig3_SKALA1.1.png 2518w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/fig3_SKALA1.1-300x132.png 300w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/fig3_SKALA1.1-1024x451.png 1024w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/fig3_SKALA1.1-768x338.png 768w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/fig3_SKALA1.1-1536x676.png 1536w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/fig3_SKALA1.1-2048x901.png 2048w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/fig3_SKALA1.1-240x106.png 240w" sizes="auto, (max-width: 2518px) 100vw, 2518px" /><figcaption class="wp-element-caption">Figure 3: Computational cost of Skala on GPU and CPU, compared with a popular metaGGA functional (r2SCAN) and two hybrid functionals (B3LYP and M06-2X). On GPU, Skala 1.1 has the same cost as r2SCAN, and the hybrid functionals become more expensive for systems with more than ~1000 orbitals. On CPU, Skala has an overhead with respect to the other functionals for smaller systems, that disappears for systems with more than ~300 orbitals. </figcaption></figure>452453454455<h2 id="acknowledgments" class="wp-block-heading">Acknowledgments</h2>456457458459<p class="wp-block-paragraph">Skala is the product of a truly collaborative effort across AI for Science, and we thank our engineering, project management, and business operations teams for making this work possible. We also thank MSR Accelerator for their partnership in advancing data generation efforts and accelerating software integrations that help bring Skala to the broader scientific community.</p>460<span id="label-external-link" class="sr-only" aria-hidden="true">Opens in a new tab</span><p>The post <a href="https://www.microsoft.com/en-us/research/blog/broadening-access-to-skala-creates-a-faster-path-to-predictive-dft/">Broadening access to Skala creates a faster path to predictive DFT </a> appeared first on <a href="https://www.microsoft.com/en-us/research">Microsoft Research</a>.</p>461]]></content:encoded>462 463 464 465 </item>466 <item>467 <title>MindTopo reveals VLMs’ spatial reasoning abilities</title>468 <link>https://www.microsoft.com/en-us/research/blog/mindtopo-reveals-vlms-spatial-reasoning-abilities/</link>469 470 <dc:creator><![CDATA[Yunfei Ge, Anbang Liu, Qineng Wang, Johnalbert Garnica, Zihan Wang, Reuben Tan, Jianfeng Gao, Ruohan Zhang, Yining Hong, Jiajun Wu, Manling Li]]></dc:creator>471 <pubDate>Wed, 12 Aug 2026 16:00:00 +0000</pubDate>472 <category><![CDATA[Research Blog]]></category>473 <guid isPermaLink="false">https://www.microsoft.com/en-us/research/blog/mindtopo-reveals-vlms-spatial-reasoning-abilities/</guid>474475 <description><![CDATA[<p>A path, a fence, a knot. MindTopo sets a new benchmark for testing how AI understands topological relationships and highlights new opportunities to strengthen spatial reasoning and planning.</p>476<p>The post <a href="https://www.microsoft.com/en-us/research/blog/mindtopo-reveals-vlms-spatial-reasoning-abilities/">MindTopo reveals VLMs’ spatial reasoning abilities</a> appeared first on <a href="https://www.microsoft.com/en-us/research">Microsoft Research</a>.</p>477]]></description>478 <content:encoded><![CDATA[479<figure class="wp-block-image size-full"><img loading="lazy" decoding="async" width="2560" height="1441" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/MindTopo-BlogHeroFeature-1400x788-1-scaled.jpg" alt="Benchmark overview showing ten spatial reasoning and planning tasks grouped into two rows. The top row, labeled “Reasoning,” includes Maze, Assembly, Bead, Sheep, and Knot. The bottom row, labeled “Planning,” includes Pipe, One Stroke, Swap, Chat Noir, and Untangle. The MindTopo logo and title are centered between the two categories." class="wp-image-1181354" srcset="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/MindTopo-BlogHeroFeature-1400x788-1-scaled.jpg 2560w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/MindTopo-BlogHeroFeature-1400x788-1-300x169.jpg 300w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/MindTopo-BlogHeroFeature-1400x788-1-1024x576.jpg 1024w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/MindTopo-BlogHeroFeature-1400x788-1-768x432.jpg 768w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/MindTopo-BlogHeroFeature-1400x788-1-1536x865.jpg 1536w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/MindTopo-BlogHeroFeature-1400x788-1-2048x1153.jpg 2048w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/MindTopo-BlogHeroFeature-1400x788-1-1066x600.jpg 1066w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/MindTopo-BlogHeroFeature-1400x788-1-655x368.jpg 655w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/MindTopo-BlogHeroFeature-1400x788-1-240x135.jpg 240w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/MindTopo-BlogHeroFeature-1400x788-1-640x360.jpg 640w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/MindTopo-BlogHeroFeature-1400x788-1-960x540.jpg 960w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/MindTopo-BlogHeroFeature-1400x788-1-1280x720.jpg 1280w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/MindTopo-BlogHeroFeature-1400x788-1-1920x1080.jpg 1920w" sizes="auto, (max-width: 2560px) 100vw, 2560px" /></figure>480481482483<div style="padding-bottom:0;padding-top:0" class="wp-block-msr-immersive-section alignfull row">484 485 <div class="container">486 <div class="wp-block-msr-immersive-section__inner wp-block-msr-immersive-section__inner--narrow">487 <div class="wp-block-columns mb-10 pb-1 pr-1 is-layout-flex wp-container-core-columns-is-layout-8f761849 wp-block-columns-is-layout-flex" style="box-shadow:var(--wp--preset--shadow--outlined)">488<div class="wp-block-column is-layout-flow wp-block-column-is-layout-flow">489<h2 id="at-a-glance" class="wp-block-heading h3">At a glance</h2>490491492493<ul class="wp-block-list">494<li>MindTopo is a new benchmark for testing topological reasoning in AI, evaluating whether multimodal models can understand concepts such as connectivity, enclosure, order, separation, and knots.</li>495496497498<li>The benchmark measures both reasoning and planning, testing not only whether models can recognize topological relationships in static images but also whether they can preserve and manipulate those relationships through a sequence of actions.</li>499500501502<li>Current multimodal models perform much better on static recognition than interactive tasks, suggesting they struggle to maintain a consistent understanding of topology over time.</li>503504505506<li>Failures often emerge during planning rather than perception, with models losing track of structural relationships as scenes change or proposing actions that violate physical constraints.</li>507508509510<li>The findings highlight an important opportunity to advance AI systems for robotics and interactive environments, where understanding what stays connected, enclosed, ordered, or knotted is essential for reliable decision-making.</li>511</ul>512</div>513</div> </div>514 </div>515516 </div>517518519520<p class="wp-block-paragraph">Can AI determine whether two rooms remain connected after a wall is added? Can it recognize whether an animal is inside a fence, distinguish a true knot from a tangled loop, or rearrange several ropes without allowing them to pass through one another?</p>521522523524<p class="wp-block-paragraph">These questions concern 3D topology, a form of spatial understanding based not on exact distances, angles, or shapes, but on structural relationships that persist as objects bend, stretch, or deform. Connectivity, enclosure, ordering, and knottedness are examples of topological properties. These properties are a foundational layer of human spatial understanding in Cognitive Science, yet they remain largely absent from how multimodal AI systems are evaluated.</p>525526527528<p class="wp-block-paragraph">In a new research study, we introduce <a class="msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall" rel="noopener noreferrer" target="_blank" href="https://mind-topo.github.io/">MindTopo<span class="sr-only"> (opens in new tab)</span></a>, a benchmark designed to evaluate whether multimodal large language models possess this kind of topological intuition. Our findings reveal a substantial gap between recognizing topology in a static image and maintaining an innate understanding of that topology while planning and acting. Current models can sometimes identify a connected path, enclosed region, or knot in a single scene, but that understanding often breaks down once the model must manipulate the scene through a sequence of actions.</p>529530531532<h2 id="how-mindtopo-defines-topological-space" class="wp-block-heading">How MindTopo defines topological space</h2>533534535536<p class="wp-block-paragraph">Most spatial evaluations for multimodal models focus on Euclidean properties such as distance, direction, size, and relative position. Inspired by Piaget and other cognitive literature’s classification of topological ability, MindTopo organizes its tasks around the following five categories:</p>537538539540<ul class="wp-block-list">541<li><strong>Continuity</strong> asks whether a path or object remains unbroken.</li>542543544545<li><strong>Separation</strong> asks whether nearby elements form one structure or distinct parts.</li>546547548549<li><strong>Order</strong> tracks how elements are arranged along a path or through a transformation.</li>550551552553<li><strong>Enclosure</strong> tests whether a boundary creates an inside and an outside.</li>554555556557<li><strong>Knots</strong> tests whether ropes are truly knotted or linked rather than merely tangled in appearance.</li>558</ul>559560561562<p class="wp-block-paragraph">Each category is evaluated at two cognitive levels. In reasoning tasks, a model examines one or more rendered scenes and answers a question about their topological structure: whether two points in a maze are connected, whether the sheep are inside the fence, whether a rope is truly knotted. In planning tasks, the model interacts with a simulated environment and selects actions that must create, preserve, or remove a particular relation, such as rotating pipe segments, drawing a separating path, rearranging blocks, trapping a moving agent, or untangling ropes. The environments enforce legal actions, so a model cannot solve a rope puzzle by passing one strand through another.</p>563564565566<figure class="wp-block-image aligncenter size-full"><img loading="lazy" decoding="async" width="2560" height="828" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Figure-1_MINDTOPO-scaled.png" alt="This figure provides an overview of MINDTOPO, a benchmark for evaluating topological reasoning in multimodal large language models. The figure illustrates two evaluation settings: reasoning, where models answer visual questions about rendered scenes, and planning, where models interact with environments to transform an initial state into a goal state. The benchmark spans five topological properties—continuity, separation, order, enclosure, and knots—with representative tasks including Maze and Pipe, IKEA and One Stroke, Bead and Swap, Sheep and Chat Noir, and Knot and Untangle. A radar chart on the right summarizes model performance across these categories and shows that current models still struggle, particularly on topological spatial reasoning that requires planning and maintaining invariants across actions." class="wp-image-1181360" srcset="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Figure-1_MINDTOPO-scaled.png 2560w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Figure-1_MINDTOPO-300x97.png 300w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Figure-1_MINDTOPO-1024x331.png 1024w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Figure-1_MINDTOPO-768x248.png 768w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Figure-1_MINDTOPO-1536x497.png 1536w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Figure-1_MINDTOPO-2048x662.png 2048w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Figure-1_MINDTOPO-240x78.png 240w" sizes="auto, (max-width: 2560px) 100vw, 2560px" /><figcaption class="wp-element-caption">Figure 1. MindTopo pairs questions about static scenes with interactive tasks that require models to preserve or change the same topological relations. </figcaption></figure>567568569570<p class="wp-block-paragraph">All scenes are generated from controlled simulators, which provide exact ground truth and adjustable difficulty. That control makes it possible to separate two failure modes that otherwise look alike: a model that fails because a scene is visually complex, and a model that fails because it cannot maintain the underlying relationship as objects move.</p>571572573574<figure class="wp-block-image aligncenter size-full"><img loading="lazy" decoding="async" width="2560" height="1548" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Figure-2_MINDTOPO-scaled.png" alt=" This figure provides an overview of the 13 MINDTOPO tasks organized by five topological properties and two cognitive levels. Continuity includes 2D Maze and 3D Maze reasoning tasks and the Pipe planning environment; Separation includes IKEA reasoning and One Stroke planning; Order includes Bead and Origami Point reasoning and Swap planning; Enclosure includes Sheep and Hole reasoning and Chat Noir planning; and Knots includes Knot reasoning and Untangle planning. Reasoning tasks pair rendered scenes with visual questions and example answers, while planning tasks, labeled “Gym Env,” show representative initial, intermediate, and final or goal states of interactive environments." class="wp-image-1181363" srcset="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Figure-2_MINDTOPO-scaled.png 2560w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Figure-2_MINDTOPO-300x181.png 300w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Figure-2_MINDTOPO-1024x619.png 1024w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Figure-2_MINDTOPO-768x464.png 768w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Figure-2_MINDTOPO-1536x929.png 1536w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Figure-2_MINDTOPO-2048x1238.png 2048w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Figure-2_MINDTOPO-240x145.png 240w" sizes="auto, (max-width: 2560px) 100vw, 2560px" /><figcaption class="wp-element-caption">Figure 2. MindTopo maps reasoning and planning tasks to continuity, separation, order, enclosure, and knots. </figcaption></figure>575576577578<h2 id="seeing-topology-is-not-the-same-as-acting-on-it" class="wp-block-heading">Seeing topology is not the same as acting on it</h2>579580581582<p class="wp-block-paragraph">Across a broad set of proprietary and open-weight models, performance was consistently stronger on static reasoning than on interactive planning, and both remained well below human performance. The contrast was especially clear when success depended on preserving a relationship across many actions.</p>583584585586<p class="wp-block-paragraph">The error patterns help locate the problem. Static mistakes usually began with perception, such as missing a wall, opening, or crossing. Planning mistakes appeared after the scene had been understood. Models followed a locally plausible move without tracking its later consequences, lost the task over multiple turns, or proposed an action that violated the environment’s dynamics.</p>587588589590 <div class="border-bottom border-top border-gray-300 mt-5 mb-5 msr-promo text-center text-md-left alignwide" data-bi-aN="promo" data-bi-id="1144027">591 592593 <p class="msr-promo__label text-gray-800 text-center text-uppercase">594 <span class="px-4 bg-white display-inline-block font-weight-semibold small">PODCAST SERIES</span>595 </p>596 597 <div class="row pt-3 pb-4 align-items-center">598 <div class="msr-promo__media col-12 col-md-5">599 <a class="bg-gray-300 display-block" href="https://www.microsoft.com/en-us/research/story/ai-testing-and-evaluation-learnings-from-science-and-industry/" aria-label="AI Testing and Evaluation: Learnings from Science and Industry" data-bi-cn="AI Testing and Evaluation: Learnings from Science and Industry" target="_blank">600 <img decoding="async" class="w-100 display-block" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2025/06/EP2-AI-TE_Hero_Feature_River_No_Text_1400x788.jpg" alt="Illustrated headshots of Daniel Carpenter, Timo Minssen, Chad Atalla, and Kathleen Sullivan for the Microsoft Research Podcast" />601 </a>602 </div>603 604 <div class="msr-promo__content p-3 px-5 col-12 col-md">605606 <h2 class="h4">AI Testing and Evaluation: Learnings from Science and Industry</h2>607 608 <p id="ai-testing-and-evaluation-learnings-from-science-and-industry" class="large">Discover how Microsoft is learning from other domains to advance evaluation and testing as a pillar of AI governance.</p>609 610 <div class="wp-block-buttons justify-content-center justify-content-md-start">611 <div class="wp-block-button">612 <a href="https://www.microsoft.com/en-us/research/story/ai-testing-and-evaluation-learnings-from-science-and-industry/" aria-describedby="ai-testing-and-evaluation-learnings-from-science-and-industry" class="btn btn-brand glyph-append glyph-append-chevron-right" data-bi-cn="AI Testing and Evaluation: Learnings from Science and Industry" target="_blank">613 Listen now </a>614 </div>615 </div>616 </div><!--/.msr-promo__content-->617 </div><!--/.msr-promo__inner-wrap-->618<span id="label-external-link" class="sr-only" aria-hidden="true">Opens in a new tab</span> </div><!--/.msr-promo-->619 620621622<h2 id="what-generative-tools-reveal" class="wp-block-heading">What generative tools reveal</h2>623624625626<p class="wp-block-paragraph">We also tested whether image and video generation could help models maintain an understanding of topological relationships. Image generation sometimes helped when the relevant relation was visible in a single frame, but it remained unreliable across a sequence of crossings or moves. Video rollouts frequently altered topology or violated task dynamics. Visual simulation appeared useful only to the extent that it preserved structural constraints over time.</p>627628629630<h2 id="building-agents-that-preserve-structure" class="wp-block-heading">Building agents that preserve structure</h2>631632633634<p class="wp-block-paragraph">MindTopo is intended as a controlled diagnostic for this gap. Robots, accessibility tools, and interactive assistants must understand not only where objects are, but also what remains connected, enclosed, ordered, or knotted as actions unfold. Closing that gap may require models that carry an explicit topological state, or world models whose predictions preserve topology by construction.</p>635<span id="label-external-link" class="sr-only" aria-hidden="true">Opens in a new tab</span><p>The post <a href="https://www.microsoft.com/en-us/research/blog/mindtopo-reveals-vlms-spatial-reasoning-abilities/">MindTopo reveals VLMs’ spatial reasoning abilities</a> appeared first on <a href="https://www.microsoft.com/en-us/research">Microsoft Research</a>.</p>636]]></content:encoded>637 638 639 640 </item>641 <item>642 <title>Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement</title>643 <link>https://www.microsoft.com/en-us/research/blog/introducing-care-x-towards-clinically-useful-radiology-vlms-with-auxiliary-supervision-reward-aligned-learning-and-tool-augmented-measurement/</link>644 645 <dc:creator><![CDATA[Mercy Ranjit, Nikhilesh E, Dr. Abhyuday Kumara Swamy, Tanuja Ganu]]></dc:creator>646 <pubDate>Tue, 11 Aug 2026 16:00:00 +0000</pubDate>647 <category><![CDATA[Research Blog]]></category>648 <guid isPermaLink="false">https://www.microsoft.com/en-us/research/blog/introducing-care-x-towards-clinically-useful-radiology-vlms-with-auxiliary-supervision-reward-aligned-learning-and-tool-augmented-measurement/</guid>649650 <description><![CDATA[<p> Radiology AI is evolving beyond report generation. CARE-X explores a unified approach that combines flexible reasoning, calibrated predictions, and measurement-based tools for chest X-ray interpretation.</p>651<p>The post <a href="https://www.microsoft.com/en-us/research/blog/introducing-care-x-towards-clinically-useful-radiology-vlms-with-auxiliary-supervision-reward-aligned-learning-and-tool-augmented-measurement/">Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement</a> appeared first on <a href="https://www.microsoft.com/en-us/research">Microsoft Research</a>.</p>652]]></description>653 <content:encoded><![CDATA[654<figure class="wp-block-image size-full"><img loading="lazy" decoding="async" width="2560" height="1441" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/CAREX-BlogHeroFeature-1400x788-1-scaled.jpg" alt="Diagram of a tool-augmented radiology vision-language model (VLM) workflow. An orchestrator routes a user’s image query to a VLM, which performs image analysis, calls measurement tools to calculate cardiac and thoracic widths and the cardiothoracic ratio (CTR), and returns diagnostic metrics with visual overlays." class="wp-image-1181346" srcset="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/CAREX-BlogHeroFeature-1400x788-1-scaled.jpg 2560w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/CAREX-BlogHeroFeature-1400x788-1-300x169.jpg 300w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/CAREX-BlogHeroFeature-1400x788-1-1024x576.jpg 1024w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/CAREX-BlogHeroFeature-1400x788-1-768x432.jpg 768w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/CAREX-BlogHeroFeature-1400x788-1-1536x865.jpg 1536w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/CAREX-BlogHeroFeature-1400x788-1-2048x1153.jpg 2048w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/CAREX-BlogHeroFeature-1400x788-1-1066x600.jpg 1066w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/CAREX-BlogHeroFeature-1400x788-1-655x368.jpg 655w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/CAREX-BlogHeroFeature-1400x788-1-240x135.jpg 240w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/CAREX-BlogHeroFeature-1400x788-1-640x360.jpg 640w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/CAREX-BlogHeroFeature-1400x788-1-960x540.jpg 960w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/CAREX-BlogHeroFeature-1400x788-1-1280x720.jpg 1280w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/CAREX-BlogHeroFeature-1400x788-1-1920x1080.jpg 1920w" sizes="auto, (max-width: 2560px) 100vw, 2560px" /></figure>655656657658<p class="wp-block-paragraph">Research Note: CARE-X is a research model and not a Microsoft product offering or medical device. It has not been cleared or approved by any regulatory authority and is not intended for clinical diagnosis, screening, or patient care. The results described below are retrospective research findings and do not establish the safety, effectiveness, or suitability of CARE-X for any clinical use. References to potential workflows describe areas for future research, not currently available capabilities or recommended uses. </p>659660661662<div style="padding-bottom:0;padding-top:0" class="wp-block-msr-immersive-section alignfull row">663 664 <div class="container">665 <div class="wp-block-msr-immersive-section__inner wp-block-msr-immersive-section__inner--narrow">666 <div class="wp-block-columns mb-10 pb-1 pr-1 is-layout-flex wp-container-core-columns-is-layout-8f761849 wp-block-columns-is-layout-flex" style="box-shadow:var(--wp--preset--shadow--outlined)">667<div class="wp-block-column is-layout-flow wp-block-column-is-layout-flow">668<h2 id="at-a-glance" class="wp-block-heading h3">At a glance</h2>669670671672<ul class="wp-block-list">673<li>The challenge: Chest X-ray interpretation spans diverse tasks that require both expressive report generation and calibrated diagnostic predictions.</li>674675676677<li>CARE-X is a unified chest X-ray VLM for diverse clinical interpretation tasks. It combines generation and structured prediction to provide both free-text reasoning and deterministic outputs.</li>678679680681<li>CARE-X uses reinforcement learning (DAPO) to reward clinical correctness in a multi-task setting.</li>682683684685<li>In a separate research experiment from CARE-X, we paired Qwen3-VL-4B-Instruct with deterministic measurement tools to evaluate whether direct computation could improve performance on measurement-dependent conditions compared with visual approximation alone. </li>686687688689<li>Validated on real-world Indian clinical data from Narayana Health, including rare ICU pathologies and CT-confirmed enlargement conditions.</li>690</ul>691</div>692</div> </div>693 </div>694695 </div>696697698699<h2 id="what-radiologists-need-task-diversity-flexibility-and-clinical-fidelity" class="wp-block-heading">What radiologists need: Task diversity, flexibility, and clinical fidelity</h2>700701702703<p class="wp-block-paragraph">A clinically useful radiology AI system must support a wide range of tasks, adapt to different workflows, and produce outputs that are medically accurate.</p>704705706707<p class="wp-block-paragraph">Radiologists and other clinicians use chest X-rays for many different purposes. A clinically useful AI system must be able to support that range of tasks. It may be asked to generate detailed findings and concise impressions for a report, answer questions about the presence, absence, or location of a finding, identify medical devices and assess their placement, or pinpoint exactly where an abnormality appears in an image.</p>708709710711<p class="wp-block-paragraph">These tasks also require different kinds of outputs, from narrative reports to calibrated diagnostic scores. And above all, they require clinical accuracy. A report could ostensibly be perfectly written yet clinically wrong if it misses a finding, reverses a negation, or misidentifies a location. Certain findings could be trivial in one context and vital to identify in another.</p>712713714715<p class="wp-block-paragraph"><a href="https://www.microsoft.com/en-us/research/publication/care-x-towards-clinically-useful-radiology-vlms-with-auxiliary-supervision-reward-aligned-learning-and-tool-augmented-measurement/">CARE-X</a> was developed as a research model to explore how a unified approach can address these diverse demands. The system combines generative and discriminative capabilities, clinically aligned optimization, and tool-based reasoning to support a broader range of radiology workflows while maintaining clinical fidelity.</p>716717718719 <div class="border-bottom border-top border-gray-300 mt-5 mb-5 msr-promo text-center text-md-left alignwide" data-bi-aN="promo" data-bi-id="999693">720 721722 <p class="msr-promo__label text-gray-800 text-center text-uppercase">723 <span class="px-4 bg-white display-inline-block font-weight-semibold small">Spotlight: Event Series</span>724 </p>725 726 <div class="row pt-3 pb-4 align-items-center">727 <div class="msr-promo__media col-12 col-md-5">728 <a class="bg-gray-300 display-block" href="https://www.microsoft.com/en-us/research/event/microsoft-research-forum/past-episodes/?OCID=msr_researchforum_MCR_Blog_Promo" aria-label="Microsoft Research Forum" data-bi-cn="Microsoft Research Forum" target="_blank">729 <img decoding="async" class="w-100 display-block" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2025/05/Research-Forum-hero_1400x788.jpg" alt="Research Forum | abstract background with colorful hexagons" />730 </a>731 </div>732 733 <div class="msr-promo__content p-3 px-5 col-12 col-md">734735 <h2 class="h4">Microsoft Research Forum</h2>736 737 <p id="microsoft-research-forum" class="large">Join us for a continuous exchange of ideas about research in the era of general AI. Watch the latest episodes on demand.</p>738 739 <div class="wp-block-buttons justify-content-center justify-content-md-start">740 <div class="wp-block-button">741 <a href="https://www.microsoft.com/en-us/research/event/microsoft-research-forum/past-episodes/?OCID=msr_researchforum_MCR_Blog_Promo" aria-describedby="microsoft-research-forum" class="btn btn-brand glyph-append glyph-append-chevron-right" data-bi-cn="Microsoft Research Forum" target="_blank">742 Watch on-demand </a>743 </div>744 </div>745 </div><!--/.msr-promo__content-->746 </div><!--/.msr-promo__inner-wrap-->747<span id="label-external-link" class="sr-only" aria-hidden="true">Opens in a new tab</span> </div><!--/.msr-promo-->748 749750751<h2 id="gaps-in-current-radiology-vision-language-models" class="wp-block-heading">Gaps in current radiology vision-language models</h2>752753754755<p class="wp-block-paragraph">Despite the impressive task breadth of recent models, critical gaps remain between what radiologists need and what current systems deliver:</p>756757758759<ol start="1" class="wp-block-list">760<li><strong>No calibrated confidence for diagnostic decisions.</strong> Generative VLMs predict diagnoses as free text, but they typically do not provide calibrated confidence scores. In clinical settings, confidence matters. Clinicians cannot tune sensitivity–specificity trade-offs across clinical contexts—an important requirement for real-world deployment. Discriminative models provide these properties but lack the flexibility of open-ended generation.</li>761762763764<li><strong>Cross-entropy loss does not optimize clinical fidelity.</strong> Standard training methods treat all token-level errors similarly, regardless of their clinical consequences. A coordinate mistake may be penalized no more than a harmless wording change. A “yes” can be flipped to a “no” even though the clinical meaning is completely different. Missing a life-threatening finding may carry the same training penalty as omitting a minor observation. As a result, models are not explicitly optimized for what matters most in patient care.</li>765766767768<li><strong>No capability for measurement-dependent findings. </strong>Some radiological findings require more than visual recognition. Radiological signs such as cardiomegaly, mediastinal widening etc. depend on precise measurements. For example, a model may correctly recognize whether a chest radiograph was acquired using an AP or PA view. But determining cardiomegaly requires measuring the cardiac and thoracic widths and determining the cardiothoracic ratio. Those quantities should be measured and computed rather than visually approximated while considering variables such as type of view, exposure, rotation of the patient etc. </li>769</ol>770771772773<p class="wp-block-paragraph">Together, these gaps call for more than a fluent generative model. The system must combine broad task coverage, structured predictions, clinically aligned optimization, and quantitative tools where direct measurement is required.</p>774775776777<hr class="wp-block-separator has-alpha-channel-opacity"/>778779780781<h2 id="care-x-one-model-flexible-outputs" class="wp-block-heading">CARE-X: One model, flexible outputs</h2>782783784785<p class="wp-block-paragraph">CARE-X brings these diverse interpretation capabilities into one model, using generative or dual inference according to the needs of each task:</p>786787788789<figure class="wp-block-table is-style-stripes"><table class="has-fixed-layout"><thead><tr><th>Task type</th><th>What CARE-X does</th><th>Inference mode</th></tr></thead><tbody><tr><td>Report generation: Findings</td><td>Produces the detailed findings section</td><td>Generative</td></tr><tr><td>Report generation: Impression</td><td>Produces the concise diagnostic impression</td><td>Generative</td></tr><tr><td>Presence and negation assessment</td><td>Determines whether a pathology is present or absent and handles negation</td><td>Dual: generative + auxiliary head</td></tr><tr><td>Disease location assessment</td><td>Identifies where an abnormality appears</td><td>Generative</td></tr><tr><td>Fine-grained multilabel disease classification</td><td>Categorizes abnormalities across multiple labels</td><td>Generative</td></tr><tr><td>Multilabel tubes and lines classification</td><td>Identifies visible medical devices</td><td>Generative</td></tr><tr><td>Abnormal placement detection of tubes and lines</td><td>Determines whether a device is positioned incorrectly</td><td>Dual: generative + auxiliary head</td></tr><tr><td>Abnormality phrase grounding</td><td>Localizes a described pathological finding</td><td>Dual: generative + auxiliary head</td></tr><tr><td>Anatomical grounding</td><td>Localizes 29 anatomical regions</td><td>Dual: generative + auxiliary head</td></tr></tbody></table><figcaption class="wp-element-caption">Table 1: CARE-X task coverage and inference modes</figcaption></figure>790791792793<p class="wp-block-paragraph"><strong>Dual inference</strong> means that a single forward pass produces both an autoregressive response and a structured auxiliary-head prediction with a confidence score. This provides free-text flexibility alongside threshold-adjustable outputs for tasks where operating-point control matters.</p>794795796797<h2 id="the-care-x-architecture-and-training-approach" class="wp-block-heading">The CARE-X architecture and training approach</h2>798799800801<p class="wp-block-paragraph">CARE-X is built on a SigLIP2-so400M vision encoder and a Phi-4-mini-instruct (3.8B) language model connected through a lightweight adapter. To support both free-text generation and structured clinical predictions, the model augments the shared language backbone with <strong>task-specific auxiliary heads</strong> for classification and visual grounding. These heads provide calibrated diagnostic predictions and spatial localization signals while sharing representations with the generative language model. Rather than being trained independently, they are co-trained with the language-modeling objective, allowing structured supervision to enrich shared representations and improve generative performance on the same tasks.</p>802803804805<p class="wp-block-paragraph"><strong>Training.</strong> CARE-X uses a three-stage supervised fine-tuning pipeline (vision pre-training, adapter/head training, and LoRA adaptation) followed by DAPO-based reinforcement learning. DAPO optimizes task-specific rewards for clinical reporting, diagnostic accuracy, and spatial grounding quality.</p>806807808809<figure class="wp-block-image aligncenter size-full"><img loading="lazy" decoding="async" width="2362" height="1824" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Figure_1_CARE-X_updated.png" alt="CARE-X architecture with a SigLIP2 vision encoder, Phi-4-mini backbone, classification and grounding auxiliary heads, language modeling, and DAPO alignment for report generation, VQA, and grounding." class="wp-image-1181260" srcset="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Figure_1_CARE-X_updated.png 2362w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Figure_1_CARE-X_updated-300x232.png 300w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Figure_1_CARE-X_updated-1024x791.png 1024w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Figure_1_CARE-X_updated-768x593.png 768w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Figure_1_CARE-X_updated-1536x1186.png 1536w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Figure_1_CARE-X_updated-2048x1582.png 2048w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Figure_1_CARE-X_updated-233x180.png 233w" sizes="auto, (max-width: 2362px) 100vw, 2362px" /><figcaption class="wp-element-caption"><em>Figure 1. The CARE-X model. (Left) Supervised fine-tuning with task-specific heads — classification, grounding, and language modeling — sharing the same Phi-4-mini-instruct backbone. The classification head outputs calibrated P(Yes)/P(No) scores; the grounding head outputs bounding box coordinate with confidence; the language modeling head generates free-text responses. (Right) DAPO with task-specific rewards for multi-task reinforcement alignment across report generation, grounding, and VQA.</em> </figcaption></figure>810811812813<h2 id="auxiliary-supervision-structured-prediction-strengthens-generation" class="wp-block-heading">Auxiliary supervision: Structured prediction strengthens generation</h2>814815816817<p class="wp-block-paragraph">A central finding of this work is that co-training discriminative auxiliary heads with a generative VLM enriches shared representations, leading to stronger generative performance on the same tasks while also providing calibrated structured predictions.</p>818819820821<h3 id="grounding-improvements" class="wp-block-heading">Grounding improvements</h3>822823824825<p class="wp-block-paragraph">The auxiliary grounding head consistently improves localization over generative decoding. On anatomical grounding (Chest ImaGenome), mAP and mIoU increase by +28.2 pp and +6.2 pp, while the largest gains occur on phrase grounding (PadChest), with +24.6 pp mAP and +14.1 pp mIoU. The composite spatial loss enhances geometric precision in shared representations.</p>826827828829<h3 id="dapo-bridges-the-gap-to-dedicated-detection-heads" class="wp-block-heading">DAPO bridges the gap to dedicated detection heads</h3>830831832833<p class="wp-block-paragraph">DAPO-trained generative output approaches or exceeds the SFT auxiliary detection head. On Anatomy grounding, CARE-X generative (0.868 mAP) surpasses the SFT detection head (0.865). This is practically significant—it demonstrates that reward-aligned learning can bring autoregressive spatial decoding to parity with structured prediction, offering clinicians a single generative inference mode without requiring auxiliary heads at test time.</p>834835836837<h3 id="calibrated-classification-with-tunable-operating-points" class="wp-block-heading">Calibrated classification with tunable operating points</h3>838839840841<p class="wp-block-paragraph">Beyond representation enrichment, the classification head offers a distinct deployment advantage: calibrated probability scores with tunable thresholds allow clinicians to shift between high-sensitivity screening and high-specificity confirmation from a single forward pass—a capability purely generative architectures cannot provide.</p>842843844845<figure class="wp-block-table"><table class="has-fixed-layout"><thead><tr><th>Model</th><th>Inference Setting</th><th>Sensitivity ↑</th><th>PPV ↑</th><th>F1 ↑</th></tr></thead><tbody><tr><td>CARE-X</td><td>Generative</td><td>0.932</td><td>0.895</td><td>0.913</td></tr><tr><td>CARE-X (Th=0.5)</td><td>Auxiliary Head</td><td><strong>0.943</strong></td><td>0.885</td><td><strong>0.913</strong></td></tr><tr><td>CARE-X (Th=0.6)</td><td>Auxiliary Head</td><td>0.855</td><td><strong>0.927</strong></td><td>0.890</td></tr><tr><td>CheXOne</td><td>Generative</td><td>0.878</td><td>0.854</td><td>0.866</td></tr><tr><td>MedGemma</td><td>Generative</td><td>0.798</td><td>0.886</td><td>0.839</td></tr></tbody></table><figcaption class="wp-element-caption">Table 2: Abnormality classification performance on Chest ImaGenome. Adjustable thresholds enable operating-point selection.</figcaption></figure>846847848849<hr class="wp-block-separator has-alpha-channel-opacity"/>850851852853<h2 id="strong-report-generation-across-four-benchmarks" class="wp-block-heading">Strong report generation across four benchmarks</h2>854855856857<p class="wp-block-paragraph">Within the paper’s comparison set, CARE-X achieves the strongest performance on most reported metrics across MIMIC-CXR, IU-Xray, CheXpert-Plus, and ReXGradient. CRIMSON, a held-out metric that evaluates abnormal findings and weights errors by clinical severity, suggests these gains reflect clinically meaningful improvements rather than reward-specific optimization.</p>858859860861<figure class="wp-block-image aligncenter size-full"><img loading="lazy" decoding="async" width="1508" height="384" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/FIG2_CARE_NOLOGO.png" alt="CRIMSON scores (↑) for CARE-X against baseline report-generation models across four chest X-ray datasets — ReXGradient, MIMIC-CXR, IU-Xray, and CheXpert-Plus. CARE-X (highlighted) achieves the highest CRIMSON score on every dataset." class="wp-image-1181277" srcset="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/FIG2_CARE_NOLOGO.png 1508w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/FIG2_CARE_NOLOGO-300x76.png 300w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/FIG2_CARE_NOLOGO-1024x261.png 1024w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/FIG2_CARE_NOLOGO-768x196.png 768w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/FIG2_CARE_NOLOGO-240x61.png 240w" sizes="auto, (max-width: 1508px) 100vw, 1508px" /><figcaption class="wp-element-caption"><em>Figure 2. CRIMSON scores (↑) for CARE-X against baseline report-generation models across four chest X-ray datasets — ReXGradient, MIMIC-CXR, IU-Xray, and CheXpert-Plus. CARE-X (highlighted) achieves the highest CRIMSON score on every dataset.</em></figcaption></figure>862863864865<h2 id="care-x-reaches-94-accuracy-on-rexvqa" class="wp-block-heading">CARE-X reaches 94% accuracy on ReXVQA</h2>866867868869<p class="wp-block-paragraph"><strong>CARE-X ranks first on the </strong><a class="msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall" href="https://rexrank.ai/" target="_blank" rel="noopener noreferrer"><strong>ReXrank RexVQA leaderboard</strong><span class="sr-only"> (opens in new tab)</span></a> as of August 2026. On the ReXVQA benchmark (41,007 question–answer pairs across five clinically relevant categories), CARE-X reaches <strong>94% overall accuracy</strong>, six percentage points above the next-best publicly reported model. </p>870871872873<figure class="wp-block-image aligncenter size-full is-resized"><img loading="lazy" decoding="async" width="2013" height="2369" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Figure_3_CARE-X_NoLogo.png" alt="Radar chart comparing ReXVQA accuracy across six categories for three models: CheXOne-R1, MedGemma, and CARE-X. " class="wp-image-1181280" style="aspect-ratio:0.8411352606831958;width:528px;height:auto" srcset="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Figure_3_CARE-X_NoLogo.png 2013w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Figure_3_CARE-X_NoLogo-255x300.png 255w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Figure_3_CARE-X_NoLogo-870x1024.png 870w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Figure_3_CARE-X_NoLogo-768x904.png 768w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Figure_3_CARE-X_NoLogo-1305x1536.png 1305w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Figure_3_CARE-X_NoLogo-1740x2048.png 1740w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Figure_3_CARE-X_NoLogo-153x180.png 153w" sizes="auto, (max-width: 2013px) 100vw, 2013px" /><figcaption class="wp-element-caption">Figure 3: ReXVQA accuracy across five findings-quality dimensions — negation, presence, location, differential diagnosis, geometric information, and overall. CARE-X consistently outperforms CheXOne-R1 and MedGemma on every axis, with the largest margins in differential diagnosis, location assessment and negation.</figcaption></figure>874875876877<hr class="wp-block-separator has-alpha-channel-opacity"/>878879880881<h2 id="tool-augmented-measurement-interleaving-perception-and-computation" class="wp-block-heading">Tool-augmented measurement: Interleaving perception and computation</h2>882883884885<p class="wp-block-paragraph">Some radiological findings depend on quantitative measurements rather than visual patterns. In a separate research experiment from CARE-X, we built an inference-time pipeline that combines Qwen3-VL-4B-Instruct with deterministic measurement tools, allowing the model to alternate between image understanding and precise computation. Qwen3-VL-4B-Instruct retains visual access to the radiograph throughout inference, invoking tools to identify anatomical landmarks, compute measurements, and evaluate diagnostic thresholds as needed. This creates a multi-turn reasoning loop that interleaves perception and measurement, enabling the model to combine visual context with exact quantitative evidence before reaching a diagnosis.</p>886887888889<figure class="wp-block-image aligncenter size-full"><img loading="lazy" decoding="async" width="2000" height="836" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Figure_4_CARE-X_updated.png" alt="Diagram illustrating a workflow for a medical assistant using a vision-language model (VLM) to analyze chest X-ray images and provide diagnostic metrics. Key components include orchestrator handling prompts and tool calls, assistant performing perception and tool calls to measure cardiac and thoracic widths, and synthesizing diagnosis with visual overlays and calculated cardiothoracic ratio (CTR) displayed in red and blue." class="wp-image-1181268" srcset="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Figure_4_CARE-X_updated.png 2000w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Figure_4_CARE-X_updated-300x125.png 300w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Figure_4_CARE-X_updated-1024x428.png 1024w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Figure_4_CARE-X_updated-768x321.png 768w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Figure_4_CARE-X_updated-1536x642.png 1536w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Figure_4_CARE-X_updated-240x100.png 240w" sizes="auto, (max-width: 2000px) 100vw, 2000px" /><figcaption class="wp-element-caption"><em>Figure 4. Tool-augmented quantitative reasoning pipeline. The orchestrator mediates a multi-turn loop: the VLM reasons over the image (perception), emits structured tool calls, receives deterministic results, and synthesizes the final diagnosis.</em></figcaption></figure>890891892893<p class="wp-block-paragraph">Despite requiring no task-specific training, this approach substantially outperforms perception-only inference across all evaluated measurement-based conditions. The results suggest that for threshold-dependent diagnoses, direct computation of clinically defined measurements is more reliable than visual approximation alone.</p>894895896897<p class="wp-block-paragraph">More broadly, this measurement-augmented approach could augment clinical workflows by expanding the set of quantitative assessments routinely derived from chest radiographs. For example, aortic dilation is not typically quantified on CXR and is often detected only incidentally on CT scans obtained for other indications. As delayed detection can contribute to adverse cardiovascular outcomes, reliable CXR-based screening could enable earlier identification and follow-up of aortic dilation.</p>898899900901<figure class="wp-block-table"><table class="has-fixed-layout"><thead><tr><th>Condition</th><th>Perception F1</th><th>Tool F1</th><th>Δ F1</th></tr></thead><tbody><tr><td>Cardiomegaly</td><td>74.56</td><td>96.00</td><td>+21.4</td></tr><tr><td>Mediastinal Widening</td><td>72.63</td><td>97.47</td><td>+24.8</td></tr><tr><td>Aortic Knob Enlargement</td><td>60.31</td><td>99.76</td><td>+39.5</td></tr><tr><td>Ascending Aorta Enlargement</td><td>39.33</td><td>100.00</td><td>+60.7</td></tr><tr><td>Descending Aorta Enlargement†</td><td>28.57</td><td>100.00</td><td>+71.4</td></tr><tr><td><strong>Average</strong></td><td></td><td></td><td><strong>+43.6</strong></td></tr></tbody></table><figcaption class="wp-element-caption">Table 3: Perception-only versus tool-augmented measurement. The average F1 improvement is 43.6 percentage points across five conditions.</figcaption></figure>902903904905<hr class="wp-block-separator has-alpha-channel-opacity"/>906907908909<h2 id="validation-on-indian-clinical-data-rare-icu-conditions-and-ct-confirmed-enlargement" class="wp-block-heading">Validation on Indian clinical data: Rare ICU conditions and CT-confirmed enlargement</h2>910911912913<p class="wp-block-paragraph"><strong>Research ethics and data use</strong>: The Narayana Health evaluations used de-identified, retrospective clinical data under applicable institutional ethics review and data-use approvals. Narayana Health approved publication of the study results described here. </p>914915916917<h3 id="study-1-inpatient-and-icu-conditions" class="wp-block-heading">Study 1: Inpatient and ICU conditions</h3>918919920921<p class="wp-block-paragraph">To assess real-world generalizability in a research setting, we evaluated CARE-X on 1,047 de-identified chest radiographs from Narayana Health, annotated for five rare, high-acuity conditions with prevalence ranging from 2.6% to 5.2%—reflecting realistic clinical distributions where missed diagnoses carry severe consequences. </p>922923924925<figure class="wp-block-table"><table class="has-fixed-layout"><thead><tr><th></th><th>Fracture</th><th>Mediastinal Shift</th><th>Pneumoperitoneum</th><th>Pneumothorax</th><th>Tubes & Lines Abnormal Placement</th></tr></thead><tbody><tr><td><strong>Model</strong></td><td>Sens / Spec</td><td>Sens / Spec</td><td>Sens / Spec</td><td>Sens / Spec</td><td>Sens / Spec</td></tr><tr><td>CheXOne</td><td>0.41 / 0.90</td><td>0.80 / 0.78</td><td>0.67 / 0.98</td><td><strong>0.85</strong> / 0.72</td><td>0.03 / 0.97</td></tr><tr><td>MedGemma</td><td>0.05 / 1.00</td><td><strong>1.00</strong> / 0.53</td><td>0.00 / 1.00</td><td>0.52 / 0.73</td><td>0.18 / 0.87</td></tr><tr><td><strong>CARE-X</strong></td><td><strong>0.62</strong> / 0.64</td><td>0.83 / 0.86</td><td><strong>0.89</strong> / 0.94</td><td>0.83 / 0.75</td><td><strong>0.66</strong> / 0.77</td></tr></tbody></table><figcaption class="wp-element-caption">Table 4: ICU pathology classification on Indian hospital data. CARE-X achieves the most balanced performance.</figcaption></figure>926927928929<p class="wp-block-paragraph">CARE-X achieves the highest sensitivity in three out of five conditions while maintaining reasonable specificity, demonstrating generalization to low-prevalence clinical settings.</p>930931932933<h3 id="study-2-ct-confirmed-enlargement-conditions" class="wp-block-heading">Study 2: CT-confirmed enlargement conditions</h3>934935936937<p class="wp-block-paragraph">In a retrospective study to measure pure recall efficacy, we evaluated measurement-dependent conditions such as mediastinal widening findings including aortic enlargement, hilar mass, and pulmonary artery enlargement on a outpatient cohort of 122 positive cases with CT-confirmed ground truth, avoiding the subjectivity of radiologist consensus on borderline enlargement findings on CXR. In the overlay setting, the VLM receives the original radiograph alongside a second image with condition-relevant anatomical segmentation masks — offering spatial guidance without direct access to measurement tools.</p>938939940941<p class="wp-block-paragraph">The tool-augmented variant reached 94.26% recall, a +10.65 percentage-point gain over the best perception-only baseline. Where CT or echocardiography access is limited, reliable triage from a widely available modality like chest X-ray can cut both unnecessary referrals and missed diagnoses.</p>942943944945<figure class="wp-block-image aligncenter size-full is-resized"><img loading="lazy" decoding="async" width="1750" height="894" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Figure_5_CARE-X_AH-1.jpg" alt="chart" class="wp-image-1181055" style="aspect-ratio:1.9579849071996738;width:701px;height:auto" srcset="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Figure_5_CARE-X_AH-1.jpg 1750w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Figure_5_CARE-X_AH-1-300x153.jpg 300w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Figure_5_CARE-X_AH-1-1024x523.jpg 1024w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Figure_5_CARE-X_AH-1-768x392.jpg 768w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Figure_5_CARE-X_AH-1-1536x785.jpg 1536w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Figure_5_CARE-X_AH-1-240x123.jpg 240w" sizes="auto, (max-width: 1750px) 100vw, 1750px" /><figcaption class="wp-element-caption">Figure 5: Recall on the CT-confirmed enlargement cohort across perception-only, overlay-assisted, and tool-augmented inference. (Study 2)</figcaption></figure>946947948949<p class="wp-block-paragraph">In a related study (accepted at EACTS conference 2026), for mild aortic dilation, the measurement-driven reasoning approach detected 40 of 43 CT-confirmed cases (93% sensitivity), compared to just 5 of 43 (12%) identified on the initial radiology reads, where aortic enlargement is usually not the primary indication for the chest X-ray. This corresponds to 35 additional mild cases that were surfaced but missed during the initial CXR interpretation. These results suggest that explicit quantitative measurements may help identify borderline enlargement that is difficult to assess through visual inspection alone. </p>950951952953<h2 id="what-this-does-and-doesn-t-show" class="wp-block-heading">What this does and doesn’t show </h2>954955956957<p class="wp-block-paragraph">These numbers are all recall, i.e., how many true positives we catch. This was the focus of the initial study because, in triage, a missed diagnosis is typically the costlier failure mode, and CT-confirmed ground truth gave us a clean way to measure it without relying on radiologist consensus for the difficult cases. </p>958959960961<p class="wp-block-paragraph">Recall, however, captures only one dimension of diagnostic performance. A model that flags everything achieves perfect recall and is useless in practice. An extended study is underway that includes CT-confirmed negative cohorts as well. Preliminary results are promising, and further studies are planned to explicitly evaluate the viability of quantitative aortic measurements on chest X-ray as a screening tool for aortic dilation. </p>962963964965<hr class="wp-block-separator has-alpha-channel-opacity"/>966967968969<h2 id="care-x-toward-clinically-useful-radiology-ai" class="wp-block-heading">CARE-X: Toward clinically useful radiology AI</h2>970971972973<p class="wp-block-paragraph">CARE-X demonstrates that discriminative and generative objectives can be effectively combined within a unified radiology AI model. By jointly training classification, grounding, and language capabilities, the model supports both flexible report generation and calibrated, threshold-adjustable predictions. The separate measurement study further highlights a practical division of labor between learned reasoning and deterministic computation: the VLM provides visual understanding and identifies relevant evidence, while measurement-dependent diagnoses are computed through transparent, tool-based calculations. Retrospective evaluation on clinically challenging Narayana Health cohorts provides encouraging evidence of the potential of this approach for real-world radiology applications. The clinical relevance of this research is underscored by the selection of the AI-based aortic dilatation screening application as a finalist for showcase at the <strong>IHF Innovation Hub, World Hospital Congress 2026</strong>, recognizing its potential to support earlier detection and clinical decision-making in cardiovascular care. </p>974975976977<p class="wp-block-paragraph">Looking ahead, CARE-X can be extended beyond its current capabilities through structured report generation, richer differential diagnosis support, and tighter integration of tools within the model itself. The framework could also benefit from incorporating broader clinical context, including laboratory results and patient history, enabling more comprehensive clinical reasoning. </p>978979980981<hr class="wp-block-separator has-alpha-channel-opacity"/>982983984985<p class="wp-block-paragraph"><em><em>CARE-X is a research model, not a Microsoft product offering or medical device. It has not been cleared or approved by any regulatory authority and is not intended or validated for clinical diagnosis, screening, patient care, or clinical decision-making. The results described are retrospective research findings and do not establish safety, effectiveness, or suitability for clinical use. </em> </em></p>986987988989<p class="wp-block-paragraph">Paper co-authors:</p>990991992993<p class="wp-block-paragraph"><a href="https://www.microsoft.com/en-us/research/people/meranjit/"><em>Mercy Ranjit</em></a>, <a class="msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall" href="https://www.linkedin.com/in/anirban-porya22/" target="_blank" rel="noopener noreferrer"><em>Anirban Porya</em><span class="sr-only"> (opens in new tab)</span></a>, <a class="msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall" href="https://www.linkedin.com/in/niharika-vadlamudi/" target="_blank" rel="noopener noreferrer"><em>Niharika Vadlamudi</em><span class="sr-only"> (opens in new tab)</span></a>, <a class="msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall" href="https://www.linkedin.com/in/nikhilesh-e-484b21255/" target="_blank" rel="noopener noreferrer"><em>Nikhilesh E</em><span class="sr-only"> (opens in new tab)</span></a>, <em><a class="msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall" href="https://www.linkedin.com/in/sathvik-joel-97524b18b/" type="link" id="https://www.linkedin.com/in/sathvik-joel-97524b18b/" target="_blank" rel="noopener noreferrer">Sathvik Joel<span class="sr-only"> (opens in new tab)</span></a>, </em><a class="msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall" href="https://www.linkedin.com/in/prasanthvv/" target="_blank" rel="noopener noreferrer"><em>Prasanth V V</em><span class="sr-only"> (opens in new tab)</span></a>, <a href="https://www.microsoft.com/en-us/research/people/taganu/"><em>Tanuja Ganu</em></a>, <a class="msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall" href="https://www.linkedin.com/in/abhyuday-swamy-105173166/" target="_blank" rel="noopener noreferrer"><em>Abhyuday Swamy</em><span class="sr-only"> (opens in new tab)</span></a>, <a class="msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall" href="https://www.linkedin.com/in/pranay97/" target="_blank" rel="noopener noreferrer"><em>Pranay Umredkar</em><span class="sr-only"> (opens in new tab)</span></a><em>, </em><a class="msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall" href="https://www.linkedin.com/in/dr-pradeep-narayan-9500101b8/" target="_blank" rel="noopener noreferrer"><em>Pradeep Narayan</em><span class="sr-only"> (opens in new tab)</span></a><em>, </em><a class="msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall" href="https://www.linkedin.com/in/vivek-rajagopal/" target="_blank" rel="noopener noreferrer"><em>Vivek Rajagopal</em><span class="sr-only"> (opens in new tab)</span></a></p>994995996997<p class="wp-block-paragraph">Collaborators: <a class="msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall" rel="noopener noreferrer" target="_blank" href="https://www.medha-analytics.ai/">Medha AI<span class="sr-only"> (opens in new tab)</span></a>, <a class="msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall" href="https://www.narayanahealth.org/" type="link" id="https://www.narayanahealth.org/" target="_blank" rel="noopener noreferrer">Narayana Health<span class="sr-only"> (opens in new tab)</span></a></p>99899910001001<p class="wp-block-paragraph"></p>1002<span id="label-external-link" class="sr-only" aria-hidden="true">Opens in a new tab</span><p>The post <a href="https://www.microsoft.com/en-us/research/blog/introducing-care-x-towards-clinically-useful-radiology-vlms-with-auxiliary-supervision-reward-aligned-learning-and-tool-augmented-measurement/">Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement</a> appeared first on <a href="https://www.microsoft.com/en-us/research">Microsoft Research</a>.</p>1003]]></content:encoded>1004 1005 1006 1007 </item>1008 <item>1009 <title>Orchard: An open framework for scalable agentic AI</title>1010 <link>https://www.microsoft.com/en-us/research/blog/orchard-an-open-framework-for-scalable-agentic-ai/</link>1011 1012 <dc:creator><![CDATA[Baolin Peng, Wenlin Yao, Qianhui Wu, Hao Cheng, Jianfeng Gao]]></dc:creator>1013 <pubDate>Mon, 03 Aug 2026 16:00:00 +0000</pubDate>1014 <category><![CDATA[Research Blog]]></category>1015 <guid isPermaLink="false"></guid>10161017 <description><![CDATA[<p>Orchard is an open-source framework for the research community to train and evaluate AI agents across task types. It reduces complexity while supporting strong performance from smaller models by enabling researchers to reuse the same infrastructure.</p>1018<p>The post <a href="https://www.microsoft.com/en-us/research/blog/orchard-an-open-framework-for-scalable-agentic-ai/">Orchard: An open framework for scalable agentic AI</a> appeared first on <a href="https://www.microsoft.com/en-us/research">Microsoft Research</a>.</p>1019]]></description>1020 <content:encoded><![CDATA[1021<figure class="wp-block-image size-full"><img loading="lazy" decoding="async" width="1400" height="788" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Orchard-BlogHeroFeature-1400x788_NEW.jpg" alt="Three Orchard framework components with benchmark results" class="wp-image-1180738" srcset="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Orchard-BlogHeroFeature-1400x788_NEW.jpg 1400w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Orchard-BlogHeroFeature-1400x788_NEW-300x169.jpg 300w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Orchard-BlogHeroFeature-1400x788_NEW-1024x576.jpg 1024w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Orchard-BlogHeroFeature-1400x788_NEW-768x432.jpg 768w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Orchard-BlogHeroFeature-1400x788_NEW-1066x600.jpg 1066w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Orchard-BlogHeroFeature-1400x788_NEW-655x368.jpg 655w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Orchard-BlogHeroFeature-1400x788_NEW-240x135.jpg 240w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Orchard-BlogHeroFeature-1400x788_NEW-640x360.jpg 640w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Orchard-BlogHeroFeature-1400x788_NEW-960x540.jpg 960w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/Orchard-BlogHeroFeature-1400x788_NEW-1280x720.jpg 1280w" sizes="auto, (max-width: 1400px) 100vw, 1400px" /></figure>1022102310241025<div style="padding-bottom:0;padding-top:0" class="wp-block-msr-immersive-section alignfull row">1026 1027 <div class="container">1028 <div class="wp-block-msr-immersive-section__inner wp-block-msr-immersive-section__inner--narrow">1029 <div class="wp-block-columns mb-10 pb-1 pr-1 is-layout-flex wp-container-core-columns-is-layout-8f761849 wp-block-columns-is-layout-flex" style="box-shadow:var(--wp--preset--shadow--outlined)">1030<div class="wp-block-column is-layout-flow wp-block-column-is-layout-flow">1031<h2 id="at-a-glance" class="wp-block-heading h3">At a glance</h2>1032103310341035<ul class="wp-block-list">1036<li>Orchard is an open-source framework for scalable and cost-effective agentic AI research, built around Orchard Env, a reusable environment service for training and evaluating agents across task domains.</li>1037103810391040<li>The same Orchard infrastructure supports software-engineering, web-navigation, and personal-assistant agents, and can train them directly inside real deployment harnesses such as Codex, OpenClaw, and ZeroClaw—letting researchers reuse environments, data pipelines, and evaluation workflows across tasks. </li>1041104210431044<li>Orchard-SWE, Orchard-GUI, and Orchard-Claw demonstrate that relatively small open-weight models can achieve strong results on complex real-world tasks. For example, Orchard-SWE reaches 69.7% on SWE-bench Verified—73.0% with value-model reranking—using only about 3 billion active parameters, approaching frontier systems using more than 10 times larger models. </li>1045</ul>1046</div>1047</div> </div>1048 </div>10491050 </div>1051105210531054<p class="wp-block-paragraph">Alongside the models and workflows, the project releases training data and evaluation methods intended to help the broader research community build and study open agentic systems. Artificial intelligence is rapidly moving beyond static question-answering toward autonomous agents that can plan, reason, and act across complex, multistep environments. These systems can fix bugs in complex codebases, navigate the web on a user’s behalf, and manage workflows involving calendars and email. </p>1055105610571058<p class="wp-block-paragraph">While there is excitement around agentic AI’s capabilities, the research community faces a persistent bottleneck. Building state-of-the-art agentic systems often requires proprietary infrastructure, including custom sandboxes, closed training pipelines, and proprietary datasets that most researchers and practitioners cannot access or reproduce.</p>1059106010611062<p class="wp-block-paragraph">To address this gap, we introduce <a class="msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall" href="https://github.com/microsoft/Orchard" target="_blank" rel="noopener noreferrer">Orchard<span class="sr-only"> (opens in new tab)</span></a>, an open-source framework for scalable agentic modeling. At the center of Orchard is Orchard Env, a lightweight, Kubernetes environment that provides reusable isolated components for running and building agents at scale—from collecting training data to reinforcement learning rollouts and evaluation. </p>1063106410651066<p class="wp-block-paragraph">Unlike many existing frameworks, Orchard Env is designed to support different agent systems and task types without modification. The same service can support software-engineering agents, web-browsing agents, and personal-assistant agents across domains. </p>1067106810691070<p class="wp-block-paragraph">To demonstrate this approach, we are releasing three domain-specific training recipes—<a class="msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall" href="https://huggingface.co/datasets/microsoft/Orchard" target="_blank" rel="noopener noreferrer">Orchard-SWE, Orchard-GUI, and Orchard-Claw.<span class="sr-only"> (opens in new tab)</span></a> We are also releasing the training data and evaluation methods used to build them.</p>1071107210731074 <div class="border-bottom border-top border-gray-300 mt-5 mb-5 msr-promo text-center text-md-left alignwide" data-bi-aN="promo" data-bi-id="1144027">1075 10761077 <p class="msr-promo__label text-gray-800 text-center text-uppercase">1078 <span class="px-4 bg-white display-inline-block font-weight-semibold small">PODCAST SERIES</span>1079 </p>1080 1081 <div class="row pt-3 pb-4 align-items-center">1082 <div class="msr-promo__media col-12 col-md-5">1083 <a class="bg-gray-300 display-block" href="https://www.microsoft.com/en-us/research/story/ai-testing-and-evaluation-learnings-from-science-and-industry/" aria-label="AI Testing and Evaluation: Learnings from Science and Industry" data-bi-cn="AI Testing and Evaluation: Learnings from Science and Industry" target="_blank">1084 <img decoding="async" class="w-100 display-block" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2025/06/EP2-AI-TE_Hero_Feature_River_No_Text_1400x788.jpg" alt="Illustrated headshots of Daniel Carpenter, Timo Minssen, Chad Atalla, and Kathleen Sullivan for the Microsoft Research Podcast" />1085 </a>1086 </div>1087 1088 <div class="msr-promo__content p-3 px-5 col-12 col-md">10891090 <h2 class="h4">AI Testing and Evaluation: Learnings from Science and Industry</h2>1091 1092 <p id="ai-testing-and-evaluation-learnings-from-science-and-industry" class="large">Discover how Microsoft is learning from other domains to advance evaluation and testing as a pillar of AI governance.</p>1093 1094 <div class="wp-block-buttons justify-content-center justify-content-md-start">1095 <div class="wp-block-button">1096 <a href="https://www.microsoft.com/en-us/research/story/ai-testing-and-evaluation-learnings-from-science-and-industry/" aria-describedby="ai-testing-and-evaluation-learnings-from-science-and-industry" class="btn btn-brand glyph-append glyph-append-chevron-right" data-bi-cn="AI Testing and Evaluation: Learnings from Science and Industry" target="_blank">1097 Listen now </a>1098 </div>1099 </div>1100 </div><!--/.msr-promo__content-->1101 </div><!--/.msr-promo__inner-wrap-->1102<span id="label-external-link" class="sr-only" aria-hidden="true">Opens in a new tab</span> </div><!--/.msr-promo-->1103 110411051106<h2 id="environment-layer-that-scales-across-types-of-tasks" class="wp-block-heading">Environment layer that scales across types of tasks</h2>1107110811091110<p class="wp-block-paragraph">The central idea behind Orchard is that the runtime environment should be a standalone, reusable service rather than infrastructure embedded inside a specific training framework. Orchard Env’s Kubernetes foundation enables it to create, manage, and remove thousands of isolated components in parallel.</p>1111111211131114<p class="wp-block-paragraph">The system is designed to work across tasks like coding, web browsing, using tools. It is also designed to work across different agent systems, along with stages of the training and evaluation process, including data distillation and reinforcement learning rollouts.</p>1115111611171118<p class="wp-block-paragraph">This flexibility makes Orchard practical at a research scale. Teams can introduce new benchmarks, agent systems, or training algorithms without rebuilding the underlying infrastructure from scratch.</p>1119112011211122<p class="wp-block-paragraph">Orchard also makes it possible to train agents inside any harness. Today’s most capable agents rarely run as a bare model. They operate through sophisticated harnesses—such as Claude Code, Codex, and OpenClaw—that manage multi-turn reasoning, tool use, and connections to external systems. Open training tools usually cannot handle these stateful, multi-process harnesses, forcing researchers to train on a simplified stand-in and then deploy in the real setting, which creates a mismatch. Orchard closes this gap: a lightweight proxy records the harness’s own model calls as training data while each rollout runs in its own container, so an agent can be trained end-to-end directly in the harness that it will be deployed with—OpenClaw, Codex, ZeroClaw, or others—and across several harnesses.</p>1123112411251126<h2 id="orchard-swe-advancing-open-source-software-engineering-agents" class="wp-block-heading">Orchard-SWE: Advancing open-source software engineering agents</h2>1127112811291130<p class="wp-block-paragraph">Software engineering is one of the most demanding settings for autonomous agents. It requires multi-step reasoning over real codebases, tool use, and the ability to recover from mistakes. Orchard-SWE is our training workflow for this domain. It is built using the Mini-SWE-Agent framework, designed to autonomously solve software engineering tasks, and evaluated on the widely used SWE-bench Verified benchmark, which tests a model’s ability to navigate, diagnose, and repair real-world codebases.</p>1131113211331134<p class="wp-block-paragraph">To train the system, we distilled 107,000 agent interactions from two advanced open-weight models (MiniMax-M2.5 and Qwen3.5-397B) covering a broad range of GitHub Issues. The training process uses credit-assignment supervised fine-tuning: rather than discarding attempts where the agent failed to fully resolve an issue, the system learns from the productive portions of those partial attempts, expanding the amount of useful training data available to the model.</p>1135113611371138<p class="wp-block-paragraph">Reinforcement learning comes next, but its feedback is sparse—an agent usually learns only whether its final patch passed or failed the hidden tests. We start with Balanced Adaptive Rollout, designed to make the most of these infrequent success signals, and then add two “dense reward” techniques for richer guidance: on-policy distillation, in which a stronger teacher model scores the agent’s decisions step by step, and a process reward model, in which an AI judge rewards sound problem-solving process—writing tests that reproduce the bug, verifying the fix, and checking that existing behavior still works—independent of whether the final tests passed. </p>1139114011411142<p class="wp-block-paragraph">Finally, we train a value model on past rollouts to rerank candidate solutions. Reinforcement learning generates many practice trajectories that are normally discarded; instead, trajectories from 20 prior experiments train a compact 4-billion-parameter value model that recognizes high-quality solutions, and at problem-solving time it scores several candidate answers and picks the best one. Together, these techniques take Orchard-SWE from a 61.4% baseline on SWE-bench Verified to 69.1% with Balanced Adaptive Rollout and 69.7% with the dense-reward techniques—a new state of the art among open-source models of comparable size (roughly 3 billion active parameters)—rising to 73% with value-model reranking, approaching frontier systems more than 10 times larger, as shown in Figure 1. </p>1143114411451146<h2 id="orchard-gui-a-lightweight-browser-agent-for-real-world-web-tasks" class="wp-block-heading">Orchard-GUI: A lightweight browser agent for real-world web tasks</h2>1147114811491150<p class="wp-block-paragraph">Web navigation presents a different set of challenges. Agents must interpret visual layouts, interact with dynamic interfaces, and complete open-ended tasks described only in natural language.</p>1151115211531154<p class="wp-block-paragraph">Orchard-GUI trains a 4-billion-parameter vision-language model as a browser agent using a relatively small amount of supervision: 400 distilled demonstrations combined with 2,200 open-ended training tasks. Despite this limited training data, the resulting model achieves strong results across several web-navigation benchmarks: 74.1% on WebVoyager, 67.0% on Online-Mind2Web, and 64.0% on DeepShop, for an average of 68.4%, as shown in Figure 1.</p>1155115611571158<figure class="wp-block-image aligncenter size-full"><img loading="lazy" decoding="async" width="1638" height="790" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/ORCHARD_Fig1_AH2.jpg" alt="On the left: Orchard-SWE (30B-A3B) reaches 67.5% on SWE-bench Verified, matching frontier systems 10—30x larger. On the right: Orchard-GUI (4B) achieves 68.4% average success across WebVoyager, Online-Mind2web, and DeepShop, making it the strongest open-source GUI agent while staying on par with proprietary systems from OpenAI and Google." class="wp-image-1184449" srcset="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/ORCHARD_Fig1_AH2.jpg 1638w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/ORCHARD_Fig1_AH2-300x145.jpg 300w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/ORCHARD_Fig1_AH2-1024x494.jpg 1024w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/ORCHARD_Fig1_AH2-768x370.jpg 768w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/ORCHARD_Fig1_AH2-1536x741.jpg 1536w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/ORCHARD_Fig1_AH2-240x116.jpg 240w" sizes="auto, (max-width: 1638px) 100vw, 1638px" /><figcaption class="wp-element-caption">Figure 1. Performance comparison. Left: Orchard-SWE (35B-A3B, ~3B active) reaches 69.7% on SWE-bench Verified—73% with value-model reranking—matching frontier systems more than 10x larger. Right: Orchard-GUI (4B) achieves 68.4% average success across WebVoyager, Online-Mind2web, and DeepShop, making it the strongest open-source GUI agent while staying on par with proprietary systems from OpenAI and Google.</figcaption></figure>1159116011611162<p class="wp-block-paragraph">These results place Orchard-GUI among the strongest open-source web agents to date while remaining competitive with larger proprietary models. The results also suggest that with the right training approach and environment, small open models can perform well on real-world web tasks.</p>1163116411651166<h2 id="orchard-claw-personal-assistant-agents-for-everyday-productivity" class="wp-block-heading">Orchard-Claw: Personal assistant agents for everyday productivity</h2>1167116811691170<p class="wp-block-paragraph">Many of the most impactful agentic applications involve everyday productivity tasks, including reading and drafting emails, managing calendars, searching for information, and coordinating across tools. Orchard-Claw focuses on personal-assistant tasks by training an agent on just 200 synthetic tasks. Evaluated on Claw-Eval, a benchmark covering realistic productivity workflows, it successfully completes 59.6% of tasks when given up to three attempts. That increases to 73.9% when paired with the stronger ZeroClaw agent system.</p>1171117211731174<p class="wp-block-paragraph">Because Orchard can train agents directly inside real deployment harnesses, Orchard-Claw is trained across several of them—including ReACT, ZeroClaw, OpenClaw, and Codex—rather than a single simplified loop. Training inside these real harnesses substantially improves the agent’s reliability; under the Codex harness, for example, its success rate rises from 18.6% for the untrained model to 51.5% after Orchard training. </p>1175117611771178<figure class="wp-block-image aligncenter size-full"><img loading="lazy" decoding="async" width="1032" height="554" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/FIG2_ORCHARD_NEW.png" alt="Diagram of the Orchard ecosystem showing three benchmark areas (Orchard‑SWE, Orchard‑GUI, Orchard‑Claw) with performance metrics, a modular training pipeline (data curation, curriculum design, SFT, RL, evaluation), and the core Orchard Env service enabling sandboxed execution, file I/O, networking, and APIs. The system supports heterogeneous environments (code, web, desktop, mobile, productivity tools) through a unified interface, emphasizing reusability across domains and efficiency features such as low latency, Kubernetes scaling, and reduced cost." class="wp-image-1180734" srcset="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/FIG2_ORCHARD_NEW.png 1032w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/FIG2_ORCHARD_NEW-300x161.png 300w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/FIG2_ORCHARD_NEW-1024x550.png 1024w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/FIG2_ORCHARD_NEW-768x412.png 768w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/FIG2_ORCHARD_NEW-710x380.png 710w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/FIG2_ORCHARD_NEW-240x129.png 240w" sizes="auto, (max-width: 1032px) 100vw, 1032px" /><figcaption class="wp-element-caption">Figure 2. Overview of the Orchard framework. Orchard Env (center) is a lightweight, Kubernetes-native environment service that provides shared capabilities such as sandbox management, command execution, file access, network controls, a REST API, and agent integration. It supports a range of task environments (bottom row) and is used across three task domains (top row): Orchard-SWE (software engineering), Orchard-GUI (browser navigation), and Orchard-Claw (AI personal assistant).</figcaption></figure>1179118011811182<h2 id="implications-and-the-road-ahead" class="wp-block-heading">Implications and the road ahead</h2>1183118411851186<p class="wp-block-paragraph">Orchard’s results reinforce a broader point: the environment layer matters. By making the underlying infrastructure open, lightweight, and reusable, Orchard lowers the cost of agentic AI research. Teams no longer need to build custom isolated environments from scratch or depend on proprietary cloud services. The same Orchard Env can be used to generate training data, run reinforcement learning rollouts, and evaluate final models without rebuilding the system each time.</p>1187118811891190<p class="wp-block-paragraph">Looking ahead, we see reusing training experience as a promising direction toward cumulative agent learning. Instead of discarding trajectories once a training run finishes, we treat them as persistent assets—for example, distilling them into reusable value models. This enables agentic experience to accumulate over time, allowing each new generation of agents to inherit and extend the knowledge acquired by previous ones, rather than starting from scratch. </p>1191119211931194<p class="wp-block-paragraph">The data efficiency demonstrated by Orchard-GUI suggests that larger-scale web agents could be trained without requiring large amounts of manually created training data. By releasing the complete Orchard stack, including the environment service, training pipelines, and training datasets, we hope to help the broader research community build more capable open agents more quickly. </p>1195119611971198<p class="wp-block-paragraph"><strong>Acknowledgements</strong></p>1199120012011202<p class="wp-block-paragraph">We thank the teams at Microsoft Research and collaborating institutions for their contributions to Orchard, as well as the open-source community whose benchmarks and tools made this research possible.</p>1203<span id="label-external-link" class="sr-only" aria-hidden="true">Opens in a new tab</span><p>The post <a href="https://www.microsoft.com/en-us/research/blog/orchard-an-open-framework-for-scalable-agentic-ai/">Orchard: An open framework for scalable agentic AI</a> appeared first on <a href="https://www.microsoft.com/en-us/research">Microsoft Research</a>.</p>1204]]></content:encoded>1205 1206 1207 1208 </item>1209 <item>1210 <title>Echoverse: Deep, evolving environments for computer-use agents</title>1211 <link>https://www.microsoft.com/en-us/research/blog/echoverse-deep-evolving-environments-for-computer-use-agents/</link>1212 1213 <dc:creator><![CDATA[Akshay Nambi, Yash Pandya, Sahil Gupta, Sarthak Harne, Kavyansh Chourasia, Yash Lara, Ahmed Awadallah]]></dc:creator>1214 <pubDate>Thu, 30 Jul 2026 17:00:00 +0000</pubDate>1215 <category><![CDATA[Research Blog]]></category>1216 <guid isPermaLink="false">https://www.microsoft.com/en-us/research/?p=1179908</guid>12171218 <description><![CDATA[<p>Computer-use AI agents struggle with multi-step workflows like email and customer support. Echoverse trains agents in realistic environments rather than simply providing more training tasks, helping them improve as the tasks, tests, and environments evolve.</p>1219<p>The post <a href="https://www.microsoft.com/en-us/research/blog/echoverse-deep-evolving-environments-for-computer-use-agents/">Echoverse: Deep, evolving environments for computer-use agents</a> appeared first on <a href="https://www.microsoft.com/en-us/research">Microsoft Research</a>.</p>1220]]></description>1221 <content:encoded><![CDATA[1222<h2 id="scaling-fidelity-over-sheer-count-targeting-the-capabilities-agents-actually-lack-and-evolving-with-the-models-they-train" class="wp-block-heading h3">Scaling fidelity over sheer count, targeting the capabilities agents actually lack, and evolving with the models they train.</h2>1223122412251226<figure class="wp-block-image size-full"><img loading="lazy" decoding="async" width="1400" height="788" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/PraxisWorld-BlogHeroFeature-1400x788-1.jpg" alt="Diagram of an iterative training loop where a model generates a world, the world produces a graded run, and feedback updates both the world and the model." class="wp-image-1180136" srcset="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/PraxisWorld-BlogHeroFeature-1400x788-1.jpg 1400w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/PraxisWorld-BlogHeroFeature-1400x788-1-300x169.jpg 300w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/PraxisWorld-BlogHeroFeature-1400x788-1-1024x576.jpg 1024w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/PraxisWorld-BlogHeroFeature-1400x788-1-768x432.jpg 768w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/PraxisWorld-BlogHeroFeature-1400x788-1-1066x600.jpg 1066w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/PraxisWorld-BlogHeroFeature-1400x788-1-655x368.jpg 655w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/PraxisWorld-BlogHeroFeature-1400x788-1-240x135.jpg 240w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/PraxisWorld-BlogHeroFeature-1400x788-1-640x360.jpg 640w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/PraxisWorld-BlogHeroFeature-1400x788-1-960x540.jpg 960w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/PraxisWorld-BlogHeroFeature-1400x788-1-1280x720.jpg 1280w" sizes="auto, (max-width: 1400px) 100vw, 1400px" /></figure>1227122812291230<div style="padding-bottom:0;padding-top:0" class="wp-block-msr-immersive-section alignfull row">1231 1232 <div class="container">1233 <div class="wp-block-msr-immersive-section__inner wp-block-msr-immersive-section__inner--narrow">1234 <div class="wp-block-columns mb-10 pb-1 pr-1 is-layout-flex wp-container-core-columns-is-layout-8f761849 wp-block-columns-is-layout-flex" style="box-shadow:var(--wp--preset--shadow--outlined)">1235<div class="wp-block-column is-layout-flow wp-block-column-is-layout-flow">1236<h2 id="at-a-glance" class="wp-block-heading h3">At a glance</h2>1237123812391240<p class="wp-block-paragraph">We built twelve training worlds for computer-use agents: ten deep domain worlds and two capability worlds, each drilling a single control rendered in many forms (date pickers and nested filters). Depth is what makes them worth training on: these worlds reproduce an application’s real behavior, come seeded with realistic data, and keep state coherent across screens and users. Trained on all twelve, a 9B model nearly doubles its base score (36.5% to 67.1%), coming within fourteen points of GPT-5.4. The experiment taught us several lessons: </p>1241124212431244<ul class="wp-block-list">1245<li><strong>High simulation fidelity is a must-have; shallow worlds hurt the agent.</strong> Trained on shallow and deep builds of the same sites, the model regressed on the shallow ones but improved on the deep ones.</li>1246124712481249<li><strong>Agents often struggle with the same challenging UI elements, like date pickers and nested filters.</strong> Drilling those controls in varied forms taught the model to operate them in domains it never saw in training.</li>1250125112521253<li><strong>Co-evolving the model, the world, and the verifier improves all of them.</strong> As the world grows more correct and its tasks grow harder, the model climbs with it.</li>1254125512561257<li><strong>Reinforcement learning against the worlds pushes the agent past imitation.</strong> Using the grounded verifier as the reward, RL lifts held-out performance and teaches the agent to reach the goal in fewer steps.</li>1258125912601261<li>We’re releasing four of the worlds with their code, data, and grounded graders, to support research on high-fidelity computer-use worlds. <br>Github: <a class="msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall" rel="noopener noreferrer" target="_blank" href="https://nam06.safelinks.protection.outlook.com/?url=https%3A%2F%2Fgithub.com%2Fmicrosoft%2FEchoverse&data=05%7C02%7Cv-amablack%40microsoft.com%7C8672491a049240e2d0cb08deecf33ea2%7C72f988bf86f141af91ab2d7cd011db47%7C1%7C0%7C639208726637682797%7CUnknown%7CTWFpbGZsb3d8eyJFbXB0eU1hcGkiOnRydWUsIlYiOiIwLjAuMDAwMCIsIlAiOiJXaW4zMiIsIkFOIjoiTWFpbCIsIldUIjoyfQ%3D%3D%7C0%7C%7C%7C&sdata=4%2F8FD0oWevHazXLVAGUwE0XnoLqVfLDC2g5TPUB94V0%3D&reserved=0"><u>microsoft/Echoverse: Deep, Evolving Environments for Computer-Use Agents</u><span class="sr-only"> (opens in new tab)</span></a> <br>Hugging Face: <a class="msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall" rel="noopener noreferrer" target="_blank" href="https://nam06.safelinks.protection.outlook.com/?url=https%3A%2F%2Fhuggingface.co%2Fdatasets%2Fmicrosoft%2FEchoverse&data=05%7C02%7Cv-amablack%40microsoft.com%7C8672491a049240e2d0cb08deecf33ea2%7C72f988bf86f141af91ab2d7cd011db47%7C1%7C0%7C639208726637691838%7CUnknown%7CTWFpbGZsb3d8eyJFbXB0eU1hcGkiOnRydWUsIlYiOiIwLjAuMDAwMCIsIlAiOiJXaW4zMiIsIkFOIjoiTWFpbCIsIldUIjoyfQ%3D%3D%7C0%7C%7C%7C&sdata=TkzGRz7uZIS97yPjvNP7a9Kp1BhmOR2WZOOxnv1GOgA%3D&reserved=0"><u>microsoft/Echoverse · Datasets at Hugging Face</u><span class="sr-only"> (opens in new tab)</span></a><br>Technical Report: <a href="https://www.microsoft.com/en-us/research/publication/echoverse-deep-evolving-environments-for-training-computer-use-agents-at-scale/" target="_blank" rel="noreferrer noopener">https://www.microsoft.com/en-us/research/publication/echoverse-deep-evolving-environments-for-training-computer-use-agents-at-scale/</a></li>1262</ul>1263</div>1264</div> </div>1265 </div>12661267 </div>1268126912701271<p class="wp-block-paragraph">A computer-use agent learns the results of what its actions do only where they have real consequences. A click changes saved state, a message reaches a real person, or a page that refuses to move tells the agent its last move did nothing. A screenshot can show what an interface looks like, but only a working world shows what an action caused.</p>1272127312741275<figure class="wp-block-video aligncenter"><video height="1200" style="aspect-ratio: 3840 / 1200;" width="3840" controls src="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/Hero-video.mp4"></video></figure>1276127712781279<p class="wp-block-paragraph">The consequences worth learning from are stateful, and most of them sit behind a login. The work people want automated lives in closed systems: email and chat, banking, health records, the internal consoles for cloud and ML. You cannot train an agent against the live versions of these. Every attempt writes to a real account, there is no reset between tries, and the true state stays hidden behind the screen. So you rebuild the system as a synthetic world where the database is yours: the state is real and changes for real, but it is safe to break, quick to reset, and graded from the data rather than a screenshot.</p>1280128112821283<p class="wp-block-paragraph">By a world we mean three things bound together: an environment (the application, its state, and the actions that change it), the tasks that set goals in it, and a verifier that grades the outcome against ground truth. The community is now good at making them: pipelines stand up an application, seed it, generate tasks, and attach verifiers, yielding hundreds of environments and thousands of checkable tasks. This work builds on that progress. However, once worlds are plentiful and its internal structure becomes the bottleneck: regardless of whether state stays coherent across users and screens, workflows keep their dependencies, a weak skill recurs in enough forms to generalize, and success is judged by outcome or by appearance.</p>1284128512861287<p class="wp-block-paragraph">Our bet, the one <strong>Echoverse </strong>tests, is that the real leverage comes less from adding worlds than from a loop that keeps improving the ones you already have. It treats building the environment and training the model as one process, not two stages: run a model in a world, find where it fails, make the world, its tasks, and its verifiers more faithful or more demanding there, train on the sharper signal, and repeat. Ordinary fine-tuning improves only the model. Here the same graded run that measures the model also improves the world that judged it, so a static benchmark saturates while the loop compounds.</p>1288128912901291<p class="wp-block-paragraph">Three levers keep that loop productive, none of them raw environment count. <strong>Depth</strong>: behaviorally faithful worlds for the domains that matter, including the closed and proprietary ones. <strong>Capability targeting</strong>: narrow worlds built around the exact interaction a model keeps failing. <strong>Co-evolution</strong>: improving the environment, its tasks, and its verifiers on every graded run, not just the model.</p>1292129312941295<figure class="wp-block-image aligncenter size-full"><img loading="lazy" decoding="async" width="2560" height="931" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-01-learning-loop-scaled.png" alt="Circular diagram of the learning loop. A model runs a task in a world and every rollout is graded against database ground truth. Two arrows branch from the graded run: surviving failures flow to model training data, while defects flow to repairs of the environment, its tasks, and its verifier, so model and world improve on the same run. " class="wp-image-1179929" srcset="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-01-learning-loop-scaled.png 2560w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-01-learning-loop-300x109.png 300w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-01-learning-loop-1024x372.png 1024w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-01-learning-loop-768x279.png 768w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-01-learning-loop-1536x558.png 1536w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-01-learning-loop-2048x745.png 2048w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-01-learning-loop-240x87.png 240w" sizes="auto, (max-width: 2560px) 100vw, 2560px" /><figcaption class="wp-element-caption">Figure 1: The learning loop: every graded run is read twice. Surviving failures become model training data, and defects in the world, its tasks, or its verifier become repairs. The same graded run that measures the model also sharpens the world.</figcaption></figure>1296129712981299<h2 id="why-synthetic-and-why-deep" class="wp-block-heading">Why synthetic, and why deep?</h2>1300130113021303<p class="wp-block-paragraph">Open, login-free sites might seem to remove the need for synthetic worlds, but they make a poor training ground for a different reason: they will not hold still. Pages get redesigned, listings and dates roll forward, and hosts throttle or block automated traffic, so a benchmark that is pinned to them drifts, and no two runs face the same site. An occasional evaluation can absorb that; training cannot, since it runs the same task thousands of times and needs the same world each time. A synthetic world is fixed in time and data: the calendar does not move, the seed data does not churn, and a task means the same thing on the thousandth rollout as on the first. We trade a little surface realism for a world we fully control.</p>1304130513061307<p class="wp-block-paragraph">Control is only the floor. A world can be perfectly stable and still be hollow, so what earns training time is depth: not its page count but how faithfully it preserves the causal structure of the work. Five properties set the bar: <strong>behavioral fidelity</strong> (controls, permissions, and errors follow the product’s logic); <strong>coherent state</strong> (a sent message appears for its recipient, a cancelled meeting clears both calendars); <strong>workflow depth</strong> (an early choice constrains what happens later); <strong>authoritative verification</strong> (application state, not pixels); and <strong>domain value</strong> (the workflow is worth improving). In the systems that matter most, the difficulty lives in permissions, shared state, and audit histories: exactly the structure a shallow clone skips. Above this bar, more environments add variety; below it, they add noise.</p>1308130913101311<h2 id="how-the-echoverse-factory-works" class="wp-block-heading">How the Echoverse factory works?</h2>1312131313141315<p class="wp-block-paragraph">Echoverse is a single pipeline with two outputs: full domain worlds that preserve workflow depth, and capability worlds that vary one diagnosed interaction. Both lean on the fact that we own the database underneath, so success is a property of the app’s own state, not a model’s read of a screenshot.</p>1316131713181319<h3 id="building-the-world" class="wp-block-heading">Building the world</h3>1320132113221323<p class="wp-block-paragraph">The pipeline expands a handful of seed scenarios into a spec, then compiles it into machine-checkable claims about routes, state, and behavior. Only then does it generate the app: a FastAPI and SQLite backend under a React interface. A fresh app is a hypothesis, not a world: the builder runs every claim against the running environment, repairing the database, backend, or frontend until each passes, then writes a readiness record that separates hard blockers from advisory risks. A world with open blockers does not advance. Depth here is not a promise in a prompt; it is the list of claims the world has been shown to pass.</p>1324132513261327<h3 id="growing-the-corpus" class="wp-block-heading">Growing the corpus</h3>1328132913301331<p class="wp-block-paragraph">A world that builds cleanly is still not training data. We reground each task on the live database, drawing goals from entities that actually exist, then send every goal through a panel of analyzers: are its entities real, is the goal plausible, does its difficulty match the work, and, the sharpest test, can an agent driving the real UI complete it? That last check runs in the browser, catching goals no interface can satisfy before a model ever sees them. A generated goal is a claim; a solve against the real app is proof.</p>1332133313341335<p class="wp-block-paragraph">Every failure becomes an issue tagged by the layer that must change: database, backend, frontend, task text, or verifier. Layer-specific fixers apply the repair, re-check it against the running app, and roll it back if it regresses. The loop re-scores against database ground truth until the pass rate stops climbing, and each surviving task is exported carrying the exact check that grades it. Those tasks become training data through one process: GPT-5.4 solves each task, a verifier keeps the trajectories that pass ground truth, and those become the supervised fine-tuning (SFT) data behind every experiment below.</p>1336133713381339<p class="wp-block-paragraph">Building the world and growing the corpus are not two stages but rather one loop: most defects belong to the world, so we re-version the environment with every iteration. Harder tasks expose gaps in the world, and a sturdier world can carry harder tasks, so each round leaves both stronger.</p>1340134113421343<figure class="wp-block-image aligncenter size-full"><img loading="lazy" decoding="async" width="2560" height="970" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-02-pipeline-scaled.png" alt="Two-phase pipeline diagram. Phase 1 expands seed scenarios into an app and repairs its database, backend, and frontend until it passes machine-checkable claims. Phase 2 regrounds tasks on live data and iterates a loop of analyzer and fixer agents, then re-scores against database ground truth until the pass rate plateaus. A dashed arrow shows many task-loop fixes landing back in the world." class="wp-image-1179930" srcset="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-02-pipeline-scaled.png 2560w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-02-pipeline-300x114.png 300w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-02-pipeline-1024x388.png 1024w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-02-pipeline-768x291.png 768w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-02-pipeline-1536x582.png 1536w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-02-pipeline-2048x776.png 2048w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-02-pipeline-240x91.png 240w" sizes="auto, (max-width: 2560px) 100vw, 2560px" /><figcaption class="wp-element-caption">Figure 2: The environment factory: the two loops behind every world. Phase 1 expands a handful of seeds into an app, then repairs the database, backend, and frontend until it passes machine-checkable claims. Phase 2 regrounds tasks on live data, runs a panel of analyzer and layer-specific fixer agents, and re-scores against database ground truth until the pass rate plateaus. Many of those fixes land in the world itself (dashed arrow). </figcaption></figure>1344134513461347<h3 id="the-verifier-is-grounded-in-the-database" class="wp-block-heading">The verifier is grounded in the database</h3>1348134913501351<p class="wp-block-paragraph">Every task carries its own answer key, a value or a state change minted from the real database by a SQL query at generation, true by construction and re-checked after the agent finishes. A <em>read</em> is graded on semantic equivalence to the stored value ($288 for $287.62 passes); a <em>write</em> on a real before/after database diff, so claiming a ticket was closed fails unless the row flipped; a <em>read_write</em> scores the lower of the two. Grading is hard to game, grounded rather than labelled, and uniform across an EchoStay booking, an EchoForge issue, and an EchoBank transfer.</p>1352135313541355<h3 id="full-domains-carry-the-workflow" class="wp-block-heading">Full domains carry the workflow</h3>1356135713581359<p class="wp-block-paragraph">The domains with the most consequential work are the hardest for public benchmarks to reach: closed, proprietary systems where the difficulty lives in permissions, shared state, and history, not layout. A faithful clone has to reproduce that. What matters is not the pixels but that an action’s consequences reach across screens and users, so a task can run a real workflow and be graded on the state it leaves behind.</p>1360136113621363<p class="wp-block-paragraph">The ten Echo domains span communication, technical work, regulated records, community, media, and travel. Where a rich public dataset exists we build on it: EchoStay is seeded from InsideAirbnb, so its listings, hosts, reviews, and amenities are real rather than invented, and EchoForum sits on a public forum corpus of 2.55 million comments. Where none exists, as with mail, calendar, banking, and health records, a seeding pipeline generates the state under strict constraints, dense and internally consistent, not a handful of placeholder rows.</p>1364136513661367<figure class="wp-block-table aligncenter"><table class="has-fixed-layout"><thead><tr><th>Workflow category</th><th>Environments</th><th>Depth the world has to carry</th></tr></thead><tbody><tr><td><strong>Communication & coordination</strong></td><td>EchoMail, EchoCalendar, EchoChat</td><td>Shared threads, schedules, participants, permissions, histories</td></tr><tr><td><strong>Technical creation & operations</strong></td><td>EchoML, EchoForge</td><td>Artifacts, configuration, dependencies, roles, multi-stage changes</td></tr><tr><td><strong>Regulated records & transactions</strong></td><td>EchoBank, EchoCare</td><td>Balances or records, authorization, audit history, consequential writes</td></tr><tr><td><strong>Community, media & travel</strong></td><td>EchoForum, EchoTunes, EchoStay</td><td>Persistent preferences, social state, search, booking, account actions</td></tr></tbody></table><figcaption class="wp-element-caption">Table 1: The ten full-domain environments of the Echo family, grouped by the work they represent. Each is a faithful stand-in for a widely used product, named for the workflow rather than the brand.</figcaption></figure>1368136913701371<p class="wp-block-paragraph">That accumulated state is what makes an action’s consequences reach across screens and users. A booking in EchoStay moves through search, listing, availability, and payment across roughly 87 routes and 23 tables, but not a single confirmation screen; an EchoMail thread carries intent from draft through delivery, reply, and label state; an EchoCare order writes each change to an audit trail. The tasks are expensive because of it, often five to twenty actions deep, and finished only when the underlying state has changed.</p>1372137313741375<figure class="wp-block-image aligncenter size-full"><img loading="lazy" decoding="async" width="2560" height="2123" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-03-domain-cards-scaled.png" alt="Grid of per-domain cards for the Echo suite. Each card names an environment and lists grounded database counts for its backend, seeded data, and feature surface, showing each is a self-contained interactive clone rather than a mockup. " class="wp-image-1179931" srcset="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-03-domain-cards-scaled.png 2560w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-03-domain-cards-300x249.png 300w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-03-domain-cards-1024x849.png 1024w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-03-domain-cards-768x637.png 768w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-03-domain-cards-1536x1274.png 1536w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-03-domain-cards-2048x1698.png 2048w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-03-domain-cards-217x180.png 217w" sizes="auto, (max-width: 2560px) 100vw, 2560px" /><figcaption class="wp-element-caption">Figure 3: Per-domain detail across the Echo suite. Each ships as a self-contained, fully-interactive clone of the app it models, with its own backend, seeded database, and feature surface. Counts are grounded database state, not mockups.</figcaption></figure>1376137713781379<h3 id="capability-worlds-isolate-one-skill" class="wp-block-heading">Capability worlds isolate one skill</h3>1380138113821383<p class="wp-block-paragraph">Not every weakness represents a missing domain; some are caused by a single control that the agent cannot reliably operate. Picture an agent booking a trip: it searches, filters, opens the right listing, then stalls at the date picker, unable to turn “the second week of March” into the right clicks on an unfamiliar calendar. Building another booking site would not fix that. The skill is learned only when the control itself appears in enough forms, and date pickers and nested filter-and-search are ubiquitous on the live web, rendered a hundred different ways, exactly the variability a single deep app cannot supply.</p>1384138513861387<p class="wp-block-paragraph">So we isolate the control and widen the interaction, mass-producing it across layouts, states, and constraints, then generating grounded tasks over each. The datepicker world renders one date control as six core widgets across 10 contexts and holds out 10 new unseen ones, from calendar heatmaps to scroll wheels and fiscal-quarter pickers; its hardest tasks turn transcription into reasoning, resolving “the last Thursday of January 2026” or “10 business days after a start date” to one exact, widget-reachable date. The nested-filter world varies 20 widget families and holds out nine compound-panel families as out-of-distribution, grading every submission by whether the filtered results actually meet the requested conditions, judged by the app’s own logic rather than by appearance.</p>1388138913901391<figure class="wp-block-image aligncenter size-full"><img loading="lazy" decoding="async" width="2560" height="1793" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-04-widget-catalog-scaled.png" alt="Plain-English catalog of the widget families the capability worlds render. Each entry is a distinct rendering of the same control (a date picker or a nested filter) re-themed across real-world verticals, with several families marked held out for evaluation only. " class="wp-image-1179932" srcset="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-04-widget-catalog-scaled.png 2560w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-04-widget-catalog-300x210.png 300w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-04-widget-catalog-1024x717.png 1024w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-04-widget-catalog-768x538.png 768w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-04-widget-catalog-1536x1076.png 1536w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-04-widget-catalog-2048x1434.png 2048w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-04-widget-catalog-240x168.png 240w" sizes="auto, (max-width: 2560px) 100vw, 2560px" /><figcaption class="wp-element-caption">Figure 4: Every widget family the two skills cover, split into training (in-distribution) and evaluation-only (held out): nested filters, 20 families plus 9 held-out compound panels; date pickers, 6 core types across 10 contexts plus 10 held-out widgets.</figcaption></figure>1392139313941395<figure class="wp-block-image aligncenter size-full"><img loading="lazy" decoding="async" width="2560" height="699" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-05-domain-coverage-scaled.png" alt="Diagram showing the capability controls re-themed across many domains: nested filters across six verticals and date pickers across ten everyday contexts, with the held-out sets reaching 36 further scenarios. " class="wp-image-1179933" srcset="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-05-domain-coverage-scaled.png 2560w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-05-domain-coverage-300x82.png 300w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-05-domain-coverage-1024x280.png 1024w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-05-domain-coverage-768x210.png 768w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-05-domain-coverage-1536x419.png 1536w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-05-domain-coverage-2048x559.png 2048w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-05-domain-coverage-240x66.png 240w" sizes="auto, (max-width: 2560px) 100vw, 2560px" /><figcaption class="wp-element-caption">Figure 5: Date pickers and nested filters themed across domains: nested filters over six verticals, from real estate to pet adoption; date pickers over ten contexts, from scheduling to insurance.</figcaption></figure>1396139713981399<h2 id="what-deeper-targeted-worlds-change" class="wp-block-heading">What deeper, targeted worlds change</h2>1400140114021403<p class="wp-block-paragraph">More trajectories do not automatically provide more training signal. What matters is depth: whether an episode carries a task through the dependent steps of a real workflow rather than just rehearsing an action in isolation. Two experiments make the difference concrete from opposite ends: one goes deeper on a whole domain, the other narrows to a single broken skill.</p>1404140514061407<h3 id="shallow-worlds-backfire-deep-worlds-transfer" class="wp-block-heading">Shallow worlds backfire; deep worlds transfer</h3>1408140914101411<p class="wp-block-paragraph">A shallow world is the cheap option. It stands up fast and looks convincing, but it only rehearses isolated, correct-looking clicks. Train on that and the model will pick up the wrong reflexes, over-stepping and looping and repeating dead actions, because nothing in the easy world ever punished them. A deep world costs more, but its trajectories carry the dependent structure that transfers to the live site.</p>1412141314141415<p class="wp-block-paragraph">To isolate that, take two live WebVoyager domains, Allrecipes and Hugging Face, and compare three checkpoints: the base model and two trained on shallow-world and deep-world trajectories built for those domains. The shallow world poses short, self-contained tasks; the deep world poses tasks that run across dependent steps, where an early action changes the state, options, and verification available later. Both give the model the same domain exposure, so only depth differs, and evaluation uses tasks from the public WebVoyager benchmark for these domains, run on the live sites outside any training world.</p>1416141714181419<p class="wp-block-paragraph">On Allrecipes, the shallow world pulls the model down, 80.0% to 75.0%; on Hugging Face it stays flat at 48.0%. Only the deep world improves both, lifting Allrecipes to 85.0% and the harder Hugging Face split to 65.0%. With exposure held equal, the gap is depth: the deep model loops less, and of the 37 Hugging Face tasks, those that exhaust their step budget fall from 15 to nine. What separated the two was not how much the model saw, but whether what it saw preserved the structure of the work.</p>1420142114221423<figure class="wp-block-image aligncenter size-full"><img loading="lazy" decoding="async" width="2560" height="955" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-06-deep-vs-shallow-scaled.png" alt="Grouped bar chart on two live WebVoyager domains, Allrecipes and Hugging Face, comparing base, shallow-world-trained, and deep-world-trained models. Shallow drops Allrecipes from 80.0% to 75.0% and leaves Hugging Face flat at 48.0%; the deep world lifts them to 85.0% and 65.0%. " class="wp-image-1179935" srcset="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-06-deep-vs-shallow-scaled.png 2560w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-06-deep-vs-shallow-300x112.png 300w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-06-deep-vs-shallow-1024x382.png 1024w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-06-deep-vs-shallow-768x287.png 768w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-06-deep-vs-shallow-1536x573.png 1536w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-06-deep-vs-shallow-2048x764.png 2048w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-06-deep-vs-shallow-240x90.png 240w" sizes="auto, (max-width: 2560px) 100vw, 2560px" /><figcaption class="wp-element-caption">Figure 6: Deep versus shallow worlds for two live WebVoyager domains, with identical domain exposure and different task depth. Deep lifts both; shallow drops below base on Allrecipes and stalls on Hugging Face.</figcaption></figure>1424142514261427<h3 id="precision-about-one-skill" class="wp-block-heading">Precision about one skill</h3>1428142914301431<p class="wp-block-paragraph">The datepicker and nested-filter worlds drill exactly the controls our evaluations flagged, and the two skills reinforce each other. Datepicker training lifts datepicker evaluations (in-distribution 60.0% to 82.6%, held-out layouts 34.0% to 54.0%); filter training lifts held-out filters 62.8% to 84.1%. Gains that hold on forms never trained on indicate that the model learned a rule, not a layout. The skills transfer across each other rather than competing: training either one alone still lifts the other, and training both is the best all-rounder on every split. Against GPT-5.4 as a frontier reference, that combined model already edges ahead on nested filters and closes most of the datepicker in-distribution gap, trailing clearly only on held-out datepickers. And the rule reaches the open web, lifting Online-Mind2Web 29.5% to 34.3% on sites it never saw. </p>1432143314341435<figure class="wp-block-image aligncenter size-full"><img loading="lazy" decoding="async" width="2560" height="1045" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-07-targeted-training_NEW-scaled.png" alt="Grouped bar chart across four capability splits (datepicker in-distribution and held-out, nested-filter in-distribution and held-out) comparing base, plus-datepicker, plus-nested-filter, and plus-both models. Training either skill lifts both controls, and training both is the best all-rounder on every split. " class="wp-image-1180128" srcset="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-07-targeted-training_NEW-scaled.png 2560w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-07-targeted-training_NEW-300x122.png 300w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-07-targeted-training_NEW-1024x418.png 1024w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-07-targeted-training_NEW-768x313.png 768w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-07-targeted-training_NEW-1536x627.png 1536w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-07-targeted-training_NEW-2048x836.png 2048w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-07-targeted-training_NEW-240x98.png 240w" sizes="auto, (max-width: 2560px) 100vw, 2560px" /><figcaption class="wp-element-caption">Figure 7: Targeted training, targeted gains: training either date pickers or nested filters lifts both controls, including held-out widgets and compositions neither was trained on, and training both is the best all-rounder on every split. Higher is better.</figcaption></figure>1436143714381439<h2 id="from-synthetic-worlds-to-the-live-web" class="wp-block-heading">From synthetic worlds to the live web</h2>1440144114421443<p class="wp-block-paragraph">Three models run through the rest of this section. Base is Qwen3.5-9B given only a handful of synthetic trajectories, just enough to align a general model to the browser action space. Our model is that same 9-billion-parameter network trained on the full synthetic corpus. GPT-5.4 is a far larger frontier model, included as a reference ceiling.</p>1444144514461447<p class="wp-block-paragraph">Does the skill survive the open web? We evaluate our model, unchanged, on WebVoyager and Online-Mind2Web, benchmarks it never trained on. They barely overlap with what we built: both are dominated by open, public sites and read-mostly browsing, while our worlds train login-gated, write-heavy workflows. A large jump was never the point; direction is. The frozen model clears base on both, WebVoyager 66.5% to 71.5% and Online-Mind2Web 40.5% to 43.4% (without BrowserBase, 50.9% to 55.6% and 29.5% to 37.2%), reported through BrowserBase because a hosted browser strips the datacenter bot-blocks and rate limits that otherwise depress every agent’s score. With no live-web data in the mix, this is transfer, not memorization.</p>1448144914501451<figure class="wp-block-image aligncenter size-full"><img loading="lazy" decoding="async" width="2560" height="954" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-08-live-transfer-scaled.png" alt="Bar chart on two live benchmarks scored through BrowserBase. The full-corpus model beats base on WebVoyager (66.5% to 71.5%) and Online-Mind2Web (40.5% to 43.4%), showing synthetic training transfers to sites it never trained on." class="wp-image-1179937" srcset="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-08-live-transfer-scaled.png 2560w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-08-live-transfer-300x112.png 300w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-08-live-transfer-1024x382.png 1024w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-08-live-transfer-768x286.png 768w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-08-live-transfer-1536x572.png 1536w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-08-live-transfer-2048x763.png 2048w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-08-live-transfer-240x89.png 240w" sizes="auto, (max-width: 2560px) 100vw, 2560px" /><figcaption class="wp-element-caption">Figure 8: Synthetic training transfers to the live web. The full-corpus model, on two benchmarks it never trained on, clears base on both; scores run through BrowserBase to remove datacenter bot-blocks.</figcaption></figure>1452145314541455<p class="wp-block-paragraph">The modest live-web gain is a coverage effect, not a ceiling: aim at a live domain and it grows. EchoForge, our code-hosting world, is the same kind of app as GitHub, one of the live sites WebVoyager tests. Add EchoForge to the training mix and the live GitHub score climbs 58.5% to 63.4%, with the overall live scores rising too (WebVoyager 50.9% to 52.9%, Online-Mind2Web 29.5% to 31.1%). The average simply reflects that most of what we built sits in domains these benchmarks never touch.</p>1456145714581459<figure class="wp-block-image aligncenter size-full"><img loading="lazy" decoding="async" width="2560" height="1442" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-09-scorecard-gap_NEW-scaled.png" alt="Dumbbell chart, per environment, of closing the gap to the frontier. For each of fourteen domains a grey dot marks base, a green dot our full-corpus model, and an amber diamond GPT-5.4; the green bar is the gain from base and the faded remainder is the distance still to GPT-5.4. A right-hand strip lists each model's exact Base, Our, and GPT score. Our model nearly doubles the base average to 67.1% and its green dot sits past the diamond on EchoBank and both nested filters, surpassing GPT-5.4. " class="wp-image-1180130" srcset="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-09-scorecard-gap_NEW-scaled.png 2560w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-09-scorecard-gap_NEW-300x169.png 300w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-09-scorecard-gap_NEW-1024x577.png 1024w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-09-scorecard-gap_NEW-768x432.png 768w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-09-scorecard-gap_NEW-1536x865.png 1536w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-09-scorecard-gap_NEW-2048x1153.png 2048w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-09-scorecard-gap_NEW-1066x600.png 1066w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-09-scorecard-gap_NEW-655x368.png 655w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-09-scorecard-gap_NEW-240x135.png 240w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-09-scorecard-gap_NEW-640x360.png 640w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-09-scorecard-gap_NEW-960x540.png 960w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-09-scorecard-gap_NEW-1280x720.png 1280w" sizes="auto, (max-width: 2560px) 100vw, 2560px" /><figcaption class="wp-element-caption">Figure 9: Closing the gap to the frontier, per environment. The green bar is the gain from base to our model; the faded remainder is the distance still to GPT-5.4. Our model surpasses GPT-5.4 on EchoBank and both nested filters and closes most of the gap elsewhere; each model’s exact score is labelled on the right.</figcaption></figure>1460146114621463<p class="wp-block-paragraph">The domains we built, most of them closed and login-gated, tell the opposite story. Across all fourteen, the model nearly doubles base, 36.5% to 67.1%, and where base was weakest it climbs three- to nine-fold, with EchoCalendar, EchoML, EchoChat, EchoCare, EchoForge, and EchoForum all moving from single or low double digits into the forties through sixties. That puts a 9-billion-parameter model within fourteen points of GPT-5.4 on the average (67.1% against 80.7%). On EchoMail, EchoBank, and both nested filters, it matches or beats the far larger frontier model outright, trailing by only a few points on in-distribution datepickers. What gets a 9B model this close is not scale but training data that is deep, targeted, and checkable, exactly what the factory is built to produce. </p>1464146514661467<h3 id="what-scaling-buys-and-what-it-doesn-t" class="wp-block-heading">What scaling buys, and what it doesn’t</h3>1468146914701471<p class="wp-block-paragraph">We scaled two axes separately: more trajectories through a fixed set of environments, drawn in equal numbers from each, and more distinct environments. They behave differently. More trajectories on the same worlds keep lifting the in-domain average, though the gains keep shrinking, while transfer to the live web flattens outright: from 6,400 to 20,000 trajectories, WebVoyager holds steady (54.8% to 55.6%) and Online-Mind2Web slips (40.1% to 37.2%). Since every point samples the worlds equally, this is no artifact: each environment holds only so much transferable skill, and once a model has drawn it out, more rollouts mostly polish what it already does. </p>1472147314741475<p class="wp-block-paragraph">Scaling environments produces the opposite result. The average keeps climbing as breadth grows, and WebVoyager reaches its best only with the full set. For generalization, the lever is diversity, not volume. A model reaches sites it never saw by training across many kinds of work, not by seeing one kind many more times. </p>1476147714781479<p class="wp-block-paragraph">Even so, scale itself is not the lever on either axis. A large trajectory budget spent on shallow worlds, or graded against the wrong answer, moves the synthetic number and goes nowhere on the live web. What travels is inside each trajectory: depth that preserves a real workflow, targeting that drills the control an agent fails, and database-grounded grading that keeps the signal honest.</p>1480148114821483<figure class="wp-block-image aligncenter size-full"><img loading="lazy" decoding="async" width="2560" height="885" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-10-scaling_NEW-scaled.png" alt="Two line charts of scaling, each plotting average synthetic score, WebVoyager, and Online-Mind2Web. Left, more trajectories through a fixed set of worlds, with the x-axis spaced by actual trajectory count so the points bunch at low counts and stretch out toward 20,000; the curves rise steeply then flatten, WebVoyager going flat and Online-Mind2Web slipping over the final stretch while the synthetic average climbs only gently. Right, more environments from two to twelve domains, where the synthetic average and WebVoyager keep climbing with breadth. The contrast shows that diversity of environments, not sheer trajectory volume, carries skill to unseen sites. " class="wp-image-1180132" srcset="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-10-scaling_NEW-scaled.png 2560w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-10-scaling_NEW-300x104.png 300w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-10-scaling_NEW-1024x354.png 1024w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-10-scaling_NEW-768x265.png 768w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-10-scaling_NEW-1536x531.png 1536w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-10-scaling_NEW-2048x708.png 2048w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-10-scaling_NEW-240x83.png 240w" sizes="auto, (max-width: 2560px) 100vw, 2560px" /><figcaption class="wp-element-caption">Figure 10: Two scaling axes, scored without BrowserBase. Left: more trajectories on a fixed set of worlds, drawn in equal numbers from each, with the x-axis spaced by actual trajectory count. The synthetic average keeps rising, but live-web transfer saturates, WebVoyager flat and Online-Mind2Web slipping past 6,400 trajectories. Right: more environments, where breadth keeps the synthetic average and WebVoyager climbing. Diversity of environments, not trajectory volume, is what carries skill to unseen sites.</figcaption></figure>1484148514861487<h2 id="the-model-is-not-the-only-thing-that-learns" class="wp-block-heading">The model is not the only thing that learns</h2>1488148914901491<p class="wp-block-paragraph">The score an agent earns is never the model alone. It comes from a coupled stack: the agent, the environment, the task, and the verifier. A zero can mean the agent failed, or the control is broken, or the requested state is impossible, or the verifier checks the wrong thing. Reading every zero as model supervision trains on defects that should have been repaired. So, we read every graded rollout as a test of the whole stack and let the whole stack learn. The environment improves as broken controls and wiring get fixed, the tasks as goals are re-grounded and made harder, the verifier is fixed when it drifts out of sync with the data. Only failures that survive all three become model curriculum.</p>1492149314941495<p class="wp-block-paragraph">EchoStay made this visible. Its failures traced to the world, not the agent: a guest-count control silently broke booking tasks, so a correct booking could never register. Fixing it raised the share of those bookings that could be completed at all from 48% to 78%, recovering 15 of the 24 that had been blocked. The same loop finds different faults elsewhere: EchoForum needed frontend fixes and a page-load speedup, which took one failing set of 37 tasks from 0 solved to 36; EchoChat’s verifier had drifted out of sync with the data, and realigning it lifted the share of gradable tasks from 34% to 99%; EchoCare needed one state-wiring fix; EchoForge had the backend logic but no UI control to reach it.</p>1496149714981499<p class="wp-block-paragraph">As the world sharpens, the model climbs with it. Re-running the loop on EchoStay across two rounds, the model trained on its corpus more than doubles, from 16.2% to 38.5%, two-thirds of the distance to GPT-5.4’s 50.4%. The model is not the only thing that learns; it is the thing that compounds once everything under it learns.</p>1500150115021503<figure class="wp-block-image aligncenter size-full"><img loading="lazy" decoding="async" width="2560" height="1317" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-11-coevolution-scaled.png" alt="Bar chart of the model's score on EchoStay before and after one co-evolution round. As the world went from v1 to v2 the model more than doubled, from 16.2% to 38.5%, shown against GPT-5.4's 50.4% reference. " class="wp-image-1179942" srcset="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-11-coevolution-scaled.png 2560w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-11-coevolution-300x154.png 300w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-11-coevolution-1024x527.png 1024w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-11-coevolution-768x395.png 768w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-11-coevolution-1536x790.png 1536w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-11-coevolution-2048x1054.png 2048w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-11-coevolution-240x123.png 240w" sizes="auto, (max-width: 2560px) 100vw, 2560px" /><figcaption class="wp-element-caption">Figure 11: Co-evolution lifts the model on EchoStay. As the world went from v1 to v2, the model trained on it more than doubled, from 16.2% to 38.5%, a separate measure from the world’s own solve rate. Higher is better. </figcaption></figure>1504150515061507<p class="wp-block-paragraph">That boundary between repairing the world and teaching the model is easy to hold inside a controlled environment, where both are inspectable. The live web erases it: there is no world to repair mid-task, so when an action lands on nothing, correctness rests entirely on the agent noticing and choosing differently. That is the last thing a world has to teach, and where the live web is least forgiving.</p>1508150915101511<h2 id="from-sft-to-rl-turning-worlds-into-rles" class="wp-block-heading">From SFT to RL: Turning worlds into RLEs</h2>1512151315141515<p class="wp-block-paragraph">Every result so far comes from imitation: the 9B model copies the trajectories GPT-5.4 got right. Imitation inherits a ceiling, though: a clean demonstration never shows how to recover from a mistake or when to stop, the failures that break agents in the wild. Reinforcement learning optimizes the outcome we grade and lets the model learn from its own trajectories, not a teacher’s.</p>1516151715181519<p class="wp-block-paragraph">But reinforcement learning needs an RL environment (RLE) it can drive at scale. Each rollout needs a reset to a known state, throughput to sample in parallel, and a reward it can trust, and a run replays the same task thousands of times. The live web is not an RLE: it will not reset, so no two rollouts begin alike; it throttles and blocks automated traffic well before RL’s scale; and it exposes no ground truth, only a screenshot a second model must judge, so the reward is as noisy as the judge and a policy learns the judge’s blind spots rather than the task. Echoverse is an RLE by construction. Every world is a self-contained app we snapshot and reset per rollout, run in parallel, and grade from its own database, so the verifier that filtered the SFT data returns a grounded, verifiable reward rather than one inferred from pixels. The same worlds that benchmark an agent train one.</p>1520152115221523<figure class="wp-block-image aligncenter size-full"><img loading="lazy" decoding="async" width="2560" height="1171" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-12-rl-echoverse-scaled.png" alt="Left-to-right block diagram of reinforcement learning on an Echoverse RL environment. A policy pi-theta, initialised from the SFT checkpoint, rolls out a group of G trajectories inside an Echoverse RLE drawn as a stack of worlds; within one world the agent repeats act and execute steps that change a database. Outside the environment, a grader, the grounded verifier, reads each rollout's final database state and returns a reward. The group of rewards updates the policy with a policy-gradient step and a KL penalty to a reference, and the loop repeats across every training world. " class="wp-image-1179943" srcset="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-12-rl-echoverse-scaled.png 2560w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-12-rl-echoverse-300x137.png 300w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-12-rl-echoverse-1024x468.png 1024w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-12-rl-echoverse-768x351.png 768w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-12-rl-echoverse-1536x702.png 1536w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-12-rl-echoverse-2048x937.png 2048w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-12-rl-echoverse-240x110.png 240w" sizes="auto, (max-width: 2560px) 100vw, 2560px" /><figcaption class="wp-element-caption">Figure 12: Reinforcement learning on an Echoverse RLE. From the SFT policy we roll out a group of trajectories in one world; each is a sequence of act and execute steps that changes the database. A grader, the same grounded verifier that filtered the SFT data, sits outside the environment and scores each rollout’s final database state into a reward. The group of rewards updates the policy, and the loop repeats across every training world.</figcaption></figure>1524152515261527<p class="wp-block-paragraph">We take the SFT model as the starting policy and run RL against five worlds: EchoBank, EchoForge, EchoForum, EchoStay, and EchoTunes. Tasks come from the harder end of each world, where the SFT policy still leaves headroom, and each update draws on several graded rollouts. Each rollout earns two rewards: a trajectory reward from our database-grounded verifier (LLM judge GPT-4.1), and a dense per-step reward from a multimodal judge that grades each screenshot (GPT-4.1 vision). We train on roughly 100 tasks per world beyond the SFT data, for two epochs. On a held-out set of 25 tasks per world, the judged score rises from 58% to 69%. The teacher taught it what to do; the world taught it when to stop, when to recover, and when to give up.</p>1528152915301531<figure class="wp-block-image aligncenter size-full"><img loading="lazy" decoding="async" width="2560" height="849" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-13-rl-training-scaled.png" alt="Two line charts of reinforcement-learning training on five worlds. Left, the1532held-out validation judge score (25 tasks per world) rises from 58.8% to a peak of153369.6% and settles near 68% over the training steps. Right, the critic's mean score,1534the RL reward signal, with its five-step moving average trends upward from1535about 0.5 to 0.6 over sixty steps." class="wp-image-1180257" srcset="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-13-rl-training-scaled.png 2560w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-13-rl-training-300x99.png 300w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-13-rl-training-1024x340.png 1024w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-13-rl-training-768x255.png 768w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-13-rl-training-1536x509.png 1536w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-13-rl-training-2048x679.png 2048w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/fig-13-rl-training-240x80.png 240w" sizes="auto, (max-width: 2560px) 100vw, 2560px" /><figcaption class="wp-element-caption">Figure 13: Reinforcement learning on five worlds, over twoepochs. Left: the held-out judge score (25 tasks per world) climbs from 58% to 69%. Right: the critic’s mean score, the RL reward signal, trends up through training. The reward sums a trajectory reward from our database-grounded verifier (LLM judge GPT-4.1) and a dense per-step reward from a multimodal judge (GPT-4.1 vision).</figcaption></figure>1536153715381539<h2 id="where-this-leaves-us" class="wp-block-heading">Where this leaves us</h2>1540154115421543<p class="wp-block-paragraph">A world is no longer a fixed benchmark you score against; it is a training surface you keep improving, where the same graded run that measures the model also sharpens the world that judged it. Deep worlds transferred where shallow clones pulled capability down; one widget rebuilt in a hundred forms taught a skill that reached the live web; co-evolution moved both sides at once; and reinforcement against the same worlds pushed the agent past imitation, lifting held-out performance and trimming wasted steps.</p>1544154515461547<p class="wp-block-paragraph">The durable advantage is not the largest inventory of synthetic websites. It is a factory that diagnoses what an agent cannot yet do, builds or repairs the world that teaches it, protects the capability already earned, and runs the loop again. The next turns scale three fronts at once. First, more deep worlds for the closed domains public benchmarks cannot reach. Second, more capability worlds for the interactions models keep failing. And, above all, more reinforcement against those grounded worlds: longer runs, harder tasks, and wider reward exploration that push the agent’s behavior and its performance further than imitation ever could. The levers compound: deeper and broader worlds make stronger RL, stronger RL boosts the agent, and every round exposes the next capability to build.</p>1548154915501551<p class="wp-block-paragraph">We are releasing a piece of the factory: environment code and graded test tasks for four worlds, two deep domains (EchoStay and EchoForge) and two capability worlds (the datepicker and nested-filter, each with an in-distribution and a held-out split). Every task carries the database-grounded verifier that scores it, so the same worlds can benchmark an agent or train one. Code and tasks: https://aka.ms/echoverse</p>1552155315541555<p class="wp-block-paragraph">When worlds grow at the frontier of an agent’s competence, evaluation stops being a scoreboard and becomes the engine that decides what to build next: worlds that keep learning alongside the agents they train.</p>1556155715581559<h2 id="acknowledgments" class="wp-block-heading">Acknowledgments</h2>1560156115621563<p class="wp-block-paragraph">We thank Alexey Taymanov, Andrew Zhao, Aravind Rajeswaran, Corby Rosset, Hussein Mozannar, Luiz Do Valle, Sara Abdali, Spencer Whitehead, Vibhav Vineet, Zach Nussbaum, Yadong Lu, Pashmina Cameron, Rafah Hosn, and Chinmay Karkar for their valuable help, insightful discussions, and continued support throughout this work.</p>1564<span id="label-external-link" class="sr-only" aria-hidden="true">Opens in a new tab</span><p>The post <a href="https://www.microsoft.com/en-us/research/blog/echoverse-deep-evolving-environments-for-computer-use-agents/">Echoverse: Deep, evolving environments for computer-use agents</a> appeared first on <a href="https://www.microsoft.com/en-us/research">Microsoft Research</a>.</p>1565]]></content:encoded>1566 1567 1568 <enclosure url="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/Hero-video.mp4" length="84092742" type="video/mp4" />15691570 </item>1571 <item>1572 <title>EvoLib: Turning experience into evolving knowledge</title>1573 <link>https://www.microsoft.com/en-us/research/blog/evolib-turning-experience-into-evolving-knowledge/</link>1574 1575 <dc:creator><![CDATA[Weijia Xu, Zelalem Gero, Michel Galley, Eric Yuan, Jianfeng Gao]]></dc:creator>1576 <pubDate>Thu, 30 Jul 2026 16:00:00 +0000</pubDate>1577 <category><![CDATA[Research Blog]]></category>1578 <guid isPermaLink="false">https://www.microsoft.com/en-us/research/?p=1179889</guid>15791580 <description><![CDATA[<p>LLMs do not get smarter just by remembering more. EvoLib turns experience into evolving knowledge, taking reusable skills and insights that help models learn and adapt across tasks long after deployment. </p>1581<p>The post <a href="https://www.microsoft.com/en-us/research/blog/evolib-turning-experience-into-evolving-knowledge/">EvoLib: Turning experience into evolving knowledge</a> appeared first on <a href="https://www.microsoft.com/en-us/research">Microsoft Research</a>.</p>1582]]></description>1583 <content:encoded><![CDATA[1584<figure class="wp-block-image size-full"><img loading="lazy" decoding="async" width="2560" height="1441" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/EvoLib-BlogHeroFeature-1400x788-1-scaled.jpg" alt="Figure 1. EvoLib transforms raw experiences into reusable skills and insights, then continually evolves them through consolidation and dynamic weighting. " class="wp-image-1180176" srcset="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/EvoLib-BlogHeroFeature-1400x788-1-scaled.jpg 2560w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/EvoLib-BlogHeroFeature-1400x788-1-300x169.jpg 300w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/EvoLib-BlogHeroFeature-1400x788-1-1024x576.jpg 1024w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/EvoLib-BlogHeroFeature-1400x788-1-768x432.jpg 768w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/EvoLib-BlogHeroFeature-1400x788-1-1536x865.jpg 1536w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/EvoLib-BlogHeroFeature-1400x788-1-2048x1153.jpg 2048w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/EvoLib-BlogHeroFeature-1400x788-1-1066x600.jpg 1066w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/EvoLib-BlogHeroFeature-1400x788-1-655x368.jpg 655w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/EvoLib-BlogHeroFeature-1400x788-1-240x135.jpg 240w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/EvoLib-BlogHeroFeature-1400x788-1-640x360.jpg 640w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/EvoLib-BlogHeroFeature-1400x788-1-960x540.jpg 960w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/EvoLib-BlogHeroFeature-1400x788-1-1280x720.jpg 1280w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/EvoLib-BlogHeroFeature-1400x788-1-1920x1080.jpg 1920w" sizes="auto, (max-width: 2560px) 100vw, 2560px" /></figure>1585158615871588<div style="padding-bottom:0;padding-top:0" class="wp-block-msr-immersive-section alignfull row">1589 1590 <div class="container">1591 <div class="wp-block-msr-immersive-section__inner wp-block-msr-immersive-section__inner--narrow">1592 <div class="wp-block-columns mb-10 pb-1 pr-1 is-layout-flex wp-container-core-columns-is-layout-8f761849 wp-block-columns-is-layout-flex" style="box-shadow:var(--wp--preset--shadow--outlined)">1593<div class="wp-block-column is-layout-flow wp-block-column-is-layout-flow">1594<h2 id="at-a-glance" class="wp-block-heading h3">At a glance</h2>1595159615971598<ul class="wp-block-list">1599<li><strong>Self-supervised.</strong> EvoLib enables large language models to learn from their own experience during inference, without requiring ground-truth labels or external feedback.</li>1600160116021603<li><strong>From experience to knowledge.</strong> EvoLib transforms past attempts into reusable skills and reflective insights that can be applied to future tasks.</li>1604160516061607<li><strong>Knowledge that evolves.</strong> Useful skills and insights are continually refined, consolidated, and reweighted, turning instance-specific observations into increasingly general knowledge over time.</li>1608160916101611<li><strong>Learning that transfers across tasks.</strong> By turning experience into reusable knowledge, EvoLib helps AI models learn from past successes and failures and evolve the knowledge that has the highest potential on improving future performance.</li>1612161316141615<li><strong>Built for today’s AI models.</strong> As EvoLib does not require model updates, it can be applied to any black-box language models and AI systems deployed through APIs.</li>1616</ul>1617</div>1618</div> </div>1619 </div>16201621 </div>1622162316241625<p class="wp-block-paragraph">Memory has become an important AI agent capability: the ability to store and retrieve past experiences. But memory alone is not learning. A collection of past conversations, reasoning traces, or action histories can quickly grow into a vast archive of experiences, making it difficult to identify the most relevant knowledge for a new task—let alone refine and evolve this knowledge to improve performance over time.</p>1626162716281629<p class="wp-block-paragraph">Humans learn differently. We do not remember every detail of our past experiences. Instead, we remember what matters: strategies that work, mistakes to avoid, and skills that transfer across situations. Over time, these lessons are refined into increasingly general and reusable knowledge. This ability to transform experience into transferable, evolving knowledge is one of the foundations of human learning.</p>1630163116321633<p class="wp-block-paragraph">In our recent paper, <a href="https://www.microsoft.com/en-us/research/publication/test-time-learning-with-an-evolving-library/" type="link" id="https://www.microsoft.com/en-us/research/publication/test-time-learning-with-an-evolving-library/"><em>Test-Time Learning with an Evolving Library</em></a>, we explore how AI systems can learn from experience in a similar way. We introduce <strong>EvoLib</strong>, a framework that transforms raw experience into an evolving library of knowledge. Rather than treating memory as a growing archive of past experiences, EvoLib extracts reusable knowledge from those experiences and continually refines it as new experiences arrive. Through the evolution of library, skills become more general, insights become more accurate, and downstream performance gets improved consistently over time. In this way, AI agents can continually learn from accumulating experience without updating the underlying model.</p>1634163516361637<h2 id="how-evolib-works" class="wp-block-heading">How EvoLib Works</h2>1638163916401641<p class="wp-block-paragraph">Unlike traditional AI memory systems that store raw experiences as static information, EvoLib is built around the idea of <strong>evolving knowledge</strong>. In EvoLib, a unit of knowledge can take the form of a reusable skill distilled from a successful solution or a reflective insight learned from mistakes. Rather than simply accumulating more memories over time, EvoLib continually refines, consolidates and reweights existing knowledge as new experiences arrive. Concretely, we design the following mechanisms around knowledge evolution:</p>1642164316441645<ul class="wp-block-list">1646<li><strong>Consolidation.</strong> As new knowledge is extracted from recent experience, EvoLib retrieves similar knowledge from the library and tries to consolidate it with the new knowledge into a more general and reusable one. This allows knowledge to move beyond individual experiences and become applicable across tasks.</li>1647164816491650<li><strong>Weighting mechanism.</strong> EvoLib continually updates the importance of each knowledge unit based not only on its immediate utility on the current task, but also on how much it contributes to generating useful knowledge on future tasks. Over time, knowledge with the greatest long-term impact naturally becomes more prominent in the library.</li>1651</ul>1652165316541655<figure class="wp-block-image aligncenter size-full"><img loading="lazy" decoding="async" width="899" height="661" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/image.gif" alt="EvoLib transforms raw experiences into reusable skills and insights, then continually evolves them through consolidation and dynamic weighting." class="wp-image-1179891"/><figcaption class="wp-element-caption">Figure 1. EvoLib transforms raw experiences into reusable skills and insights, then continually evolves them through consolidation and dynamic weighting.</figcaption></figure>1656165716581659<h2 id="key-results" class="wp-block-heading">Key Results</h2>1660166116621663<p class="wp-block-paragraph">To evaluate EvoLib, we tested it across a diverse set of challenging tasks with different types of experiences and demands for learning:</p>1664166516661667<ul class="wp-block-list">1668<li>Solving mathematical reasoning problems</li>1669167016711672<li>Writing code to perform the given tasks under efficiency constraints</li>1673167416751676<li>Making decisions to explore and interact with an environment to perform long-horizon tasks</li>1677</ul>1678167916801681<p class="wp-block-paragraph">Across these tasks, EvoLib consistently outperforms the top retrieval-based memory approaches and other abstract memory mechanisms with more efficient token usage.</p>1682168316841685<p class="wp-block-paragraph">We also evaluated how effectively EvoLib converts test-time compute into performance gains through continually evolving knowledge. Figure 2 compares EvoLib against both compute scaling methods that perform each task in isolation and strong memory-based learning approaches. Each curve shows how performance improves as the amount of test-time compute increases.</p>1686168716881689<p class="wp-block-paragraph">Across all three benchmarks, EvoLib achieves higher performance throughout most of the compute range and improves performance more rapidly with increasing compute.</p>1690169116921693<p class="wp-block-paragraph">These results suggest that the key to better learning may not simply be storing more memories or spending more compute. Instead, the greatest gains come from transforming experience into reusable knowledge that can be continually refined and applied across tasks.</p>1694169516961697<figure class="wp-block-image aligncenter size-full"><img loading="lazy" decoding="async" width="1696" height="375" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/EvoLib_Fig2.jpg" alt="Across all tasks, EvoLib converts test-time compute into performance gains more efficiently than existing methods. " class="wp-image-1179901" srcset="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/EvoLib_Fig2.jpg 1696w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/EvoLib_Fig2-300x66.jpg 300w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/EvoLib_Fig2-1024x226.jpg 1024w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/EvoLib_Fig2-768x170.jpg 768w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/EvoLib_Fig2-1536x340.jpg 1536w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/EvoLib_Fig2-240x53.jpg 240w" sizes="auto, (max-width: 1696px) 100vw, 1696px" /><figcaption class="wp-element-caption">Figure 2. Across all tasks, EvoLib converts test-time compute into performance gains more effectively than existing methods. </figcaption></figure>1698169917001701<h2 id="robustness-to-random-task-order" class="wp-block-heading">Robustness to random task order</h2>1702170317041705<p class="wp-block-paragraph">A natural question is whether such learning depends heavily on the order in which tasks are encountered. In the real world, an AI system may face diverse types of tasks in arbitrary order, and a useful learning framework should be robust to the randomness in task order. To evaluate this, we measured the task performance on the same set of heterogeneous tasks but with different task orders. We found that EvoLib consistently improves over existing memory-based learning approaches and maintains stable performance across different orderings. This indicates that EvoLib can continually learn from diverse tasks even when they are interleaved, suggesting its practical advantage in real-world scenarios where an agent must handle and learn from a mixed stream of heterogeneous user requests without relying on a structured curriculum.</p>1706170717081709<p class="wp-block-paragraph">As AI systems take on longer-running and more complex tasks, learning from experience will become increasingly important. The future of AI may depend not only on larger models and more computation, but also on mechanisms that allow systems to continually accumulate, refine, and reuse knowledge.</p>1710171117121713<p class="wp-block-paragraph">EvoLib is one step toward that vision. By transforming experience into evolving knowledge, it enables AI systems to continually improve and adapt after deployment. Rather than repeatedly starting from scratch, future AI systems may be able to build upon an evolving library of reusable skills and insights, much like humans do.</p>1714171517161717<p class="wp-block-paragraph">Code and experiment results are available on <a class="msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall" href="https://github.com/microsoft/EvoLib" target="_blank" rel="noopener noreferrer">GitHub<span class="sr-only"> (opens in new tab)</span></a> to support future research on memory and knowledge evolution in AI systems.</p>1718<span id="label-external-link" class="sr-only" aria-hidden="true">Opens in a new tab</span><p>The post <a href="https://www.microsoft.com/en-us/research/blog/evolib-turning-experience-into-evolving-knowledge/">EvoLib: Turning experience into evolving knowledge</a> appeared first on <a href="https://www.microsoft.com/en-us/research">Microsoft Research</a>.</p>1719]]></content:encoded>1720 1721 1722 1723 </item>1724 <item>1725 <title>Verifying Rust cryptography in SymCrypt, from standards to code</title>1726 <link>https://www.microsoft.com/en-us/research/blog/verifying-rust-cryptography-in-symcrypt-from-standards-to-code/</link>1727 1728 <dc:creator><![CDATA[Son Ho, Cédric Fournet, Antoine Delignat-Lavaud, Samuel Lee, Jason Fisher, Jessica Krynitsky]]></dc:creator>1729 <pubDate>Mon, 13 Jul 2026 16:00:00 +0000</pubDate>1730 <category><![CDATA[Research Blog]]></category>1731 <guid isPermaLink="false">https://www.microsoft.com/en-us/research/blog/verifying-rust-cryptography-in-symcrypt-from-standards-to-code/</guid>17321733 <description><![CDATA[<p>Cryptographic code supports vital protections in modern computing systems. Learn how a new method helps verify code as developers write it while preserving speed and adaptability as it gets implemented and evolves.</p>1734<p>The post <a href="https://www.microsoft.com/en-us/research/blog/verifying-rust-cryptography-in-symcrypt-from-standards-to-code/">Verifying Rust cryptography in SymCrypt, from standards to code</a> appeared first on <a href="https://www.microsoft.com/en-us/research">Microsoft Research</a>.</p>1735]]></description>1736 <content:encoded><![CDATA[1737<h2 id="how-rust-lean-aeneas-and-ai-agents-are-helping-scale-formal-verification-for-production-cryptographic-algorithms" class="wp-block-heading h3">How Rust, Lean, Aeneas, and AI agents are helping scale formal verification for production cryptographic algorithms</h2>1738173917401741<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="576" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/RustSymCrypt-BlogHeroFeature-1400x788-1-1024x576.jpg" alt="Diagram showing the process of verifying cryptographic code. An algorithm from a standard is converted into a formal specification, while Rust code is converted into a code model. The specification and code model are then compared through proof and verification steps." class="wp-image-1178491" srcset="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/RustSymCrypt-BlogHeroFeature-1400x788-1-1024x576.jpg 1024w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/RustSymCrypt-BlogHeroFeature-1400x788-1-300x169.jpg 300w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/RustSymCrypt-BlogHeroFeature-1400x788-1-768x432.jpg 768w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/RustSymCrypt-BlogHeroFeature-1400x788-1-1536x865.jpg 1536w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/RustSymCrypt-BlogHeroFeature-1400x788-1-2048x1153.jpg 2048w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/RustSymCrypt-BlogHeroFeature-1400x788-1-1066x600.jpg 1066w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/RustSymCrypt-BlogHeroFeature-1400x788-1-655x368.jpg 655w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/RustSymCrypt-BlogHeroFeature-1400x788-1-240x135.jpg 240w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/RustSymCrypt-BlogHeroFeature-1400x788-1-640x360.jpg 640w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/RustSymCrypt-BlogHeroFeature-1400x788-1-960x540.jpg 960w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/RustSymCrypt-BlogHeroFeature-1400x788-1-1280x720.jpg 1280w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/RustSymCrypt-BlogHeroFeature-1400x788-1-1920x1080.jpg 1920w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /></figure>1742174317441745<div style="padding-bottom:0;padding-top:0" class="wp-block-msr-immersive-section alignfull row">1746 1747 <div class="container">1748 <div class="wp-block-msr-immersive-section__inner wp-block-msr-immersive-section__inner--narrow">1749 <div class="wp-block-columns mb-10 pb-1 pr-1 is-layout-flex wp-container-core-columns-is-layout-8f761849 wp-block-columns-is-layout-flex" style="box-shadow:var(--wp--preset--shadow--outlined)">1750<div class="wp-block-column is-layout-flow wp-block-column-is-layout-flow">1751<h2 id="at-a-glance" class="wp-block-heading h3">At a glance</h2>1752175317541755<ul class="wp-block-list">1756<li>SymCrypt develops new verified cryptography using Rust, Aeneas, and Lean to provide higher security assurance.</li>1757175817591760<li>We prove that their code safely and correctly implements standard algorithms, notably for post-quantum cryptography.</li>1761176217631764<li>We are releasing verified code, specs, properties, and proofs initially for SHA-3 and ML-KEM. </li>1765176617671768<li>Aeneas allows verifying a large subset of Rust code and provides efficient automation in Lean to support the proof effort.</li>1769177017711772<li>Agents allow scaling automation by writing proofs that are independently-verifiable.</li>1773</ul>1774</div>1775</div> </div>1776 </div>17771778 </div>1779178017811782<h2 id="introduction-and-motivation-for-formal-verification" class="wp-block-heading">Introduction and motivation for formal verification</h2>1783178417851786<p class="wp-block-paragraph">Cryptographic code sits at the foundation of modern computing. It protects operating systems, cloud services, firmware, messaging systems, and the protocols that connect them. Small mistakes can have outsized consequences: a single arithmetic slip, missing bounds check, or incorrect state transition can undermine the security of an otherwise sound design.</p>1787178817891790<p class="wp-block-paragraph">Testing and auditing remain essential, but they are not enough on their own. Cryptographic implementations are often optimized, constant-time, architecture-specific, and deliberately low level. The code that ships rarely looks like the clean algorithm in a standard: it contains reductions, bit manipulations, SIMD intrinsics, carefully shaped loops, and portability layers for many environments.</p>1791179217931794<p class="wp-block-paragraph">Formal verification addresses this gap by deploying machine-checked proofs instead of relying on testing alone. Rather than merely checking that the code usually behaves correctly, verification implements a precise mathematical specification for all inputs that satisfy the stated preconditions.</p>1795179617971798<p class="wp-block-paragraph">In June last year, Microsoft announced we would <a href="https://www.microsoft.com/en-us/research/blog/rewriting-symcrypt-in-rust-to-modernize-microsofts-cryptographic-library/">formally verify new algorithms written in Rust in SymCrypt</a>, the cryptographic provider used across products and services including Windows and Azure. New cryptographic implementations are being written in safe Rust, then verified in the <a class="msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall" href="https://lean-lang.org/" target="_blank" rel="noopener noreferrer">Lean<span class="sr-only"> (opens in new tab)</span></a> formal proof framework using the <a class="msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall" href="https://github.com/AeneasVerif/aeneas" target="_blank" rel="noopener noreferrer">Aeneas<span class="sr-only"> (opens in new tab)</span></a> toolchain. This applies in particular to post-quantum cryptography, which require fast secure implementations of complex algorithms. This combination gives us two layers of assurance: Rust rules out broad classes of memory-safety bugs, while Lean proofs establish functional correctness against formal specifications derived from standards.</p>1799180018011802<p class="wp-block-paragraph">The result is a new verification methodology for production cryptography: verify code as developers write it, preserve performance-oriented implementation choices, and make the proof process scalable enough to keep up with an evolving codebase.</p>1803180418051806<figure class="wp-block-image aligncenter size-full is-resized"><img loading="lazy" decoding="async" width="847" height="489" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/image-1.png" alt="Agents (stochastic, in blue) and tools (algorithmic, in green) for software verification. Human effort focuses on reviewing formalization of standards and main properties. Agents write proofs and intermediate properties. Compilation, code extraction, and proof verification are deterministic, not agentic." class="wp-image-1178356" style="width:667px;height:auto" srcset="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/image-1.png 847w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/image-1-300x173.png 300w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/image-1-768x443.png 768w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/image-1-240x139.png 240w" sizes="auto, (max-width: 847px) 100vw, 847px" /><figcaption class="wp-element-caption">Figure 1. Agents (stochastic, in blue) and tools (algorithmic, in green) for software verification. Human effort focuses on reviewing formalization of standards and main properties. Agents write proofs and intermediate properties. Compilation, code extraction, and proof verification are deterministic, not agentic.</figcaption></figure>1807180818091810<h2 id="status-of-verification-in-symcrypt" class="wp-block-heading">Status of verification in SymCrypt</h2>1811181218131814<p class="wp-block-paragraph">We have open sourced a <a class="msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall" href="https://github.com/microsoft/SymCrypt/tree/feature/verifiedcrypto" target="_blank" rel="noopener noreferrer">SymCrypt branch<span class="sr-only"> (opens in new tab)</span></a> that includes formal specifications and proofs. This public branch makes the proof artifacts available alongside the Rust algorithm implementations they validate, showing how the methodology applies to production cryptographic code. SymCrypt is not a standalone research prototype; it is Microsoft’s open-source cryptographic library used across products and services including Windows and Azure Linux.</p>1815181618171818<p class="wp-block-paragraph">This first release includes complete proofs for the Rust ML-KEM and SHA3 code that is being used in insiders builds of Windows today. SymCrypt is extending the same Rust, Lean, and Aeneas-based workflow to more Rust-native algorithms and integrating them into production versions for Windows and Linux, including for instance verified Rust code for, e.g., AES-GCM, FrodoKEM, and ML-DSA. The rest of this post uses this SymCrypt work as a concrete example, starting with how public standards become executable Lean specifications.</p>1819182018211822<h2 id="turning-standards-into-formal-lean-specifications" class="wp-block-heading">Turning standards into formal Lean specifications</h2>1823182418251826<p class="wp-block-paragraph">The first step is to formalize what the algorithm is supposed to do. For cryptographic primitives, the source of truth is usually a public standard: a NIST specification, an IETF RFC, or another carefully reviewed algorithm description.</p>1827182818291830<p class="wp-block-paragraph">In our approach, the Lean specification is designed to stay close to the standard. When the standard describes a loop, an array update, or a mathematical operation, the Lean model follows the same structure wherever possible. This syntactic proximity matters: it makes the formal specification easier to audit because reviewers can compare the standard and the Lean side by side.</p>1831183218331834<p class="wp-block-paragraph">Lean also lets us write executable specifications. That means we can run the formal model against official test vectors to catch transcription errors, off-by-one mistakes, or misunderstandings of the standard. For algorithms such as ML-KEM, we can go further and prove high-level mathematical properties, such as showing that the formal model of the number-theoretic transform corresponds to the intended operation over the relevant polynomial ring.</p>1835183618371838<p class="wp-block-paragraph">A representative example is the number-theoretic transform (NTT) from ML-KEM. The standard describes the algorithm as an in-place transformation over 256 coefficients modulo q, with three nested loops that update pairs of coefficients using successive powers of the constant ζ (= 17). </p>1839184018411842<script src="https://gist.github.com/fournet/49988a5ea726483f06d315d68a08c074.js"></script>1843184418451846<p class="wp-block-paragraph">Here is a direct translation of the NIST standard in Lean, trying to stick as close as possible to the original syntax:</p>1847184818491850<script src="https://gist.github.com/fournet/a1e50e3a280d59a86d43975e2aac3cff.js"></script>1851185218531854<p class="wp-block-paragraph">The Lean version deliberately mirrors the structure of the standard: the same loop nest, the same zeta selection, and the same coefficient updates, allowing easy line-by-line human review. At the same time, it is executable and uses mathematical types, so it can be tested against known vectors and connected to higher-level theorems about the NTT’s algebraic meaning. In summary, the Lean specification is a concise, executable, mathematically meaningful model that tracks the standard closely enough to be reviewed by cryptographers and proof engineers alike.</p>1855185618571858<h2 id="connecting-the-formal-specification-to-the-code" class="wp-block-heading">Connecting the formal specification to the code</h2>1859186018611862<p class="wp-block-paragraph">Once the specification is formalized, the next challenge is to connect it to the implementation. We do not ask developers to rewrite production cryptographic code in a verification-oriented language, nor do we generate code that product teams must then own. Instead, we verify the Rust code that engineers write, exactly as they write it.</p>1863186418651866<p class="wp-block-paragraph">Aeneas makes this possible by translating Rust’s mid-level representation into a pure Lean model. Rust’s ownership and borrowing discipline are crucial here. They let Aeneas safely eliminate much of the reasoning about pointer aliasing, liveness, and mutation that makes verification of C-style code so expensive.</p>1867186818691870<p class="wp-block-paragraph">For example, a Rust function that updates an array in place becomes, in Lean, a function that explicitly takes and returns a functional array. Mutable borrows are translated into value transformations. This preserves the behaviour that matters while presenting proof engineers with a functional model that is far easier to reason about.</p>1871187218731874<p class="wp-block-paragraph">Once in Lean, the function can be equipped with a theorem that states that it refines a formal specification. In other words, for every input satisfying the required bounds and well-formedness conditions, the implementation function returns the same mathematical result as the standard-derived Lean specification.</p>1875187618771878<p class="wp-block-paragraph">This style keeps responsibilities cleanly separated. Software engineers continue to write idiomatic, performant Rust. Verification engineers work against generated Lean models and prove theorems about them. The Rust code and the proofs live side by side, but the proof burden does not shape the code into something unnatural.</p>1879188018811882<p class="wp-block-paragraph">Going back to the NTT example, its Rust implementation is a function fn ntt(&mut [u16; 256]) that uses a mutable borrow to update an array in-place. The Lean translation purifies it into a function ntt : Array U16 256#usize → Result (Array U16 256#usize) that directly outputs the updated array, while wrapping it into a Result type to explicitly capture the fact that Rust functions may panic.</p>1883188418851886<p class="wp-block-paragraph">In this case, the theorem states that, if the array satisfies a well-formedness invariant (ensuring it represents a valid polynomial), then running the Rust model ntt returns the well-formed representation of the result of the mathematical specification Spec.ntt, modulo conversion from low-level arrays to high-level polynomials.</p>1887188818891890<script src="https://gist.github.com/fournet/7f67e3604972d43b0f06a493305d1897.js"></script>1891189218931894<p class="wp-block-paragraph">Scaling this to every function in real cryptographic code required substantial automation. Lean’s extensibility lets us build a gradient of automation with tactics for symbolic execution, arithmetic, arrays, and bit-vector reasoning. The experience becomes closer to debugging: automation handles the routine proof obligations, while engineers can inspect and refine the proof when a goal does not close automatically.</p>1895189618971898<h2 id="supporting-intrinsics-and-multiple-architectures" class="wp-block-heading">Supporting intrinsics and multiple architectures</h2>1899190019011902<p class="wp-block-paragraph">Production cryptography cannot ignore hardware. SymCrypt must run across environments ranging from embedded and kernel contexts to cloud services. It also needs to take advantage of platform-specific instructions when they are available, including SIMD intrinsics and architecture-specific optimized paths.</p>1903190419051906<p class="wp-block-paragraph">A verification story that only works for a portable reference implementation is therefore incomplete. We need to verify the code that actually ships: dispatch logic, optimized routines, and target-specific variants included.</p>1907190819091910<p class="wp-block-paragraph">The code below is adapted from the ntt_layer function that is internally used by the NTT. This function is compiled differently for x86-64 and aarch64, allowing dynamic dispatch to target-specific or portable implementations. On x86-64, it checks the availability of SSE2 instructions, while on aarch64 it checks for Neon.</p>1911191219131914<style data-wp-block-html="css">1915.gist .gist-meta {1916 display: none !important;1917}1918</style>19191920<div class="gist" data-gist-id="your-gist-id"><script src="https://gist.github.com/fournet/8f94607c600e0d78d640148003f021f6.js"></script></div>1921192219231924<p class="wp-block-paragraph">As rustc’s output is inherently target specific, our toolchain compiles the code several times, one per compilation target for which verification is required, before merging the corresponding models. In effect, this merge operation turns the static dispatch permitted by the cfg attributes in the Rust code into a first layer of dynamic dispatch between x86-64 and aarch64 in the Lean model. Following what the Rust code does, these target specific models then themselves dynamically dispatch to the models of the XMM, Neon, and generic implementations.</p>1925192619271928<p class="wp-block-paragraph">Intrinsics require a slightly different treatment. Some low-level wrappers, especially those that manipulate raw pointers or expose platform instructions, are modelled by small, carefully reviewed Lean specifications. Others can be modelled using Rust code, which can be tested against hardware reference documentation, then translated and verified. The surrounding safe Rust code is then verified against those models. This keeps the trusted surface narrow while preserving the performance benefits of hardware acceleration.</p>1929193019311932<p class="wp-block-paragraph">The important point is that verification does not require giving up optimization. The methodology is designed to preserve the complexities of production code – including intrinsics, dispatch, and platform-specific implementations – while still proving a single, auditable correctness statement.</p>1933193419351936<h2 id="reflecting-formal-guarantees-to-the-code-developer" class="wp-block-heading">Reflecting formal guarantees to the code developer</h2>1937193819391940<p class="wp-block-paragraph">Formal verification only scales in an engineering organization if developers can understand what has been proved. It is not enough for a proof to exist in a repository; the guarantee must be visible, reviewable, and synchronized to the code that engineers maintain.</p>1941194219431944<p class="wp-block-paragraph">To support this, we expose verification results through automatically generated dashboards. These dashboards summarize theorems in developer-facing terms: preconditions, postconditions, covered functions, trusted models, and remaining assumptions. Engineers do not need to open Lean to see what has been verified. For instance, below is the page displayed by the dashboard for our ntt function.</p>1945194619471948<figure class="wp-block-image aligncenter size-full"><img loading="lazy" decoding="async" width="1048" height="585" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/FIG3_SymCrypt.png" alt="Screenshot of a verified formal specification for the symcrust::mlkem::ntt function. The page shows a green “Verified” badge, links to the Lean model and source code, and a specification stating the mathematical conditions the NTT implementation must satisfy." class="wp-image-1178444" srcset="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/FIG3_SymCrypt.png 1048w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/FIG3_SymCrypt-300x167.png 300w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/FIG3_SymCrypt-1024x572.png 1024w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/FIG3_SymCrypt-768x429.png 768w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/FIG3_SymCrypt-240x134.png 240w" sizes="auto, (max-width: 1048px) 100vw, 1048px" /><figcaption class="wp-element-caption">Figure 2. Dashboard page for the theorem that shows the Rust function mlkem.ntt correctly implements the NTT specified in the NIST standard.</figcaption></figure>1949195019511952<p class="wp-block-paragraph">The specification clearly presents the theorem statement included in the Lean formal development: it separates the function input and preconditions from the post-condition by putting them above a horizontal line, and use fully qualified names with links to navigate to Rust and Lean definitions.</p>1953195419551956<p class="wp-block-paragraph">This feedback loop is especially useful for reviewing assumptions around intrinsics, target-specific code, and boundary conditions. A cryptographic developer can for example check whether the theorem fully captures what they expect their code to guarantee, and notice a formal statement is too weak, or a precondition is wrong.</p>1957195819591960<p class="wp-block-paragraph">The dashboards also aligns verification with continuous development. As Rust code changes, Lean models and proofs can be regenerated and replayed. When a proof breaks, that failure becomes a signal: either the implementation changed in a way that needs a proof update, or the change has exposed a real discrepancy with the specification.</p>1961196219631964<p class="wp-block-paragraph">This turns formal verification from a one-time research artifact into part of the engineering workflow.</p>1965196619671968<h2 id="agentic-proofs" class="wp-block-heading">Agentic proofs</h2>1969197019711972<p class="wp-block-paragraph">The final ingredient is automation beyond traditional tactics: AI agents. Lean is well suited to this because proofs are machine-checked by a small trusted kernel. An agent may propose a proof script, but Lean independently verifies whether the proof is valid.</p>1973197419751976<p class="wp-block-paragraph">We use agents in two places. First, they help translate standards into Lean specifications. Because the resulting specification is executable, aligned to the original standard, tested against official vectors, supported by mathematical theorems, and much simpler than an implementation, it can be thoroughly audited even when an agent helped draft it.</p>1977197819791980<p class="wp-block-paragraph">Second, agents help write and maintain proofs. With the right libraries, tactics, examples, and documentation, agents can handle large amounts of proof work: unfolding generated models, applying specifications for helper functions, discharging arithmetic obligations, and repairing proofs after refactors.</p>1981198219831984<p class="wp-block-paragraph">This is particularly powerful because the Rust code and Lean proofs are separated. Agents do not need to annotate or modify the production Rust implementation to make a proof go through. They operate on the proof side, and the result is accepted only if Lean validates it and the final theorem states the desired guarantee without introducing unreviewed assumptions.</p>1985198619871988<p class="wp-block-paragraph">In practice, this changes the economics of verification. Work that previously required months of specialist effort can be accelerated dramatically. The proof engineer’s role shifts from writing every proof by hand to designing specifications, curating automation, reviewing theorem statements, and steering agents to complete their proofs.</p>1989199019911992<h2 id="conclusion" class="wp-block-heading">Conclusion</h2>1993199419951996<p class="wp-block-paragraph">Verified cryptography has often faced a difficult trade-off: the strongest guarantees came from specialized toolchains, generated code, and workflows that were hard for product teams to adopt. Rust, Lean, Aeneas, and agentic proof automation let us revisit that tradeoff.</p>1997199819992000<p class="wp-block-paragraph">By verifying Rust as written, deriving auditable specifications from standards, supporting optimized multi-architecture implementations, and reflecting proof results back to developers, formal verification can become part of normal cryptographic engineering rather than an after-the-fact research exercise.</p>2001200220032004<p class="wp-block-paragraph">That is the long-term promise: cryptographic code that remains fast, portable, maintainable, and developer-owned, while carrying machine-checked evidence that it implements the standards it is meant to realize.</p>2005<span id="label-external-link" class="sr-only" aria-hidden="true">Opens in a new tab</span><p>The post <a href="https://www.microsoft.com/en-us/research/blog/verifying-rust-cryptography-in-symcrypt-from-standards-to-code/">Verifying Rust cryptography in SymCrypt, from standards to code</a> appeared first on <a href="https://www.microsoft.com/en-us/research">Microsoft Research</a>.</p>2006]]></content:encoded>2007 2008 2009 2010 </item>2011 <item>2012 <title>Aurora 1.5: Extending open foundation models for weather and Earth-system applications</title>2013 <link>https://www.microsoft.com/en-us/research/blog/aurora-1-5-extending-open-foundation-models-for-weather-and-earth-system-applications/</link>2014 2015 <dc:creator><![CDATA[Kenji Takeda, Haiyu Dong, Jonathan Weyn, Amit Misra, Matt Corey, Kevin White, Shannon Monroe, Juan M. Lavista Ferres, Ashley Llorens, Bonnie Kruft]]></dc:creator>2016 <pubDate>Thu, 09 Jul 2026 16:46:22 +0000</pubDate>2017 <category><![CDATA[Research Blog]]></category>2018 <guid isPermaLink="false">https://www.microsoft.com/en-us/research/blog/aurora-1-5-extending-open-foundation-models-for-weather-and-earth-system-applications/</guid>20192020 <description><![CDATA[<p>Aurora 1.5 adds 22 more variables, hourly temporal resolution, and probabilistic ensemble forecasting to the Aurora foundation model, making it more useful for real-world weather, climate, and energy applications.</p>2021<p>The post <a href="https://www.microsoft.com/en-us/research/blog/aurora-1-5-extending-open-foundation-models-for-weather-and-earth-system-applications/">Aurora 1.5: Extending open foundation models for weather and Earth-system applications</a> appeared first on <a href="https://www.microsoft.com/en-us/research">Microsoft Research</a>.</p>2022]]></description>2023 <content:encoded><![CDATA[2024<figure class="wp-block-image size-full"><img loading="lazy" decoding="async" width="1400" height="788" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/05/AuroraUpdate-BlogHeroFeature-1400x788-1.jpg" alt="Aurora 1.5 | three white line icons on an abstract blue and purple background: globe, thunder cloud, tree" class="wp-image-1173975" srcset="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/05/AuroraUpdate-BlogHeroFeature-1400x788-1.jpg 1400w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/05/AuroraUpdate-BlogHeroFeature-1400x788-1-300x169.jpg 300w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/05/AuroraUpdate-BlogHeroFeature-1400x788-1-1024x576.jpg 1024w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/05/AuroraUpdate-BlogHeroFeature-1400x788-1-768x432.jpg 768w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/05/AuroraUpdate-BlogHeroFeature-1400x788-1-1066x600.jpg 1066w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/05/AuroraUpdate-BlogHeroFeature-1400x788-1-655x368.jpg 655w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/05/AuroraUpdate-BlogHeroFeature-1400x788-1-240x135.jpg 240w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/05/AuroraUpdate-BlogHeroFeature-1400x788-1-640x360.jpg 640w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/05/AuroraUpdate-BlogHeroFeature-1400x788-1-960x540.jpg 960w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/05/AuroraUpdate-BlogHeroFeature-1400x788-1-1280x720.jpg 1280w" sizes="auto, (max-width: 1400px) 100vw, 1400px" /></figure>2025202620272028<div style="padding-bottom:0;padding-top:0" class="wp-block-msr-immersive-section alignfull row">2029 2030 <div class="container">2031 <div class="wp-block-msr-immersive-section__inner wp-block-msr-immersive-section__inner--narrow">2032 <div class="wp-block-columns mb-10 pb-1 pr-1 is-layout-flex wp-container-core-columns-is-layout-8f761849 wp-block-columns-is-layout-flex" style="box-shadow:var(--wp--preset--shadow--outlined)">2033<div class="wp-block-column is-layout-flow wp-block-column-is-layout-flow">2034<h2 id="at-a-glance" class="wp-block-heading h3">At a glance</h2>2035203620372038<ul class="wp-block-list">2039<li>Aurora 1.5 is a major extension of Microsoft’s Aurora Earth System foundation model that adds 22 more weather variables relevant to energy, agriculture, transport, and climate risk, along with hourly temporal resolution and probabilistic ensemble forecasting.</li>2040204120422043<li>Released as open source on GitHub with model checkpoints on Hugging Face, Aurora 1.5 enables researchers and developers to use, evaluate, and build on the model.</li>2044204520462047<li>Aurora 1.5 connects open research to Microsoft Weather services, linking the model with data, infrastructure, managed access, and operational use for weather and Earth-system applications.</li>2048</ul>2049</div>2050</div> </div>2051 </div>20522053 </div>2054205520562057<p class="wp-block-paragraph">Aurora 1.5 is a major update to the open Aurora Earth-system foundation model, adding 22 new weather variables for a broader view of atmospheric conditions, hourly forecasts, and probabilistic ensemble forecasting. Developed by Microsoft Weather as an extension of the original model from Microsoft Research AI for Science, Aurora 1.5 shows how frontier research can move into broader use: open for researchers and developers to evaluate and extend, and designed to support customers where additional data, infrastructure, and operational assurance is needed. As climate and weather-related risks continue to affect communities, infrastructure, and economies worldwide, advances in Earth-system forecasting can help improve preparedness and decision-making.</p>2058205920602061<h2 id="what-is-aurora" class="wp-block-heading">What is Aurora?</h2>2062206320642065<p class="wp-block-paragraph">Aurora is a foundation model for the Earth system developed by Microsoft Research AI for Science, first introduced in 2024 and <a class="msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall" href="https://www.nature.com/articles/s41586-025-09005-y" target="_blank" rel="noopener noreferrer">published in Nature<span class="sr-only"> (opens in new tab)</span></a> in 2025. It showed that a single model could be adapted to medium-range weather, ocean waves, atmospheric chemistry, and emerging climate applications, including high-resolution weather forecasting through fine-tuning. Its growing use has reinforced the value of an open, collaborative model that is easier to adapt, evaluate, and put to use. </p>2066206720682069<p class="wp-block-paragraph">This <a class="msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall" href="https://www.bing.com/ck/a?!&&p=f9c93e7b19f62b3737c7c3282badddf1233badf5058fb7d7861ef84db05d08e0JmltdHM9MTc4MjQzMjAwMA&ptn=3&ver=2&hsh=4&fclid=24451b10-f799-6468-1027-0c47f6ba6571&psq=microsoft+aurora+ai+weather+2024&u=a1aHR0cHM6Ly9ibG9ncy5taWNyb3NvZnQuY29tL29uLXRoZS1pc3N1ZXMvMjAyNS8xMS8xMy90aGUtbmV4dC1waGFzZS1vZi1hdXJvcmEtb3Blbi1hbmQtY29sbGFib3JhdGl2ZS1haS1mb3Itd2VhdGhlci1hbmQtY2xpbWF0ZS1mb3JlY2FzdGluZy8" target="_blank" rel="noopener noreferrer">next phase of Aurora<span class="sr-only"> (opens in new tab)</span></a> builds on that foundation by making the model openly available for the global community to adapt, extend, and build on. </p>2070207120722073<h2 id="what-is-new-in-aurora-1-5" class="wp-block-heading">What is new in Aurora 1.5?</h2>2074207520762077<p class="wp-block-paragraph">Aurora 1.5 advances the broader effort to make open weather foundation models practical and scalable for organizations that rely on atmospheric and Earth-system intelligence. Alongside new variables and higher temporal resolution, Aurora 1.5 adds one of the most requested capabilities from users: ensemble forecasting. Because forecasts are sensitive to initial conditions and model uncertainty, ensembles run multiple simulations to show the range and likelihood of possible outcomes. Aurora 1.5 builds on Microsoft Research’s scientific foundation with new product engineering, cloud infrastructure, managed access, and decision-support capabilities. Together, these advances make Aurora 1.5 a valuable enterprise-grade weather solution for organizations. </p>2078207920802081<figure class="wp-block-image aligncenter size-full"><img loading="lazy" decoding="async" width="2133" height="2414" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/aurora_1.5_demo_ensemble_forecast.png" alt="Aurora 1.5 ensemble forecast example showing mean and ensemble uncertainty for total cloud cover and surface solar radiation (SSRD) over the Atlantic and Europe region at a 2–3 day forecast range. Four globe maps display the ensemble mean and standard deviation for each variable, illustrating Aurora's ability to predict both expected conditions and forecast uncertainty for cloud cover and solar radiation. " class="wp-image-1178260" srcset="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/aurora_1.5_demo_ensemble_forecast.png 2133w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/aurora_1.5_demo_ensemble_forecast-265x300.png 265w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/aurora_1.5_demo_ensemble_forecast-905x1024.png 905w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/aurora_1.5_demo_ensemble_forecast-768x869.png 768w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/aurora_1.5_demo_ensemble_forecast-1357x1536.png 1357w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/aurora_1.5_demo_ensemble_forecast-1810x2048.png 1810w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/aurora_1.5_demo_ensemble_forecast-159x180.png 159w" sizes="auto, (max-width: 2133px) 100vw, 2133px" /><figcaption class="wp-element-caption">Figure 1: Illustration of the capabilities of Aurora 1.5 ensemble for predicting new impactful parameters such as total cloud cover and solar radiation. Ensemble mean and standard deviation are shown<em><em>.</em> </em></figcaption></figure>2082208320842085<p class="wp-block-paragraph">The breadth update adds 22 new variables to Aurora’s original 4, including representative surface, pressure-level, wind, temperature, humidity, precipitation, and radiation fields. That broader coverage makes the model more relevant for sectors that depend on integrated Earth-system signals, from energy and agriculture to transport and resilience planning. </p>2086208720882089<p class="wp-block-paragraph">The update to hourly temporal resolution enables fine-grained detail for precision operational guidance, such as the onset of precipitation, trade decisions, or a landfalling tropical cyclone. </p>2090209120922093<blockquote class="wp-block-quote is-style-spectrum is-layout-flow wp-block-quote-is-layout-flow">2094<p class="wp-block-paragraph"><em>“Aurora 1.5 is a meaningful step toward making weather foundation models more open, useful, and practical. By releasing the model openly, we give researchers, developers, and organizations a clearer path to evaluate it, adapt it, and understand where it can help. Microsoft Weather’s role is to connect that open research foundation with the data, infrastructure, and applied workflows required by enterprises to use weather intelligence responsibly and with confidence.”</em></p>2095<cite><strong>Sridhar Iyer, Corporate Vice President, Microsoft AI</strong></cite></blockquote>2096209720982099 <div class="border-bottom border-top border-gray-300 mt-5 mb-5 msr-promo text-center text-md-left alignwide" data-bi-aN="promo" data-bi-id="1144028">2100 21012102 <p class="msr-promo__label text-gray-800 text-center text-uppercase">2103 <span class="px-4 bg-white display-inline-block font-weight-semibold small">PODCAST SERIES</span>2104 </p>2105 2106 <div class="row pt-3 pb-4 align-items-center">2107 <div class="msr-promo__media col-12 col-md-5">2108 <a class="bg-gray-300 display-block" href="https://www.microsoft.com/en-us/research/story/the-ai-revolution-in-medicine-revisited/" aria-label="The AI Revolution in Medicine, Revisited" data-bi-cn="The AI Revolution in Medicine, Revisited" target="_blank">2109 <img decoding="async" class="w-100 display-block" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2025/06/Episode7-PeterBillSebastien-AIRevolution_Hero_Feature_River_No_Text_1400x788.jpg" alt="Illustrated headshot of Bill Gates, Peter Lee, and Sébastien Bubeck" />2110 </a>2111 </div>2112 2113 <div class="msr-promo__content p-3 px-5 col-12 col-md">21142115 <h2 class="h4">The AI Revolution in Medicine, Revisited</h2>2116 2117 <p id="the-ai-revolution-in-medicine-revisited" class="large">Join Microsoft’s Peter Lee on a journey to discover how AI is impacting healthcare and what it means for the future of medicine.</p>2118 2119 <div class="wp-block-buttons justify-content-center justify-content-md-start">2120 <div class="wp-block-button">2121 <a href="https://www.microsoft.com/en-us/research/story/the-ai-revolution-in-medicine-revisited/" aria-describedby="the-ai-revolution-in-medicine-revisited" class="btn btn-brand glyph-append glyph-append-chevron-right" data-bi-cn="The AI Revolution in Medicine, Revisited" target="_blank">2122 Listen now </a>2123 </div>2124 </div>2125 </div><!--/.msr-promo__content-->2126 </div><!--/.msr-promo__inner-wrap-->2127<span id="label-external-link" class="sr-only" aria-hidden="true">Opens in a new tab</span> </div><!--/.msr-promo-->2128 212921302131<h2 id="ensemble-forecasting-in-aurora-1-5-unlocks-more-confident-decisions-in-the-face-of-weather-uncertainty" class="wp-block-heading">Ensemble Forecasting in Aurora 1.5 Unlocks More Confident Decisions in the Face of Weather Uncertainty</h2>2132213321342135<p class="wp-block-paragraph">The ensemble version of Aurora 1.5 introduces stochastic perturbations to represent model uncertainty, allowing the generation of multiple forecast members to estimate the spread of possible futures. For a multitude of applications including power systems, transport, agriculture, extreme-weather planning, and climate risk, the model distribution matters as much as the best estimate. </p>2136213721382139<p class="wp-block-paragraph">This ensemble capability was developed through multi-stage fine-tuning on top of the original Aurora model. After expanding the variable set and adding hourly temporal resolution, the team introduced controlled perturbations into the model’s latent conditioning pathway and optimized the ensemble for probabilistic forecast quality. A final round of auto-regressive fine-tuning on ECMWF High Resolution (HRES) analysis data from 2018 to 2023 improved rollout behavior and stability.</p>2140214121422143<figure class="wp-block-image aligncenter size-full"><img loading="lazy" decoding="async" width="2560" height="1020" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/ensemble_scorecard-scaled.png" alt="Heat maps comparing Aurora 1.5 and ECMWF ensemble forecast skill. Aurora 1.5 achieves lower probabilistic forecast error across most variables and forecast lead times. " class="wp-image-1178119" srcset="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/ensemble_scorecard-scaled.png 2560w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/ensemble_scorecard-300x120.png 300w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/ensemble_scorecard-1024x408.png 1024w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/ensemble_scorecard-768x306.png 768w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/ensemble_scorecard-1536x612.png 1536w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/ensemble_scorecard-2048x816.png 2048w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/ensemble_scorecard-240x96.png 240w" sizes="auto, (max-width: 2560px) 100vw, 2560px" /><figcaption class="wp-element-caption">Figure 2. Comparing Aurora 1.5’s probabilistic forecasts with the ECMWF ensemble forecast. The shading shows relative probabilistic forecast error, using ECMWF ENS as the baseline: blue areas indicate where Aurora 1.5 performs better, and red areas indicate where it performs worse. Across upper-air geopotential, temperature, and humidity, together with five surface variables, Aurora 1.5 outperforms ECMWF ENS on 88.9% of the evaluated variable-and-lead-time targets. </figcaption></figure>2144214521462147<p class="wp-block-paragraph">Aurora’s ensemble approach summarizes uncertainty across multiple model runs. Its probabilistic forecasts outperform those of the state-of-the-art ECWMF dynamical ensemble on 88.9% of evaluated targets (Figure 1). In evaluations on all 2024–2025 tropical cyclones, Aurora 1.5 substantially reduced track errors, including roughly one-third lower track error when comparing the ensemble median to the original Aurora. An example for the devastating Hurricane Helene shows how Aurora 1.5’s skill translates to high-impact weather applications. </p>2148214921502151<figure class="wp-block-image aligncenter size-full"><img loading="lazy" decoding="async" width="600" height="967" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/05/helene_aurora15_2024092400_forecast_600px.png" alt="Aurora 1.5 ensemble forecasts for Hurricane Helene compared with operational and observed storm tracks. The ensemble forecasts closely follow the observed path while representing uncertainty through multiple plausible trajectories. " class="wp-image-1173971" srcset="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/05/helene_aurora15_2024092400_forecast_600px.png 600w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/05/helene_aurora15_2024092400_forecast_600px-186x300.png 186w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/05/helene_aurora15_2024092400_forecast_600px-112x180.png 112w" sizes="auto, (max-width: 600px) 100vw, 600px" /><figcaption class="wp-element-caption">Figure 3. Hurricane Helene ensemble forecast from Aurora 1.5, showing multiple plausible storm tracks starting at 0 UTC on September 24, 2024. The probabilistic ensemble forecast envelops the verified track, effectively capturing uncertainty in the storm’s progression.</figcaption></figure>2152215321542155<figure class="wp-block-image aligncenter size-full"><img loading="lazy" decoding="async" width="1400" height="938" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/aurora15_vs_original_merged.png" alt="Track-error reductions for Aurora 1.5 relative to the original Aurora model. Error decreases across all forecast lead times, with the largest improvements from the ensemble median forecast. " class="wp-image-1178271" srcset="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/aurora15_vs_original_merged.png 1400w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/aurora15_vs_original_merged-300x201.png 300w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/aurora15_vs_original_merged-1024x686.png 1024w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/aurora15_vs_original_merged-768x515.png 768w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/aurora15_vs_original_merged-240x161.png 240w" sizes="auto, (max-width: 1400px) 100vw, 1400px" /><figcaption class="wp-element-caption">Figure 4. Aurora 1.5 reduces track error relative to the original model across lead times. Ensemble mean and median tracks are used for diagnostics, with the median showing the strongest gains, reaching roughly one-third lower error by day 5. Results reflect track position only. </figcaption></figure>2156215721582159<h2 id="beyond-weather-aurora-as-an-earth-system-foundation" class="wp-block-heading">Beyond weather: Aurora as an Earth-system foundation</h2>2160216121622163<p class="wp-block-paragraph">Beyond medium-range weather applications, Terradot – part of the Microsoft Climate Innovation Fund portfolio—is working with the <a class="msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall" href="https://iclr.cc/virtual/2026/10014507" target="_blank" rel="noopener noreferrer">AI for Good Lab<span class="sr-only"> (opens in new tab)</span></a> and the Microsoft Research Accelerator on <a class="msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall" href="https://iclr.cc/virtual/2026/10014507" target="_blank" rel="noopener noreferrer">TerraNova, using Aurora-derived weather representations<span class="sr-only"> (opens in new tab)</span></a> to estimate and optimize carbon dioxide removal from enhanced rock weathering under real field conditions. Sasankh Munukutla, Co-Founder of Terradot, highlights<em>, “By building on Aurora, we’re significantly advancing our R&D timelines and accelerating our path towards gigaton-scale carbon removal.” </em>This work shows how Earth-system foundation models can support climate mitigation and public-interest science beyond forecasting, including settings where rigorous evaluation and responsible deployment matter.</p>2164216521662167<p class="wp-block-paragraph">Aurora is also being explored with partners such as the UK Met Office, exploring how foundation models can work alongside established physics-based systems to tackle problems from weather to climate time scales. The aim is faster, more flexible forecasts that support decision-making without replacing the science behind trusted prediction. </p>2168216921702171<blockquote class="wp-block-quote is-style-spectrum is-layout-flow wp-block-quote-is-layout-flow">2172<p class="wp-block-paragraph"><em>“Microsoft’s Aurora model is an exciting and promising tool, enabling Met Office scientists to bring their data and expertise to help solve climate problems and provide new kinds of climate information. Met Office and Microsoft scientists and engineers are working together every day to translate lessons from AI weather prediction into the climate information space, sharing expertise in data science and climate science. Aurora is a great platform for learning how to translate these tools for use in climate projection to make the AI climate models of the future.”</em></p>2173<cite>— Doug McNeall, Science lead for Data-Driven Climate Modelling, Met Office Hadley Centre </cite></blockquote>2174217521762177<h2 id="connecting-open-models-to-operational-use" class="wp-block-heading">Connecting open models to operational use</h2>2178217921802181<p class="wp-block-paragraph">Microsoft connects open research, product engineering, responsible deployment, and partner ecosystems so that models can move from scientific advance to evaluated operational use. As an example, Aurora began in Microsoft Research AI for Science and is now being built on for operational use by Microsoft Weather, with AI for Good helping to evaluate public-interest applications. The platform path brings <a href="https://www.microsoft.com/en/customers/story/26785-bkw-fmb-energie-ag-foundry-models" target="_blank" rel="noreferrer noopener">Aurora into Microsoft Foundry and Planetary Computer Pro</a>, alongside Agent skills and Azure services that connect models with geospatial data, scalable infrastructure, and applied workflows. <a href="https://www.microsoft.com/en/customers/story/26785-bkw-fmb-energie-ag-foundry-models" target="_blank" rel="noreferrer noopener">BKW provides an early proof point</a>: the company is using Aurora 1.5 alongside existing operational Microsoft Weather models to support energy operations where weather-dependent generation, infrastructure planning, and environmental data need to come together. </p>2182218321842185<blockquote class="wp-block-quote is-style-spectrum is-layout-flow wp-block-quote-is-layout-flow">2186<p class="wp-block-paragraph"><em>“This collaboration demonstrates how advanced AI capabilities and robust cloud infrastructure can be applied to one of the most strategic domains — energy, where weather plays a fundamental role. In a time of accelerated transformation, it supports our ambition to operate increasingly renewable-based systems, where generation is inherently weather-dependent, and to better anticipate and manage this variability with greater confidence and precision.”</em> </p>2187<cite>Farhat Quiñones Yamshid, Lead, AI and Technology, BKW </cite></blockquote>2188218921902191<h2 id="from-open-research-to-broader-impact" class="wp-block-heading">From open research to broader impact</h2>2192219321942195<p class="wp-block-paragraph">Aurora’s open-source availability is intended to help researchers, agencies, companies, and civil society evaluate, apply, and extend the model. Microsoft Weather is building on that open foundation to deliver easier access to Aurora forecasts through managed services, integrations, and responsible deployment paths for organizations that depend on weather and Earth-system intelligence.</p>2196219721982199<p class="wp-block-paragraph">Foundation models should complement—not replace—physics-based models and domain expertise. The opportunity is to use them responsibly, with careful evaluation and transparency, and to invite researchers, agencies, companies, and public-interest partners to test where Aurora and related Microsoft Weather capabilities can improve forecasting, planning, and climate resilience in their own settings.</p>2200220122022203<h2 id="about-microsoft-weather" class="wp-block-heading">About Microsoft Weather </h2>2204220522062207<p class="wp-block-paragraph">Microsoft Weather is the AI-based forecasting team behind weather experiences across Windows, Bing, Copilot, Edge, and MSN, reaching more than a billion devices across 180 countries. The team has been applying AI to operational weather forecasting for more than seven years and has built a proven track record of delivering high-quality forecasts at global scale. Microsoft Weather has won multiple forecasting competitions and was ranked the world’s most accurate global forecast provider by an independent third party for three consecutive years from 2022 to 2024. Building on today’s Aurora 1.5 announcement, the team plans to extend this work in the coming months with additional fit-for-purpose AI weather models designed for enterprise scenarios where forecast quality, speed, uncertainty, and operational decision support matter most.</p>2208220922102211<p class="wp-block-paragraph">If you are interested in exploring Aurora and Microsoft Weather solutions for commercial or organizational applications, please contact us at <a href="mailto:AIWeatherClimate@microsoft.com" target="_blank" rel="noreferrer noopener">AIWeatherClimate@microsoft.com</a> </p>2212221322142215<div class="wp-block-buttons is-content-justification-center is-content-justification-center is-layout-flex wp-container-core-buttons-is-layout-fe48e5de wp-block-buttons-is-layout-flex">2216<div class="wp-block-button"><a data-bi-type="button" class="wp-block-button__link wp-element-button" href="https://ai.azure.com/catalog/models/Aurora-1.5">Aurora 1.5 on Microsoft Foundry</a></div>2217221822192220<div class="wp-block-button is-style-fill-github"><a data-bi-type="button" class="wp-block-button__link wp-element-button" href="https://github.com/microsoft/aurora">Aurora 1.5 on GitHub</a></div>2221222222232224<div class="wp-block-button is-style-fill"><a data-bi-type="button" class="wp-block-button__link wp-element-button" href="https://www.microsoft.com/en-us/research/publication/aurora-1-5-fine-tuning-a-foundation-model-for-medium-range-ensemble-weather-prediction/">Aurora 1.5 paper</a></div>2225222622272228<div class="wp-block-button is-style-fill-github"><a data-bi-type="button" class="wp-block-button__link wp-element-button" href="https://github.com/microsoft/vibe-kit/tree/main/skills/msresearch-aurora" target="_blank" rel="noreferrer noopener">Agent Skills for adapting Aurora to new applications</a></div>2229</div>2230223122322233<p class="wp-block-paragraph"></p>2234<span id="label-external-link" class="sr-only" aria-hidden="true">Opens in a new tab</span><p>The post <a href="https://www.microsoft.com/en-us/research/blog/aurora-1-5-extending-open-foundation-models-for-weather-and-earth-system-applications/">Aurora 1.5: Extending open foundation models for weather and Earth-system applications</a> appeared first on <a href="https://www.microsoft.com/en-us/research">Microsoft Research</a>.</p>2235]]></content:encoded>2236 2237 2238 2239 </item>2240 <item>2241 <title>Flint: A visualization language for the AI era</title>2242 <link>https://www.microsoft.com/en-us/research/blog/flint-a-visualization-language-for-the-ai-era/</link>2243 2244 <dc:creator><![CDATA[Chenglong Wang, Alper Sarikaya, Scott Tsukamaki, Michel Galley, Jianfeng Gao]]></dc:creator>2245 <pubDate>Wed, 08 Jul 2026 16:00:00 +0000</pubDate>2246 <category><![CDATA[Research Blog]]></category>2247 <guid isPermaLink="false">https://www.microsoft.com/en-us/research/?p=1177589</guid>22482249 <description><![CDATA[<p>Short chart specifications are easy to write, but often produce uninspiring results. Flint is an open-source visualization language that offers a middle path, letting AI agents create expressive charts from compact, human-editable specifications.</p>2250<p>The post <a href="https://www.microsoft.com/en-us/research/blog/flint-a-visualization-language-for-the-ai-era/">Flint: A visualization language for the AI era</a> appeared first on <a href="https://www.microsoft.com/en-us/research">Microsoft Research</a>.</p>2251]]></description>2252 <content:encoded><![CDATA[2253<figure class="wp-block-image size-full"><img loading="lazy" decoding="async" width="1400" height="788" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/Flint-BlogHeroFeature-1400x788-1.jpg" alt="Flint blog | three white line icons on an abstract green background; bar chart icon, connected nodes icon, flowchart icon" class="wp-image-1177592" srcset="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/Flint-BlogHeroFeature-1400x788-1.jpg 1400w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/Flint-BlogHeroFeature-1400x788-1-300x169.jpg 300w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/Flint-BlogHeroFeature-1400x788-1-1024x576.jpg 1024w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/Flint-BlogHeroFeature-1400x788-1-768x432.jpg 768w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/Flint-BlogHeroFeature-1400x788-1-1066x600.jpg 1066w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/Flint-BlogHeroFeature-1400x788-1-655x368.jpg 655w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/Flint-BlogHeroFeature-1400x788-1-240x135.jpg 240w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/Flint-BlogHeroFeature-1400x788-1-640x360.jpg 640w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/Flint-BlogHeroFeature-1400x788-1-960x540.jpg 960w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/Flint-BlogHeroFeature-1400x788-1-1280x720.jpg 1280w" sizes="auto, (max-width: 1400px) 100vw, 1400px" /></figure>2254225522562257<div style="padding-bottom:0;padding-top:0" class="wp-block-msr-immersive-section alignfull row">2258 2259 <div class="container">2260 <div class="wp-block-msr-immersive-section__inner wp-block-msr-immersive-section__inner--narrow">2261 <div class="wp-block-columns mb-10 pb-1 pr-1 is-layout-flex wp-container-core-columns-is-layout-8f761849 wp-block-columns-is-layout-flex" style="box-shadow:var(--wp--preset--shadow--outlined)">2262<div class="wp-block-column is-layout-flow wp-block-column-is-layout-flow">2263<h2 id="at-a-glance" class="wp-block-heading h3">At a glance</h2>2264226522662267<ul class="wp-block-list">2268<li><strong>Polished charts from simple specs</strong>. Flint allows AI agents to reliably generate expressive, visually polished charts from simple, human-editable specifications.</li>2269227022712272<li><strong>Semantic types guide design</strong>. Flint leverages semantic data types to express meanings of data. They help the compiler choose appropriate scales, baselines, formatting, and color schemes.</li>2273227422752276<li><strong>Layouts adapt to the data</strong>. Flint automatically manages sizing, spacing, labels, and layout so charts remain readable as cardinality and density change, without explicit user configurations.</li>2277227822792280<li><strong>One spec can target multiple backends</strong>. A single Flint specification can compile to Vega-Lite, Apache ECharts, or Chart.js without rewriting the chart from scratch.</li>2281228222832284<li><strong>Built for agent workflows</strong>. The open-source project includes the <em>flint-chart library</em> and the <em>flint-chart-mcp server</em>, so agents can create, validate, and render charts directly in chat or coding environments.</li>2285</ul>2286</div>2287</div> </div>2288 </div>22892290 </div>2291229222932294<div style="height:20px" aria-hidden="true" class="wp-block-spacer"></div>2295229622972298<figure class="wp-block-image aligncenter size-full"><img loading="lazy" decoding="async" width="2560" height="764" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/chartwall_FLINT-scaled.png" alt="A dense grid displaying a diverse gallery of data visualizations. The collection showcases over twenty different chart types, including stacked area charts, line graphs, sunburst charts, stacked bar charts, treemaps, radar charts, Sankey diagrams, dense heatmaps, diverging bar charts, candlestick charts, violin plots, a choropleth map of the United States, scatter plots, grouped bar charts, waterfall charts, and parallel coordinate plots." class="wp-image-1178156" srcset="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/chartwall_FLINT-scaled.png 2560w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/chartwall_FLINT-300x90.png 300w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/chartwall_FLINT-1024x306.png 1024w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/chartwall_FLINT-768x229.png 768w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/chartwall_FLINT-1536x458.png 1536w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/chartwall_FLINT-2048x611.png 2048w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/chartwall_FLINT-240x72.png 240w" sizes="auto, (max-width: 2560px) 100vw, 2560px" /><figcaption class="wp-element-caption">Figure 1. Flint supports a diverse collection of visualizations with its simple spec, which can be rendered with visualization libraries like Vega-Lite, Echarts, and Chart.js.</figcaption></figure>2299230023012302<p class="wp-block-paragraph">Creating a good chart requires many design decisions: how dates should be parsed, whether a scale should start at zero, how values should be formatted, how much room labels need, and which colors make the data easier to read. Modern visualization libraries such as Vega-Lite, Apache ECharts, and Chart.js expose these controls, but there is a trade-off: Short specifications that rely on system defaults often produce uninspiring charts, while polished visualizations require detailed specifications with purposely chosen parameters that are often verbose, fragile, and error-prone.</p>2303230423052306<p class="wp-block-paragraph">This trade-off becomes sharper as large language models (LLMs) and AI agents take on more visualization work. Agents are especially prone to errors when they must manage complex, low-level specification details, and the resulting fragile code can be difficult for people to inspect, repair, or reuse. Ideally, we need something in between: a compact specification that agents can produce reliably, people can edit directly, and a system can compile into a well-designed chart.</p>2307230823092310<p class="wp-block-paragraph">To address this challenge, we introduce <a class="msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall" href="https://microsoft.github.io/flint-chart/" type="link" id="https://microsoft.github.io/flint-chart/" target="_blank" rel="noopener noreferrer">Flint<span class="sr-only"> (opens in new tab)</span></a>, a visualization intermediate language for AI-driven chart creation. Flint helps AI agents create expressive, attractive charts from simple, human-editable chart specs. Instead of requiring verbose low-level parameters for scales, axes, spacing, and layout, the Flint compiler derives optimized chart settings from the data, semantic types, chart type, and encodings. The same Flint spec can render through multiple backends, including Vega-Lite, Apache ECharts, and Chart.js.</p>2311231223132314<figure class="wp-block-image aligncenter size-full"><img loading="lazy" decoding="async" width="2560" height="1260" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/compile-demo_FLINT-scaled.png" alt="A three-step diagram illustrating the Flint workflow from left to right. It starts with a short JSON code snippet labeled "FLINT SPEC," which flows into a significantly longer, more detailed JSON code snippet labeled "COMPILED SPEC (VEGA-LITE)." This final specification then points to the end result under "VISUALIZATION," displaying a rendered heatmap chart showing data across different games and periods." class="wp-image-1178168" srcset="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/compile-demo_FLINT-scaled.png 2560w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/compile-demo_FLINT-300x148.png 300w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/compile-demo_FLINT-1024x504.png 1024w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/compile-demo_FLINT-768x378.png 768w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/compile-demo_FLINT-1536x756.png 1536w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/compile-demo_FLINT-2048x1008.png 2048w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/compile-demo_FLINT-240x118.png 240w" sizes="auto, (max-width: 2560px) 100vw, 2560px" /><figcaption class="wp-element-caption">Figure 2. Flint compiles a compact, human-editable chart specification into a complete backend-native specification and rendered visualization. In this heatmap example, the Flint spec names semantic types (period as YearMonth, newUsers as Profit) and maps fields to visual channels. The compiler derives the Vega-Lite details, including temporal parsing, axis formatting, color scale, cell sizing, legend configuration, and layout.</figcaption></figure>2315231623172318<h2 id="how-flint-works" class="wp-block-heading">How Flint works</h2>2319232023212322<p class="wp-block-paragraph">Figure 2 illustrates the how the Flint compiler turns a compact chart specification into a refined heatmap.</p>2323232423252326<p class="wp-block-paragraph">To produce a high-quality heatmap, traditionally, we need to explicitly tell the system with low-level chart properties about how to process the period field, how to properly label MonthYear values, size individual heatmap cells, and choose a color scale that appropriately represents positive and negative newUsers values. Without these configurations, visualization libraries must guess from field names and raw values, which can lead to charts that are technically valid but potentially misleading. While they are important, hard-coding these details can be difficult and error-prone, and they make specification fragile and hard for users to understand or adapt.</p>2327232823292330<p class="wp-block-paragraph">In Flint, these low-level details are systematically managed, where the compiler infers them from high-level data and chart specifications. Here, the <strong>data specification</strong> captures semantic types and optional metadata, and the <strong>chart specification</strong> defines the chart type and maps fields to visual channels such as x, y, color, size, or facet. From this information, the compiler derives the parsing rules, scales, axes, aggregations, formatting, color schemes, layout, and generates the backend-native specification, which is used to render the final polished visualization. This frees users from explicitly setting fragile and error-prone low-level details.</p>2331233223332334<p class="wp-block-paragraph">Furthermore, because the intermediate representation is separate from any single rendering library, Flint can target backends with very different APIs and programming models. Users can keep the same compact chart intent while compiling to Vega-Lite, ECharts, or Chart.js, and choose the backend whose capabilities best fit the visualization.</p>2335233623372338 <div class="border-bottom border-top border-gray-300 mt-5 mb-5 msr-promo text-center text-md-left alignwide" data-bi-aN="promo" data-bi-id="1144027">2339 23402341 <p class="msr-promo__label text-gray-800 text-center text-uppercase">2342 <span class="px-4 bg-white display-inline-block font-weight-semibold small">PODCAST SERIES</span>2343 </p>2344 2345 <div class="row pt-3 pb-4 align-items-center">2346 <div class="msr-promo__media col-12 col-md-5">2347 <a class="bg-gray-300 display-block" href="https://www.microsoft.com/en-us/research/story/ai-testing-and-evaluation-learnings-from-science-and-industry/" aria-label="AI Testing and Evaluation: Learnings from Science and Industry" data-bi-cn="AI Testing and Evaluation: Learnings from Science and Industry" target="_blank">2348 <img decoding="async" class="w-100 display-block" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2025/06/EP2-AI-TE_Hero_Feature_River_No_Text_1400x788.jpg" alt="Illustrated headshots of Daniel Carpenter, Timo Minssen, Chad Atalla, and Kathleen Sullivan for the Microsoft Research Podcast" />2349 </a>2350 </div>2351 2352 <div class="msr-promo__content p-3 px-5 col-12 col-md">23532354 <h2 class="h4">AI Testing and Evaluation: Learnings from Science and Industry</h2>2355 2356 <p id="ai-testing-and-evaluation-learnings-from-science-and-industry" class="large">Discover how Microsoft is learning from other domains to advance evaluation and testing as a pillar of AI governance.</p>2357 2358 <div class="wp-block-buttons justify-content-center justify-content-md-start">2359 <div class="wp-block-button">2360 <a href="https://www.microsoft.com/en-us/research/story/ai-testing-and-evaluation-learnings-from-science-and-industry/" aria-describedby="ai-testing-and-evaluation-learnings-from-science-and-industry" class="btn btn-brand glyph-append glyph-append-chevron-right" data-bi-cn="AI Testing and Evaluation: Learnings from Science and Industry" target="_blank">2361 Listen now </a>2362 </div>2363 </div>2364 </div><!--/.msr-promo__content-->2365 </div><!--/.msr-promo__inner-wrap-->2366<span id="label-external-link" class="sr-only" aria-hidden="true">Opens in a new tab</span> </div><!--/.msr-promo-->2367 236823692370<h2 id="flint-for-ai-assisted-visualization" class="wp-block-heading">Flint for AI-assisted visualization</h2>2371237223732374<p class="wp-block-paragraph">Flint is well suited to LLM-based chart generation because semantic types are often easier for models to infer than the full set of low-level visualization parameters. Field names, value patterns, and common data knowledge can help an agent recognize whether a column represents a date, price, percentage, country, ranking, or correlation. Once those meanings are explicit, the compiler can handle many design decisions that would otherwise appear as brittle, library-specific code.</p>2375237623772378<p class="wp-block-paragraph">In our research study, we compared Flint with DirectVL, a baseline that asks the model to directly generate full (more complex) Vega-Lite specifications in a LLM self-evaluation pipeline. Across three tested models based on testing data from Tidy Tuesdays, Flint received higher overall LLM-judge scores: 16.27 vs. 15.91 with GPT-5.1, 16.16 vs. 15.60 with GPT-5-mini, and 15.91 vs. 15.34 with GPT-4.1. In fact, Flint has been so powerful and reliable that it is now used to power <a class="msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall" href="https://github.com/microsoft/data-formulator" type="link" id="https://github.com/microsoft/data-formulator" target="_blank" rel="noopener noreferrer">Data Formulator<span class="sr-only"> (opens in new tab)</span></a>, a Microsoft Research project for AI-assisted data analysis and visualization.</p>2379238023812382<p class="wp-block-paragraph">To make Flint easy for your agents to access, we also release <strong><em>flint-chart-mcp</em></strong>, a Model Context Protocol (MCP) server that allows agents to create, validate, and render charts inside a chat or coding environment. MCP calls can embed data inline or read configured local files, and the server can open an interactive chart view so users can inspect and refine the results.</p>2383238423852386<figure class="wp-block-image aligncenter size-full"><img loading="lazy" decoding="async" width="1628" height="1338" src="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/flint-mcp-experience_FLINT.png" alt="A mockup of an AI agent chat interface. A user sends the message, "Show me quarterly revenue by region as a grouped bar chart." The AI responds with an interactive widget labeled "Flint Chart MCP APP" displaying a grouped bar chart of revenue by quarter, color-coded by region (North, South, West). Below the chart are interactive UI controls for adjusting the corner radius, sorting, toggling values on or off, and a button to "Copy spec to chat."" class="wp-image-1178170" srcset="https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/flint-mcp-experience_FLINT.png 1628w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/flint-mcp-experience_FLINT-300x247.png 300w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/flint-mcp-experience_FLINT-1024x842.png 1024w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/flint-mcp-experience_FLINT-768x631.png 768w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/flint-mcp-experience_FLINT-1536x1262.png 1536w, https://www.microsoft.com/en-us/research/wp-content/uploads/2026/07/flint-mcp-experience_FLINT-219x180.png 219w" sizes="auto, (max-width: 1628px) 100vw, 1628px" /><figcaption class="wp-element-caption">Figure 3. Once you set up the flint-chart-mcp with your favorite AI client, the agent can generate interactive visualizations powered by Flint to answer your data exploration questions. </figcaption></figure>2387238823892390<h2 id="try-flint" class="wp-block-heading">Try Flint</h2>2391239223932394<p class="wp-block-paragraph">Flint is open source and ready to use:</p>2395239623972398<ul class="wp-block-list">2399<li>Project site: <a class="msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall" href="https://microsoft.github.io/flint-chart/" target="_blank" rel="noopener noreferrer">https://microsoft.github.io/flint-chart/<span class="sr-only"> (opens in new tab)</span></a></li>2400240124022403<li>GitHub: <a class="msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall" href="https://github.com/microsoft/flint-chart" target="_blank" rel="noopener noreferrer">https://github.com/microsoft/flint-chart<span class="sr-only"> (opens in new tab)</span></a></li>2404240524062407<li>Flint MCP server instruction: <a class="msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall" href="https://microsoft.github.io/flint-chart/#/mcp" target="_blank" rel="noopener noreferrer">https://microsoft.github.io/flint-chart/#/mcp<span class="sr-only"> (opens in new tab)</span></a></li>2408</ul>2409241024112412<p class="wp-block-paragraph">Flint points toward a shared semantic layer for visualization, where people and AI agents can work with compact chart intent while a compiler handles the careful low-level details. We invite the community to explore the project and build on it.</p>2413<span id="label-external-link" class="sr-only" aria-hidden="true">Opens in a new tab</span><p>The post <a href="https://www.microsoft.com/en-us/research/blog/flint-a-visualization-language-for-the-ai-era/">Flint: A visualization language for the AI era</a> appeared first on <a href="https://www.microsoft.com/en-us/research">Microsoft Research</a>.</p>2414]]></content:encoded>2415 2416 2417 2418 </item>2419 </channel>2420</rss>2421