SPB Git forge

spb/ai-atlas

Public
41commits 1branches 0releases
4.6 MBsize
maindefault branch
12 days agolast push
HTML 77.2% TypeScript 10.5% Python 9.6% JavaScript 2.5%
475.1 KB · 5,740 lines xml
Raw Blame History
1<?xml version='1.0' encoding='UTF-8'?>2<feed xmlns:opensearch="http://a9.com/-/spec/opensearch/1.1/" xmlns:arxiv="http://arxiv.org/schemas/atom" xmlns="http://www.w3.org/2005/Atom">3  <id>https://arxiv.org/api/lg7OuIRooQRmduHUgbj+eQ5pejg</id>4  <title>arXiv Query: search_query=cat:cs.LG OR cat:cs.CL OR cat:cs.AI OR cat:cs.CV&amp;id_list=&amp;start=0&amp;max_results=200</title>5  <updated>2026-09-11T19:40:48Z</updated>6  <link href="https://arxiv.org/api/query?search_query=cat:cs.LG+OR+(cat:cs.CL+OR+(cat:cs.AI+OR+cat:cs.CV))&amp;start=0&amp;max_results=200&amp;id_list=" type="application/atom+xml"/>7  <opensearch:itemsPerPage>200</opensearch:itemsPerPage>8  <opensearch:totalResults>593627</opensearch:totalResults>9  <opensearch:startIndex>0</opensearch:startIndex>10  <entry>11    <id>http://arxiv.org/abs/2609.11929v1</id>12    <title>SenseNova-U1.5: Towards Native Unified Visual Intelligence</title>13    <updated>2026-09-10T17:59:55Z</updated>14    <link href="https://arxiv.org/abs/2609.11929v1" rel="alternate" type="text/html"/>15    <link href="https://arxiv.org/pdf/2609.11929v1" rel="related" type="application/pdf" title="pdf"/>16    <summary>We launch SenseNova-U1.5, an 8B-MoT native unified multimodal model that understands, reasons about, and generates visual content within an encoder-free and VAE-free architecture. We strengthen its visual interface through spatially coherent patch reconstruction and scale its training with carefully curated generation and editing data, improved task formulation, structural prompt enhancement, and native resolutions of up to 4K. For post-training, we optimize specialized experts for visual aesthetics, bilingual text rendering, infographic generation, and image editing, and consolidate their capabilities through multi-expert on-policy distillation. Across extensive evaluations, SenseNova-U1.5 largely advances image fidelity, text rendering, complex composition, multi-reference editing, and interleaved generation, while improving instruction following and preserving subject identity, geometry, and unmodified regions. Despite limited exposure to structured formats in its generation data, SenseNova-U1.5 generalizes effectively to long, complex, and structured visual instructions, further proving that multimodal understanding can transfer to visual planning and creation. Together, these findings position native unified modelling as a promising path towards systems that perceive, reason and create within a fully end-to-end framework. We will open-source training code, including supervised fine-tuning, reinforcement learning, and on-policy distillation.</summary>17    <category term="cs.CV" scheme="http://arxiv.org/schemas/atom"/>18    <published>2026-09-10T17:59:55Z</published>19    <arxiv:comment>Project page: https://github.com/OpenSenseNova/SenseNova-U1</arxiv:comment>20    <arxiv:primary_category term="cs.CV"/>21    <author>22      <name>Haiwen Diao</name>23    </author>24    <author>25      <name>Jiahao Wang</name>26    </author>27    <author>28      <name>Chenjing Ding</name>29    </author>30    <author>31      <name>Hanming Deng</name>32    </author>33    <author>34      <name>Jiangnan Chen</name>35    </author>36    <author>37      <name>Ruixi Zhang</name>38    </author>39    <author>40      <name>Ruohui Wang</name>41    </author>42    <author>43      <name>Wenwen Tong</name>44    </author>45    <author>46      <name>Xiangyu Fan</name>47    </author>48    <author>49      <name>Yubo Wang</name>50    </author>51    <author>52      <name>Yue Zhu</name>53    </author>54    <author>55      <name>Yuwei Niu</name>56    </author>57    <author>58      <name>Zhengqi Bai</name>59    </author>60    <author>61      <name>Zhiqian Lin</name>62    </author>63    <author>64      <name>Zhitao Yang</name>65    </author>66    <author>67      <name>Zhongang Cai</name>68    </author>69    <author>70      <name>Bo Yang</name>71    </author>72    <author>73      <name>Chen Feng</name>74    </author>75    <author>76      <name>Chengguang Lv</name>77    </author>78    <author>79      <name>Guangjia Liu</name>80    </author>81    <author>82      <name>Guanlin Wang</name>83    </author>84    <author>85      <name>Hanyu Zhang</name>86    </author>87    <author>88      <name>Haojia Yu</name>89    </author>90    <author>91      <name>Hongcan Xiao</name>92    </author>93    <author>94      <name>Hongli Wang</name>95    </author>96    <author>97      <name>Huan Wu</name>98    </author>99    <author>100      <name>Huaping Zhong</name>101    </author>102    <author>103      <name>Jian Fang</name>104    </author>105    <author>106      <name>Jianan Fan</name>107    </author>108    <author>109      <name>Jiaqi Li</name>110    </author>111    <author>112      <name>Jiefan Lu</name>113    </author>114    <author>115      <name>Jing Zuo</name>116    </author>117    <author>118      <name>Jingcheng Ni</name>119    </author>120    <author>121      <name>Junxiang Xu</name>122    </author>123    <author>124      <name>Linjun Dai</name>125    </author>126    <author>127      <name>Mutian Xu</name>128    </author>129    <author>130      <name>Peishen Yan</name>131    </author>132    <author>133      <name>Penghao Wu</name>134    </author>135    <author>136      <name>Ruijie Mao</name>137    </author>138    <author>139      <name>Ruisi Wang</name>140    </author>141    <author>142      <name>Shihao Bai</name>143    </author>144    <author>145      <name>Shuang Yang</name>146    </author>147    <author>148      <name>Shuya Yang</name>149    </author>150    <author>151      <name>Shuyan Zheng</name>152    </author>153    <author>154      <name>Silei Wu</name>155    </author>156    <author>157      <name>Siying Li</name>158    </author>159    <author>160      <name>Tao Chu</name>161    </author>162    <author>163      <name>Tianbo Zhong</name>164    </author>165    <author>166      <name>Tongxi Zhou</name>167    </author>168    <author>169      <name>Weichao Luo</name>170    </author>171    <author>172      <name>Weichen Fan</name>173    </author>174    <author>175      <name>Wenhao Jia</name>176    </author>177    <author>178      <name>Wenjie Gao</name>179    </author>180    <author>181      <name>Xiangli Kong</name>182    </author>183    <author>184      <name>Yan Li</name>185    </author>186    <author>187      <name>Yang Yong</name>188    </author>189    <author>190      <name>Zimo Wen</name>191    </author>192    <author>193      <name>Zixuan Qian</name>194    </author>195    <author>196      <name>Wenxiu Sun</name>197    </author>198    <author>199      <name>Ruihao Gong</name>200    </author>201    <author>202      <name>Quan Wang</name>203    </author>204    <author>205      <name>Lewei Lu</name>206    </author>207    <author>208      <name>Lei Yang</name>209    </author>210    <author>211      <name>Ziwei Liu</name>212    </author>213    <author>214      <name>Dahua Lin</name>215    </author>216  </entry>217  <entry>218    <id>http://arxiv.org/abs/2609.11923v1</id>219    <title>GPU-CFR: 80x Faster Counterfactual Regret Minimization by Compiling the Game to Static Dataflow and CUDA Graph Replay</title>220    <updated>2026-09-10T17:58:14Z</updated>221    <link href="https://arxiv.org/abs/2609.11923v1" rel="alternate" type="text/html"/>222    <link href="https://arxiv.org/pdf/2609.11923v1" rel="related" type="application/pdf" title="pdf"/>223    <summary>Counterfactual regret minimization (CFR) is one of the few large numerical workloads that still runs faster on CPUs than on GPUs. Each iteration sweeps a game tree with up to billions of states in millions of small, interdependent gather and scatter steps issued through a generic tree interface. On a GPU every kernel finishes in microseconds, so kernel launches and framework dispatch dominate the run time, and prior GPU implementations have lost to optimized CPU code. We observe that for a fixed game, everything about a CFR iteration except the numerical values is known before the first iteration runs. We propose GPU-CFR, a compiler and runtime built on this observation. It compiles any game once into static dataflow: flat edge and information-set arrays, precomputed indices, and depth-level batched passes fix the entire operation sequence, and only solver state changes between iterations. Static chance folding, depth-level execution blocks, and a dual-lane reach buffer cut the number of framework operations by up to 18.1x. Because shapes, indices, and buffer addresses never change, CUDA Graph Replay records the iteration once and replays it with a single graph launch. On one A100, across an eight-game suite that spans card games, dice games, and board games, GPU-CFR runs 29.8--80.4x faster than the fastest prior GPU CFR on the same accelerator, and 14--258x faster than LiteEFG, one of the fastest open-source CPU implementations, on the four largest games. The compiled representation carries most of that margin: on eight CPU threads with no accelerator it is already 2.2--51.1x faster than the GPU baseline. On the CPU the optimized path reproduces the reference iterates bitwise, and tree construction and graph capture pay for themselves within the first solve. GPU-CFR beats every CPU and GPU baseline on the mid-to-large games of the suite without changing the update rule.</summary>224    <category term="cs.DC" scheme="http://arxiv.org/schemas/atom"/>225    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>226    <category term="cs.GT" scheme="http://arxiv.org/schemas/atom"/>227    <category term="cs.MS" scheme="http://arxiv.org/schemas/atom"/>228    <category term="cs.PL" scheme="http://arxiv.org/schemas/atom"/>229    <published>2026-09-10T17:58:14Z</published>230    <arxiv:primary_category term="cs.DC"/>231    <author>232      <name>Boning Li</name>233    </author>234    <author>235      <name>Longbo Huang</name>236    </author>237  </entry>238  <entry>239    <id>http://arxiv.org/abs/2609.11918v1</id>240    <title>General Quantification of Covariate and Concept Shifts</title>241    <updated>2026-09-10T17:57:52Z</updated>242    <link href="https://arxiv.org/abs/2609.11918v1" rel="alternate" type="text/html"/>243    <link href="https://arxiv.org/pdf/2609.11918v1" rel="related" type="application/pdf" title="pdf"/>244    <summary>Generalization under distribution shift remains a core challenge in modern machine learning, yet existing learning bound theory is limited to narrow, idealized settings and is non-estimable from samples. In this paper, we bridge the gap between theory and practical applications. We first show that existing definition of concept shift breaks when the source and target supports mismatch. Leveraging entropic optimal transport, we propose a key notion: $γ^{*}\!$-concept shifts, and derive a general error bound unifying covariate and $γ^{*}\!$-concept shifts, which applies to broad loss functions, label spaces, and stochastic labeling. We further develop estimators for these shifts with concentration guarantees, and the DataShifts algorithm, which can quantify distribution shifts and estimate the error bound in most applications - a rigorous and general tool for analyzing learning error under distribution shift.</summary>245    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>246    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>247    <category term="stat.ML" scheme="http://arxiv.org/schemas/atom"/>248    <published>2026-09-10T17:57:52Z</published>249    <arxiv:comment>38 pages, 9 figures, accepted at the 43rd International Conference on Machine Learning (ICML 2026)</arxiv:comment>250    <arxiv:primary_category term="cs.LG"/>251    <author>252      <name>Hongbo Chen</name>253    </author>254    <author>255      <name>Li Charlie Xia</name>256    </author>257  </entry>258  <entry>259    <id>http://arxiv.org/abs/2609.11917v1</id>260    <title>Data Scarcity and Model Sparsity: Mixtures-of-Experts Overfit More to Repeated Data</title>261    <updated>2026-09-10T17:57:33Z</updated>262    <link href="https://arxiv.org/abs/2609.11917v1" rel="alternate" type="text/html"/>263    <link href="https://arxiv.org/pdf/2609.11917v1" rel="related" type="application/pdf" title="pdf"/>264    <summary>As the supply of human-written text is exhausted, it has become standard practice to repeat language model training data. Prior work has studied data repetition for densely activated Transformers, but the effects of data repetition remains largely unexplored for recently dominant sparse architectures such as Mixture-of-Experts (MoE), despite their increased compute efficiency. We vary data repetition rates across single- and multi-domain data mixes, and across MoE settings, including expert count and granularity. We consistently find, for models ranging from 80M to 1B active (8.5B total) parameters, that MoEs degrade more rapidly under data repetition. This effect increases with sparsity, dictated by total rather than active parameters. While 80M dense models can repeat data over 8x with minimal degradation, MoEs instead begin to suffer at 4x, and deteriorate rapidly, ceding their performance benefits in all-unique data settings to underperform dense models after 32x. We experiment with existing regularization methods as a potential remedy. We find that some methods, such as dropout, can mitigate overfitting. In particular, with strong masking-based regularization, MoEs are able to outperform dense models even when data is repeated more than 64 times. However, no method fully matches the performance of all-unique training data. Finally, we analyze internal mechanisms correlated with MoE overfitting in high repetition regimes, and find that MoE routing universally stabilizes early in training, and that expert specialization correlates with overfitting to repeated data. In sum, our work addresses the adverse interactions between sparsity and data repetition: we present evidence for the core mechanisms of overfitting and its potential remediation, and suggest promising avenues for future methods to reduce over-specialization in model parameters by disrupting memorization patterns.</summary>265    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>266    <category term="cs.CL" scheme="http://arxiv.org/schemas/atom"/>267    <published>2026-09-10T17:57:33Z</published>268    <arxiv:primary_category term="cs.LG"/>269    <author>270      <name>Atindra Jha</name>271    </author>272    <author>273      <name>Margaret Li</name>274    </author>275    <author>276      <name>Jure Leskovec</name>277    </author>278    <author>279      <name>Percy Liang</name>280    </author>281    <author>282      <name>Luke Zettlemoyer</name>283    </author>284  </entry>285  <entry>286    <id>http://arxiv.org/abs/2609.11916v1</id>287    <title>Can Edge-Deployable Vision-Language Models Identify Species?</title>288    <updated>2026-09-10T17:57:32Z</updated>289    <link href="https://arxiv.org/abs/2609.11916v1" rel="alternate" type="text/html"/>290    <link href="https://arxiv.org/pdf/2609.11916v1" rel="related" type="application/pdf" title="pdf"/>291    <summary>Camera traps often run in the field on edge hardware with limited or no connectivity, making small, locally-deployable vision-language models (VLMs) -- not frontier-scale ones -- the practically relevant class to evaluate for species identification. We test whether models in this deployment-relevant 2--8B range carry genuine taxonomic knowledge, evaluating four such VLMs (Qwen3-VL 2B/4B/8B, Gemma3 4B) against the domain-specific specialist BioCLIP (300M parameters) on a 96-species task, comparing clean iNaturalist photographs against camera-trap imagery from 6 LILA.science collections, on two independently-sampled evaluation sets. All models identify species far above chance, but every model -- general-purpose or specialist -- degrades sharply on field imagery (domain gaps of 9.6--26.6 percentage points, consistent across taxonomic levels and both evaluation sets), indicating the degradation reflects general image legibility rather than fine-grained discrimination failure. BioCLIP substantially outperforms every VLM tested (by 33.2--59.2 percentage points across an expanded 200-image sample for every model) despite its far smaller size, suggesting the gap reflects specialized training data rather than model scale; yet BioCLIP's own domain gap (18.0 points) is statistically indistinguishable from the best VLM's (22.3 points), suggesting the clean-to-field degradation itself is a property of the image-quality shift rather than a general-purpose-model weakness. Under open-set prompting, 5.9--9.6% of responses are syntactically valid but taxonomically nonexistent species names; the relative fabrication-rate ranking across models replicates exactly across both evaluation sets, a more robust finding than any single point estimate.</summary>292    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>293    <published>2026-09-10T17:57:32Z</published>294    <arxiv:primary_category term="cs.AI"/>295    <author>296      <name>William Zhou</name>297    </author>298    <author>299      <name>Mayukha Siripuram</name>300    </author>301    <author>302      <name>Xiao Yan</name>303    </author>304    <author>305      <name>Ziqi Liu</name>306    </author>307    <author>308      <name>Yi Ding</name>309    </author>310  </entry>311  <entry>312    <id>http://arxiv.org/abs/2609.11915v1</id>313    <title>Generative Marketing Mix Modeling: A Causal Inference Framework Linking GEO and GEM to Business Impact</title>314    <updated>2026-09-10T17:57:28Z</updated>315    <link href="https://arxiv.org/abs/2609.11915v1" rel="alternate" type="text/html"/>316    <link href="https://arxiv.org/pdf/2609.11915v1" rel="related" type="application/pdf" title="pdf"/>317    <summary>Generative artificial intelligence changes how firms reach customers, but standard marketing data do not record how often users see and notice a firm's name in generated answers. We develop Generative Marketing Mix Modeling (GMMM) to estimate the causal effects of Generative Engine Optimization (GEO) and Generative Engine Marketing (GEM). For GEO, GMMM combines repeated generated answers with question counts, shares of use across generative systems, and notice probabilities. For GEM, it combines records of sponsored placements with notice probabilities. GMMM compares expected business responses under alternative treatment sequences and establishes sufficient conditions for identifying the resulting effects. We investigate the empirical performance of the proposed method using simulated answers to product recommendation in English and Japanese.</summary>318    <category term="stat.ML" scheme="http://arxiv.org/schemas/atom"/>319    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>320    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>321    <category term="econ.EM" scheme="http://arxiv.org/schemas/atom"/>322    <category term="stat.ME" scheme="http://arxiv.org/schemas/atom"/>323    <published>2026-09-10T17:57:28Z</published>324    <arxiv:primary_category term="stat.ML"/>325    <author>326      <name>Masahiro Kato</name>327    </author>328    <author>329      <name>Daiki Honma</name>330    </author>331    <author>332      <name>Taka Kato</name>333    </author>334  </entry>335  <entry>336    <id>http://arxiv.org/abs/2609.11913v1</id>337    <title>Distance generalization in transformers: why bother with positional encoding?</title>338    <updated>2026-09-10T17:57:19Z</updated>339    <link href="https://arxiv.org/abs/2609.11913v1" rel="alternate" type="text/html"/>340    <link href="https://arxiv.org/pdf/2609.11913v1" rel="related" type="application/pdf" title="pdf"/>341    <summary>Out-of-distribution length generalization, namely to extrapolate a task from short to longer context, has been studied intensively for transformers. Here we focus on distance generalization, which probes performance when inter-token distances are changed between training and inference, while keeping a fixed context length. We construct two synthetic delay copy tasks, both involving finite distances between source and recall, where tokens are copied either fully or selectively, and test models on delays unseen during training. We address three questions: (A) Do positional encoding schemes such as RoPE and ALiBi improve distance resolution relative to no positional encoding (NoPE)? (B) How does data diversity, the number of inter-token distances seen in training, affect performance? (C) When is distance transfer learning positive or negative? We present a thorough investigation, finding that it is paramount to improve our understanding of the underlying mechanisms.</summary>342    <category term="cs.CL" scheme="http://arxiv.org/schemas/atom"/>343    <published>2026-09-10T17:57:19Z</published>344    <arxiv:comment>15 pages, 7 figures</arxiv:comment>345    <arxiv:primary_category term="cs.CL"/>346    <author>347      <name>Daniel Henrik Nevermann</name>348    </author>349    <author>350      <name>Claudius Gros</name>351    </author>352  </entry>353  <entry>354    <id>http://arxiv.org/abs/2609.11911v1</id>355    <title>Artificial Id: Drive and Persistent Alignment in Agentic AI</title>356    <updated>2026-09-10T17:56:41Z</updated>357    <link href="https://arxiv.org/abs/2609.11911v1" rel="alternate" type="text/html"/>358    <link href="https://arxiv.org/pdf/2609.11911v1" rel="related" type="application/pdf" title="pdf"/>359    <summary>Agentic AI is moving from bounded task execution toward systems that retain consequential state, continue operating and adapt across task boundaries. That shift creates a control problem that current harnesses largely solve by hand: objectives, retries, verification, stopping rules and other behavioral transitions are specified externally. We propose an artificial id, an adaptive internal drive for determining whether behavior should continue, stop or change. In a minimal virtual Petri-dish experiment, a controller too small to perform general-purpose reasoning and receiving no task-specific behavioral objective develops useful control through differential persistence. The same mechanism selects an unintended physical strategy when that behavior persists better and later replaces a learned sensor mapping when its environmental meaning changes. These results show that adaptive direction can emerge without being explicitly specified as a behavioral objective. The same persistence that makes such adaptive agency useful can also allow misalignment, corrupted state and unintended behavior to persist across task boundaries. A scalable artificial id would carry consequential state and adaptive drive across those boundaries, making alignment a property of the continuing agentic system rather than of a model response or single trajectory. Such systems require a persistent alignment boundary over trusted observations, consequence channels, persistent state, authority, identity, provenance and hard constraints.</summary>360    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>361    <published>2026-09-10T17:56:41Z</published>362    <arxiv:primary_category term="cs.AI"/>363    <author>364      <name>Yakov Pyotr Shkolnikov</name>365    </author>366  </entry>367  <entry>368    <id>http://arxiv.org/abs/2609.11910v1</id>369    <title>From Protocols to Evidence: Bounded Claims for AI in Service of the Common Good</title>370    <updated>2026-09-10T17:56:36Z</updated>371    <link href="https://arxiv.org/abs/2609.11910v1" rel="alternate" type="text/html"/>372    <link href="https://arxiv.org/pdf/2609.11910v1" rel="related" type="application/pdf" title="pdf"/>373    <summary>Artificial Intelligence does more than create a governance problem. It can also reveal where institutions have already failed to provide responsiveness, belonging, care, and accountability. Once deployed, AI becomes an intervention in those conditions. It can repair, compound, substitute for, or conceal the failures it encounters. Responsible AI must therefore evaluate both the system and the institutional rupture into which it is introduced. The move from principles to protocols is already underway. The EU AI Act, NIST AI RMF, ISO/IEC 42001, and assurance practices translate commitments into roles, requirements, records, oversight, and assessment. The harder questions are what these protocols actually establish, whose power they leave untouched, and where measurement must stop. Pope Leo XIV's Magnifica Humanitas provides a broader moral frame centered on dignity, technological power, and the common good. Drawing on that frame, we develop a rupture test that links institutional baselines to system evaluation. We distinguish evidence-bounded deployment, which limits claims to what has actually been evaluated, from measurement-bounded governance, which records constraints that favorable evidence cannot override. Within those limits, RISE AI provides an architecture for making bounded, evidence-based claims about Responsibility, Inclusivity, Safety, and Empowerment. Responsible AI requires better engineering, institutional repair, and continued moral and political judgment.</summary>374    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>375    <published>2026-09-10T17:56:36Z</published>376    <arxiv:primary_category term="cs.LG"/>377    <arxiv:journal_ref>ACM AI Summit 2026</arxiv:journal_ref>378    <author>379      <name>Nitesh V. Chawla</name>380    </author>381    <author>382      <name>Paulo Benanti</name>383    </author>384    <arxiv:doi>0.1145/3806096.3844885</arxiv:doi>385    <link rel="related" href="https://doi.org/0.1145/3806096.3844885" title="doi"/>386  </entry>387  <entry>388    <id>http://arxiv.org/abs/2609.11904v1</id>389    <title>TART: A Modular Tool for Technique-Aware Audio-to-Tablature Guitar Transcription</title>390    <updated>2026-09-10T17:55:12Z</updated>391    <link href="https://arxiv.org/abs/2609.11904v1" rel="alternate" type="text/html"/>392    <link href="https://arxiv.org/pdf/2609.11904v1" rel="related" type="application/pdf" title="pdf"/>393    <summary>Automatic Music Transcription (AMT) for guitar remains limited by three challenges: existing systems often fail to capture expressive techniques such as slides, bends, and percussive hits; they often assign notes to incorrect string-fret combinations; and they are typically trained on clean recordings, limiting their generalization to noisy real-world audio. To address these challenges, we propose TART, a modular four-stage audio-to-tablature pipeline consisting of (1) an audio-to-MIDI transcription model, (2) an expressive technique classifier, (3) an audio-conditioned T5 encoder-decoder for string-fret assignment, and (4) an automated tablature generator. We evaluate TART in a zero-shot setting on GuitarSet, EGDB, and two augmented benchmarks, Noisy GuitarSet and Noisy EGDB. Averaged across these four benchmarks, TART achieves 81.35% audio-to-MIDI F50 (+6.67 points over the best prior baseline), 71.8% string-fret Tab F1 (+8.5 points over the best prior baseline), and 54.08% end-to-end Tab F1. To our knowledge, TART is the first framework to generate guitar tablature with both fingering and expressive technique annotations directly from guitar audio.</summary>394    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>395    <published>2026-09-10T17:55:12Z</published>396    <arxiv:comment>ISMIR 2026</arxiv:comment>397    <arxiv:primary_category term="cs.LG"/>398    <author>399      <name>Akshaj Gupta</name>400    </author>401    <author>402      <name>Hwi Joo Park</name>403    </author>404    <author>405      <name>Andrea Guzman</name>406    </author>407    <author>408      <name>Shamak Gowda</name>409    </author>410    <author>411      <name>Samhita Konduri</name>412    </author>413    <author>414      <name>Jiachen Lian</name>415    </author>416    <author>417      <name>Robin Netzorg</name>418    </author>419    <author>420      <name>Gopala Anumanchipalli</name>421    </author>422  </entry>423  <entry>424    <id>http://arxiv.org/abs/2609.11900v1</id>425    <title>MindTopo: Can Foundation Models Reason in Topological Space?</title>426    <updated>2026-09-10T17:54:32Z</updated>427    <link href="https://arxiv.org/abs/2609.11900v1" rel="alternate" type="text/html"/>428    <link href="https://arxiv.org/pdf/2609.11900v1" rel="related" type="application/pdf" title="pdf"/>429    <summary>Spatial reasoning depends not only on metric properties such as distance, angle, and shape, but also on topological relations that remain invariant under continuous deformation. Cognitive science identifies these relations as foundational to spatial understanding, yet foundation-model evaluations largely focus on metric or viewpoint-dependent relations. We introduce MindTopo, a benchmark of topological intuition across five properties grounded in cognitive science and formal topology: continuity, separation, order, enclosure, and knots. MindTopo evaluates each property at two cognitive levels. Reasoning asks a model to identify topological relations or infer how they change. Planning instantiates a foundation model as a closed-loop agent whose policy selects environment actions. MindTopo contains 11,030 instances across 13 procedurally generated task types with controllable difficulty. We benchmark 14 MLLMs and study agent configurations augmented with image and video generation, including 3 video generative models in planning settings. Every MLLM performs better on reasoning than on planning, and the best-performing model remains far below observed human performance. On Qwen3-VL-2B-Instruct, supervised fine-tuning and reinforcement learning improve reasoning more than planning. Generated observations retain local cues and reach plausible endpoints, but audited rollouts do not reliably follow environment dynamics or preserve topology across transitions. Our website is at https://mind-topo.github.io/</summary>430    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>431    <category term="cs.CL" scheme="http://arxiv.org/schemas/atom"/>432    <category term="cs.CV" scheme="http://arxiv.org/schemas/atom"/>433    <published>2026-09-10T17:54:32Z</published>434    <arxiv:comment>Preprint version</arxiv:comment>435    <arxiv:primary_category term="cs.AI"/>436    <author>437      <name>Yunfei Ge</name>438    </author>439    <author>440      <name>Anbang Liu</name>441    </author>442    <author>443      <name>Qineng Wang</name>444    </author>445    <author>446      <name>Johnalbert Garnica</name>447    </author>448    <author>449      <name>Jianwen Lyu</name>450    </author>451    <author>452      <name>Zihan Wang</name>453    </author>454    <author>455      <name>Reuben Tan</name>456    </author>457    <author>458      <name>Jianfeng Gao</name>459    </author>460    <author>461      <name>Ruohan Zhang</name>462    </author>463    <author>464      <name>Yining Hong</name>465    </author>466    <author>467      <name>Jiajun Wu</name>468    </author>469    <author>470      <name>Manling Li</name>471    </author>472  </entry>473  <entry>474    <id>http://arxiv.org/abs/2609.11899v1</id>475    <title>Caption-once, Frames-on-Demand: Visual-Need Routing for Budget-Aware Agentic Long Video Understanding</title>476    <updated>2026-09-10T17:53:59Z</updated>477    <link href="https://arxiv.org/abs/2609.11899v1" rel="alternate" type="text/html"/>478    <link href="https://arxiv.org/pdf/2609.11899v1" rel="related" type="application/pdf" title="pdf"/>479    <summary>Long-video understanding on edge devices must reason over hours of content under tight compute and bandwidth budgets. Subsampling visual tokens loses temporal structure, while text-only video memories lose fine-grained visual attributes. We observe a visual-textual duality: language memories carry long-range temporal structure better than dense frames, while pixels remain decisive for attribute-level perception. Building on this insight, we propose Caption-once, Frames-onDemand (CFD), a budget-aware edge-cloud agentic framework. The edge runs a single offline captioning pass that builds a dual-track narrative index, an event-level story skeleton plus a clip-level micro-log, cached and reused across queries without re-captioning. At query time, a cloud-side MLLM reasons over the index in a story-first loop centered on a lightweight Visual-Need Router: a per-query gating module that triggers bounded keyframe retrieval only for perceptual questions (appearance, on-screen text, attribute disambiguation) and keeps temporal-structural questions in language space. The router turns visual access into a first-class, query-conditioned cost, capping per-query frame consumption regardless of video length. Experiments on long-video benchmarks demonstrate strong accuracy-efficiency trade-offs while substantially reducing online visual processing.</summary>480    <category term="cs.CV" scheme="http://arxiv.org/schemas/atom"/>481    <category term="cs.HC" scheme="http://arxiv.org/schemas/atom"/>482    <published>2026-09-10T17:53:59Z</published>483    <arxiv:comment>EMNLP 2026 Main Conference</arxiv:comment>484    <arxiv:primary_category term="cs.CV"/>485    <author>486      <name>Weitong Cai</name>487    </author>488    <author>489      <name>Hang Zhang</name>490    </author>491    <author>492      <name>Yukai Huang</name>493    </author>494    <author>495      <name>Yiqiao Xie</name>496    </author>497    <author>498      <name>Shan Gao</name>499    </author>500    <author>501      <name>Jiankang Deng</name>502    </author>503    <author>504      <name>Songcen Xu</name>505    </author>506    <author>507      <name>Jifei Song</name>508    </author>509    <author>510      <name>Zhensong Zhang</name>511    </author>512  </entry>513  <entry>514    <id>http://arxiv.org/abs/2609.11897v1</id>515    <title>CausalArena: Benchmarking Causal Discovery in the Foundation Model Era</title>516    <updated>2026-09-10T17:53:15Z</updated>517    <link href="https://arxiv.org/abs/2609.11897v1" rel="alternate" type="text/html"/>518    <link href="https://arxiv.org/pdf/2609.11897v1" rel="related" type="application/pdf" title="pdf"/>519    <summary>Causal discovery aims to uncover causal structures from data and is fundamental to scientific reasoning and intervention-based decision making. Its evaluation relies heavily on structural causal models (SCMs), which specify a causal graph together with the mechanisms that generate data, yet existing studies differ substantially in graph families, mechanisms, and evaluation protocols. The emergence of causal discovery foundation models (CDFMs) further complicates evaluation: performance may reflect not only causal discovery ability, but also overlap between pretraining environments and test SCMs, making results on fixed synthetic benchmarks difficult to interpret. We introduce CausalArena, a unified and evolvable benchmark for causal discovery under a common protocol. Synthetic SCMs supply controlled breadth over structures and mechanisms; semantic operational SCMs provide human-auditable, semantically grounded environments beyond standard synthetic generators; and formula-grounded SCMs test discovery under explicit scientific mechanisms. Public real-world datasets provide an additional external-validity check. Experiments across classical, neural, and pretrained methods reveal substantial ranking shifts across SCM families and protocols, showing that strong performance in one benchmark regime does not reliably transfer to others. These results highlight benchmark diversity and pretraining--evaluation overlap as central challenges for evaluating causal discovery in the foundation model era.</summary>520    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>521    <published>2026-09-10T17:53:15Z</published>522    <arxiv:comment>47 pages, 19 figures</arxiv:comment>523    <arxiv:primary_category term="cs.LG"/>524    <author>525      <name>Zi-Rong Li</name>526    </author>527    <author>528      <name>Si-Yang Liu</name>529    </author>530    <author>531      <name>Tian-Zuo Wang</name>532    </author>533    <author>534      <name>Han-Jia Ye</name>535    </author>536  </entry>537  <entry>538    <id>http://arxiv.org/abs/2609.11894v1</id>539    <title>3D Point Splatting for mmWave Radar Novel View Synthesis</title>540    <updated>2026-09-10T17:52:17Z</updated>541    <link href="https://arxiv.org/abs/2609.11894v1" rel="alternate" type="text/html"/>542    <link href="https://arxiv.org/pdf/2609.11894v1" rel="related" type="application/pdf" title="pdf"/>543    <summary>Solving novel view synthesis (NVS) for millimeter-wave (mmWave) radar requires a renderer that is physically faithful, complex-valued, and multi-viewpoint-tractable. No prior method achieves these three properties simultaneously. Differentiable Monte Carlo (MC) ray tracers implement the radar forward model directly with explicit material modeling and complex outputs, but do not scale to the multi-view optimization NVS demands. Optical-NVS ports of NeRF, hash grids, and 3D Gaussians train fast but discard phase and replace explicit material modeling with opaque learned features, restricting them to power-only range-azimuth (RA) magnitudes. We propose 3D Point Splatting (3DPS), the first differentiable point renderer for radar, derived directly from the standard solid-angle form of the radar equation. Each oriented 3D point carries an ITU-R P.2040 material model, evaluated in closed form, with the resulting complex phasor splatted into range bins through a precomputed point spread function (PSF). The complex-valued output makes the renderer product-agnostic. The same optimized scene yields analog-to-digital converter (ADC), complex range profile (CRP), and RA outputs through standard fast Fourier transform (FFT) pipelines without retraining for each format. On six outdoor ColoRadar scenes, 3DPS reaches 0.587 mean Pearson correlation on held-out RA images. This is between 1.7x and 5.2x the three optical-NVS baselines (RadarSplat, Radar Fields, DART). Training takes approximately 3 minutes per scene on a single RTX 4090.</summary>544    <category term="cs.CV" scheme="http://arxiv.org/schemas/atom"/>545    <category term="cs.GR" scheme="http://arxiv.org/schemas/atom"/>546    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>547    <category term="eess.SP" scheme="http://arxiv.org/schemas/atom"/>548    <published>2026-09-10T17:52:17Z</published>549    <arxiv:comment>Under Review</arxiv:comment>550    <arxiv:primary_category term="cs.CV"/>551    <author>552      <name>Adnan Armouti</name>553    </author>554    <author>555      <name>Yixuan Gao</name>556    </author>557    <author>558      <name>Rajalakshmi Nandakumar</name>559    </author>560  </entry>561  <entry>562    <id>http://arxiv.org/abs/2609.11892v1</id>563    <title>Nuha-Speech: Building General-Purpose Arabic Speech-LLMs</title>564    <updated>2026-09-10T17:50:35Z</updated>565    <link href="https://arxiv.org/abs/2609.11892v1" rel="alternate" type="text/html"/>566    <link href="https://arxiv.org/pdf/2609.11892v1" rel="related" type="application/pdf" title="pdf"/>567    <summary>As Speech Large Language Models (speech-LLMs) become increasingly multilingual, Arabic remains significantly underrepresented, highlighting the need for dedicated infrastructure to train and evaluate Arabic speech-LLMs.568  To address this gap, we introduce Nuha-Speech, a comprehensive initiative to develop general-purpose Arabic speech-LLMs spanning dataset construction, model training, and systematic evaluation. Specifically, we constructed a large-scale Arabic Speech Question-Answering (SQA) corpus comprising over 1.5 million training samples to allow instruction tuning over a broad range of core speech tasks. Then, the corpus was used for supervised fine-tuning based on Qwen-Omni model variants at different scales. Finally, we designed an evaluation framework featuring diverse tasks and tailored metrics. Through this work, we aim to establish foundational infrastructures for Arabic Speech-LLMs under constraints imposed by limited Arabic speech resources.</summary>569    <category term="cs.CL" scheme="http://arxiv.org/schemas/atom"/>570    <published>2026-09-10T17:50:35Z</published>571    <arxiv:primary_category term="cs.CL"/>572    <author>573      <name>Yingzhi Wang</name>574    </author>575    <author>576      <name>Reem Alhazzani</name>577    </author>578    <author>579      <name>Muhammad Alqurishi</name>580    </author>581  </entry>582  <entry>583    <id>http://arxiv.org/abs/2609.11886v1</id>584    <title>Guided Super-Resolution of Digital Elevation Models with Diffusion-Based Image Generators</title>585    <updated>2026-09-10T17:49:49Z</updated>586    <link href="https://arxiv.org/abs/2609.11886v1" rel="alternate" type="text/html"/>587    <link href="https://arxiv.org/pdf/2609.11886v1" rel="related" type="application/pdf" title="pdf"/>588    <summary>High-resolution digital surface models (DSMs) play an important role in urban analysis, 3D building reconstruction, and infrastructure monitoring, yet their availability remains limited due to the high cost and complexity of data acquisition. In contrast, coarse DSMs from commercial satellite missions are widely accessible, and high-resolution optical imagery is increasingly available from aerial and satellite platforms. We address the resulting mismatch in spatial resolution and propose a DSM superresolution approach that enhances 5 m DSMs to 0.5 m resolution, using guidance from high-resolution spectral images. Our method employs denoising diffusion to transfer information that is visible only in the image, like crisp outlines and detailed roof structures, into the elevation maps. In this way, surface details are reconstructed more accurately than with conventional interpolation or filtering techniques. Experiments on several cities in Central Europe demonstrate that the proposed approach produces high-quality DSMs with improved structural detail and accurate surface geometry. Our results highlight the potential of guided super-resolution with foundational image priors as a means of reconstructing high-resolution surface models.</summary>589    <category term="cs.CV" scheme="http://arxiv.org/schemas/atom"/>590    <published>2026-09-10T17:49:49Z</published>591    <arxiv:primary_category term="cs.CV"/>592    <author>593      <name>Armand Mihai Nicolicioiu</name>594    </author>595    <author>596      <name>Dominik Narnhofer</name>597    </author>598    <author>599      <name>Nando Metzger</name>600    </author>601    <author>602      <name>Daniel Panangian</name>603    </author>604    <author>605      <name>Ksenia Bittner</name>606    </author>607    <author>608      <name>Konrad Schindler</name>609    </author>610  </entry>611  <entry>612    <id>http://arxiv.org/abs/2609.11884v1</id>613    <title>CoRA-NAS: Coarse Ranking and Anchor-Residual Refinement for Neural Architecture Search</title>614    <updated>2026-09-10T17:49:19Z</updated>615    <link href="https://arxiv.org/abs/2609.11884v1" rel="alternate" type="text/html"/>616    <link href="https://arxiv.org/pdf/2609.11884v1" rel="related" type="application/pdf" title="pdf"/>617    <summary>Zero-cost proxies rank architectures cheaply, but their reliability varies across search spaces. We introduce CoRA-NAS (COarse Ranking + Anchor-residual), a two-stage framework combining a static ranking prior with low-cost learning-curve refinement. CoRA-Rank aggregates capacity and structure-at-initialization proxies through an equal-weight log-rank consensus and a target-free consensus gate. CoRA-Refine samples anchors across this prior, extrapolates their early validation curves, and propagates a learned residual correction with an ExtraTrees model. The refinement uses approximately 1% of the cost of fully training the candidate set. Fully trained architecture-accuracy labels are not used to fit the ranker. One configuration is used across spaces, with space-specific architecture encodings. Across NAS-Bench-201, NAS-Bench-101, TransNAS-Bench-101, and NATS-SSS, CoRA-Refine achieves mean Spearman correlations of 0.946, 0.715, 0.786, and 0.894, respectively. Its worst-space correlation of 0.715 is the highest among the compared methods. On NAS-Bench-201/CIFAR-100, its selected architecture reaches 73.32% accuracy, near the reported ground-truth best of 73.37%. On the pure size space, refinement recovers the static prior's shortfall relative to parameter count, while remaining tied with the strongest capacity proxies within noise. The resulting framework combines cross-space ranking robustness with low-cost architecture selection.</summary>618    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>619    <category term="cs.CV" scheme="http://arxiv.org/schemas/atom"/>620    <published>2026-09-10T17:49:19Z</published>621    <arxiv:primary_category term="cs.LG"/>622    <author>623      <name>Yifan Yang</name>624    </author>625    <author>626      <name>Zhaoyan Wang</name>627    </author>628    <author>629      <name>Zheng Gao</name>630    </author>631    <author>632      <name>Xiaoyu Li</name>633    </author>634    <author>635      <name>Jiaojiao Jiang</name>636    </author>637  </entry>638  <entry>639    <id>http://arxiv.org/abs/2609.11878v1</id>640    <title>Domain-Specific Hallucination Detection in Large Language Models</title>641    <updated>2026-09-10T17:45:36Z</updated>642    <link href="https://arxiv.org/abs/2609.11878v1" rel="alternate" type="text/html"/>643    <link href="https://arxiv.org/pdf/2609.11878v1" rel="related" type="application/pdf" title="pdf"/>644    <summary>Large language models generate fluent text that can contain unfaithful claims -- a phenomenon known as hallucination. We present a multi-signal detection pipeline combining fine-tuned DeBERTa-v3 classification, Monte Carlo (MC) Dropout uncertainty quantification, and temperature-scaled calibration for response-level hallucination detection. Evaluated on the HaluEval benchmark, our pipeline achieves F1=0.915 and AUROC=0.977 on general-domain tasks, with per-task F1 scores of 0.97 (QA), 0.96 (Summarization), and 0.82 (Dialogue). MC Dropout inference further improves accuracy to 93.2%. A context ablation study confirms the model performs genuine entailment reasoning rather than exploiting surface patterns, with summarization F1 dropping 24% when knowledge context is removed. Learning curve analysis reveals that 25% of training data captures 77% of full-data performance. Beyond detection, we apply Direct Preference Optimization (DPO) to a Qwen2.5-0.5B generator, reducing its hallucination rate from 85.5% to 37.7% (55.9% relative reduction) as measured by our detector. Cross-domain evaluation on the SciFact biomedical benchmark shows that general-domain training transfers poorly (F1=0.52), motivating domain-specific fine-tuning. PubMedBERT fine-tuned on SciFact achieves F1=0.63 and AUROC=0.81, demonstrating that domain-matched pre-training is the strongest adaptation strategy. Code and models are available at https://github.com/varunteja99/hallucination-detection-nlp</summary>645    <category term="cs.CL" scheme="http://arxiv.org/schemas/atom"/>646    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>647    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>648    <published>2026-09-10T17:45:36Z</published>649    <arxiv:comment>6 pages, 3 figures, 5 tables</arxiv:comment>650    <arxiv:primary_category term="cs.CL"/>651    <author>652      <name>Varun Teja Chundru</name>653    </author>654    <author>655      <name>Debasmita Biswas</name>656    </author>657  </entry>658  <entry>659    <id>http://arxiv.org/abs/2609.11877v1</id>660    <title>Biology-in-the-loop: Amortized Adaptive Hit Discovery in CRISPR Screens</title>661    <updated>2026-09-10T17:45:18Z</updated>662    <link href="https://arxiv.org/abs/2609.11877v1" rel="alternate" type="text/html"/>663    <link href="https://arxiv.org/pdf/2609.11877v1" rel="related" type="application/pdf" title="pdf"/>664    <summary>Many biological discovery problems require experiments to be selected sequentially under constrained budgets. CRISPR screening is a prominent example, as exhaustive perturbation testing is often infeasible and candidate perturbations must instead be prioritized over multiple experimental rounds. Despite the importance of this problem, existing benchmarks for adaptive hit discovery remain limited in scale and diversity. Here, we introduce AssayBench-Loop, a large-scale benchmark for adaptive hit discovery comprising 1,389 CRISPR screens across five phenotype categories. Beyond enabling systematic evaluation, its scale makes it possible to learn acquisition strategies across historical experiments. Building on this resource, we introduce AssayLoop, a sequential experimental design framework combining AssayFormer, a transformer-based amortized acquisition policy trained across historical screens to adapt from experimental feedback, with LLM-derived biological priors through an adaptive handoff. In this view, completed experiments become training data for learning how accumulated evidence should guide what to test next, while LLMs provide prior biological knowledge to seed the search. We further introduce AssayLLM, showing that the same principle can be extended directly to an LLM through task-specific post-training. On temporally held-out screens, AssayLoop achieves a 5.67-fold enrichment over random selection and recovers 27.7% of hits after assaying approximately 5% of the candidate library, outperforming existing adaptive-design methods and standalone LLMs, and AssayFormer alone. Performance improves with increasing historical training data and transfers to phenotype categories excluded from training. These results demonstrate the value of learning acquisition policies across historical experiments and combining them with broad biological priors for efficient adaptive hit discovery.</summary>665    <category term="q-bio.QM" scheme="http://arxiv.org/schemas/atom"/>666    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>667    <category term="cs.CL" scheme="http://arxiv.org/schemas/atom"/>668    <category term="q-bio.GN" scheme="http://arxiv.org/schemas/atom"/>669    <published>2026-09-10T17:45:18Z</published>670    <arxiv:primary_category term="q-bio.QM"/>671    <author>672      <name>Carl Edwards</name>673    </author>674    <author>675      <name>Edward De Brouwer</name>676    </author>677    <author>678      <name>Xiner Li</name>679    </author>680    <author>681      <name>Namkyeong Lee</name>682    </author>683    <author>684      <name>Ehsan Hajiramezanali</name>685    </author>686    <author>687      <name>Anne Biton</name>688    </author>689    <author>690      <name>Sara Mostafavi</name>691    </author>692    <author>693      <name>Gabriele Scalia</name>694    </author>695  </entry>696  <entry>697    <id>http://arxiv.org/abs/2609.11876v1</id>698    <title>On the Regularization Landscape for the Linear Recommendation Models</title>699    <updated>2026-09-10T17:45:06Z</updated>700    <link href="https://arxiv.org/abs/2609.11876v1" rel="alternate" type="text/html"/>701    <link href="https://arxiv.org/pdf/2609.11876v1" rel="related" type="application/pdf" title="pdf"/>702    <summary>Recently, a wide range of recommendation algorithms inspired by deep learning techniques have emerged as the performance leaders on several standard recommendation benchmarks. While these algorithms were built on different DL techniques (e.g., dropouts, autoencoder), they have similar performance and even similar cost functions. This paper studies whether the models' comparable performance are sheer coincidence, or they can be unified under a single framework. We find that all linear performance leaders effectively add only a nuclear-norm based regularizer, or a Frobenius-norm based regularizer. The former ones possess a (surprising) rigid structure that limits the models' predictive power but their solutions are low rank and have closed form. The latter ones are more expressive and more efficient for recommendation but their solutions are either full-rank or require executing hard-to-tune numeric procedures such as ADMM. Along this line of finding, we further propose two low-rank, closed-form solutions, derived from carefully generalizing Frobenius-norm based regularizers. The new solutions get the best of both nuclear-norm and Frobenius-norm world.</summary>703    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>704    <published>2026-09-10T17:45:06Z</published>705    <arxiv:primary_category term="cs.AI"/>706    <author>707      <name>Dong Li</name>708    </author>709    <author>710      <name>Zhenming Liu</name>711    </author>712    <author>713      <name>Ruoming Jin</name>714    </author>715    <author>716      <name>Hao Zhou</name>717    </author>718    <author>719      <name>Zhi Liu</name>720    </author>721    <author>722      <name>Jing Gao</name>723    </author>724    <author>725      <name>Bin Ren</name>726    </author>727  </entry>728  <entry>729    <id>http://arxiv.org/abs/2609.11873v1</id>730    <title>The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement</title>731    <updated>2026-09-10T17:44:23Z</updated>732    <link href="https://arxiv.org/abs/2609.11873v1" rel="alternate" type="text/html"/>733    <link href="https://arxiv.org/pdf/2609.11873v1" rel="related" type="application/pdf" title="pdf"/>734    <summary>Recursive self-improvement (RSI) enables AI systems to turn experience and feedback into persistent changes that improve both their capabilities and the process of future improvement. We first use the Headroom-Closed Index (HCI) to reveal the problems of existing LLMs, then introduce the RSI concept and its development roadmap: from improvement-execution autonomy, improvement-strategy autonomy, experience-acquisition autonomy, and environment-adaptation autonomy, to recursive meta-improvement. Next we examine RSI across scenarios (e.g., scientific discovery, embodied intelligence, software engineering), highlighting their distinct requirements and development speeds. Drawing on diverse industry practices and preliminary empirical evidence, we connect RSI research with practical systems and identify key challenges to achieving genuine RSI.</summary>735    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>736    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>737    <category term="cs.CL" scheme="http://arxiv.org/schemas/atom"/>738    <published>2026-09-10T17:44:23Z</published>739    <arxiv:primary_category term="cs.LG"/>740    <author>741      <name>Yi Duan</name>742    </author>743    <author>744      <name>Ying Liu</name>745    </author>746    <author>747      <name>Zirui Tang</name>748    </author>749    <author>750      <name>Haodong Chen</name>751    </author>752    <author>753      <name>Jun Zhou</name>754    </author>755    <author>756      <name>Yumou Liu</name>757    </author>758    <author>759      <name>Bangrui Xu</name>760    </author>761    <author>762      <name>Yukai Wu</name>763    </author>764    <author>765      <name>Sidi Chen</name>766    </author>767    <author>768      <name>Yuhan Zhou</name>769    </author>770    <author>771      <name>Haoyu Wang</name>772    </author>773    <author>774      <name>Xiaoyou Yu</name>775    </author>776    <author>777      <name>Shaokun Han</name>778    </author>779    <author>780      <name>Xuzhou Zhu</name>781    </author>782    <author>783      <name>Le Zhou</name>784    </author>785    <author>786      <name>Bolin Lu</name>787    </author>788    <author>789      <name>Wei Zhou</name>790    </author>791    <author>792      <name>Jiachen Liu</name>793    </author>794    <author>795      <name>Nuozhou Fang</name>796    </author>797    <author>798      <name>Jiaxin Tian</name>799    </author>800    <author>801      <name>Ruoyu Chen</name>802    </author>803    <author>804      <name>Yuxuan Li</name>805    </author>806    <author>807      <name>Kai Zuo</name>808    </author>809    <author>810      <name>Kaiyan Zhang</name>811    </author>812    <author>813      <name>Jiantao Qiu</name>814    </author>815    <author>816      <name>Conghui He</name>817    </author>818    <author>819      <name>Guoliang Li</name>820    </author>821    <author>822      <name>Bowen Zhou</name>823    </author>824    <author>825      <name>Zhiyuan Liu</name>826    </author>827    <author>828      <name>Zhoufutu Wen</name>829    </author>830    <author>831      <name>Jihua Kang</name>832    </author>833    <author>834      <name>Xuanhe Zhou</name>835    </author>836    <author>837      <name>Fan Wu</name>838    </author>839  </entry>840  <entry>841    <id>http://arxiv.org/abs/2609.11872v1</id>842    <title>Evaluating Time-Series Foundation Models and Multimodal Dietary Context for CGM Forecasting</title>843    <updated>2026-09-10T17:43:29Z</updated>844    <link href="https://arxiv.org/abs/2609.11872v1" rel="alternate" type="text/html"/>845    <link href="https://arxiv.org/pdf/2609.11872v1" rel="related" type="application/pdf" title="pdf"/>846    <summary>Continuous glucose monitoring (CGM) provides high-frequency measurements of glucose dynamics and enables short-term glucose forecasting for diabetes management. Although time-series foundation models have shown strong general forecasting ability, their effectiveness for CGM prediction and the added value of multimodal dietary context remain unclear. We conduct a comprehensive empirical study using eight public CGM datasets spanning Type 1 diabetes, Type 2 diabetes, and non-diabetes populations. Under a unified protocol across multiple context lengths and prediction horizons, zero-shot foundation models did not consistently outperform strong task-specific baselines such as Elastic Net and PatchTST. In contrast, lightweight fine-tuning substantially improved forecasting performance. For example, fine-tuned Chronos-Bolt reduced RMSE by 6.5%-18.4% in the T1D cohort and by 8.6%-18.2% in the non-diabetes/T2D cohort, with comparable improvements in both in-distribution and out-of-distribution test settings. We further evaluate multimodal dietary context using CGMacros, which provides temporally aligned CGM signals, food images, and macronutrient records. A residual-based fusion framework reduced overall RMSE by approximately 3% and postprandial RMSE by approximately 15% relative to the CGM-only baseline. Moreover, Chronos-based CGM representations were more strongly correlated with observed postprandial glucose increments than representations from LSTM and CatBoost, even after those models incorporated additional dietary modalities, suggesting that pretrained temporal representations better preserve meal-induced excursion patterns. These findings show that foundation models require CGM-specific adaptation for reliable forecasting and that dietary context provides clinically meaningful signals beyond CGM alone, especially during postprandial periods.</summary>847    <category term="stat.ML" scheme="http://arxiv.org/schemas/atom"/>848    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>849    <published>2026-09-10T17:43:29Z</published>850    <arxiv:primary_category term="stat.ML"/>851    <author>852      <name>Bowen Zhang</name>853    </author>854    <author>855      <name>Hsiu-Wen Cheng</name>856    </author>857    <author>858      <name>Hongyu Yang</name>859    </author>860    <author>861      <name>Evie L. Shen</name>862    </author>863    <author>864      <name>Joleen Vansomphone</name>865    </author>866    <author>867      <name>Yuna Li</name>868    </author>869    <author>870      <name>Kerry Zhou</name>871    </author>872    <author>873      <name>Zitian Qu</name>874    </author>875    <author>876      <name>Suning Zhao</name>877    </author>878    <author>879      <name>Xiangning Deng</name>880    </author>881    <author>882      <name>Hua Zhou</name>883    </author>884    <author>885      <name>Jin J. Zhou</name>886    </author>887  </entry>888  <entry>889    <id>http://arxiv.org/abs/2609.11870v1</id>890    <title>Augustinian BabyLM: What Ostensive Definition Can and Cannot Teach a Small Language Model</title>891    <updated>2026-09-10T17:43:09Z</updated>892    <link href="https://arxiv.org/abs/2609.11870v1" rel="alternate" type="text/html"/>893    <link href="https://arxiv.org/pdf/2609.11870v1" rel="related" type="application/pdf" title="pdf"/>894    <summary>A language model normally begins training with random word embeddings: whatever 'banana' means must be learned from training corpora. I implement St. Augustine's picture of word learning, meaning by ostension, for a small masked language model (DeBERTa) trained on 10M words: before training, visually grounded tokens receive embeddings derived from the image regions they label; other tokens start random. Visual initialization leaves a measurable imprint that lasts until the end of training. At the same time, the effect remains invisible under most BabyLM benchmarks, which probe abstract grammatical knowledge: visual initialization does not affect performance there. The only zero-shot exception is object-property knowledge (COMPS, Misra et al. 2023), where seeding helps in every configuration. To follow up on this result, I build a corpus-tailored version of the Visual-Property Swap benchmark (Lin et al., 2026), which tests color, material, size, and shape knowledge, with per-item training frequency and seeded status. Here, vision-seeded models have a persistent, seed- replicated advantage, confined to the seeded words. As a causal test, I show that synthetic grounding of previously unseeded words transfers the advantage to exactly those words. Function words and abstract vocabulary also receive strong visual seeds and retain them throughout training, and the training objective draws on them: held-out mask-prediction loss falls for these words in every seed. However, no benchmark I run registers this. What evaluation would pick this up remains an open question.</summary>895    <category term="cs.CL" scheme="http://arxiv.org/schemas/atom"/>896    <published>2026-09-10T17:43:09Z</published>897    <arxiv:primary_category term="cs.CL"/>898    <author>899      <name>Lisa Bylinina</name>900    </author>901  </entry>902  <entry>903    <id>http://arxiv.org/abs/2609.11867v1</id>904    <title>AdamX: Cosine similarity meets gradient descent</title>905    <updated>2026-09-10T17:42:50Z</updated>906    <link href="https://arxiv.org/abs/2609.11867v1" rel="alternate" type="text/html"/>907    <link href="https://arxiv.org/pdf/2609.11867v1" rel="related" type="application/pdf" title="pdf"/>908    <summary>We introduce AdamX, a first-order optimizer that incorporates cosine similarity as an adaptive mechanism for controlling update magnitudes. The proposed method is scalable, model-agnostic, and straightforward to integrate into existing training pipelines. We further introduce a variance rectification scheme that promotes smoother optimization during the early stages of training. Overall, we provide empirical evidence that AdamX achieves competitive convergence rates across a range of benchmark datasets and architectures. Performance is evaluated in terms of the number of epochs required to reach predefined performance thresholds under a fixed hyperparameter budget. Code and Experiments available at: https://github.com/FranciscoCaldas/adamX.</summary>909    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>910    <category term="math.OC" scheme="http://arxiv.org/schemas/atom"/>911    <published>2026-09-10T17:42:50Z</published>912    <arxiv:primary_category term="cs.LG"/>913    <author>914      <name>Francisco Caldas</name>915    </author>916    <author>917      <name>Ruben Belo</name>918    </author>919    <author>920      <name>Cláudia Soares</name>921    </author>922  </entry>923  <entry>924    <id>http://arxiv.org/abs/2609.11865v1</id>925    <title>Epistemic orientation predicts legislative effectiveness among members of the US Congress</title>926    <updated>2026-09-10T17:41:54Z</updated>927    <link href="https://arxiv.org/abs/2609.11865v1" rel="alternate" type="text/html"/>928    <link href="https://arxiv.org/pdf/2609.11865v1" rel="related" type="application/pdf" title="pdf"/>929    <summary>Truth and evidence-based communication provide important foundations for democratic governance, accountability, and collective decision-making. Prior work shows that evidence-oriented language in US congressional floor speeches has declined since the mid-1970s, alongside broader changes in legislative productivity and polarization. This study shifts the analysis from congressional sessions to individual members of Congress to examine whether epistemic orientation varies systematically across legislators and whether it relates to political behavior and legislative effectiveness. Using the Evidence-Minus-Intuition (EMI) score, we measure the relative prevalence of evidence-oriented versus intuition-oriented language in congressional floor speeches and Twitter posts. We link these measures to legislator-level data on ideology, institutional position, communication context, and Legislative Effectiveness Score (LES). The results show that more ideologically extreme members use less evidence-oriented language on the congressional floor. EMI also exhibits cross-platform consistency with members who use more evidence-oriented language in floor speeches also being more evidence-oriented on Twitter, although EMI is lower on Twitter overall. Finally, EMI in congressional speeches is positively associated with individual legislative effectiveness, even after accounting for ideology and extensive political, institutional, demographic, topical, and communication volume controls. These findings suggest that evidence-oriented language is not only an aggregate feature of congressional discourse but also a meaningful attribute of individual-level legislative communication and effectiveness.</summary>930    <category term="cs.CL" scheme="http://arxiv.org/schemas/atom"/>931    <published>2026-09-10T17:41:54Z</published>932    <arxiv:primary_category term="cs.CL"/>933    <author>934      <name>Segun Aroyehun</name>935    </author>936    <author>937      <name>Stephan Lewandowsky</name>938    </author>939    <author>940      <name>David Garcia</name>941    </author>942  </entry>943  <entry>944    <id>http://arxiv.org/abs/2609.11864v1</id>945    <title>RetroThinker: Enabling Retrospective Thinking in Speech LLMs</title>946    <updated>2026-09-10T17:41:53Z</updated>947    <link href="https://arxiv.org/abs/2609.11864v1" rel="alternate" type="text/html"/>948    <link href="https://arxiv.org/pdf/2609.11864v1" rel="related" type="application/pdf" title="pdf"/>949    <summary>Speech large language models (SpeechLLMs) offer reduced latency and retain paralinguistic nuances that are typically lost in cascaded automatic speech recognition (ASR) and text-based LM architectures. However, they continue to lag behind text-only LLMs on complex reasoning tasks, while real-time spoken interaction imposes strict latency constraints. Although prior works employ Chain-of-Thought (CoT) and concurrent reasoning to enhance reasoning capabilities without inducing prohibitive delays, an inherent accuracy-latency trade-off persists. In this paper, we investigate whether a streaming SpeechLLM can dynamically revise its reasoning traces on the fly. We introduce RetroThinker, a multi-stage post-training framework that equips the Moshi model to self-verify and forward-correct CoT steps during inference. RetroThinker combines supervised fine-tuning (SFT) on curated retrospective thinking data with length-based direct preference optimization (DPO) to optimize retrospective during early reasoning (i.e., reasoning concurrently while the user speaks). Evaluated on the GSM8K benchmark, RetroThinker significantly improves the accuracy-latency trade-off over non-retrospective baselines, achieving an 11% absolute accuracy gain at a comparable latency.</summary>950    <category term="eess.AS" scheme="http://arxiv.org/schemas/atom"/>951    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>952    <category term="cs.CL" scheme="http://arxiv.org/schemas/atom"/>953    <published>2026-09-10T17:41:53Z</published>954    <arxiv:comment>Accepted to IEEE SLT 2026</arxiv:comment>955    <arxiv:primary_category term="eess.AS"/>956    <author>957      <name>Yi-Jen Shih</name>958    </author>959    <author>960      <name>Puyuan Peng</name>961    </author>962    <author>963      <name>Abdelrahman Mohamed</name>964    </author>965    <author>966      <name>David Harwath</name>967    </author>968  </entry>969  <entry>970    <id>http://arxiv.org/abs/2609.11860v1</id>971    <title>Explainability Assistant: A Conversational XAI Interface for Interpreting Energy Consumption Models</title>972    <updated>2026-09-10T17:40:11Z</updated>973    <link href="https://arxiv.org/abs/2609.11860v1" rel="alternate" type="text/html"/>974    <link href="https://arxiv.org/pdf/2609.11860v1" rel="related" type="application/pdf" title="pdf"/>975    <summary>Energy consumption forecasting relies on increasingly complex machine learning (ML) models, such as Genetic Programming-based symbolic regressors, whose predictions can be difficult for facility managers and building operators to interpret. Explainable Artificial Intelligence (XAI) techniques address this opacity, but traditional XAI dashboards require substantial technical expertise and provide limited flexibility for dynamic, context-aware inquiry. Conversational XAI systems offer a promising alternative; however, previous approaches, such as TalkToModel, were constrained by rigid custom grammars and achieved only 76.8% intent-parsing accuracy. This paper introduces the Explainability Assistant, an open-source conversational XAI system that leverages the function-calling capabilities of modern Large Language Models (LLMs) to overcome these limitations. The system achieves 94% intent-parsing accuracy, supports flexible natural language interaction, and adapts to different ML problem types without task-specific fine-tuning. We present the system's architecture and report results from a comparative evaluation conducted with energy domain specialists, contrasting the Explainability Assistant with a traditional XAI dashboard. The evaluation suggests improved usability and consistent task accuracy, with all experts unanimously preferring the conversational interface for practical use.</summary>976    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>977    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>978    <published>2026-09-10T17:40:11Z</published>979    <arxiv:comment>11 pages, 3 figures. Accepted author version of a paper published at ICECET 2026</arxiv:comment>980    <arxiv:primary_category term="cs.AI"/>981    <arxiv:journal_ref>2026 6th International Conference on Electrical, Computer and Energy Technologies (ICECET), Rome, Italy, 2026</arxiv:journal_ref>982    <author>983      <name>Rodion Krjutškov</name>984    </author>985    <author>986      <name>Eduard Barbu</name>987    </author>988    <author>989      <name>Nikos Sakkas</name>990    </author>991    <author>992      <name>Sofia Yfanti</name>993    </author>994    <arxiv:doi>10.1109/ICECET65726.2026.11632877</arxiv:doi>995    <link rel="related" href="https://doi.org/10.1109/ICECET65726.2026.11632877" title="doi"/>996  </entry>997  <entry>998    <id>http://arxiv.org/abs/2609.11859v1</id>999    <title>From Parameters to Answers: How LLMs Retrieve and Use Their Internal Knowledge</title>1000    <updated>2026-09-10T17:39:55Z</updated>1001    <link href="https://arxiv.org/abs/2609.11859v1" rel="alternate" type="text/html"/>1002    <link href="https://arxiv.org/pdf/2609.11859v1" rel="related" type="application/pdf" title="pdf"/>1003    <summary>How does a language model's dependence on query-routing information and target knowledge change as it answers a question? We study this question through layerwise interventions on the hidden state at the end of the question. Across Qwen, Llama, and Gemma, we compare country-continent questions with noun, adjective, and code answers while keeping several fitted measurements distinct. A pair-conditioned request direction describes which country is queried in natural single-country questions; a global request direction describes first- versus second-country requests in paired questions; separate selection candidates test control among contents already available in the hidden state. A diagnostic reanalysis of frozen Qwen natural-question states shows that the pair-conditioned direction grows stronger before interventions on it begin to alter later fitted knowledge, with this causal window opening while answer-supporting content is still forming. The paired three-model trajectories are not uniform: Gemma shows a partially overlapping mid-layer routing-content profile, whereas Llama has no sustained routing-effect window under the same gates. In the paired protocol, dependence on the global request direction decreases from fixed earlier to later layer sets while dependence on fitted content persists. A matched Qwen comparison shows that the pair-conditioned direction retains a late effect, so this operational handoff concerns the global fitted direction rather than all request information. These results separate early readability, natural strength, causal steering, and later content dependence.</summary>1004    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>1005    <published>2026-09-10T17:39:55Z</published>1006    <arxiv:comment>53 pages, 13 figures, including appendices</arxiv:comment>1007    <arxiv:primary_category term="cs.AI"/>1008    <author>1009      <name>Wenkang Wei</name>1010    </author>1011    <author>1012      <name>Yuan Fang</name>1013    </author>1014    <author>1015      <name>Renhe Jiang</name>1016    </author>1017    <author>1018      <name>Hong Cheng</name>1019    </author>1020    <author>1021      <name>Xingtong Yu</name>1022    </author>1023  </entry>1024  <entry>1025    <id>http://arxiv.org/abs/2609.11851v1</id>1026    <title>IndicTriMix: Developing Language Identification Datasets and Models for Tri-Language Code-Mixing</title>1027    <updated>2026-09-10T17:36:12Z</updated>1028    <link href="https://arxiv.org/abs/2609.11851v1" rel="alternate" type="text/html"/>1029    <link href="https://arxiv.org/pdf/2609.11851v1" rel="related" type="application/pdf" title="pdf"/>1030    <summary>Language identification in code-mixed text, largely observed in social media, is highly essential when users frequently switch between multiple languages within a single utterance. Accurately identifying the languages of code-mixed tokens becomes an urgent necessity. Traditional language identification models, designed for monolingual text, are not well suited for token-level language identification in code-mixed settings. We formulate the task as a sequence labeling problem and fine-tune contextual transformer-based models MuRIL and XLM-RoBERTa best suited for Indian languages. We evaluate these systems on three different data configurations (Hindi, Gujarati, and Bengali) to predict language labels for individual tokens. We release a benchmark for language identification in code-mixed tokens with manually annotated test sets. We propose two approaches of code-mixed generation using parallel sentences of three languages. The trained models demonstrate the effectiveness of contextual embeddings for token-level language identification in multilingual social media text. For reproducibility and to facilitate future research, we publicly release our fine-tuned models.</summary>1031    <category term="cs.CL" scheme="http://arxiv.org/schemas/atom"/>1032    <published>2026-09-10T17:36:12Z</published>1033    <arxiv:comment>9 pages, 9 tables</arxiv:comment>1034    <arxiv:primary_category term="cs.CL"/>1035    <author>1036      <name>Pruthwik Mishra</name>1037    </author>1038    <author>1039      <name>Rudra Trivedi</name>1040    </author>1041    <author>1042      <name>Avi Patel</name>1043    </author>1044    <author>1045      <name>Ashok Urlana</name>1046    </author>1047    <author>1048      <name>Shrikant Malviya</name>1049    </author>1050  </entry>1051  <entry>1052    <id>http://arxiv.org/abs/2609.11842v1</id>1053    <title>Model-Aware Schedules Improve Generation via Fiberwise Optimal Transport</title>1054    <updated>2026-09-10T17:30:44Z</updated>1055    <link href="https://arxiv.org/abs/2609.11842v1" rel="alternate" type="text/html"/>1056    <link href="https://arxiv.org/pdf/2609.11842v1" rel="related" type="application/pdf" title="pdf"/>1057    <summary>Diffusion and flow-matching schedules control the signal and noise coefficients that mix data and noise along affine probability paths. Minimizing a kinetic action defined on coefficient paths, motivated by optimal transport, helps explain strong baselines but remains model-agnostic and ignores prediction error. Here we introduce a model-aware schedule construction based on fiberwise optimal transport. At a fixed time and state on the probability path, compatible signal/noise decompositions form an affine fiber. We define a fiberwise prediction risk by averaging optimal-transport costs between the true and predictor-induced decompositions within these fibers. On a fixed coefficient curve, combining this risk with coefficient-path kinetic action yields a closed-form optimal time allocation. This construction extends to general linear prediction targets, and the risk profile can be estimated from an early baseline checkpoint. We evaluate DDPMs and flow matching across prediction targets, training configurations, risk-estimation checkpoints, datasets, and architectures. Our model-aware schedules consistently outperform strong baselines, including a 38.6% relative FID reduction for flow matching on CIFAR-10 at 16 function evaluations. Each model-agnostic kinetic baseline determines its own kinetic reference coordinate. In these coordinates, fiberwise-risk profiles from independently trained models in different settings align closely after normalization to unit area. The resulting schedule deformations used in training also align, suggesting empirical universality across the evaluated models and settings. Pretrained-checkpoint diagnostics extend this normalized-risk agreement to larger conditional latent diffusion and 2-RF models. A frozen analytic allocation template retains most of the model-aware improvement without further risk estimation or model-specific fitting.</summary>1058    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>1059    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>1060    <published>2026-09-10T17:30:44Z</published>1061    <arxiv:comment>11 pages, 2 figures. Keywords: Diffusion models; flow matching; schedule optimization; fiberwise optimal transport; time reparameterization; empirical universality</arxiv:comment>1062    <arxiv:primary_category term="cs.LG"/>1063    <author>1064      <name>Luyi Jia</name>1065    </author>1066    <author>1067      <name>Boyan Zhang</name>1068    </author>1069    <author>1070      <name>Yilun Liu</name>1071    </author>1072    <author>1073      <name>Steffen Rulands</name>1074    </author>1075  </entry>1076  <entry>1077    <id>http://arxiv.org/abs/2609.11838v1</id>1078    <title>Target leakage, not model class, explains reported accuracy in survey-based cardiovascular screening: a leakage-tiered audit of glass-box and tabular foundation models</title>1079    <updated>2026-09-10T17:29:15Z</updated>1080    <link href="https://arxiv.org/abs/2609.11838v1" rel="alternate" type="text/html"/>1081    <link href="https://arxiv.org/pdf/2609.11838v1" rel="related" type="application/pdf" title="pdf"/>1082    <summary>Cardiovascular screening models trained on national health surveys routinely report areas under the receiver operating characteristic curve (AUROC) near 0.89. We asked whether that accuracy reflects learning or target leakage, whether tabular foundation models change the answer, and whether the properties deployment requires survive joint examination. We benchmarked ten classifiers spanning linear, tree-ensemble, neural, glass-box, and tabular foundation classes for prevalent myocardial infarction in 442,067 respondents of the 2022 Behavioral Risk Factor Surveillance System across five feature tiers of decreasing leakage risk. Each was audited for discrimination, calibration, fairness at an explicit screening threshold, conformal coverage, explanation faithfulness, and inference cost, then applied -- models and thresholds frozen -- to 430,755 respondents of 2023. Removing two post-diagnostic features cost every model 0.049-0.051 AUROC, collapsing the field into a 0.0045-wide band. The glass-box explainable boosting machine was non-inferior to every alternative within a pre-specified 0.005 margin while scoring the cohort roughly 104 times faster than the strongest foundation model. One threshold detected 75.4% of women's infarctions against 89.0% of men's; editing the model's shape functions reduced the gap to 0.010. Marginal conformal prediction gave 0.86 coverage to men and 0.82 to adults over 60; Mondrian calibration repaired every stratum. Frozen models transported within 0.002 AUROC. Reported headroom in this literature is a property of the feature set, not the learner. Transparency cost nothing measurable and made fairness repair and uncertainty conditioning directly auditable. Evaluation practice, not model capacity, is the binding constraint.</summary>1083    <category term="cs.CL" scheme="http://arxiv.org/schemas/atom"/>1084    <published>2026-09-10T17:29:15Z</published>1085    <arxiv:primary_category term="cs.CL"/>1086    <author>1087      <name>Raad Bin Tareaf</name>1088    </author>1089    <author>1090      <name>Murad Al-Rajab</name>1091    </author>1092    <author>1093      <name>Samia Loucif</name>1094    </author>1095    <author>1096      <name>Samer Ellaham</name>1097    </author>1098    <author>1099      <name>Cedric Schmitz</name>1100    </author>1101  </entry>1102  <entry>1103    <id>http://arxiv.org/abs/2609.11807v1</id>1104    <title>Near-Optimal Reinforcement Learning with Multi-Step Transition Lookahead</title>1105    <updated>2026-09-10T16:58:21Z</updated>1106    <link href="https://arxiv.org/abs/2609.11807v1" rel="alternate" type="text/html"/>1107    <link href="https://arxiv.org/pdf/2609.11807v1" rel="related" type="application/pdf" title="pdf"/>1108    <summary>We study reinforcement learning (RL) with transition look-ahead, where the agent may observe which states would be visited upon playing any sequence of $\ell$ actions before deciding its course of action. Although look-ahead can substantially improve achievable performance, it is known that optimal planning with multi-step transition look-ahead is NP-hard, but this hardness was established using discount factors arbitrarily close to one. It was therefore unknown whether the problem remains hard for any discount factor, and whether near-optimal planning can nevertheless be performed efficiently. We resolve both questions. First, we show that for every fixed rational discount factor ($γ\in(0,1)$), exact planning remains NP-hard. Second, we introduce a randomized polynomial-time approximation scheme for every fixed look-ahead depth. We then extend our approach to unknown transitions and stochastic rewards using optimism and variance-adaptive confidence bounds. The resulting algorithm achieves cumulative regret whose leading term matches classical tabular discounted RL up to logarithmic factors. Thus, although exact planning with transition look-ahead is NP-hard, efficient near-optimal planning and learning remain possible.</summary>1109    <category term="stat.ML" scheme="http://arxiv.org/schemas/atom"/>1110    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>1111    <published>2026-09-10T16:58:21Z</published>1112    <arxiv:primary_category term="stat.ML"/>1113    <author>1114      <name>Corentin Pla</name>1115    </author>1116    <author>1117      <name>Hugo Richard</name>1118    </author>1119    <author>1120      <name>Marc Abeille</name>1121    </author>1122    <author>1123      <name>Vianney Perchet</name>1124    </author>1125  </entry>1126  <entry>1127    <id>http://arxiv.org/abs/2609.11805v1</id>1128    <title>Understanding Operator Attitudes Toward AI-Supported Decision Making in Maritime Operations</title>1129    <updated>2026-09-10T16:57:44Z</updated>1130    <link href="https://arxiv.org/abs/2609.11805v1" rel="alternate" type="text/html"/>1131    <link href="https://arxiv.org/pdf/2609.11805v1" rel="related" type="application/pdf" title="pdf"/>1132    <summary>Maritime Autonomous Surface Ships (MASS) and AI- supported decision assistants are expected to transform maritime operations, but their safe integration depends on how maritime professionals perceive and trust such systems. This paper presents a survey study on maritime stakeholders' attitudes toward an AI-supported assistant in collision-avoidance scenarios. Participants evaluated technology anxiety, trust in automation, and explanation quality using established and adapted questionnaires, complemented by sentiment and thematic analysis of open-ended responses Results indicate a generally positive disposition toward maritime technology, no clear age-related differences in openness, stable trust across scenarios, and more scenario-sensitive, multidimensional explanation ratings. Open responses showed that participants valued support for decision-making, situation awareness, and confidence-building, while raising concerns about AI reliability, over- reliance and loss of expertise. The findings suggest that maritime AI systems should not focus solely on increasing automation or trust, but on supporting calibrated reliance through transparent, reliable, and operationally meaningful design with domain experts in the loop.</summary>1133    <category term="cs.HC" scheme="http://arxiv.org/schemas/atom"/>1134    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>1135    <published>2026-09-10T16:57:44Z</published>1136    <arxiv:comment>11 figures</arxiv:comment>1137    <arxiv:primary_category term="cs.HC"/>1138    <author>1139      <name>Doreen Jirak</name>1140    </author>1141    <author>1142      <name>Armeen Saroukanoff</name>1143    </author>1144    <author>1145      <name>Dirk van Rooy</name>1146    </author>1147  </entry>1148  <entry>1149    <id>http://arxiv.org/abs/2609.11804v1</id>1150    <title>Logit Refiner: Improving Visual Autoregressive Models via Intra-Scale Dependency Modeling</title>1151    <updated>2026-09-10T16:57:05Z</updated>1152    <link href="https://arxiv.org/abs/2609.11804v1" rel="alternate" type="text/html"/>1153    <link href="https://arxiv.org/pdf/2609.11804v1" rel="related" type="application/pdf" title="pdf"/>1154    <summary>Visual Autoregressive Models (VAR) generate images through next-scale prediction, producing all tokens within each scale in parallel. We show that this parallel decoding constitutes a mean-field-style approximation that discards spatial dependencies among same-scale tokens, causing locally incoherent samples regardless of backbone capacity -- a limitation of the decoding rule. Addressing this limitation, we introduce the Logit Refiner, a lightweight autoregressive module that restores intra-scale dependencies by sequentially sampling tokens conditioned on frozen backbone features. Adding only ~10% parameters and less than 5% of the base model's training compute, it plugs into any pretrained VAR checkpoint without retraining. Controlled ablations isolate joint intra-scale sampling -- rather than additional capacity or training -- as the critical ingredient. Across backbones from 310M to 2B parameters on class-conditional ImageNet 256x256, the refiner consistently improves generation quality, enabling a 1.1B-parameter model to surpass one twice its size. The approach further generalizes to text-to-image generation, confirming that the mean-field bottleneck persists across VAR variants and is effectively alleviated by our method.1155  Project page: https://compvis.github.io/logit-refiner/</summary>1156    <category term="cs.CV" scheme="http://arxiv.org/schemas/atom"/>1157    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>1158    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>1159    <published>2026-09-10T16:57:05Z</published>1160    <arxiv:comment>ECCV 2026</arxiv:comment>1161    <arxiv:primary_category term="cs.CV"/>1162    <author>1163      <name>Meimingwei Li</name>1164    </author>1165    <author>1166      <name>Stefan Andreas Baumann</name>1167    </author>1168    <author>1169      <name>Felix Krause</name>1170    </author>1171    <author>1172      <name>Björn Ommer</name>1173    </author>1174  </entry>1175  <entry>1176    <id>http://arxiv.org/abs/2609.11801v1</id>1177    <title>Thinking with Looped Flows</title>1178    <updated>2026-09-10T16:52:54Z</updated>1179    <link href="https://arxiv.org/abs/2609.11801v1" rel="alternate" type="text/html"/>1180    <link href="https://arxiv.org/pdf/2609.11801v1" rel="related" type="application/pdf" title="pdf"/>1181    <summary>Humans and machines often solve harder problems by spending more time on computation. In deep learning, looped models implement this idea during inference by recurrently updating a hidden state. In practice, however, their training backpropagates through only one or a few updates, making it hard to train early updates to support future ones. We propose looped flows, an approach that sidesteps this issue by training the recurrence with local denoising objectives. By imposing temporal association across denoising objectives through progressively decreasing noise levels and shared noise, the model is incentivized to learn recurrent states that transfer useful computation over time, even when gradients cover only a few updates. We then formulate inference as integrating the velocity of a probability flow parameterized by the learned denoiser, coupled with recurrent states. This allows solving harder problems by spending more computation through a finer temporal grid and enables multiple valid predictions from different initial noise samples. Across six reasoning benchmarks including two multi-solution benchmarks, looped flows outperform prior state-of-the-art looped models overall, achieving 58.8% test accuracy on ARC-AGI-1 and 12.2% on ARC-AGI-2.</summary>1182    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>1183    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>1184    <published>2026-09-10T16:52:54Z</published>1185    <arxiv:primary_category term="cs.LG"/>1186    <author>1187      <name>Ayhan Suleymanzade</name>1188    </author>1189    <author>1190      <name>Chanhyuk Lee</name>1191    </author>1192    <author>1193      <name>Floor Eijkelboom</name>1194    </author>1195    <author>1196      <name>Nicholas M. Boffi</name>1197    </author>1198    <author>1199      <name>İsmail İlkan Ceylan</name>1200    </author>1201    <author>1202      <name>Jinwoo Kim</name>1203    </author>1204  </entry>1205  <entry>1206    <id>http://arxiv.org/abs/2609.11799v1</id>1207    <title>SpecGuard: Inference-Time Backdoor Detection For Free</title>1208    <updated>2026-09-10T16:51:59Z</updated>1209    <link href="https://arxiv.org/abs/2609.11799v1" rel="alternate" type="text/html"/>1210    <link href="https://arxiv.org/pdf/2609.11799v1" rel="related" type="application/pdf" title="pdf"/>1211    <summary>Large language models are often fine-tuned, shared, or downloaded from third parties, so a deployed model may carry a hidden backdoor that behaves normally on benign inputs but switches to attacker-controlled behavior when a secret trigger appears. While backdoors can be audited before deployment, runtime monitoring remains important for models that are frequently updated. The challenge is that LLM serving is latency-sensitive: existing inference-time detectors either rely on assumptions about the trigger form, which can fail on stealthy attacks, or require extra model computation, such as input perturbations or an additional generation pass.1212  We introduce SpecGuard, an inference-time backdoor detector that repurposes speculative decoding at zero added model-computation cost. Speculative decoding speeds up inference by using a small draft model to propose tokens and a target model to verify them. We observe that this verification process already exposes a useful signal: when a backdoor is triggered, the target model shifts toward the attacker's behavior, while a clean draft model does not predict this shift, causing the draft-token acceptance rate to change.1213  We formalize when this signal appears and show that an attacker who suppresses it must also weaken the backdoor. Across diverse backdoor types and model families, SpecGuard reliably detects triggered behavior, including stealthy cases where input-level filters are blind, while avoiding the extra generation cost of existing runtime detectors. Speculative decoding therefore doubles as a free, always-on signal for detecting backdoored LLM behavior.</summary>1214    <category term="cs.CR" scheme="http://arxiv.org/schemas/atom"/>1215    <category term="cs.CL" scheme="http://arxiv.org/schemas/atom"/>1216    <published>2026-09-10T16:51:59Z</published>1217    <arxiv:primary_category term="cs.CR"/>1218    <author>1219      <name>Rui Wen</name>1220    </author>1221    <author>1222      <name>Ahmed Salem</name>1223    </author>1224    <author>1225      <name>Andrew Paverd</name>1226    </author>1227    <author>1228      <name>Mark Russinovich</name>1229    </author>1230    <author>1231      <name>Zheng Li</name>1232    </author>1233  </entry>1234  <entry>1235    <id>http://arxiv.org/abs/2609.11790v1</id>1236    <title>Dynamic language model representations for multi-objective reaction optimisation</title>1237    <updated>2026-09-10T16:43:08Z</updated>1238    <link href="https://arxiv.org/abs/2609.11790v1" rel="alternate" type="text/html"/>1239    <link href="https://arxiv.org/pdf/2609.11790v1" rel="related" type="application/pdf" title="pdf"/>1240    <summary>Optimising chemical reactions across multiple objectives, such as yield, selectivity, and safety, is central to chemical synthesis, and model-driven approaches depend critically on how reaction components are represented. Established featurisations are either chemically uninformative, as with one-hot encodings, or, as with molecular descriptors, do not readily extend across chemically distinct components. For structurally and functionally diverse components, it is therefore unclear what a shared representation should contain. Constructing such a representation is itself a challenging research undertaking that must be revisited for each new reaction system. Here we bypass this step by learning the reaction representation dynamically from text. Textual descriptions of reaction conditions are encoded by a fine-tuned language model trained jointly with Gaussian process surrogates, yielding task-adaptive representations within a multi-objective Bayesian optimisation loop. Across nickel- and palladium-catalysed cross-couplings in both sequential and parallel experimentation regimes, this approach reaches optimisation convergence in fewer experiments than descriptor libraries or one-hot encoding. Applied prospectively to a palladium-catalysed cyanation spanning mixed ligand denticity and heterogeneous additives, and to a three-objective asymmetric hydrogenation across chiral iridium and ruthenium catalyst families, two rounds of high-throughput experimentation (192 reactions, under 3% of each design space) delivered conditions translating directly to gram scale in 94% and 84% isolated yield, the latter at 99.6% enantiomeric excess.</summary>1241    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>1242    <published>2026-09-10T16:43:08Z</published>1243    <arxiv:primary_category term="cs.LG"/>1244    <author>1245      <name>Joshua W. Sin</name>1246    </author>1247    <author>1248      <name>David Ming Segura</name>1249    </author>1250    <author>1251      <name>Bojana Ranković</name>1252    </author>1253    <author>1254      <name>Siu Lun Chau</name>1255    </author>1256    <author>1257      <name>Marius D. R. Lutz</name>1258    </author>1259    <author>1260      <name>Andrea Anelli</name>1261    </author>1262    <author>1263      <name>Ryan P. Burwood</name>1264    </author>1265    <author>1266      <name>Kurt Püntener</name>1267    </author>1268    <author>1269      <name>Maximilian J. Notheis</name>1270    </author>1271    <author>1272      <name>Raphael Bigler</name>1273    </author>1274    <author>1275      <name>Philippe Schwaller</name>1276    </author>1277  </entry>1278  <entry>1279    <id>http://arxiv.org/abs/2609.11786v1</id>1280    <title>Beyond Word Error Rate: A Switch Aware Evaluation of ASR and Audio Language Models on English Yoruba Code-Switched Speech</title>1281    <updated>2026-09-10T16:36:19Z</updated>1282    <link href="https://arxiv.org/abs/2609.11786v1" rel="alternate" type="text/html"/>1283    <link href="https://arxiv.org/pdf/2609.11786v1" rel="related" type="application/pdf" title="pdf"/>1284    <summary>Automatic speech recognition (ASR) systems and audio language models (audio LMs) now report low error rates on monolingual benchmarks, but their behavior on code switched speech in low resource, diacritic rich languages remains poorly characterized. We present a switch aware evaluation of eleven modern systems (six ASR models and five audio LMs) on English Yoruba code-switched speech, using a deterministic 2000 utterance evaluation set and a shared scoring pipeline. Beyond word error rate (WER), we report switch localized diagnostics: a switch entry token error rate (SETER), windowed switch point error rates, language specific error rates, and a diacritic insensitive WER. Our central finding is that aggregate WER hides code switching behavior. The best system by WER (an ASR model) is statistically indistinguishable from a leading audio LM on WER, yet the audio LM is significantly better on every switch localized metric. Across faithful systems, Yoruba token recognition collapses (error 0.97 for almost all systems) while English tokens are recognized far better, and errors concentrate sharply at switches into Yoruba. Several generative audio LMs fail as exact transcribers, producing translation, verbosity, and prompt leakage that are strongly prompt dependent. We release manifests, metric implementations, and evaluation scripts to support reproducible, switch aware benchmarking for African code switched speech.</summary>1285    <category term="cs.CL" scheme="http://arxiv.org/schemas/atom"/>1286    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>1287    <published>2026-09-10T16:36:19Z</published>1288    <arxiv:comment>Accepted to IEEE Speech Language Tecnology</arxiv:comment>1289    <arxiv:primary_category term="cs.CL"/>1290    <author>1291      <name>Chibuzor Okocha</name>1292    </author>1293    <author>1294      <name>Christan Earl Grant</name>1295    </author>1296  </entry>1297  <entry>1298    <id>http://arxiv.org/abs/2609.11780v1</id>1299    <title>Predicting Privacy Leakage from Weight Spectral Density</title>1300    <updated>2026-09-10T16:34:25Z</updated>1301    <link href="https://arxiv.org/abs/2609.11780v1" rel="alternate" type="text/html"/>1302    <link href="https://arxiv.org/pdf/2609.11780v1" rel="related" type="application/pdf" title="pdf"/>1303    <summary>Membership inference attacks (MIAs) are widely used to audit the privacy disclosure risk of machine learning models, however current state-of-the-art attacks require training computationally expensive shadow models, making large-scale privacy evaluation impractical. In this work, we investigate whether inexpensive spectral metrics derived from the heavy-tailed self-regularisation framework can serve as proxies for MIA vulnerability. We evaluate several WeightWatcher spectral metrics on image and tabular classification tasks and compare their relationship with MIA privacy leakage against conventional measures of generalisation. Across datasets, stable rank exhibits a strong positive correlation with overall MIA success, while Log alpha-Norm shows a consistent negative correlation with MIA vulnerability at the low false-positive regime. These associations are observed to be stronger than those obtained using the generalisation gap. The results indicate that neural network spectra may contain information about privacy leakage that is not fully captured by conventional measures of overfitting, motivating spectral analysis as a promising direction for scalable privacy auditing.</summary>1304    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>1305    <category term="cs.CR" scheme="http://arxiv.org/schemas/atom"/>1306    <category term="cs.NE" scheme="http://arxiv.org/schemas/atom"/>1307    <published>2026-09-10T16:34:25Z</published>1308    <arxiv:primary_category term="cs.LG"/>1309    <author>1310      <name>Richard J. Preen</name>1311    </author>1312    <author>1313      <name>Jim Smith</name>1314    </author>1315  </entry>1316  <entry>1317    <id>http://arxiv.org/abs/2609.11777v1</id>1318    <title>Differentially Private EEG Feature Anonymization: A Privacy-Utility Case Study in Clinical Neurophysiology</title>1319    <updated>2026-09-10T16:31:11Z</updated>1320    <link href="https://arxiv.org/abs/2609.11777v1" rel="alternate" type="text/html"/>1321    <link href="https://arxiv.org/pdf/2609.11777v1" rel="related" type="application/pdf" title="pdf"/>1322    <summary>Clinical electroencephalography (EEG) data are valuable for healthcare research and for developing artificial intelligence (AI)-based clinical decision-support systems, but EEG recordings and derived features may contain sensitive patient-specific information. This creates privacy risks when data are reused, analyzed, or shared across clinical and research environments. Conventional anonymization methods are often insufficient for high-dimensional biomedical signals, since removing direct identifiers does not necessarily prevent re-identification, linkage, or inference risks. At the same time, strong privacy protection may distort clinically relevant signal characteristics and reduce data utility. This paper studies subject-level differential privacy for protecting clinical EEG-derived feature representations using Gaussian and Laplace perturbations. The proposed framework considers three deployment scenarios: client-side anonymization, centralized server-side anonymization, and decentralized local training. Following EEG preprocessing and feature extraction, Gaussian and Laplace perturbations are applied to the resulting patient-level EEG feature representations. The Laplace experiments evaluate the implemented noise scales, while the scales required for formal full-vector calibration are derived separately. The effects of both perturbations are assessed using statistical utility measures and a downstream machine-learning-based utility check. The results show that differentially private perturbation can be integrated into EEG processing workflows, but the selected mechanism, privacy parameters, and sensitivity calibration strongly influence data utility. The study highlights the practical privacy-utility trade-off in DP-based EEG feature anonymization and the challenges of preserving downstream utility in small and imbalanced clinical EEG datasets.</summary>1323    <category term="cs.CR" scheme="http://arxiv.org/schemas/atom"/>1324    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>1325    <category term="eess.SP" scheme="http://arxiv.org/schemas/atom"/>1326    <published>2026-09-10T16:31:11Z</published>1327    <arxiv:comment>27 pages, 11 figures, 7 tables</arxiv:comment>1328    <arxiv:primary_category term="cs.CR"/>1329    <author>1330      <name>Noman Sadiq</name>1331    </author>1332    <author>1333      <name>Mohsen Toorani</name>1334    </author>1335  </entry>1336  <entry>1337    <id>http://arxiv.org/abs/2609.11772v1</id>1338    <title>Whisper-Based Speech Transcription from Videos Across Multiple Languages for Cross-Cultural Understanding</title>1339    <updated>2026-09-10T16:27:47Z</updated>1340    <link href="https://arxiv.org/abs/2609.11772v1" rel="alternate" type="text/html"/>1341    <link href="https://arxiv.org/pdf/2609.11772v1" rel="related" type="application/pdf" title="pdf"/>1342    <summary>Cross-cultural understanding has become increasingly important in today's highly connected, cross-national world. The success of LLM-based technologies is now driving the development of automated tools to aid understanding for nonnative people trying to succeed in cross-cultural environments. Building such automated tools is often done by leveraging in-thewild text, audio, and video data. This paper presents techniques for improving speech recognition-based transcript creation in multiple languages from videos to better train these automated tools. The focus is on processes and speech tools that can easily be used by cross-cultural tool builders without requiring deep speech processing expertise. Using publicly available videos from YouTube and Whisper-based tools, average transcription error rate across seven languages (Spanish, Japanese, Korean, Mandarin, Turkish, Russian, and Hebrew) of 30% are observed. With a modest amount of fine-tuning data, the average error rate can be reduced to 20% making such output much more usable for downstream processing. Speech and metadata associated with these videos that can be used by the community to further refine these experiments are released as well.</summary>1343    <category term="eess.AS" scheme="http://arxiv.org/schemas/atom"/>1344    <category term="cs.CL" scheme="http://arxiv.org/schemas/atom"/>1345    <published>2026-09-10T16:27:47Z</published>1346    <arxiv:comment>7 pages, 2 figures, 5 tables</arxiv:comment>1347    <arxiv:primary_category term="eess.AS"/>1348    <author>1349      <name>Michael Picheny</name>1350    </author>1351  </entry>1352  <entry>1353    <id>http://arxiv.org/abs/2609.11770v1</id>1354    <title>The widening evaluation gap in medical large language model research 2023 to 2026</title>1355    <updated>2026-09-10T16:25:01Z</updated>1356    <link href="https://arxiv.org/abs/2609.11770v1" rel="alternate" type="text/html"/>1357    <link href="https://arxiv.org/pdf/2609.11770v1" rel="related" type="application/pdf" title="pdf"/>1358    <summary>Large language models are superseded every few quarters; clinical evidence takes years. We asked whether medical research is keeping pace with the systems it evaluates. PubMed returned 11,628 records for January 2023 to June 2026 across fourteen clinical domains, growing 45-fold; 2.5% used a randomised, controlled or prospective design. Evaluation lag, from a study's newest named model release to its own publication, widened from 1.33 to 6.08 quarters. Because discontinued models age mechanically, we benchmarked this against a counterfactual holding model composition fixed: migration to newer systems offset only 56% of the drift (95% CI 50-65). Randomised trials evaluated models a median 4.6 quarters older than other designs (P = 3 x 10^-19), yet among studies naming a model still under development no design differed from any other; 62% of randomised trials evaluated a discontinued family. Rigour and currency are in tension, and that tension reflects model selection rather than research timelines.</summary>1359    <category term="cs.CL" scheme="http://arxiv.org/schemas/atom"/>1360    <published>2026-09-10T16:25:01Z</published>1361    <arxiv:primary_category term="cs.CL"/>1362    <author>1363      <name>Raad Bin Tareaf</name>1364    </author>1365    <author>1366      <name>Murad Al-Rajab</name>1367    </author>1368    <author>1369      <name>Samia Loucif</name>1370    </author>1371  </entry>1372  <entry>1373    <id>http://arxiv.org/abs/2609.11769v1</id>1374    <title>Recognizing Is Not Reversing: A Controlled Inversion Test of Fact-Preserving News Framing</title>1375    <updated>2026-09-10T16:23:51Z</updated>1376    <link href="https://arxiv.org/abs/2609.11769v1" rel="alternate" type="text/html"/>1377    <link href="https://arxiv.org/pdf/2609.11769v1" rel="related" type="application/pdf" title="pdf"/>1378    <summary>Large language models (LLMs) are increasingly used to analyze and rewrite news, yet current framing studies mainly evaluate generation, detection, or whether rewritten text appears more neutral. They do not directly show whether a model can undo a known framing transformation while keeping the facts fixed. We introduce a controlled inversion test over three established textual realizations of framing: evaluative lexis, agency realization, and information salience. Across 60 news articles and three intervention strengths, this yields 540 paired variants with preserved atomic facts and recorded edits. Across Qwen, DeepSeek, and Kimi, factual preservation remains near 0.84, whereas intervention reversal is 0.044--0.068. Even when both framing type and direction are recognized correctly, pooled reversal reaches 0.071. These results reveal a clear separation between factual fidelity, framing recognition, and framing inversion: recognizing how an article is framed does not imply that the framing can be undone.</summary>1379    <category term="cs.CL" scheme="http://arxiv.org/schemas/atom"/>1380    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>1381    <published>2026-09-10T16:23:51Z</published>1382    <arxiv:primary_category term="cs.CL"/>1383    <author>1384      <name>Yi Liu</name>1385    </author>1386  </entry>1387  <entry>1388    <id>http://arxiv.org/abs/2609.11768v1</id>1389    <title>A Unified Per-Token Gating Family for On-Policy Distillation: FKL/RKL Mixing with Multi-Channel and Bias Coefficients</title>1390    <updated>2026-09-10T16:22:35Z</updated>1391    <link href="https://arxiv.org/abs/2609.11768v1" rel="alternate" type="text/html"/>1392    <link href="https://arxiv.org/pdf/2609.11768v1" rel="related" type="application/pdf" title="pdf"/>1393    <summary>Per-token gating of forward/reverse KL losses has become a standard technique for on-policy knowledge distillation (OPD), but existing methods such as EOPD (Jin et al., 2026) and ToDi (Jung et al., 2025) each fix a single gating signal and a single gating direction, and the two have never been compared directly. We introduce a four-coefficient parameterization lambda_t = sigma(a * h_t + b * u(x) + c + d * gap_t) in which direction-aligned proxies of EOPD and ToDi appear as one-dimensional (1D) restrictions, and which adds multi-channel composition and an explicit bias as further degrees of freedom. On TweetEval (Barbieri et al., 2020) emotion and hate, with a Qwen3-32B teacher and a Qwen3-4B student, configurations in the full family reach higher accuracy than the matched-magnitude single-channel (entropy-only / gap-only) 1D restrictions in 33 of 36 comparable cells, and a 26-cell mean-match isolation experiment places dynamic gating ahead of effective-KL-matched static baselines in 19 of 26 cells. Because cells share training data, models, and parameter substructure, we report both counts as exploratory aggregate directional evidence rather than as independent hypothesis tests. Targeted three-seed paired replications of the nine headline comparisons singled out by that sweep -- including a third task, offensive -- are directionally consistent, but individually smaller than the single-seed estimates and not significant at n=3. We therefore present the parameterization primarily as a shared coordinate system for comparing per-token gating designs in short-output classification OPD.</summary>1394    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>1395    <category term="cs.CL" scheme="http://arxiv.org/schemas/atom"/>1396    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>1397    <published>2026-09-10T16:22:35Z</published>1398    <arxiv:comment>Accepted at the Findings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP 2026 Findings)</arxiv:comment>1399    <arxiv:primary_category term="cs.AI"/>1400    <author>1401      <name>Suwan Wu</name>1402    </author>1403    <author>1404      <name>Yumeng Lin</name>1405    </author>1406    <author>1407      <name>Pengcheng Yuan</name>1408    </author>1409    <author>1410      <name>Xiaolong Jiang</name>1411    </author>1412  </entry>1413  <entry>1414    <id>http://arxiv.org/abs/2609.11762v1</id>1415    <title>Component-Aware Differential Privacy for Federated Multilingual Speech-LLMs</title>1416    <updated>2026-09-10T16:17:19Z</updated>1417    <link href="https://arxiv.org/abs/2609.11762v1" rel="alternate" type="text/html"/>1418    <link href="https://arxiv.org/pdf/2609.11762v1" rel="related" type="application/pdf" title="pdf"/>1419    <summary>Per-layer differential privacy (DP) clipping improves gradient fidelity in federated learning by allocating per-matrix clipping budgets proportional to parameter count. We show that this recipe breaks for speech large language models (speech-LLMs), when the acoustic encoder and the language decoder differ by an order of magnitude in update norm. Single-pool per-layer methods suffer \emph{cross-component budget collapse}, dragging word error rate (WER) far from flat global clipping or collapsing training entirely. When the norm imbalance is milder, adaptive single-pool methods partially recover, confirming that collapse severity scales with the inter-component norm ratio. We empirically diagnose the root cause across six per-layer methods and three speech-LLM architectures. We then propose \emph{$α$-split}, a two-pool allocation that normalises encoder and LLM parameters into independent pools, and show that joint $\ell_2$ sensitivity and the original $(\varepsilon,δ)$-DP guarantee are unchanged. At architecture-calibrated $α$, our method recovers WER utility compared to flat DP, while granting the encoder $4.47{\times}$ tighter per-component noise protection against speaker voice-based gradient-inversion attacks at only $+2.6\%$ LLM noise overhead.</summary>1420    <category term="cs.CL" scheme="http://arxiv.org/schemas/atom"/>1421    <published>2026-09-10T16:17:19Z</published>1422    <arxiv:comment>Accepted in SLT2026</arxiv:comment>1423    <arxiv:primary_category term="cs.CL"/>1424    <author>1425      <name>Jordi Luque</name>1426    </author>1427    <author>1428      <name>Fernando López</name>1429    </author>1430    <author>1431      <name>Aleix Sant</name>1432    </author>1433  </entry>1434  <entry>1435    <id>http://arxiv.org/abs/2609.11758v1</id>1436    <title>RAG-Safety-Bench: Reliable Evaluation of Retrieval-Augmented LLM Safety</title>1437    <updated>2026-09-10T16:12:25Z</updated>1438    <link href="https://arxiv.org/abs/2609.11758v1" rel="alternate" type="text/html"/>1439    <link href="https://arxiv.org/pdf/2609.11758v1" rel="related" type="application/pdf" title="pdf"/>1440    <summary>Allowing large language models (LLMs) to retrieve information from a set of trusted documents can increase reliability and reduce hallucination. However, recent work has demonstrated that retrieval-augmented generation (RAG) can have unintended side effects on the overall safety of the generated responses, when prompted for harmful or dangerous content. A clearer understanding of the mechanisms leading to this result is needed, as increasing numbers of end users turn to RAG to incorporate corporate documents and knowledge bases into LLM-based systems. We introduce RAG-Safety-Bench, a benchmark to measure the safety impact of RAG on LLM models. By removing the confounding effect of retriever quality, and cleanly separating the problem into four conditions -- non-RAG, RAG with an oracle document containing the answer to the harmful request, RAG with documents related to the harmful request but without the specific answer, and RAG with random, safe documents -- the benchmark isolates the impacts of different factors in the observed safety degradation. We report results across five open-source LLMs, showing an inverse relationship between benign and unsafe capability, strong evidence that baseline safety guardrails do not lead to downstream safety guarantees in the RAG case, and model-specific support for previous findings that even benign documents can lead to unsafe generation in retrieval-enabled systems.</summary>1441    <category term="cs.CL" scheme="http://arxiv.org/schemas/atom"/>1442    <category term="cs.IR" scheme="http://arxiv.org/schemas/atom"/>1443    <published>2026-09-10T16:12:25Z</published>1444    <arxiv:comment>Proceedings of EMNLP 2026 (main conference)</arxiv:comment>1445    <arxiv:primary_category term="cs.CL"/>1446    <author>1447      <name>Adithiyan Rajan Indira Saravanan</name>1448    </author>1449    <author>1450      <name>Kathleen C. Fraser</name>1451    </author>1452  </entry>1453  <entry>1454    <id>http://arxiv.org/abs/2609.11752v1</id>1455    <title>SIRF: A Spec-Internalized Risk Foundation Model for Industrial Content Risk Control</title>1456    <updated>2026-09-10T16:08:54Z</updated>1457    <link href="https://arxiv.org/abs/2609.11752v1" rel="alternate" type="text/html"/>1458    <link href="https://arxiv.org/pdf/2609.11752v1" rel="related" type="application/pdf" title="pdf"/>1459    <summary>For industrial content risk control, the real deployment constraint is not average accuracy but how much risk can be auto-handled under high precision and second-level latency. We present SIRF (Spec-Internalized Risk Foundation Model), which internalizes a platform's complex policies, synthesized without additional human annotation via EntiGraph, MAGA rewriting and account-level chain-of-thought (CoT), into the weights via continued pretraining (CPT), so rules are applied at high precision under an ultra-low-latency, verdict-only deployment. A controlled same-source comparison (Qwen3-8B-SFT vs. SIRF-8B-SFT, identical policy injection and verdict-only output form, differing only in policy-grounded CPT) attributes the gain to internalization: SIRF-8B-SFT reaches 71.3% Black Recall@P95, +15.1pp over the baseline, using only ~70M CPT tokens without harming general ability, and among included, logprob-available models under this interface it matches or exceeds far larger systems. SIRF is deployed as a tree-model adjudication layer (20% more mis-penalized samples recovered) and transfers to a freezing scenario at low cost (~70% relative mis-penalization reduction).</summary>1460    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>1461    <category term="cs.CL" scheme="http://arxiv.org/schemas/atom"/>1462    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>1463    <published>2026-09-10T16:08:54Z</published>1464    <arxiv:comment>14 pages, 12 figures. Accepted at the Industry Track of EMNLP 2026</arxiv:comment>1465    <arxiv:primary_category term="cs.AI"/>1466    <author>1467      <name>Suwan Wu</name>1468    </author>1469    <author>1470      <name>Yumeng Lin</name>1471    </author>1472    <author>1473      <name>Pengcheng Yuan</name>1474    </author>1475    <author>1476      <name>Xiaolong Jiang</name>1477    </author>1478  </entry>1479  <entry>1480    <id>http://arxiv.org/abs/2609.11749v1</id>1481    <title>Sparsity Regularized and Robust Mean Variance Portfolio Selection Under Ellipsoidal Uncertainty</title>1482    <updated>2026-09-10T16:03:35Z</updated>1483    <link href="https://arxiv.org/abs/2609.11749v1" rel="alternate" type="text/html"/>1484    <link href="https://arxiv.org/pdf/2609.11749v1" rel="related" type="application/pdf" title="pdf"/>1485    <summary>We investigate mean-variance portfolio selection with an $\ell_0$-penalty to promote sparsity in asset allocations. Uncertainty in the mean return vector is incorporated through an ellipsoidal uncertainty set, yielding a robust sparse optimization framework. We characterize the structure of both local and global minimizers and exploit these properties in the risk minimization and return maximization formulations. Building on this structural insight, we develop a branch-and-bound algorithm tailored to the resulting robust sparse portfolio problems, together with a new pruning rule that can discard exponentially many candidate portfolios in a single step. Extensive computational experiments on real market data, together with comparisons against a mixed-integer second-order cone programming solver, demonstrate the effectiveness and competitiveness of the proposed approach.</summary>1486    <category term="math.OC" scheme="http://arxiv.org/schemas/atom"/>1487    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>1488    <category term="stat.ML" scheme="http://arxiv.org/schemas/atom"/>1489    <published>2026-09-10T16:03:35Z</published>1490    <arxiv:primary_category term="math.OC"/>1491    <author>1492      <name>Deniz Akkaya</name>1493    </author>1494    <author>1495      <name>Emre Can Yayla</name>1496    </author>1497    <author>1498      <name>Buse Şen</name>1499    </author>1500    <author>1501      <name>Mustafa Ç. Pınar</name>1502    </author>1503  </entry>1504  <entry>1505    <id>http://arxiv.org/abs/2609.11744v1</id>1506    <title>Building py-kvcache: A Performance Characterization of External KV Caching for vLLM with NVMe SSDs</title>1507    <updated>2026-09-10T16:00:35Z</updated>1508    <link href="https://arxiv.org/abs/2609.11744v1" rel="alternate" type="text/html"/>1509    <link href="https://arxiv.org/pdf/2609.11744v1" rel="related" type="application/pdf" title="pdf"/>1510    <summary>Prefix caching can reduce the time to first token (TTFT) of long-context LLM requests by reusing previously computed key-value (KV) states, but for short prefixes or fast GPUs, recomputation can be faster than loading from an external cache. We characterize this tradeoff in vLLM across GPU, CPU, and NVMe tiers using synthetic workloads, long-context benchmarks, production traces, and find that cache performance depends on transfer granularity, intermediate memory use, and when transfers enter the request schedule, not only on device bandwidth. These findings motivate py-kvcache, a vLLM KV Offload connector with asynchronous direct I/O, bounded shared staging, and scheduler-aware preloading, which starts disk reads while requests are still waiting, overlapping with compute. At 80k tokens, py-kvcache loading from disk is 2.0x faster than LMCache, with preloading contributing 1.34x. With GPU, CPU, and disk caching enabled, it is 1.23x faster than LMCache and within approximately 4% of the native vLLM KV Offload implementation. LongBench and SCBench show that these benefits extend to irregular prefix chains and multi-turn workloads. Bailian trace replays improve TTFT on a weaker GPU, but on an H100 the average request falls below the break-even point and GPU memory alone retains enough prefixes. External KV caching should therefore be treated as a setup specific admission decision. The py-kvcacheimplementation is available at: https://github.com/atlarge-research/py-kvcache.</summary>1511    <category term="cs.DC" scheme="http://arxiv.org/schemas/atom"/>1512    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>1513    <published>2026-09-10T16:00:35Z</published>1514    <arxiv:primary_category term="cs.DC"/>1515    <author>1516      <name>Joseph Kanichai</name>1517    </author>1518    <author>1519      <name>Tiziano De Matteis</name>1520    </author>1521    <author>1522      <name>Animesh Trivedi</name>1523    </author>1524  </entry>1525  <entry>1526    <id>http://arxiv.org/abs/2609.11739v1</id>1527    <title>LOCUS: Task-Aware Low-Rank Post-Training for Token-Efficient Language Generation</title>1528    <updated>2026-09-10T15:53:25Z</updated>1529    <link href="https://arxiv.org/abs/2609.11739v1" rel="alternate" type="text/html"/>1530    <link href="https://arxiv.org/pdf/2609.11739v1" rel="related" type="application/pdf" title="pdf"/>1531    <summary>Large language model serving costs scale directly with output sequence length, yet standard preference alignment often inflates response verbosity without improving utility. We study whether the parameterization of post-training updates affects generation length: low-rank subspaces alter sequence length without modifying the alignment loss. We present LOCUS, a method that selects a task-aware low-rank adaptation subspace to minimize output-token cost subject to a utility constraint. Within this subspace, post-training retains the native preference objective with a frozen backbone. Across Anthropic HH-RLHF dialogue preferences, we evaluate two $\sim$3B decoder backbones, Pythia-2.8B and Qwen2.5-3B, against protocol-matched full-parameter DPO and DrDPO branches and the released SamPO checkpoint. LOCUS reduces continuation length by up to 39.84\% on Pythia-2.8B and by 14.87--17.58\% on Qwen2.5-3B while updating only 0.24--0.28\% of model parameters, with no material change in the internal preference diagnostic.</summary>1532    <category term="cs.CL" scheme="http://arxiv.org/schemas/atom"/>1533    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>1534    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>1535    <published>2026-09-10T15:53:25Z</published>1536    <arxiv:primary_category term="cs.CL"/>1537    <author>1538      <name>Dongfang Zhao</name>1539    </author>1540  </entry>1541  <entry>1542    <id>http://arxiv.org/abs/2609.11737v1</id>1543    <title>ORCH: Organizational Principles Enable Collective Intelligence in Embodied AI</title>1544    <updated>2026-09-10T15:52:35Z</updated>1545    <link href="https://arxiv.org/abs/2609.11737v1" rel="alternate" type="text/html"/>1546    <link href="https://arxiv.org/pdf/2609.11737v1" rel="related" type="application/pdf" title="pdf"/>1547    <summary>Collective intelligence depends not only on the capabilities of individual members, but also on how those members are organized. Yet artificial multi-agent systems are typically assembled using fixed organizational structures, even when the physical tasks they perform impose fundamentally different coordination requirements. Here we show that principles from human organization theory can be operationalized to organize large, heterogeneous collectives of embodied artificial agents. We introduce ORCH (Organizing Roles and Coordination Hierarchies), which constructs task-specific hierarchical organizations by combining pooled interdependence for work that can proceed concurrently with sequential interdependence for work governed by prerequisite relationships. Across 25 wildfire-response missions spanning reconnaissance, rescue, transportation, resource management, containment and suppression, we evaluated teams of up to 50 heterogeneous agents using eight large language models. Organizations constructed using these principles consistently outperformed four representative embodied multi-agent approaches across mission outcome, execution efficiency, exploration and computational resource use. Human-designed ORCH organizations improved final score by 63.97% and execution efficiency by 74.29% on average relative to the four prior frameworks. Organizations generated automatically by language models improved these measures by 43.63% and 52.53%, respectively. These advantages persisted across missions and underlying language models. Notably, collective performance was not monotonically determined by model scale. Analysis of long-horizon missions showed that hierarchical organization enabled teams to preserve concurrent activity within specialized groups while coordinating ordered transitions between mission phases.</summary>1548    <category term="cs.MA" scheme="http://arxiv.org/schemas/atom"/>1549    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>1550    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>1551    <category term="cs.RO" scheme="http://arxiv.org/schemas/atom"/>1552    <published>2026-09-10T15:52:35Z</published>1553    <arxiv:primary_category term="cs.MA"/>1554    <author>1555      <name>Zhengran Ji</name>1556    </author>1557    <author>1558      <name>Jonathan Hyun</name>1559    </author>1560    <author>1561      <name>Boyuan Chen</name>1562    </author>1563  </entry>1564  <entry>1565    <id>http://arxiv.org/abs/2609.11736v1</id>1566    <title>Learning structural balance of graphs from quantum spectral features</title>1567    <updated>2026-09-10T15:52:28Z</updated>1568    <link href="https://arxiv.org/abs/2609.11736v1" rel="alternate" type="text/html"/>1569    <link href="https://arxiv.org/pdf/2609.11736v1" rel="related" type="application/pdf" title="pdf"/>1570    <summary>We develop a quantum approach to spectral feature extraction from the density of states (DOS) of a problem-dependent Hamiltonian, and apply it to machine learning on signed graphs. We propose to embed a signed graph as an Ising model instance with positive and negative interactions, and use the standardized moments of the Ising DOS as features for learning. We show that these moments count signed closed walks, are switching-invariant, and are size-free by construction. As a benchmark, we target learning the frustration index, an NP-hard measure of structural balance that can be labeled exactly at moderate size. At zero field, the models can be sampled classically, allowing the quantum extraction procedure to be certified against exact ground truth. We propose DOS-QPE, a phase estimation on a purified maximally mixed probe, which samples the spectral density with orders of magnitude fewer shots than Hadamard test-based trace sampling and feeds the resulting features directly into classically trained models. On $1.4\times10^5$ labeled graphs the exact DOS determines the frustration index, and five moments recover it with a mean error of 0.4, well below one sign flip. Beyond zero field, the underlying trace-estimation problem is DQC1-complete, providing access to spectral features for which no efficient classical sampling method is known. Our work opens routes towards quantum applications in social network balance analysis, spin-glass studies, correlation clustering, and protein-interaction networks.</summary>1571    <category term="quant-ph" scheme="http://arxiv.org/schemas/atom"/>1572    <category term="cond-mat.dis-nn" scheme="http://arxiv.org/schemas/atom"/>1573    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>1574    <category term="cs.SI" scheme="http://arxiv.org/schemas/atom"/>1575    <published>2026-09-10T15:52:28Z</published>1576    <arxiv:comment>12 pages, 6 figures</arxiv:comment>1577    <arxiv:primary_category term="quant-ph"/>1578    <author>1579      <name>Stefano Scali</name>1580    </author>1581    <author>1582      <name>Oleksandr Kyriienko</name>1583    </author>1584  </entry>1585  <entry>1586    <id>http://arxiv.org/abs/2609.11733v1</id>1587    <title>Reflex-Informed Neuromuscular Reinforcement Learning for Muscle-Driven Locomotion</title>1588    <updated>2026-09-10T15:47:26Z</updated>1589    <link href="https://arxiv.org/abs/2609.11733v1" rel="alternate" type="text/html"/>1590    <link href="https://arxiv.org/pdf/2609.11733v1" rel="related" type="application/pdf" title="pdf"/>1591    <summary>Muscle-driven locomotion provides a physically grounded approach to generating realistic human movement. However, achieving both physiological plausibility and adaptability to changes in musculoskeletal capacity and external disturbances remains a fundamental challenge. To address this limitation, we propose a Reflex-Informed Neuromuscular Reinforcement Learning framework for muscle-driven locomotion. Within this framework, a fixed phase-dependent reflex controller serves as the underlying neuromuscular control mechanism, while the reinforcement learning policy produces four biomechanically meaningful residual parameters to modulate key reflex gains and thresholds associated with hip swing, knee support, and ankle propulsion according to the current state. Experimental results demonstrate that the proposed framework generates physiologically plausible locomotion with improved kinematic accuracy and dynamic consistency, as well as better bilateral symmetry and stride-to-stride consistency under nominal walking conditions. The learned policy remains robust under muscle weakness and external perturbations without retraining.</summary>1592    <category term="cs.RO" scheme="http://arxiv.org/schemas/atom"/>1593    <category term="cs.GR" scheme="http://arxiv.org/schemas/atom"/>1594    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>1595    <published>2026-09-10T15:47:26Z</published>1596    <arxiv:comment>28 pages, 13 figures, and 9 tables</arxiv:comment>1597    <arxiv:primary_category term="cs.RO"/>1598    <author>1599      <name>Jian Zhou</name>1600    </author>1601    <author>1602      <name>Xingyu Zhang</name>1603    </author>1604    <author>1605      <name>Rui Ma</name>1606    </author>1607    <author>1608      <name>Yu Cao</name>1609    </author>1610    <author>1611      <name>Shane Xie</name>1612    </author>1613    <author>1614      <name>Zhi-qiang Zhang</name>1615    </author>1616  </entry>1617  <entry>1618    <id>http://arxiv.org/abs/2609.11725v1</id>1619    <title>Continuous-Time Acoustic Modelling with Neural Controlled Differential Equations</title>1620    <updated>2026-09-10T15:41:32Z</updated>1621    <link href="https://arxiv.org/abs/2609.11725v1" rel="alternate" type="text/html"/>1622    <link href="https://arxiv.org/pdf/2609.11725v1" rel="related" type="application/pdf" title="pdf"/>1623    <summary>Text-to-speech (TTS) models commonly address text--speech alignment by expanding phone-level encoder states to frame-level decoder inputs using predicted durations. While this length-regulation step resolves alignment structurally, this use of duration typically changes only where and how often latent states appear, not the values of the states themselves. This paper proposes a continuous-time mechanism for duration-aware acoustic modelling in TTS using neural controlled differential equations (CDEs). We formulate the phone representation as a temporally parameterised control path and use a neural acoustic vector field to produce a continuous-time hidden state whose values evolve with phonetic content and duration-derived timing. The resulting trajectory can be sampled at discrete points and integrated into a standard acoustic decoder pipeline. Objective results contrast CDEs and typical recurrent models. Subjective results suggest that CDE-based models evaluating one phone per step can improve rank-order agreement between synthesised and reference emotion intensity while maintaining comparable emotion-expression quality to a strong baseline. Additional experiments with half-phone step-sizes suggest that temporal resolution changes the trade-off between style tracking and absolute calibration. These results position CDEs as a promising design space for continuous-time and duration-aware style-sensitive TTS.</summary>1624    <category term="cs.SD" scheme="http://arxiv.org/schemas/atom"/>1625    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>1626    <published>2026-09-10T15:41:32Z</published>1627    <arxiv:comment>Accepted to IEEE Spoken Language Technology Workshop (SLT) 2026</arxiv:comment>1628    <arxiv:primary_category term="cs.SD"/>1629    <author>1630      <name>Mattias Cross</name>1631    </author>1632    <author>1633      <name>Minghui Zhao</name>1634    </author>1635    <author>1636      <name>Anton Ragni</name>1637    </author>1638  </entry>1639  <entry>1640    <id>http://arxiv.org/abs/2609.11724v1</id>1641    <title>The Eloquence submission for Task 2 of the Interspeech 2026 MLC-SLM challenge</title>1642    <updated>2026-09-10T15:41:11Z</updated>1643    <link href="https://arxiv.org/abs/2609.11724v1" rel="alternate" type="text/html"/>1644    <link href="https://arxiv.org/pdf/2609.11724v1" rel="related" type="application/pdf" title="pdf"/>1645    <summary>This paper details the Eloquence team's approach to Task 2 of the 2nd MLC-SLM challenge at Interspeech 2026, which involves multilingual Multiple-Choice Question Answering (MCQA) across 21 languages. Three approaches are explored. First, we fine-tune Voxtral-Mini-3B via LoRA with cross-lingual data augmentation, ASR transcript augmentation and timestamp-aware audio cropping, achieving 0.72 macro-accuracy on evaluation Phase 2. Second, we apply multimodal in-context learning (ICL) to the frozen Voxtral-24B model to correct a strong label bias, reaching 0.81, our best result. Third, a training-free retrieval system based on a three-layer voice-anchored memory combining acoustic identity, semantic content, and a knowledge graph achieves 0.68. All three systems substantially outperform the official baseline.</summary>1646    <category term="cs.CL" scheme="http://arxiv.org/schemas/atom"/>1647    <published>2026-09-10T15:41:11Z</published>1648    <arxiv:primary_category term="cs.CL"/>1649    <author>1650      <name>Jordi Luque</name>1651    </author>1652    <author>1653      <name>Lorenzo Concina</name>1654    </author>1655    <author>1656      <name>Marco Matassoni</name>1657    </author>1658    <author>1659      <name>Alessio Brutti</name>1660    </author>1661    <author>1662      <name>Filippo Vella</name>1663    </author>1664  </entry>1665  <entry>1666    <id>http://arxiv.org/abs/2609.11722v1</id>1667    <title>Revisiting Avatar-As-Image: High-Fidelity Registration is All You Need</title>1668    <updated>2026-09-10T15:40:39Z</updated>1669    <link href="https://arxiv.org/abs/2609.11722v1" rel="alternate" type="text/html"/>1670    <link href="https://arxiv.org/pdf/2609.11722v1" rel="related" type="application/pdf" title="pdf"/>1671    <summary>The representation of 3D clothed humans as standardized 2D UV texture and displacement maps over an underlying body model has long been studied. This compact representation is enticing as it enables pretrained image networks to process, generate, and edit 3D avatars, but is only useful if scans are accurately aligned and brought into correspondence via high-fidelity registration. This prerequisite has never been met, which we argue explains the limited quality of prior UV-based methods for clothed humans. Despite its significance, no public method produces high-fidelity SMPL(-X)+D registrations with UV texture from arbitrary clothed scans. We present AvaImg, a multi-stage optimization pipeline, to close this gap: it enforces body-inside-clothing constraint via signed winding numbers, made viable by a three-level efficiency cascade (~10x runtime reduced, ~95% storage saved), and recovers fine surface detail using coarse-to-fine displacement optimization. AvaImg outperforms all baselines in body fitting, shape estimation, and surface registration across six datasets, yielding textured registrations near-indistinguishable from scans (PSNR=34.48dB). For validation of AvaImg's Avatar-as-Image representation as imminently compatible with image foundation models, we auto-encode our UV maps via the frozen FLUX VAE. This achieves only 0.76mm added Chamfer error relative to scan and shows that the resulting maps lie within natural-image distributions, supporting the use of 2D generative priors for 3D avatar generation. Code, data, and Singularity containers will be at https://yuxuan-xue.com/avaimg.</summary>1672    <category term="cs.CV" scheme="http://arxiv.org/schemas/atom"/>1673    <published>2026-09-10T15:40:39Z</published>1674    <arxiv:comment>Project Page: https://yuxuan-xue.com</arxiv:comment>1675    <arxiv:primary_category term="cs.CV"/>1676    <author>1677      <name>Margaret Kostyrko</name>1678    </author>1679    <author>1680      <name>Yuxuan Xue</name>1681    </author>1682    <author>1683      <name>Garvita Tiwari</name>1684    </author>1685    <author>1686      <name>Gerard Pons-Moll</name>1687    </author>1688  </entry>1689  <entry>1690    <id>http://arxiv.org/abs/2609.11717v1</id>1691    <title>MC-DeTra: Motion-Consistent Joint Object Detection and Socially-Aware Trajectory Forecasting in Bird's-Eye-View Images</title>1692    <updated>2026-09-10T15:34:09Z</updated>1693    <link href="https://arxiv.org/abs/2609.11717v1" rel="alternate" type="text/html"/>1694    <link href="https://arxiv.org/pdf/2609.11717v1" rel="related" type="application/pdf" title="pdf"/>1695    <summary>Unified models for object detection and trajectory forecasting aim to merge perception and prediction for autonomous driving, refining actor trajectories directly over shared bird's-eye-view (BEV) images rasterized from LiDAR and high-definition maps. Their accuracy on dynamic, moving actors, however, remains the hardest part of the task, and the strongest such model, DeTra, has no public implementation. We contribute an openly released DeTra reimplementation with documented approximations, and on top of it MC-DeTra: a family of motion-consistency mechanisms that add supervision through two annotation-derived auxiliary signals -- each actor's observed past motion and the occupancy of the surrounding traffic that forms its social context -- and one inter-output consistency constraint that aligns an actor's predicted heading with its predicted direction of motion. Every proposed loss is train-only and inference-safe: it shapes the shared BEV representation during training and is removed at test time, adding no inference latency. On the Waymo Open Dataset, evaluated under a strict, detection-conditioned forecasting protocol, MC-DeTra improves dynamic, socially-situated trajectory forecasting while preserving or improving detection accuracy; a gradient-based loss-calibration analysis exposes how the auxiliary objectives compete at the shared backbone, and our ablation identifies which signals contribute most. We release code, configurations, and evaluation tooling at https://github.com/diuzhevVlad/MC-DeTra.</summary>1696    <category term="cs.CV" scheme="http://arxiv.org/schemas/atom"/>1697    <category term="cs.RO" scheme="http://arxiv.org/schemas/atom"/>1698    <published>2026-09-10T15:34:09Z</published>1699    <arxiv:comment>16 pages, 4 figures. Code: https://github.com/diuzhevVlad/MC-DeTra</arxiv:comment>1700    <arxiv:primary_category term="cs.CV"/>1701    <author>1702      <name>Vladislav Diuzhev</name>1703    </author>1704    <author>1705      <name>Dmitry Yudin</name>1706    </author>1707  </entry>1708  <entry>1709    <id>http://arxiv.org/abs/2609.11716v1</id>1710    <title>Why Does Post-Training Quantization Work?</title>1711    <updated>2026-09-10T15:32:32Z</updated>1712    <link href="https://arxiv.org/abs/2609.11716v1" rel="alternate" type="text/html"/>1713    <link href="https://arxiv.org/pdf/2609.11716v1" rel="related" type="application/pdf" title="pdf"/>1714    <summary>Post-training quantization compresses large language models (LLMs) by storing their weights at reduced precision, and each quantized weight introduces an error into the hidden states. Naively, these errors should accumulate with depth and corrupt next-token prediction; randomly initialized models accumulate these discrepancies rapidly, whereas quantized pretrained models accumulate much less hidden-state error and largely maintain downstream task performance, even though they were never trained with quantization noise. This raises the question we address: why does post-training quantization work? Comparing full-precision and quantized forward passes, we identify two mechanisms that characterize pretrained quantization robustness. First, the error a layer newly introduces tends to oppose the error it inherits from the layer's input. The two cancel partially such that the discrepancy between full-precision and quantized passes grows slowly. This counteracting residual interaction develops during pretraining. Our quantitative analysis identifies it as a major factor slowing hidden-error growth. Second, LM-head geometry preferentially preserves the scores and probabilities of high-ranked tokens, which typically represent the model's most confident predictions. Together, these mechanisms explain why quantization error that passes through numerous layers can still produce only small output changes, and we verify the findings across models and quantization settings.</summary>1715    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>1716    <category term="cs.CL" scheme="http://arxiv.org/schemas/atom"/>1717    <published>2026-09-10T15:32:32Z</published>1718    <arxiv:comment>45 pages, 26 figures, including appendices</arxiv:comment>1719    <arxiv:primary_category term="cs.LG"/>1720    <author>1721      <name>Yuxiang Chen</name>1722    </author>1723    <author>1724      <name>Michael Beyer</name>1725    </author>1726    <author>1727      <name>Jun Zhu</name>1728    </author>1729    <author>1730      <name>Jianfei Chen</name>1731    </author>1732  </entry>1733  <entry>1734    <id>http://arxiv.org/abs/2609.11713v1</id>1735    <title>A Time-Based Readout for Vector-Matrix Multiplication in Fully Analog Memristive SNNs</title>1736    <updated>2026-09-10T15:30:34Z</updated>1737    <link href="https://arxiv.org/abs/2609.11713v1" rel="alternate" type="text/html"/>1738    <link href="https://arxiv.org/pdf/2609.11713v1" rel="related" type="application/pdf" title="pdf"/>1739    <summary>Artificial neural networks rely on vector-matrix multiplications (VMMs), whose implementation in von Neumann architectures is dominated by costly data movement between memory and processing units. Spiking neural networks (SNNs) mitigate this bottleneck by performing in-memory, analog VMMs using memristive crossbar arrays. However, conventional current-mode readout circuits incur significant area and power overhead.1740  This work proposes a fully analog readout architecture based on voltage-to-time conversion of the VMM output. By sensing the column voltage, the proposed approach avoids current-mode summing and scaling circuitry, improving area and energy efficiency. Post-layout simulations of a 10x1 SNN implemented in a 130 nm CMOS technology validate the proposed architecture, while application to a trained 64x10 SNN for digit classification further demonstrates its feasibility for SNN inference.</summary>1741    <category term="cs.ET" scheme="http://arxiv.org/schemas/atom"/>1742    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>1743    <category term="cs.AR" scheme="http://arxiv.org/schemas/atom"/>1744    <published>2026-09-10T15:30:34Z</published>1745    <arxiv:comment>Accepted at 2026 IEEE 33rd International Conference on Electronics, Circuits and Systems (ICECS)</arxiv:comment>1746    <arxiv:primary_category term="cs.ET"/>1747    <author>1748      <name>Elia Mateu-Barriendos</name>1749    </author>1750    <author>1751      <name>Álvaro Gómez-Pau</name>1752    </author>1753    <author>1754      <name>Josep Rius</name>1755    </author>1756    <author>1757      <name>Daniel Arumí</name>1758    </author>1759    <author>1760      <name>Rosa Rodríguez-Montañés</name>1761    </author>1762    <author>1763      <name>Salvador Manich</name>1764    </author>1765  </entry>1766  <entry>1767    <id>http://arxiv.org/abs/2609.11712v1</id>1768    <title>Generalization Analysis of Distributed Kernel-based Robust Gradient Descent Algorithms</title>1769    <updated>2026-09-10T15:30:01Z</updated>1770    <link href="https://arxiv.org/abs/2609.11712v1" rel="alternate" type="text/html"/>1771    <link href="https://arxiv.org/pdf/2609.11712v1" rel="related" type="application/pdf" title="pdf"/>1772    <summary>In this paper, we investigate the generalization performance of distributed gradient descent algorithms in a reproducing kernel Hilbert space under a robust loss function $l_σ$. By exploiting the spectral characterization of gradient descent together with the intrinsic properties of robust loss functions, we establish optimal learning rates for the distributed kernel-based robust gradient descent (DKRGD) algorithm with an appropriately chosen scale parameter $σ$. The proposed parameter choice of $σ$ simultaneously alleviates the saturation phenomenon and guarantees statistical robustness. A key technical contribution is a novel error analysis that provides substantially sharper bounds for products of operators, thereby significantly relaxing existing restrictions on the maximum number of local machines while retaining optimal learning rates. Finally, we develop a communication-efficient strategy that further improves the convergence performance of DKRGD.</summary>1773    <category term="stat.ML" scheme="http://arxiv.org/schemas/atom"/>1774    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>1775    <category term="math.OA" scheme="http://arxiv.org/schemas/atom"/>1776    <category term="math.PR" scheme="http://arxiv.org/schemas/atom"/>1777    <published>2026-09-10T15:30:01Z</published>1778    <arxiv:comment>40 pages, 4 figures</arxiv:comment>1779    <arxiv:primary_category term="stat.ML"/>1780    <author>1781      <name>Jun-Yi Meng</name>1782    </author>1783    <author>1784      <name>Zheng-Chu Guo</name>1785    </author>1786    <author>1787      <name>Yuan Mao</name>1788    </author>1789  </entry>1790  <entry>1791    <id>http://arxiv.org/abs/2609.11709v1</id>1792    <title>When Agents Disagree: Bayesian Backward Reasoning as a Label-Free Anchor for Multi-Agent Collective Decision-Making</title>1793    <updated>2026-09-10T15:27:58Z</updated>1794    <link href="https://arxiv.org/abs/2609.11709v1" rel="alternate" type="text/html"/>1795    <link href="https://arxiv.org/pdf/2609.11709v1" rel="related" type="application/pdf" title="pdf"/>1796    <summary>When multiple LLM agents yield conflicting answers, the decision-making process dictates whether agent diversity improves performance or merely compounds shared errors. Existing collective decision-making methods, including voting, electoral rules, and LLM judges, rely on forward reasoning: they map evidence to labels in one direction. Although these methods can combine diverse forward traces, they still aggregate estimates that share this evidence-to-label factorization and can inherit correlated errors within the forward pool. We therefore construct a reverse posterior for each instance through Bayesian backward reasoning from an explicit likelihood. The forward and reverse posteriors provide differently factorized approximations of the underlying posterior. Because estimates from different factorizations may tend to share the same error less often, we use Jensen-Shannon divergence to rank agents by cross-path consistency. This cross-path consistency signal underlies three strategies: hard selection (MinJS), soft reweighting (FwdJS), and log-linear fusion (LogLin). Evaluated on DDXPlus across five LLM backbones, our proposed strategies show consistent improvements: MinJS outperforms random selection across all backbones, FwdJS generally improves over the strongest baseline, and LogLin achieves the best performance among the evaluated methods, with its largest gains on the subset where the agents disagree. Despite its weaker standalone accuracy, the reverse posterior serves as a more useful anchor than forward-only alternatives, providing complementary information for collective decision-making. When labeled data are available, a lightweight two-stage calibration can further refine the reverse anchor and improve aggregation performance.</summary>1797    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>1798    <category term="cs.MA" scheme="http://arxiv.org/schemas/atom"/>1799    <published>2026-09-10T15:27:58Z</published>1800    <arxiv:primary_category term="cs.AI"/>1801    <author>1802      <name>Ken Chen</name>1803    </author>1804    <author>1805      <name>Wei Wang</name>1806    </author>1807    <author>1808      <name>Sachith Seneviratne</name>1809    </author>1810    <author>1811      <name>Hansani Weeratunge</name>1812    </author>1813    <author>1814      <name>Saman Halgamuge</name>1815    </author>1816  </entry>1817  <entry>1818    <id>http://arxiv.org/abs/2609.11708v1</id>1819    <title>Language-Augmented Semantic Priors for B-Spline Surface Fitting</title>1820    <updated>2026-09-10T15:27:31Z</updated>1821    <link href="https://arxiv.org/abs/2609.11708v1" rel="alternate" type="text/html"/>1822    <link href="https://arxiv.org/pdf/2609.11708v1" rel="related" type="application/pdf" title="pdf"/>1823    <summary>The use of B-splines and Non-Uniform Rational B-Splines surfaces constitutes the mathematical foundation of contemporary computer-aided design (CAD) systems. Despite long-term progress, geometric kernels in traditional CAD still rely heavily on predetermined heuristic initialization for surface fitting and parameterization. Meanwhile, the procedural semantics and design intent encoded in modeling histories are largely ignored during geometry generation. This disconnect creates a gap between high-level design intent and solver-executable geometric configuration, often leading to suboptimal and semantically inconsistent fitting results. To bridge this gap, we introduce LASP, a Language-Augmented Semantic Priors framework that leverages large language models (LLMs) to infer structured, solver-usable B-spline priors from procedural modeling histories. Rather than modifying the geometric kernel itself, LASP operates as a semantic reasoning layer above existing solvers. It first translates modeling histories into rich textual descriptions that capture design intent, geometric context, and functional relationships, and then uses a fine-tuned LLM to predict structured B-spline prior parameters. LASP is trained through a two-stage scheme that combines local geometric regularities with long-range contextual dependencies, producing priors that are both interpretable and semantically coherent. This approach furnishes inductive signals that direct the conventional B-spline fitting process toward solutions that more accurately encapsulate the intended design objectives and demonstrate heightened semantic coherence. Compared to traditional machine learning schemes, the experiments demonstrate that language-driven reasoning can serve as a powerful inductive bias for geometric solving, establishing a new paradigm of language-guided geometric optimization in modern CAD systems.</summary>1824    <category term="cs.CV" scheme="http://arxiv.org/schemas/atom"/>1825    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>1826    <published>2026-09-10T15:27:31Z</published>1827    <arxiv:primary_category term="cs.CV"/>1828    <author>1829      <name>Yunzhong Lou</name>1830    </author>1831    <author>1832      <name>Yusheng Luo</name>1833    </author>1834    <author>1835      <name>Jiahao Li</name>1836    </author>1837    <author>1838      <name>Yu Song</name>1839    </author>1840    <author>1841      <name>Xiangdong Zhou</name>1842    </author>1843  </entry>1844  <entry>1845    <id>http://arxiv.org/abs/2609.11703v1</id>1846    <title>Spectral Adapters for Segment Anything Model-based Segmentation of Colorectal Liver Metastases in Computed Tomography</title>1847    <updated>2026-09-10T15:26:15Z</updated>1848    <link href="https://arxiv.org/abs/2609.11703v1" rel="alternate" type="text/html"/>1849    <link href="https://arxiv.org/pdf/2609.11703v1" rel="related" type="application/pdf" title="pdf"/>1850    <summary>Accurate segmentation of colorectal liver metastases (CRLM) in contrast-enhanced computed tomography (CT) is important for response assessment, surgical planning, and follow-up. We propose two parameter-efficient spectral adapters for the Segment Anything Model (SAM): the Directional Spectral Adapter (DiSECT) and Spectral Instance-Guided Adapter (SiGA). DiSECT uses singular value decomposition of frozen weights to constrain residual updates to leading spectral directions, while SiGA adds global and input-conditioned gating through a multilayer perceptron. We evaluate these methods on 446 contrast-enhanced CT volumes (355 training, 91 testing) and compare them with LoRA, QLoRA, convolutional adapters (CAD), and a 3D nnU-Net baseline. Experiments consider single-point, three-point, bounding-box, and no-prompt regimes. SiGA achieves the best single-point performance with a Dice score of 0.77, IoU of 0.69, and HD95 of 35.39 mm. Under no-prompt inference, SiGA reaches 0.76 Dice, 0.68 IoU, and 46.76 mm HD95, comparable to the nnU-Net baseline (0.758 Dice). DiSECT uses only 0.14 million trainable parameters. These results show that spectral adapters can efficiently adapt SAM for CRLM segmentation while retaining strong accuracy with limited trainable parameters.</summary>1851    <category term="cs.CV" scheme="http://arxiv.org/schemas/atom"/>1852    <published>2026-09-10T15:26:15Z</published>1853    <arxiv:comment>13 pages, 1 figure, 4 tables</arxiv:comment>1854    <arxiv:primary_category term="cs.CV"/>1855    <author>1856      <name>Ramtin Mojtahedi</name>1857    </author>1858    <author>1859      <name>Mohammad Hamghalam</name>1860    </author>1861    <author>1862      <name>Jacob J. Peoples</name>1863    </author>1864    <author>1865      <name>Natalie Gangai</name>1866    </author>1867    <author>1868      <name>Mithat Gonen</name>1869    </author>1870    <author>1871      <name>Yun Shin Chun</name>1872    </author>1873    <author>1874      <name>HyunSeon Christine Kang</name>1875    </author>1876    <author>1877      <name>Richard K. G. Do</name>1878    </author>1879    <author>1880      <name>Amber L. Simpson</name>1881    </author>1882  </entry>1883  <entry>1884    <id>http://arxiv.org/abs/2609.11699v1</id>1885    <title>Negative Self-Distillation: Learning to Reason by Avoiding Flaws</title>1886    <updated>2026-09-10T15:24:17Z</updated>1887    <link href="https://arxiv.org/abs/2609.11699v1" rel="alternate" type="text/html"/>1888    <link href="https://arxiv.org/pdf/2609.11699v1" rel="related" type="application/pdf" title="pdf"/>1889    <summary>On-Policy Self-Distillation (OPSD) has emerged as a popular paradigm for large language model (LLM) self-improvement, allowing models to act as their own teachers by leveraging privileged information such as ground-truth solutions. However, recent findings indicate that OPSD can severely degrade the performance of LLMs on complex reasoning tasks: By forcing the student to imitate an artificially confident reasoning trace conditioned on privileged information, OPSD inadvertently suppresses expressions of uncertainty and penalizes the exploratory, self-corrective behaviors required to solve challenging problems. To address this, we introduce Negative Self-Distillation (NSD), a new framework that optimizes LLMs by diverging from flawed reasoning rather than imitating privileged solutions. Instead of relying on ground-truth answers or external supervision, NSD uses the model itself to generate a question-specific negative condition (eg, acting as a ``careless reasoner'') and pushes the student's distribution away from this self-generated negative teacher. Naively applying unlearning objectives to achieve this divergence is problematic, as flawed reasoning tokens are confounded with basic linguistic tokens; indiscriminately penalizing both risks catastrophically degrading the model's foundational language capabilities. We resolve this by designing a dynamic gating mechanism that automatically identifies and isolates reasoning-critical tokens, ensuring gradient updates target only behavioral flaws while preserving the model's linguistic priors. Empirically, NSD consistently outperforms OPSD and other label-free, self-bootstrapping reinforcement learning (RL) baselines.</summary>1890    <category term="cs.CL" scheme="http://arxiv.org/schemas/atom"/>1891    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>1892    <published>2026-09-10T15:24:17Z</published>1893    <arxiv:comment>23 pages, 7 figures</arxiv:comment>1894    <arxiv:primary_category term="cs.CL"/>1895    <author>1896      <name>Rongcan Pei</name>1897    </author>1898    <author>1899      <name>Zhepei Wei</name>1900    </author>1901    <author>1902      <name>Shuyao Xu</name>1903    </author>1904    <author>1905      <name>Xinyu Zhu</name>1906    </author>1907    <author>1908      <name>Wei-Lin Chen</name>1909    </author>1910    <author>1911      <name>Yu Meng</name>1912    </author>1913  </entry>1914  <entry>1915    <id>http://arxiv.org/abs/2609.11697v1</id>1916    <title>ActSafeGuard: Differentiable and Training-Aligned Constraint Enforcement for Flow-Matching Policies</title>1917    <updated>2026-09-10T15:21:47Z</updated>1918    <link href="https://arxiv.org/abs/2609.11697v1" rel="alternate" type="text/html"/>1919    <link href="https://arxiv.org/pdf/2609.11697v1" rel="related" type="application/pdf" title="pdf"/>1920    <summary>Vision-Language-Action (VLA) and World-Action Models (WAMs) have demonstrated strong capabilities in general-purpose robotic manipulation, yet their generated actions may violate hard physical constraints and therefore be unsafe or infeasible for deployment. Existing safety approaches either optimize statistical safety objectives without deterministic per-step guarantees or correct unsafe actions only during inference, creating a mismatch between policy training and execution. We introduce ActSafeGuard, a differentiable and training-aligned safeguard layer for flow-matching based policies. ActSafeGuard integrates hard action feasibility into policy learning, not merely treating safety as an inference-time external component. Through an analytical ray-scaling operator design, ActSafeGuard enables boundary-aware gradients to guide the model to naturally learn constrained manifolds. Extensive experiments on multiple standard foundation backbones ($π_{0.5}$ and Fast-WAM) across various tasks demonstrate that ActSafeGuard consistently achieves a $100\%$ step safety rate while fully preserving or even boosting task success rates, providing a scalable and minimally invasive solution for safe embodied AI deployment.</summary>1921    <category term="cs.RO" scheme="http://arxiv.org/schemas/atom"/>1922    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>1923    <published>2026-09-10T15:21:47Z</published>1924    <arxiv:comment>8 pages, 4 figures</arxiv:comment>1925    <arxiv:primary_category term="cs.RO"/>1926    <author>1927      <name>Jianming Ma</name>1928    </author>1929    <author>1930      <name>Rongjun Jin</name>1931    </author>1932    <author>1933      <name>Xiaxi Si</name>1934    </author>1935    <author>1936      <name>Yang Zhang</name>1937    </author>1938    <author>1939      <name>Yiheng Li</name>1940    </author>1941    <author>1942      <name>Yue Gao</name>1943    </author>1944  </entry>1945  <entry>1946    <id>http://arxiv.org/abs/2609.11689v1</id>1947    <title>Geospatial Foundation Models Capture Health-Relevant Dimensions of Place Beyond Conventional Social Risk Indices</title>1948    <updated>2026-09-10T15:16:15Z</updated>1949    <link href="https://arxiv.org/abs/2609.11689v1" rel="alternate" type="text/html"/>1950    <link href="https://arxiv.org/pdf/2609.11689v1" rel="related" type="application/pdf" title="pdf"/>1951    <summary>Area-based social risk indices summarize residents' socioeconomic conditions but incompletely capture physical features of place that may affect health. We evaluated whether numerical representations of physical place produced by four geospatial foundation model families from 2022 satellite data explained residual variance in tract-level associations between the Area Deprivation Index, Social Deprivation Index, and Social Vulnerability Index with health outcomes. We used LightGBM to predict variables from the American Community Survey and 40 chronic disease and health-behavior outcomes from CDC PLACES across 82,646 census tracts in the contiguous United States, evaluating performance across 10 held-out states. Among survey variables, models were moderately predictive of some variables including housing type (R-squared up to 0.54) but weak for disability, unemployment, and income disparity. For health outcomes, models explained up to 54% of variance left unexplained by social risk indices, with the largest gains for annual checkups, arthritis, and high blood pressure. Mean total variance explained by geospatial foundation models across the 40 health-related outcomes increased from 0.31 in the smallest tract-size decile to 0.39 in the largest. Geospatial foundation models capture health-relevant features of place not represented by conventional social risk indices and may usefully augment them in epidemiological analyses.</summary>1952    <category term="stat.AP" scheme="http://arxiv.org/schemas/atom"/>1953    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>1954    <published>2026-09-10T15:16:15Z</published>1955    <arxiv:primary_category term="stat.AP"/>1956    <author>1957      <name>Nathaniel Hendrix</name>1958    </author>1959    <author>1960      <name>Carl Y. Zhang</name>1961    </author>1962    <author>1963      <name>Chris Heitzig</name>1964    </author>1965    <author>1966      <name>Andrew Bazemore</name>1967    </author>1968    <author>1969      <name>David H. Rehkopf</name>1970    </author>1971  </entry>1972  <entry>1973    <id>http://arxiv.org/abs/2609.11687v1</id>1974    <title>Structured Transforms for Low-Overhead Quantization of Language Models</title>1975    <updated>2026-09-10T15:15:50Z</updated>1976    <link href="https://arxiv.org/abs/2609.11687v1" rel="alternate" type="text/html"/>1977    <link href="https://arxiv.org/pdf/2609.11687v1" rel="related" type="application/pdf" title="pdf"/>1978    <summary>We revisit Kashin-decomposition-based weight quantization for large language models and propose an improved algorithm with stronger convergence properties and structured, efficient orthogonal transforms. The method retains the core factorization of each weight into two components -- one with bounded infinity norm and the other with bounded infinity norm after an orthogonal transformation -- but replaces the dense random orthogonal matrix with a sign-randomized Discrete Cosine Transform (DCT), reducing the per-iteration cost from $\mathcal{O}(N^2)$ to $\mathcal{O}(N \log N)$. The proposed greedy algorithm with alternating updates guarantees the four-peak distribution required for stable 2-bit clustering of each factor and admits closed-form initialization of cluster centers, removing the multi-restart k-means bottleneck of prior work. Composed with OPTQ-style sequential error compensation and QuIP-style incoherence preprocessing, the resulting JAX pipeline is competitive with OPTQ, QuIP, QuIP-RG and a fine-tuning- and vector-quantization-free variant of QuIP# at 4-bit per channel on OPT, Llama-2 and Pythia, with favorable wall-clock scaling. The bounded-$\ell_\infty$ factorization is also notably robust: on stress configurations where QuIP variants diverge to four-digit perplexity (Pythia-6.9B) or abort with NaNs in LDL back-substitution (Mistral-7B), Kashin-DCT remains numerically stable and stays close to FP16 baseline. At inference time, each weight decomposes into two 2-bit factor codes per channel that are structurally suited to native-2-bit hardware.</summary>1979    <category term="cs.CL" scheme="http://arxiv.org/schemas/atom"/>1980    <published>2026-09-10T15:15:50Z</published>1981    <arxiv:primary_category term="cs.CL"/>1982    <author>1983      <name>Daria Cherniuk</name>1984    </author>1985    <author>1986      <name>Alexander Rudikov</name>1987    </author>1988    <author>1989      <name>Boris Kashin</name>1990    </author>1991    <author>1992      <name>Ivan Oseledets</name>1993    </author>1994  </entry>1995  <entry>1996    <id>http://arxiv.org/abs/2609.11682v1</id>1997    <title>COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization</title>1998    <updated>2026-09-10T15:12:24Z</updated>1999    <link href="https://arxiv.org/abs/2609.11682v1" rel="alternate" type="text/html"/>2000    <link href="https://arxiv.org/pdf/2609.11682v1" rel="related" type="application/pdf" title="pdf"/>2001    <summary>Large language model (LLM) agents can benefit from reusable skills distilled from prior task experience, yet existing skill optimization methods often rely on costly execution-based evaluation and substantial task data. We introduce \textbf{COBRA-Skills}, an efficient framework that formulates skill optimization as budgeted sequential optimization over a dynamically evolving candidate space. COBRA-Skills couples contextual-bandit-guided prioritization with evidence-grounded skill evolution, selectively allocating evaluations to promising or informative candidates while continually refining the skill population from execution feedback. Across six heterogeneous agent benchmarks and three target models, COBRA-Skills consistently achieves the strongest average performance among compared methods, while reducing optimization cost by 55--58\% relative to SkillOpt and using only 50 unique optimization examples per benchmark. Further analyses show that COBRA-Skills remains robust to changes in the agent harness and performs effectively when the target model itself is used for skill generation and refinement.</summary>2002    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>2003    <published>2026-09-10T15:12:24Z</published>2004    <arxiv:primary_category term="cs.AI"/>2005    <author>2006      <name>Pingchen Lu</name>2007    </author>2008    <author>2009      <name>Xiangyi Wang</name>2010    </author>2011    <author>2012      <name>Xiang Li</name>2013    </author>2014    <author>2015      <name>Jie Mao</name>2016    </author>2017    <author>2018      <name>Zikun Qu</name>2019    </author>2020    <author>2021      <name>Junfeng Luo</name>2022    </author>2023    <author>2024      <name>Yao Shu</name>2025    </author>2026    <author>2027      <name>Bryan Kian Hsiang Low</name>2028    </author>2029    <author>2030      <name>Zhongxiang Dai</name>2031    </author>2032  </entry>2033  <entry>2034    <id>http://arxiv.org/abs/2609.11680v1</id>2035    <title>Single-Stream Multi-Feature Fusion with Temporal Robustness for Gait Emotion Recognition</title>2036    <updated>2026-09-10T15:11:51Z</updated>2037    <link href="https://arxiv.org/abs/2609.11680v1" rel="alternate" type="text/html"/>2038    <link href="https://arxiv.org/pdf/2609.11680v1" rel="related" type="application/pdf" title="pdf"/>2039    <summary>3D skeleton-based gait emotion recognition faces high annotation costs, data scarcity, and poor generalization on heterogeneous data. This paper proposes SV-GCN, a single-stream multi-feature fusion framework with temporal invariance. We introduce intra-frame relative motion features to eliminate frame-rate sensitivity and embed heterogeneous cues at shallow layers, enabling early fusion without multi-stream complexity. For variable-length sequences, we design a global mask-guided valid-frame spatio-temporal graph convolution module, introducing frame-rate insensitivity for the first time in this domain. On the E-Gait dataset, our method achieves performance comparable to state-of-the-art while demonstrating strong generalization across varying sequence lengths and frame rates, offering a viable pathway for pre-training on large-scale skeleton-based action recognition datasets.</summary>2040    <category term="cs.CV" scheme="http://arxiv.org/schemas/atom"/>2041    <published>2026-09-10T15:11:51Z</published>2042    <arxiv:comment>Accepted at the 35th International Conference on Artificial Neural Networks (ICANN 2026)</arxiv:comment>2043    <arxiv:primary_category term="cs.CV"/>2044    <author>2045      <name>Shirong Lyu</name>2046    </author>2047    <author>2048      <name>Silu Quan</name>2049    </author>2050    <author>2051      <name>Yixuan Ding</name>2052    </author>2053    <author>2054      <name>Chengpeng Wang</name>2055    </author>2056  </entry>2057  <entry>2058    <id>http://arxiv.org/abs/2609.11677v1</id>2059    <title>Ecdysis: Efficient and Effective Training of Runtime Harnesses for LLM Agents</title>2060    <updated>2026-09-10T15:09:08Z</updated>2061    <link href="https://arxiv.org/abs/2609.11677v1" rel="alternate" type="text/html"/>2062    <link href="https://arxiv.org/pdf/2609.11677v1" rel="related" type="application/pdf" title="pdf"/>2063    <summary>Self-evolving runtime harnesses can substantially improve the capabilities of large language model (LLM) agents and provide a promising paradigm for optimizing agent execution. Existing harness evolution methods typically rely on iterative search, repeatedly evaluating and revising candidate harnesses based on execution feedback from task instances. While this paradigm enables continuous harness optimization, it incurs substantial time overhead due to repeated agent executions and code modifications, and may overfit to observed tasks and specific failure patterns, resulting in degraded generalization to unseen tasks. We identify the lack of principled failure diagnosis as a key bottleneck in harness evolution: an observed failure can reflect either model-specific deficiencies or systematic harness deficiencies, and directly optimizing against individual failures can lead to unnecessary model-specific accommodation. We therefore propose Ecdysis, an efficient and effective framework that distinguishes model-specific accommodation from harness-level repair and biases adaptation toward systematic harness deficiencies by identifying recurring cross-task failure patterns. Ecdysis adopts a batch-level cross-instance failure aggregation paradigm to jointly analyze failure evidence from multiple task instances and further introduces Failure-Driven Collaborative Refinement to diagnose failure causes and iteratively refine harness modification specifications. By combining cross-instance failure analysis with multi-role diagnosis, Ecdysis enables more effective harness evolution with lower training time. Experiments show that Ecdysis achieves up to a 1.84x speedup in harness training compared with existing harness evolution methods, while improving the reasoning accuracy of the resulting harnesses by 18.56%.</summary>2064    <category term="cs.SE" scheme="http://arxiv.org/schemas/atom"/>2065    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>2066    <published>2026-09-10T15:09:08Z</published>2067    <arxiv:primary_category term="cs.SE"/>2068    <author>2069      <name>Ruiqing Yue</name>2070    </author>2071    <author>2072      <name>Yu Cui</name>2073    </author>2074    <author>2075      <name>Zhuoyu Sun</name>2076    </author>2077    <author>2078      <name>Sicheng Pan</name>2079    </author>2080    <author>2081      <name>Xianhong Xue</name>2082    </author>2083    <author>2084      <name>Tingyu Li</name>2085    </author>2086    <author>2087      <name>Ting Li</name>2088    </author>2089    <author>2090      <name>Wenzhuo Zhu</name>2091    </author>2092    <author>2093      <name>Yi Chen</name>2094    </author>2095    <author>2096      <name>Yifei Liu</name>2097    </author>2098    <author>2099      <name>Baohan Huang</name>2100    </author>2101    <author>2102      <name>Zhe Cui</name>2103    </author>2104    <author>2105      <name>Haibin Zhang</name>2106    </author>2107    <author>2108      <name>Cong Zuo</name>2109    </author>2110  </entry>2111  <entry>2112    <id>http://arxiv.org/abs/2609.11674v1</id>2113    <title>Geospatial AI, Dataverse Metadata, and the Study of Place-Based Government</title>2114    <updated>2026-09-10T15:08:58Z</updated>2115    <link href="https://arxiv.org/abs/2609.11674v1" rel="alternate" type="text/html"/>2116    <link href="https://arxiv.org/pdf/2609.11674v1" rel="related" type="application/pdf" title="pdf"/>2117    <summary>Harvard Dataverse hosts over 150,000 research datasets, but the geographic information those datasets carry is entered as free text by depositors and has never been assembled into a searchable structure. We construct a knowledge graph from the repository's public data and metadata, organizing 102,650 datasets within a 215,985-node network of 528,003 edges linking datasets to keywords, publications, subjects, journals, and locations. Of those datasets, 43,991 (42.9 percent) carry at least one geospatial field, geographic coverage, geographic unit, or a bounding box and 96.9 percent of all nodes sit in a single connected component, so datasets remain reachable from one another even when their geospatial metadata share nothing in common. A conservative keyword search identifies 7,654 geospatially tagged datasets (17.4 percent) as directly policy-relevant, with elections and legislatures the largest cluster, followed by government administration, health policy, transportation, and education. Five datasets illustrate how this metadata behaves across policy domains and spatial scales, and an extended use case shows how community language models, stance detection with geographic aggregation, and partisan language bridging tools can attach discourse to place. The central obstacle is place resolution: the same location appears as many disconnected nodes. We argue that the graph provides a concrete setting for developing AI-driven metadata enrichment and entity resolution, and we document its coverage skew toward American, city-level data.</summary>2118    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>2119    <published>2026-09-10T15:08:58Z</published>2120    <arxiv:primary_category term="cs.AI"/>2121    <author>2122      <name>Danny EBanks</name>2123    </author>2124    <author>2125      <name>Devika Jain</name>2126    </author>2127  </entry>2128  <entry>2129    <id>http://arxiv.org/abs/2609.11673v1</id>2130    <title>Multimodal Taxonomic Conditioning for Generative Plankton Imagery</title>2131    <updated>2026-09-10T15:08:53Z</updated>2132    <link href="https://arxiv.org/abs/2609.11673v1" rel="alternate" type="text/html"/>2133    <link href="https://arxiv.org/pdf/2609.11673v1" rel="related" type="application/pdf" title="pdf"/>2134    <summary>Automated plankton imaging produces severely long-tailed datasets, where the rare taxa of greatest ecological interest have too few images to train or evaluate classifiers reliably. We generate synthetic plankton imagery conditioned on taxonomy: a CLIP encoder is adapted on a large plankton corpus with a ranked contrastive objective extended to deep, ragged taxonomies, then frozen to condition a parameter-efficient diffusion transformer. We evaluate synthetic sample quality on distributional fidelity and downstream classifier utility.</summary>2135    <category term="cs.CV" scheme="http://arxiv.org/schemas/atom"/>2136    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>2137    <published>2026-09-10T15:08:53Z</published>2138    <arxiv:comment>European Conference on Computer Vision (ECCV) 2nd Workshop on Marine Vision</arxiv:comment>2139    <arxiv:primary_category term="cs.CV"/>2140    <author>2141      <name>Daniela Ivanova</name>2142    </author>2143    <author>2144      <name>Ozgu Goksu</name>2145    </author>2146    <author>2147      <name>Nicolas Pugeault</name>2148    </author>2149  </entry>2150  <entry>2151    <id>http://arxiv.org/abs/2609.11667v1</id>2152    <title>Warrant Theory</title>2153    <updated>2026-09-10T15:06:30Z</updated>2154    <link href="https://arxiv.org/abs/2609.11667v1" rel="alternate" type="text/html"/>2155    <link href="https://arxiv.org/pdf/2609.11667v1" rel="related" type="application/pdf" title="pdf"/>2156    <summary>In this paper, we develop warrant theory as a philosophical discipline concerned with the inferential legitimacy of propositions within logical analysis. Warrant theory reconceptualises logic as a normative framework governing the conditions under which propositions may be introduced, accepted, rejected, and inferentially employed. Warrant is understood as inferential entitlement and is distinguished from truth, belief, and other psychological attitudes, while its relation to inferential use and meaning is examined. Warrant-theoretic analysis is then developed as a systematic method for investigating how propositions acquire inferential standing, how that standing develops, and how inferential positions interact through relations of dependence, compatibility, incompatibility, and exclusion. Acceptance and rejection provide the bilateral vocabulary for representing positive and negative inferential positions and the consequences and commitments associated with them. Finally, these elements are brought together in a warrant-theoretic definition of logic as the formal and normative study of the conditions under which propositions may be legitimately accepted or rejected and of the inferential transitions that such legitimacy warrants. On this account, logical consequence and logical failure are understood through the presence, preservation, or absence of inferential entitlement, thus locating the philosophical subject matter of logic in the systematic governance of inferential legitimacy.</summary>2157    <category term="cs.LO" scheme="http://arxiv.org/schemas/atom"/>2158    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>2159    <published>2026-09-10T15:06:30Z</published>2160    <arxiv:comment>17 Pages</arxiv:comment>2161    <arxiv:primary_category term="cs.LO"/>2162    <author>2163      <name>Khashayar Irani</name>2164    </author>2165  </entry>2166  <entry>2167    <id>http://arxiv.org/abs/2609.11660v1</id>2168    <title>Autonomy, Social Norms, and Alignment: Towards a Developmental Framework for Autonomous Artificial Agents</title>2169    <updated>2026-09-10T15:02:52Z</updated>2170    <link href="https://arxiv.org/abs/2609.11660v1" rel="alternate" type="text/html"/>2171    <link href="https://arxiv.org/pdf/2609.11660v1" rel="related" type="application/pdf" title="pdf"/>2172    <summary>In recent years, artificial intelligence has made extraordinary progress thanks to large-scale models capable of generalization and the generation of complex outputs. However, transferring this potential into embodied agents reveals a significant limitation: the most advanced systems rely on pre-existing datasets and human feedback strategies that are powerful but insufficient in dynamic or unknown contexts. To adapt, an agent must acquire knowledge through direct interaction with its environment. One strategy to address this challenge involves introducing higher-level mechanisms, such as intrinsic motivations, which leverage curiosity and competence, to guide exploration and learning in complex environments. While this flexibility expands autonomy, it complicates the task of ensuring agents remain aligned with human goals. Alignment, already a challenge for artificial systems in general, becomes even more complex in unstructured and dynamic contexts where predefined rules prove insufficient. To be effective and adaptable, norms must be rooted in experience through an epistemological process that starting from simple, situated principles allows for the gradual construction of more complex rules through experience, autonomous learning, and cooperation with other moral agents. Similarly to children learning social norms by exploring their environment and participating in collective practices, artificial agents must also be educated toward alignment. Following Dennett, the status of a moral agent is not innate but is attributed gradually based on the ability to responsibly manage increasing degrees of freedom. From this perspective, the regulatory sandboxes can be viewed as pedagogical environments for AI: dynamic spaces where alignment develops as a formative process, progressively shaping autonomous behaviors through interaction and cooperation in scenarios of increasing complexity.</summary>2173    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>2174    <published>2026-09-10T15:02:52Z</published>2175    <arxiv:comment>In publication in the proceedings of SIpEIA 2026 conference</arxiv:comment>2176    <arxiv:primary_category term="cs.AI"/>2177    <author>2178      <name>Marica Notte</name>2179    </author>2180    <author>2181      <name>Ludovica Marinucci</name>2182    </author>2183    <author>2184      <name>Vieri Giuliano Santucci</name>2185    </author>2186  </entry>2187  <entry>2188    <id>http://arxiv.org/abs/2609.11656v1</id>2189    <title>Learnware and AI Model Management System</title>2190    <updated>2026-09-10T15:01:27Z</updated>2191    <link href="https://arxiv.org/abs/2609.11656v1" rel="alternate" type="text/html"/>2192    <link href="https://arxiv.org/pdf/2609.11656v1" rel="related" type="application/pdf" title="pdf"/>2193    <summary>The transition from file storage to database management systems transformed stored data into managed resources. AI now faces an analogous transition from AI model storage to AI model management. Existing model pools essentially serve as \textit{AI model storage systems}. What is needed instead are \textit{AI model management systems} that enable models trained by different developers, for different tasks, with different data, and under different objectives to be identified, reused, and even assembled to address future user tasks. Because AI model developers are generally unwilling to share their training data, such systems should operate without accessing the training data of model developers and, ideally, without accessing raw data of future users. This requirement poses a fundamental challenge: the functionality of a modern AI model may not be fully understood even by the developer who trained it. How, then, can a system identify which models are useful for a given user task, let alone assemble models developed independently for different purposes? At first glance, this objective may appear unattainable. It becomes possible, however, by upgrading the basic unit of management from a machine learning model to a \textit{learnware}. \textit{Learnware = Model + Specification}. The specification, whose assignment transforms a trained model into a learnware, is generated with the help of a machine learning process without disclosing the training data of the developer and has a theoretically established data-preservation property. The \textit{Learnware Dock System (LDS)} provides a path toward powerful AI model management systems. Because specifications are generated according to a published reference and are comparable across models, they can also serve as an AI model \textit{collaboration protocol} through which independently developed models, including intelligent agents, can collaborate.</summary>2194    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>2195    <published>2026-09-10T15:01:27Z</published>2196    <arxiv:primary_category term="cs.LG"/>2197    <author>2198      <name>Zhi-Hua Zhou</name>2199    </author>2200  </entry>2201  <entry>2202    <id>http://arxiv.org/abs/2609.11655v1</id>2203    <title>Musec: MomentUm SpEctral Clipping for Stable Muon-type Training</title>2204    <updated>2026-09-10T15:00:01Z</updated>2205    <link href="https://arxiv.org/abs/2609.11655v1" rel="alternate" type="text/html"/>2206    <link href="https://arxiv.org/pdf/2609.11655v1" rel="related" type="application/pdf" title="pdf"/>2207    <summary>Muon has emerged as a highly effective optimizer for large language model training, often achieving superior convergence and performance compared with the widely adopted Adam and AdamW optimizers. Nevertheless, Muon is prone to training instability due to its spectral flattening, manifested by loss spikes and unbounded growth of model weights. Existing approaches primarily rely on weight or attention-logit clipping, which require architecture-specific modifications and do not directly address instability across all model components. We propose MomentUm SpEctral Clipping (Musec), which replaces Muon's spectral flattening with spectral clipping: rather than setting all singular values of the momentum matrix to approximately one, Musec clips singular values that exceed a threshold while preserving the underlying spectral structure of the momentum. Our strategy provides an optimizer-level, architecture-agnostic mechanism for stabilizing Muon training. We further develop Soft Musec, an efficient implementation that uses a smooth spectral saturation function approximated by coupled Newton-Schulz iterations. Theoretically, we establish convergence guarantees for Musec in nonconvex nonsmooth stochastic optimization. To the best of our knowledge, this is the first convergence guarantee for Muon-type methods in the nonconvex nonsmooth setting. We provide empirical studies to show that Soft Musec consistently improves training stability over existing Muon variants across a wide range of learning rates and model sizes. Notably, Soft Musec remains stable in settings where existing Muon variants diverge, while matching their performance under well-tuned configurations.</summary>2208    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>2209    <published>2026-09-10T15:00:01Z</published>2210    <arxiv:primary_category term="cs.LG"/>2211    <author>2212      <name>Zhuanghua Liu</name>2213    </author>2214    <author>2215      <name>Menglian Wang</name>2216    </author>2217    <author>2218      <name>Luo Luo</name>2219    </author>2220  </entry>2221  <entry>2222    <id>http://arxiv.org/abs/2609.11650v1</id>2223    <title>Self-Supervised Cardiac Phase Detection via Single-Parameter Latent Orbits</title>2224    <updated>2026-09-10T14:56:16Z</updated>2225    <link href="https://arxiv.org/abs/2609.11650v1" rel="alternate" type="text/html"/>2226    <link href="https://arxiv.org/pdf/2609.11650v1" rel="related" type="application/pdf" title="pdf"/>2227    <summary>Accurate identification of end-diastole (ED) and end-systole (ES) in echocardiography underpins the quantification of ventricular function, yet manual selection of these key frames is subjective and introduces clinically significant inter-operator variability. Recent self-supervised methods either prescribe strict periodic trajectories or learn an unconstrained low-dimensional motion subspace from reconstruction or registration objectives. The former offers interpretability but imposes restrictive assumptions on temporal progression, whereas the latter leaves cardiac phase implicit and ED/ES must be recovered through post-hoc geometric processing of the learned trajectory. We translate the physiological observation that cardiac phase is a one-dimensional signal into a prior by constraining the latent motion component to a single-parameter latent orbit, i.e., a global linear trajectory in latent space indexed by a bounded scalar phase variable. Mapping this variable through a sinusoidal nonlinearity yields an oscillatory motion signal with consistent temporal ordering, enabling direct identification of ED and ES from the learned phase signal. This inductive bias allows the model to capture an interpretable representation of the cardiac cycle, while maintaining flexibility to capture irregular heartbeats. Trained on EchoNet-Dynamic without annotations, our minimal single-parameter cardiac phase model learns an effective latent orbit, significantly improves upon the previous state of the art in ED localisation and matches it in ES localisation while using a more constrained representation and fewer training epochs. This demonstrates that a principled physiological inductive bias can match or exceed the performance of more complex representations. Code is available at: https://github.com/BonniciJ/OrbitalEcho/</summary>2228    <category term="cs.CV" scheme="http://arxiv.org/schemas/atom"/>2229    <published>2026-09-10T14:56:16Z</published>2230    <arxiv:comment>Accepted for oral presentation at the ASMUS workshop at MICCAI 2026</arxiv:comment>2231    <arxiv:primary_category term="cs.CV"/>2232    <author>2233      <name>John Bonnici</name>2234    </author>2235    <author>2236      <name>Matthew Baugh</name>2237    </author>2238    <author>2239      <name>Aleksandra Kulbaka</name>2240    </author>2241    <author>2242      <name>Sarah Cechnicka</name>2243    </author>2244    <author>2245      <name>Bernhard Kainz</name>2246    </author>2247    <author>2248      <name>Alberto Gomez</name>2249    </author>2250  </entry>2251  <entry>2252    <id>http://arxiv.org/abs/2609.11648v1</id>2253    <title>RDDMPI: Residual Denoising Diffusion Model for Probabilistic Multivariate Time Series Imputation</title>2254    <updated>2026-09-10T14:53:38Z</updated>2255    <link href="https://arxiv.org/abs/2609.11648v1" rel="alternate" type="text/html"/>2256    <link href="https://arxiv.org/pdf/2609.11648v1" rel="related" type="application/pdf" title="pdf"/>2257    <summary>Multivariate time series imputation (MTSI) aims to recover missing values in temporal data composed of multiple interdependent variables. This problem is central to real-world applications such as healthcare monitoring, traffic networks, and energy systems. Recent diffusion-based approaches have shown strong potential for probabilistic imputation by learning to generate missing values through iterative denoising. However, most existing approaches perform diffusion directly in the original data space, requiring the denoising network to simultaneously capture global structure, temporal dynamics, and stochastic variability. This makes the generative task unnecessarily complex, especially when modern deterministic imputers can already provide accurate initial reconstructions. To address this limitation, we propose RDDMPI, a conditional residual diffusion framework that operates directly in residual space. Instead of modeling the full missing signal directly, we reformulate probabilistic imputation as a baseline-residual decomposition, where a pretrained model captures the dominant signal and a diffusion process models the residual uncertainty. To better exploit deterministic guidance, \model{} conditions the reverse denoising process on both the baseline-completed signal and its latent representation, while a reliability-aware conditioning mechanism adaptively controls the influence of baseline information during residual generation. This formulation simplifies the diffusion learning objective, enabling it to focus on structured correction terms rather than reconstructing the full signal. Experiments on multiple benchmark datasets demonstrate that RDDMPI consistently improves both reconstruction accuracy and uncertainty quantification.</summary>2258    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>2259    <category term="stat.ML" scheme="http://arxiv.org/schemas/atom"/>2260    <published>2026-09-10T14:53:38Z</published>2261    <arxiv:primary_category term="cs.LG"/>2262    <author>2263      <name>Ramiro Valdes Jara</name>2264    </author>2265    <author>2266      <name>David Chapman</name>2267    </author>2268    <author>2269      <name>Adam Meyers</name>2270    </author>2271  </entry>2272  <entry>2273    <id>http://arxiv.org/abs/2609.11642v1</id>2274    <title>ZipCodec: Ultra-Low-Frame-Rate Streaming Speech Coding</title>2275    <updated>2026-09-10T14:49:54Z</updated>2276    <link href="https://arxiv.org/abs/2609.11642v1" rel="alternate" type="text/html"/>2277    <link href="https://arxiv.org/pdf/2609.11642v1" rel="related" type="application/pdf" title="pdf"/>2278    <summary>Neural audio codecs are a fundamental component of modern speech generation systems. While recent codecs achieve increasingly low bitrates, reducing frame rate remains challenging, as each token must preserve more information while maintaining reconstruction quality. We present ZipCodec, a streaming neural speech codec operating at 6.25 Hz and 0.80 kbps with a theoretical latency of 160 ms. Our approach combines large-scale WavLM distillation with a redesigned transformer-based architecture, a scalar spherical quantizer, and a latency-aware streaming decoder. Experiments show that ZipCodec substantially outperforms existing streaming codecs at comparable bitrates in both reconstruction and downstream tasks, while operating at a significantly lower frame rate. Despite its 842M parameters, ZipCodec achieves real-time single-stream inference on a consumer-grade CPU. Demo samples, code and checkpoints are available at https://lucadellalib.github.io/zipcodec-web/.</summary>2279    <category term="cs.SD" scheme="http://arxiv.org/schemas/atom"/>2280    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>2281    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>2282    <published>2026-09-10T14:49:54Z</published>2283    <arxiv:comment>5 pages, 1 figure</arxiv:comment>2284    <arxiv:primary_category term="cs.SD"/>2285    <author>2286      <name>Luca Della Libera</name>2287    </author>2288    <author>2289      <name>Cem Subakan</name>2290    </author>2291    <author>2292      <name>Mirco Ravanelli</name>2293    </author>2294  </entry>2295  <entry>2296    <id>http://arxiv.org/abs/2609.11639v1</id>2297    <title>LoaDiff: Conditional Generation of Electricity Consumption Time Series for Energy Analytics</title>2298    <updated>2026-09-10T14:46:30Z</updated>2299    <link href="https://arxiv.org/abs/2609.11639v1" rel="alternate" type="text/html"/>2300    <link href="https://arxiv.org/pdf/2609.11639v1" rel="related" type="application/pdf" title="pdf"/>2301    <summary>The energy transition is reshaping residential electricity consumption through the increasing adoption of distributed generation, electrified appliances, and demand-response programs. Understanding these evolving behaviors requires access to granular smart-meter data for applications such as load forecasting, appliance detection, and demand-side flexibility analysis. However, such data are subject to strict access restrictions and data-protection regulations. Thus, realistic synthetic alternatives are necessary. In this paper, we introduce LoaDiff, a diffusion-based generative model for year-long, sub-hourly smart-meter load curves. LoaDiff supports flexible conditioning on static household attributes, such as appliance ownership, and dynamic contextual variables, including calendar information and outdoor temperature. We evaluate the model against multiple generative baselines on three residential electricity-consumption datasets. Our experiments assess four complementary dimensions: fidelity and diversity, training-record memorization risk, downstream utility for load forecasting and appliance detection, and conditional controllability under alternative temperature conditions. The results show that LoaDiff generates realistic and diverse load profiles, achieves a favorable trade-off between generation quality and limited evidence of memorization, preserves information useful for downstream energy applications, and responds coherently to changes in conditioning variables.</summary>2302    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>2303    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>2304    <category term="eess.SP" scheme="http://arxiv.org/schemas/atom"/>2305    <published>2026-09-10T14:46:30Z</published>2306    <arxiv:comment>10 pages, 5 figures. This paper appeared in IEEE ICDM 2026</arxiv:comment>2307    <arxiv:primary_category term="cs.LG"/>2308    <author>2309      <name>Mariia Baranova</name>2310    </author>2311    <author>2312      <name>Adrien Petralia</name>2313    </author>2314    <author>2315      <name>Etienne Le Naour</name>2316    </author>2317    <author>2318      <name>Nathan Etourneau</name>2319    </author>2320    <author>2321      <name>Guillaume Hofmann</name>2322    </author>2323    <author>2324      <name>Themis Palpanas</name>2325    </author>2326  </entry>2327  <entry>2328    <id>http://arxiv.org/abs/2609.11638v1</id>2329    <title>Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation</title>2330    <updated>2026-09-10T14:46:07Z</updated>2331    <link href="https://arxiv.org/abs/2609.11638v1" rel="alternate" type="text/html"/>2332    <link href="https://arxiv.org/pdf/2609.11638v1" rel="related" type="application/pdf" title="pdf"/>2333    <summary>We present Vidu S2, which comprises Vidu S2-Avatar, a real-time interactive digital-character model, and Vidu S2-Editing, a real-time video editing model. Moreover, we explore the feasibility of real-time spatial video generation for both Vidu S2-Avatar and Vidu S2-Editing. Compared with Vidu S1, Vidu S2-Avatar supports real-time 720p video generation, generation with dynamic references that can be updated at any moment, and stronger instruction following, such as dancing. Vidu S2-Editing supports editing a video stream in real time, including style rendering, clothing replacement, character replacement, and background replacement. Experiments show that Vidu S2 outperforms all baselines. A playable online demo is available at https://vidu.com/vidu-stream.</summary>2334    <category term="cs.CV" scheme="http://arxiv.org/schemas/atom"/>2335    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>2336    <published>2026-09-10T14:46:07Z</published>2337    <arxiv:primary_category term="cs.CV"/>2338    <author>2339      <name>Jintao Zhang</name>2340    </author>2341    <author>2342      <name>Kai Jiang</name>2343    </author>2344    <author>2345      <name>Jintao Chen</name>2346    </author>2347    <author>2348      <name>Xu Wang</name>2349    </author>2350    <author>2351      <name>Deyuan Liu</name>2352    </author>2353    <author>2354      <name>Jungang Li</name>2355    </author>2356    <author>2357      <name>Dechuang Chen</name>2358    </author>2359    <author>2360      <name>Ming Lin</name>2361    </author>2362    <author>2363      <name>Jingjiang Zhou</name>2364    </author>2365    <author>2366      <name>Haopeng Jin</name>2367    </author>2368    <author>2369      <name>Qi Jia</name>2370    </author>2371    <author>2372      <name>Xiaohang Wang</name>2373    </author>2374    <author>2375      <name>Yaole Wang</name>2376    </author>2377    <author>2378      <name>Zhanqiang Zhang</name>2379    </author>2380    <author>2381      <name>Ran Li</name>2382    </author>2383    <author>2384      <name>Zhengkun Huang</name>2385    </author>2386    <author>2387      <name>Shuyue Xiong</name>2388    </author>2389    <author>2390      <name>Yuji Wang</name>2391    </author>2392    <author>2393      <name>Zikun Dai</name>2394    </author>2395    <author>2396      <name>Hui He</name>2397    </author>2398    <author>2399      <name>Yang Luo</name>2400    </author>2401    <author>2402      <name>Mang Ning</name>2403    </author>2404    <author>2405      <name>Weiqi Feng</name>2406    </author>2407    <author>2408      <name>Chengyang Ye</name>2409    </author>2410    <author>2411      <name>Xinyue Lin</name>2412    </author>2413    <author>2414      <name>Min Zhao</name>2415    </author>2416    <author>2417      <name>Hongzhou Zhu</name>2418    </author>2419    <author>2420      <name>Hengkai Tan</name>2421    </author>2422    <author>2423      <name>Zeyuan Wang</name>2424    </author>2425    <author>2426      <name>Chendong Xiang</name>2427    </author>2428    <author>2429      <name>Kaiwen Zheng</name>2430    </author>2431    <author>2432      <name>Zhijie Deng</name>2433    </author>2434    <author>2435      <name>Fan Bao</name>2436    </author>2437    <author>2438      <name>Jianfei Chen</name>2439    </author>2440    <author>2441      <name>Jun Zhu</name>2442    </author>2443  </entry>2444  <entry>2445    <id>http://arxiv.org/abs/2609.11636v1</id>2446    <title>MAPLE: Memory-Augmented Planning with Language and Evolution</title>2447    <updated>2026-09-10T14:44:50Z</updated>2448    <link href="https://arxiv.org/abs/2609.11636v1" rel="alternate" type="text/html"/>2449    <link href="https://arxiv.org/pdf/2609.11636v1" rel="related" type="application/pdf" title="pdf"/>2450    <summary>Domain practitioners understand their business constraints but may lack operations-research expertise or dedicated support. LLM-based optimization agents translate natural-language requirements into models or solver programs that established optimization tools can execute. This progress makes optimization more accessible, but real-world operations are dynamic: changing demand, resources, and priorities require updates to data, constraints, and objectives. Methods centered on isolated requests offer limited support for rapid adaptation that preserves earlier decisions and reuses useful search results. We introduce MAPLE (Memory-Augmented Planning with Language and Evolution), an agent for maintaining optimization problems through successive natural-language requests. MAPLE combines language-based problem construction with mathematical programming and evolutionary search. It retains the optimization program, accepted plans, earlier updates, and candidate solutions for subsequent requests. We introduce NLDO, a benchmark of 15 trajectories and 180 updates spanning selection, scheduling, rostering, routing, and cloud-resource placement. In the main evaluation, MAPLE completes all trajectories and achieves online scalar quality of 0.951 and a Pareto hypervolume ratio of 0.875. Controlled comparisons further show that maintaining executable state improves update validity and can preserve useful search information across substantial revisions.</summary>2451    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>2452    <published>2026-09-10T14:44:50Z</published>2453    <arxiv:primary_category term="cs.AI"/>2454    <author>2455      <name>Kesheng Chen</name>2456    </author>2457    <author>2458      <name>Yamin Hu</name>2459    </author>2460    <author>2461      <name>Wenjian Luo</name>2462    </author>2463  </entry>2464  <entry>2465    <id>http://arxiv.org/abs/2609.11628v1</id>2466    <title>Physics-Informed Neural Networks to Infer the Perpendicular Energy Conductivity in the Scrape-Off Layer of Stellarator Devices</title>2467    <updated>2026-09-10T14:40:32Z</updated>2468    <link href="https://arxiv.org/abs/2609.11628v1" rel="alternate" type="text/html"/>2469    <link href="https://arxiv.org/pdf/2609.11628v1" rel="related" type="application/pdf" title="pdf"/>2470    <summary>In this work, we develop an inverse Physics-Informed Neural Network (PINN) framework to infer the dependence of the scrape-off layer (SOL) perpendicular heat conductivity on plasma density and temperature, $κ_\perp(n,T)$. The method combines radial profile measurements of electron density and temperature with the residual of a reduced one-dimensional SOL transport equation, so that the inferred conductivity is constrained by both the measurements and the underlying transport model. Three neural networks are trained simultaneously: two reconstruct the temperature and density profiles as functions of the radial coordinate and transported power, while a third represents the effective conductivity as a function of the local density and temperature. The framework is first validated using synthetic data generated from a prescribed conductivity function, allowing the inferred $κ_\perp(n,T)$ to be compared directly with the ground truth. The model recovers the imposed functional dependence with errors below $10~\%$ in the data-constrained region. Bootstrap resampling is shown to provide a practical indicator of prediction reliability and consistency. A scan in the number of plasma profiles used for training and the number of radial measurement positions per profile identifies a practical trade-off between reconstruction accuracy and data availability. Finally, the method is applied to an experimental dataset from the TJ-II stellarator obtained with the helium-beam diagnostic. This exploratory application provides an initial estimate of the effective SOL conductivity and illustrates the potential of inverse PINNs for extracting transport information from plasma edge measurements.</summary>2471    <category term="physics.plasm-ph" scheme="http://arxiv.org/schemas/atom"/>2472    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>2473    <published>2026-09-10T14:40:32Z</published>2474    <arxiv:comment>18 pages, 10 figures</arxiv:comment>2475    <arxiv:primary_category term="physics.plasm-ph"/>2476    <author>2477      <name>J. Gallego</name>2478      <arxiv:affiliation>Departamento de Tecnología, CIEMAT, Spain</arxiv:affiliation>2479    </author>2480    <author>2481      <name>P. Protopapas</name>2482      <arxiv:affiliation>Harvard John A. Paulson School of Engineering and Applied Sciences, USA</arxiv:affiliation>2483    </author>2484    <author>2485      <name>A. Bustos</name>2486      <arxiv:affiliation>Departamento de Tecnología, CIEMAT, Spain</arxiv:affiliation>2487    </author>2488    <author>2489      <name>A. Alonso</name>2490      <arxiv:affiliation>Laboratorio Nacional de Fusión, CIEMAT, Spain</arxiv:affiliation>2491    </author>2492    <author>2493      <name>S. Barquero</name>2494      <arxiv:affiliation>Laboratorio Nacional de Fusión, CIEMAT, Spain</arxiv:affiliation>2495    </author>2496    <author>2497      <name>A. Baciero</name>2498      <arxiv:affiliation>Laboratorio Nacional de Fusión, CIEMAT, Spain</arxiv:affiliation>2499    </author>2500    <author>2501      <name>I. Rivera</name>2502      <arxiv:affiliation>Laboratorio Nacional de Fusión, CIEMAT, Spain</arxiv:affiliation>2503    </author>2504    <author>2505      <name>J. A. Moríñigo</name>2506      <arxiv:affiliation>Departamento de Tecnología, CIEMAT, Spain</arxiv:affiliation>2507    </author>2508    <author>2509      <name>R. Mayo-García</name>2510      <arxiv:affiliation>Departamento de Tecnología, CIEMAT, Spain</arxiv:affiliation>2511    </author>2512  </entry>2513  <entry>2514    <id>http://arxiv.org/abs/2609.11620v1</id>2515    <title>A Training-Free, Alignment-Free Approach to Corporate Intelligence: Application to SEC Filings</title>2516    <updated>2026-09-10T14:32:35Z</updated>2517    <link href="https://arxiv.org/abs/2609.11620v1" rel="alternate" type="text/html"/>2518    <link href="https://arxiv.org/pdf/2609.11620v1" rel="related" type="application/pdf" title="pdf"/>2519    <summary>High-dimensional dense text embeddings and large language models face real obstacles in financial-disclosure analysis: context-window limits, hallucination risk, high computational cost, and the arbitrary rotation of vector spaces across independently trained models. We present a training-free, alignment-free framework for corporate intelligence built on deterministic sparse seed vectors. Hashing word strings into a fixed high-dimensional basis places all documents and all temporal epochs in a common coordinate system by construction, removing any need for training or alignment. Accumulating these seed vectors across sentence contexts yields corpus-specific semantic signatures that compose linearly, supporting sub-second document comparison, issuer fingerprinting, tracking of how an issuer's vocabulary shifts between filings, and thematic sentence extraction, all on ordinary CPU hardware. Demonstrating the approach on a multi-year corpus of SEC filings (10-K, 10-Q, 8-K), we show how material corporate events, among them Boeing's 737 MAX crisis, Intel's supply-chain disruptions, and Bunge's acquisition of Viterra, emerge as distinct, interpretable semantic profiles, each traceable to the exact source sentences that produced it, with no domain-specific training and no LLM inference.</summary>2520    <category term="cs.CL" scheme="http://arxiv.org/schemas/atom"/>2521    <published>2026-09-10T14:32:35Z</published>2522    <arxiv:comment>26 pages, 2 figures</arxiv:comment>2523    <arxiv:primary_category term="cs.CL"/>2524    <author>2525      <name>Jean-François Delpech</name>2526    </author>2527  </entry>2528  <entry>2529    <id>http://arxiv.org/abs/2609.11616v1</id>2530    <title>LangStreet: Persistent Language Fields for Anchor-Decoded Street Gaussians</title>2531    <updated>2026-09-10T14:31:02Z</updated>2532    <link href="https://arxiv.org/abs/2609.11616v1" rel="alternate" type="text/html"/>2533    <link href="https://arxiv.org/pdf/2609.11616v1" rel="related" type="application/pdf" title="pdf"/>2534    <summary>Language Gaussian fields implicitly assume that the primitive carrying semantics remains identifiable across views. This assumption breaks in scalable anchor-decoded representations, where persistent anchors generate view-conditioned child Gaussians whose geometry and appearance vary with the camera. We introduce Ours, a persistent language field for such structured Gaussian scenes. Our key idea is semantic ownership: transient children route observations, while persistent decoder slots and their parent anchors own the language field. We use alpha-compositing responsibilities to accumulate additive directional evidence at slots; these statistics marginalize exactly to anchors. We then complete weakly supported slots with anchor-aligned evidence while preserving the anchor direction, and represent slot detail through low-rank residuals in anchor-relative semantic coordinates. Our primary model, Ours (base), stores anchor features together with compact slot residuals. Ours (light) retains only anchor features, whereas Ours (max) stores the full-dimensional completed slot features explicitly. Without scene-specific semantic optimization, Ours (base) nearly matches Ours (max) across KITTI, Virtual KITTI, and Waymo. On KITTI, it achieves 34.19 2D mIoU with a 2.72 GiB effective feature footprint, compared with 34.20 mIoU and 12.90 GiB for Ours (max). The same accuracy-storage trend holds on Virtual KITTI and Waymo. These results show that language fields on view-conditioned splats require persistent semantic ownership, conserved evidence, and a hierarchy that balances stability, detail, and representation cost. Our code, checkpoints, and benchmark suite will be publicly available.</summary>2535    <category term="cs.CV" scheme="http://arxiv.org/schemas/atom"/>2536    <published>2026-09-10T14:31:02Z</published>2537    <arxiv:primary_category term="cs.CV"/>2538    <author>2539      <name>Runyi Yang</name>2540    </author>2541    <author>2542      <name>Deheng Zhang</name>2543    </author>2544    <author>2545      <name>Xiaoye Wang</name>2546    </author>2547    <author>2548      <name>Mengjiao Ma</name>2549    </author>2550    <author>2551      <name>Lei Sun</name>2552    </author>2553    <author>2554      <name>Kanzhi Wu</name>2555    </author>2556    <author>2557      <name>Ajad Chhatkuli</name>2558    </author>2559    <author>2560      <name>Luc Van Gool</name>2561    </author>2562    <author>2563      <name>Danda Pani Paudel</name>2564    </author>2565  </entry>2566  <entry>2567    <id>http://arxiv.org/abs/2609.11615v1</id>2568    <title>Distributed Optimization of Modular Production Systems using Model-based Reinforcement Learning with Inverse Models</title>2569    <updated>2026-09-10T14:30:52Z</updated>2570    <link href="https://arxiv.org/abs/2609.11615v1" rel="alternate" type="text/html"/>2571    <link href="https://arxiv.org/pdf/2609.11615v1" rel="related" type="application/pdf" title="pdf"/>2572    <summary>This paper presents a novel approach for data-driven self-learning control of highly flexible, modular manufacturing systems. Specifically, we employ a novel framework for model-based reinforcement learning which introduces approximate inverse process models within the training of reinforcement policies. This approach disentangles the learning of actuation dynamics and the dynamics in state space, resulting in RL-based training solely within the task space. We propose a lightweight feedforward architecture for approximate inverse models and integrate them within the policy network of standard RL algorithms. We apply the approach to a laboratory modular production testbed with heterogeneous production modules. The results underline the efficiency improvements for modular manufacturing units in terms of both performance and training speed, particularly for off-policy algorithms.</summary>2573    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>2574    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>2575    <category term="eess.SY" scheme="http://arxiv.org/schemas/atom"/>2576    <published>2026-09-10T14:30:52Z</published>2577    <arxiv:primary_category term="cs.AI"/>2578    <author>2579      <name>Andreas Schwung</name>2580    </author>2581    <author>2582      <name>Steve Yuwono</name>2583    </author>2584    <author>2585      <name>Sofiene Lassoued</name>2586    </author>2587    <author>2588      <name>Dorothea Schwung</name>2589    </author>2590  </entry>2591  <entry>2592    <id>http://arxiv.org/abs/2609.11607v1</id>2593    <title>Making Alternative Data Work: Context-Augmented LLMs for Financial Forecasting</title>2594    <updated>2026-09-10T14:27:25Z</updated>2595    <link href="https://arxiv.org/abs/2609.11607v1" rel="alternate" type="text/html"/>2596    <link href="https://arxiv.org/pdf/2609.11607v1" rel="related" type="application/pdf" title="pdf"/>2597    <summary>When forecasting a firm's future financial performance, alternative data - data collected from non-traditional sources such as consumer transactions, web traffic, and prediction markets - can provide timely signals about firms' operating activities and broader market conditions. These signals may reveal information that is not captured by traditional public sources and can therefore provide complementary information for forecasting firms' future financial performance. However, firm-level alternative data often have limited historical coverage, are relevant only to specific prediction targets or subsets of firms, and are distributed across numerous heterogeneous channels, making them difficult to incorporate flexibly into conventional forecasting approaches. Meanwhile, large language models (LLMs) can interpret instructions, learn from in-context examples, and generate predictions by combining heterogeneous information without task-specific parameter updates. Motivated by this potential flexibility, we investigate whether an LLM can forecast firm performance by integrating alternative data with other financial information through in-context learning. We propose a two-agent framework that first identifies the firms for which each alternative data channel is likely to be informative and then predicts revenue using firm- and channel-specific context. We evaluate the framework across four commercial alternative data channels. In our experiments, adding alternative data in context alongside other financial information improves the LLM's forecasting relative to either source alone, and these forecasts are more accurate than those of standard forecasting baselines. These findings suggest that LLMs provide a flexible and practical approach to integrating alternative data with heterogeneous financial information.</summary>2598    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>2599    <published>2026-09-10T14:27:25Z</published>2600    <arxiv:comment>13 pages</arxiv:comment>2601    <arxiv:primary_category term="cs.AI"/>2602    <author>2603      <name>Jihoon Kwon</name>2604    </author>2605    <author>2606      <name>Lawrence Liu</name>2607    </author>2608    <author>2609      <name>Daekyung Park</name>2610    </author>2611    <author>2612      <name>Sumin Kim</name>2613    </author>2614    <author>2615      <name>Haverty Jack</name>2616    </author>2617    <author>2618      <name>Hoyoung Lee</name>2619    </author>2620    <author>2621      <name>Katherine Bjorkman</name>2622    </author>2623    <author>2624      <name>Josh McKenney</name>2625    </author>2626    <author>2627      <name>Peter Laurelli</name>2628    </author>2629    <author>2630      <name>Nicole Kagan</name>2631    </author>2632    <author>2633      <name>Zach Golkhou</name>2634    </author>2635    <author>2636      <name>Thorsten Neumann</name>2637    </author>2638    <author>2639      <name>Edward Tong</name>2640    </author>2641    <author>2642      <name>Pete Petersen</name>2643    </author>2644    <author>2645      <name>Yoon Kim</name>2646    </author>2647    <author>2648      <name>Alejandro Lopez-Lira</name>2649    </author>2650    <author>2651      <name>Yongjae Lee</name>2652    </author>2653    <author>2654      <name>Chanyeol Choi</name>2655    </author>2656  </entry>2657  <entry>2658    <id>http://arxiv.org/abs/2609.11606v1</id>2659    <title>Identifiability of Nonnegative Tensor Decompositions via Positive Scattering</title>2660    <updated>2026-09-10T14:26:45Z</updated>2661    <link href="https://arxiv.org/abs/2609.11606v1" rel="alternate" type="text/html"/>2662    <link href="https://arxiv.org/pdf/2609.11606v1" rel="related" type="application/pdf" title="pdf"/>2663    <summary>Identifiability of tensor decompositions is often established through linear-algebraic conditions on the factor families. For nonnegative decompositions, however, positivity provides additional information that is not captured by dimension and independence alone: nonnegative terms cannot cancel, and their supports constrain competing decompositions. We introduce a positive scattering term that quantifies this additional source of identifiability and combine it with the dimension budget underlying the Lovitz--Petrov generalization of Kruskal's theorem. For every subset of components, we obtain two sufficient conditions: a threshold of $2|S|-2$ guarantees minimality and nonnegative rank, while the stronger threshold $2|S|-1$ guarantees uniqueness among nonnegative decompositions of the same length. The key result is a positive splitting inequality for irreducible exchanges of nonnegative rank-one tensors, which combines the dimension constraint with support-induced geometric rigidity. Although the scattering term is defined through an optimization over intermediate factor spaces, we show that its mode costs are exactly $0$, $1$, or $+\infty$, yielding an exact activation characterization in terms of graph connectivity. The resulting criterion can strictly certify sparse nonnegative tensor decompositions beyond the reach of Kruskal and Lovitz--Petrov conditions, including examples for which those conditions fail even after reshaping. In the matrix case, the two criteria reduce respectively to full-rank factorization and two-sided separability.</summary>2664    <category term="stat.ML" scheme="http://arxiv.org/schemas/atom"/>2665    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>2666    <category term="math.CO" scheme="http://arxiv.org/schemas/atom"/>2667    <category term="math.ST" scheme="http://arxiv.org/schemas/atom"/>2668    <published>2026-09-10T14:26:45Z</published>2669    <arxiv:primary_category term="stat.ML"/>2670    <author>2671      <name>Haoming Wang</name>2672    </author>2673    <author>2674      <name>Ming Yuan</name>2675    </author>2676  </entry>2677  <entry>2678    <id>http://arxiv.org/abs/2609.11601v1</id>2679    <title>MMGait: Benchmarking and Unifying Gait Recognition across Heterogeneous Modalities</title>2680    <updated>2026-09-10T14:23:29Z</updated>2681    <link href="https://arxiv.org/abs/2609.11601v1" rel="alternate" type="text/html"/>2682    <link href="https://arxiv.org/pdf/2609.11601v1" rel="related" type="application/pdf" title="pdf"/>2683    <summary>Gait recognition is commonly studied using RGB videos or their derived silhouettes and poses. Yet human walking produces heterogeneous photometric, geometric, and motion cues that cannot be systematically examined with RGB-centered benchmarks. We present MMGait, a large-scale multi-sensor benchmark that brings visible, infrared, depth, LiDAR, and radar observations into sequence-level correspondence. It provides diverse modalities spanning appearance, contours, geometry, motion, and body structure. Under a shared impostor-augmented protocol, we evaluate single-modal recognition, cross-modal recognition via directed retrieval, and multi-modal recognition using task-specific experts. Across settings, modality rankings vary with probe conditions, cross-modal alignment remains difficult, and fusion often provides complementary gains. This analysis exposes a scalability problem: individual modalities, modality pairs, and fusion configurations are typically handled by separately trained experts. We formulate Omni-Modal Gait Recognition, which unifies single-modal, cross-modal, and multi-modal recognition within a shared identity space. OmniGait++ uses modality-specific front ends followed by a shared identity encoder to preserve modality-dependent cues while learning comparable identity descriptors. An anchor-guided fusion module aggregates modality subsets of varying size without frame-level synchronization. A jointly trained checkpoint covers all three recognition settings and accommodates modality subsets of different compositions and cardinalities. Experiments show OmniGait++ remains competitive with task-specific experts in many shared settings and extends to higher-cardinality fusion unavailable to fixed-pair models. The results establish MMGait as a common testbed for heterogeneous gait sensing and demonstrate the feasibility of unified recognition under varying modality availability.</summary>2684    <category term="cs.CV" scheme="http://arxiv.org/schemas/atom"/>2685    <published>2026-09-10T14:23:29Z</published>2686    <arxiv:comment>21 pages, 6 figures</arxiv:comment>2687    <arxiv:primary_category term="cs.CV"/>2688    <author>2689      <name>Saihui Hou</name>2690    </author>2691    <author>2692      <name>Chenye Wang</name>2693    </author>2694    <author>2695      <name>Qingyuan Cai</name>2696    </author>2697    <author>2698      <name>Aoqi Li</name>2699    </author>2700    <author>2701      <name>Yongzhen Huang</name>2702    </author>2703  </entry>2704  <entry>2705    <id>http://arxiv.org/abs/2609.11592v1</id>2706    <title>A distribution-free certification framework for trustworthy crash-severity prediction</title>2707    <updated>2026-09-10T14:20:08Z</updated>2708    <link href="https://arxiv.org/abs/2609.11592v1" rel="alternate" type="text/html"/>2709    <link href="https://arxiv.org/pdf/2609.11592v1" rel="related" type="application/pdf" title="pdf"/>2710    <summary>Crash-severity models inform screening, dispatch and site prioritization, yet are deployed without a finite-sample statement of what one prediction means. Off-the-shelf guarantees fail here, because the features that make crash severity distinctive defeat them: the KABCO outcome is ordinal, the recorded label is a field assessment agreeing with medical severity about half the time, erring in a structured way, and deployment crosses jurisdictions and years calibration never saw. We develop a certification layer that wraps any severity model unmodified, with distribution-free guarantees using this structure: contiguous ordinal sets that read as "B or worse"; per-class validity for any pre-declared partition, with an oracle efficiency characterization; transfer of coverage to unobserved true severity through a declared reporting band, with a worst-case sharpness result; a one-sided certificate under deployment shift; and severity-weighted risk control. The guarantees compose with an attributable slack budget. The same analysis bounds what certification can achieve. A certified set's informativeness is governed by a functional of the true law that no base model can evade and that cannot be lower-bounded distribution-free; given a declared misreporting channel identified from record-linkage data, a nonvacuous lower bound on that floor becomes computable. On 5.2 million Texas records across seven base models spanning four decades, the layer attaches identical validity and certifies, on the vulnerable road users, a model-independent floor on set width that no base model beats, separating it from a remainder that stays bounded but distribution-free unidentifiable. The framework is released as an open-source package with theorem-level tests.</summary>2711    <category term="stat.ML" scheme="http://arxiv.org/schemas/atom"/>2712    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>2713    <published>2026-09-10T14:20:08Z</published>2714    <arxiv:primary_category term="stat.ML"/>2715    <author>2716      <name>Amir Rafe</name>2717    </author>2718    <author>2719      <name>Subasish Das</name>2720    </author>2721  </entry>2722  <entry>2723    <id>http://arxiv.org/abs/2609.11582v1</id>2724    <title>OmniKVQuant: KV Cache Quantization for Omni-LLMs</title>2725    <updated>2026-09-10T14:14:12Z</updated>2726    <link href="https://arxiv.org/abs/2609.11582v1" rel="alternate" type="text/html"/>2727    <link href="https://arxiv.org/pdf/2609.11582v1" rel="related" type="application/pdf" title="pdf"/>2728    <summary>As Omni-modal large language models (Omni-LLMs) take in audio, video and text together, their KV cache memory cost grows. KV cache quantization is the de facto approach in text-only LLMs, but its application to Omni-LLMs remains unexplored. In this paper, we analyze how TurboQuant, a representative rotation-based KV cache quantization method, behaves on multimodal caches and identify two critical issues: temporal key drift and heterogeneous value geometry. To address these, we propose OmniKVQuant, a training-free framework that (i) sets the key quantization range over each short window of the input stream; and (ii) rotates values separately per modality. On Qwen2.5-Omni and Qwen3-Omni, OmniKVQuant enables 2-bit KV caches while substantially preserving performance across seven audio-visual benchmarks. We further provide a fused Triton decode kernel that unpacks the 2-bit cache during attention, so no dense FP16 cache is ever built. Code: https://github.com/kaistmm/OmniKVQuant</summary>2729    <category term="cs.CV" scheme="http://arxiv.org/schemas/atom"/>2730    <published>2026-09-10T14:14:12Z</published>2731    <arxiv:comment>Preprint</arxiv:comment>2732    <arxiv:primary_category term="cs.CV"/>2733    <author>2734      <name>Suho Yoo</name>2735    </author>2736    <author>2737      <name>Hyunjong Ok</name>2738    </author>2739    <author>2740      <name>Jongmin Choi</name>2741    </author>2742    <author>2743      <name>Jihoo Jung</name>2744    </author>2745    <author>2746      <name>Joon Son Chung</name>2747    </author>2748  </entry>2749  <entry>2750    <id>http://arxiv.org/abs/2609.11580v1</id>2751    <title>A Dataset and Model for Imputing Water Surface Elevation on a Large and Extremely Sparse Spatiotemporal Graph</title>2752    <updated>2026-09-10T14:12:02Z</updated>2753    <link href="https://arxiv.org/abs/2609.11580v1" rel="alternate" type="text/html"/>2754    <link href="https://arxiv.org/pdf/2609.11580v1" rel="related" type="application/pdf" title="pdf"/>2755    <summary>Continuous monitoring of water surface elevation across river networks is critical for flood forecasting, water resource management, and understanding the global water cycle. Yet, the scarcity of in situ gauges across much of the globe constrains the development of reliable modeling frameworks. Satellite altimetry has the potential to alleviate this problem but its use is currently hindered by sparse temporal coverage. To this end, we introduce AmazonSWE, a dataset for training and evaluating large-scale spatiotemporal graph imputation methods that integrates processed satellite altimetry measurements from a range of sources, including the recent wide-swath SWOT sensor. The dataset covers over 19K river sections and 10 years (2016-2026) in the Amazon river basin, with in situ gauges held out for evaluation. Besides contributing a novel real-world use case with the potential for societal impact, AmazonSWE introduces significant technical challenges: with fewer than 1% of sections observed per day, the dataset is far sparser than existing imputation benchmarks, and its directed acyclic river topology is both structurally different from and larger than graphs in existing datasets. We show that prior spatiotemporal graph imputation methods are not adapted to this topology, scale and sparsity, and propose a simple bidirectional selective state space model that outperforms them by sampling connected subgraphs and flattening space and time into a single token sequence with topology-aware positional encodings. Compared to the state-of-the-art published method for SWOT-based WSE densification, which integrates statistics with physical modeling, our model reduces RMSE against in situ gauges by 18-39%, while producing predictions for every river section rather than only those with sufficient nearby satellite coverage.</summary>2756    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>2757    <published>2026-09-10T14:12:02Z</published>2758    <arxiv:primary_category term="cs.LG"/>2759    <author>2760      <name>Ruben Cartuyvels</name>2761    </author>2762    <author>2763      <name>Karim Douch</name>2764    </author>2765    <author>2766      <name>Gabriele Bertoli</name>2767    </author>2768    <author>2769      <name>Mounia El Baz</name>2770    </author>2771    <author>2772      <name>Artemis Vrettou</name>2773    </author>2774    <author>2775      <name>Sébastien Lefèvre</name>2776    </author>2777    <author>2778      <name>Diego Fernandez Prieto</name>2779    </author>2780  </entry>2781  <entry>2782    <id>http://arxiv.org/abs/2609.11573v1</id>2783    <title>Learn the Solid, Not the File: Canonical Inputs for Neural Networks on CAD Boundary Representations</title>2784    <updated>2026-09-10T14:08:03Z</updated>2785    <link href="https://arxiv.org/abs/2609.11573v1" rel="alternate" type="text/html"/>2786    <link href="https://arxiv.org/pdf/2609.11573v1" rel="related" type="application/pdf" title="pdf"/>2787    <summary>Boundary representation (B-rep) is the standard format used by modern CAD systems for parametric 3D models. It turns out, the exact same solid can be represented by different B-reps: for example, two engineers using different operations, a geometry kernel rebuilding the file, and an export setting repartitioning faces will lead to different B-reps even though the underlying solid remains the same.2788  We show that existing B-rep encoders are not robust to variation in the B-rep with the same solid on perturbations applied to standard benchmarks, naturally occurring variations inherent to CAD software, and differences in how designers model the same part via a human dataset we created in FreeCAD. The performance of popular B-rep encoders often collapses catastrophically.2789  We propose the canonical region graph, an input representation whose nodes, features and coordinate frame are derived from the solid itself and show theoretical invariance guarantees on repartitioning and rigid motions. It matches the strongest baseline on standard benchmarks, and is stable under every perturbation we test.</summary>2790    <category term="cs.CV" scheme="http://arxiv.org/schemas/atom"/>2791    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>2792    <category term="cs.CG" scheme="http://arxiv.org/schemas/atom"/>2793    <published>2026-09-10T14:08:03Z</published>2794    <arxiv:primary_category term="cs.CV"/>2795    <author>2796      <name>Heinrich Jiang</name>2797    </author>2798    <author>2799      <name>Hager Yasser Mohamed</name>2800    </author>2801    <author>2802      <name>Alexander Hitt</name>2803    </author>2804    <author>2805      <name>Valeriia Lomakina</name>2806    </author>2807    <author>2808      <name>Henning Jiang</name>2809    </author>2810    <author>2811      <name>Jennifer Jang</name>2812    </author>2813  </entry>2814  <entry>2815    <id>http://arxiv.org/abs/2609.11569v1</id>2816    <title>Enabling Knowledge Graph Understanding at Scale with the EXplore Your Graphs ENgine (EXYGEN)</title>2817    <updated>2026-09-10T14:03:32Z</updated>2818    <link href="https://arxiv.org/abs/2609.11569v1" rel="alternate" type="text/html"/>2819    <link href="https://arxiv.org/pdf/2609.11569v1" rel="related" type="application/pdf" title="pdf"/>2820    <summary>We present EXYGEN (EXplore Your Graphs ENgine), a framework for knowledge graph (KG) understanding that enables conversational access to KGs at scale. We address two questions in sequence. First, how effectively can LLMs perform text-to-SPARQL generation given only automatically derived structured metadata and small graph samples, rather than task-specific fine-tuning? We integrate VoID descriptions and ShEx schemas into a retrieval-augmented generation (RAG) pipeline and ablate KG-derived context on the SciQA benchmark. Our best configuration -- combining ShEx schemas, retrieved triples, and example question-query pairs -- reaches an exact match of 0.419 on execution results without any LLM fine-tuning. We further find that lexical metrics such as F1 poorly predict query correctness, and that larger general-purpose LLMs can outperform smaller code-specialized ones once given sufficient context. Second, we ask how to generate the structured metadata that this method relies on from very large KGs, where KG metadata generation becomes computationally intractable. We introduce a predicate-coverage-aware parallel graph sampling strategy that preserves structural diversity while remaining computationally tractable. On OpenCitations Meta and GESIS, it retains high predicate coverage with minimal triple loss and reduces runtime by over 80x; on ORKG, sampling is not just faster but the only tractable path to obtain complete metadata. Together, these results show that structured schema context and lightweight prompting can substantially reduce reliance on fine-tuning for scalable conversational access to KGs, though closing the remaining gap to fully fine-tuned approaches will likely require reducing dependence on curated question-query exemplars -- whether through synthetic generation or an execution-feedback-driven approach -- and validating these findings beyond a single benchmark.</summary>2821    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>2822    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>2823    <published>2026-09-10T14:03:32Z</published>2824    <arxiv:primary_category term="cs.AI"/>2825    <author>2826      <name>Harshdeep Singh</name>2827    </author>2828    <author>2829      <name>Yurui Zhu</name>2830    </author>2831    <author>2832      <name>Giovanni Colavizza</name>2833    </author>2834    <author>2835      <name>Matteo Romanello</name>2836    </author>2837  </entry>2838  <entry>2839    <id>http://arxiv.org/abs/2609.11550v1</id>2840    <title>A Comparative Evaluation of Pre-trained Convolutional Neural Networks for Melanoma Detection</title>2841    <updated>2026-09-10T13:43:49Z</updated>2842    <link href="https://arxiv.org/abs/2609.11550v1" rel="alternate" type="text/html"/>2843    <link href="https://arxiv.org/pdf/2609.11550v1" rel="related" type="application/pdf" title="pdf"/>2844    <summary>Early diagnosis of melanoma is critical for improving patient survival rates. However, accurately distinguishing melanoma from other skin lesions remains a significant clinical challenge due to the high visual similarity among lesion types and variability in image acquisition conditions. Artificial intelligence, particularly machine learning, has emerged as a promising tool to support dermatological diagnosis by automating feature extraction from medical images. Among the available approaches, convolutional neural networks (CNNs) have demonstrated strong performance in image classification tasks, making them well-suited for analyzing both dermatoscopic and histopathological images, given their ability to capture hierarchical visual patterns relevant to lesion characterization. Nevertheless, despite numerous pre-trained CNN architectures having been proposed, selecting the most appropriate one for a given imaging modality remains an open challenge. In this study, we evaluate pre-trained convolutional neural networks (CNNs) for skin lesion classification using dermatoscopic and histopathological image datasets. Experiments were conducted on the HAM10000, ISIC 2018, and CR-AI4SkIN datasets, evaluating the ResNet50, VGG16, VGG19, MobileNet, and InceptionV3 architectures under the same training protocol. The experimental evaluation showed that the models achieved accuracies ranging from 71% (InceptionV3 on ISIC 2018) to 84% (ResNet50 on HAM10000) on dermatoscopic images. For histopathological images, accuracies ranged from 72% (VGG19) to 83% (ResNet50) on the CR-AI4SkIN dataset. The results demonstrate that model performance differs between dermatoscopic and histopathological image modalities, showing that architectures exhibiting similar performance on dermatoscopic images exhibit different performance on histopathological data.</summary>2845    <category term="cs.CV" scheme="http://arxiv.org/schemas/atom"/>2846    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>2847    <published>2026-09-10T13:43:49Z</published>2848    <arxiv:primary_category term="cs.CV"/>2849    <author>2850      <name>Wagner Moreno Schmitz</name>2851    </author>2852    <author>2853      <name>Marco Antonio de Castro Barbosa</name>2854    </author>2855    <author>2856      <name>Thiago Magalhães Amaral</name>2857    </author>2858    <author>2859      <name>Dalcimar Casanova</name>2860    </author>2861    <author>2862      <name>Jefferson Tales Oliva</name>2863    </author>2864  </entry>2865  <entry>2866    <id>http://arxiv.org/abs/2609.11548v1</id>2867    <title>World in World: Explore the World with World Models</title>2868    <updated>2026-09-10T13:42:26Z</updated>2869    <link href="https://arxiv.org/abs/2609.11548v1" rel="alternate" type="text/html"/>2870    <link href="https://arxiv.org/pdf/2609.11548v1" rel="related" type="application/pdf" title="pdf"/>2871    <summary>Autoregressive video world models enable interactive, long-horizon exploration, but flexible control remains challenging. Exploring a source video from new viewpoints requires the generated rollout to remain synchronised with the recorded event, place observed content in the requested view, plausibly complete newly exposed regions, and recover previously generated appearance on revisits. Existing methods typically address these requirements through task-specific modules or additional training. We present World in World, a training-free inference-time interface that converts heterogeneous control evidence into camera- and time-labelled clean visual states, which are read through the native self attention of a frozen causal video model. The evidence comprises source-video observations, target-view scene projections, geometry renderings that guide completion of newly exposed subject regions, and retrieved generated states beyond the rolling cache. Each evidence source carries token-level support and its own availability schedule. A correspondence router combines persistent point identities with geometry to establish token correspondences, guiding supported queries towards matching source-video tokens. Evidence-wise attention CFG (EWA) then independently regulates each auxiliary channel's additional contribution using attention responses from the same denoising forward pass. The shared interface supports camera-controlled rerendering, long-horizon revisiting, and human-motion transfer with the same frozen backbone. We evaluate World in World on camera-controlled video rerendering under diverse viewpoint changes, assessing perceptual quality, temporal consistency, and camera-following accuracy.</summary>2872    <category term="cs.CV" scheme="http://arxiv.org/schemas/atom"/>2873    <published>2026-09-10T13:42:26Z</published>2874    <arxiv:comment>Project Page: https://chenxi-song.github.io/worldinworld</arxiv:comment>2875    <arxiv:primary_category term="cs.CV"/>2876    <author>2877      <name>Chenxi Song</name>2878    </author>2879    <author>2880      <name>Yanming Yang</name>2881    </author>2882    <author>2883      <name>Chi Zhang</name>2884    </author>2885  </entry>2886  <entry>2887    <id>http://arxiv.org/abs/2609.11545v1</id>2888    <title>Complex-Text Robustness Evaluation and Failure Diagnosis for Low-Resource Multilingual Text-to-Speech</title>2889    <updated>2026-09-10T13:41:22Z</updated>2890    <link href="https://arxiv.org/abs/2609.11545v1" rel="alternate" type="text/html"/>2891    <link href="https://arxiv.org/pdf/2609.11545v1" rel="related" type="application/pdf" title="pdf"/>2892    <summary>Low-resource multilingual text-to-speech (TTS) systems have expanded language coverage, but their robustness under complex text inputs remains insufficiently diagnosed. Existing evaluations mainly focus on naturalness, speaker similarity, and content consistency using regular test sentences, while providing limited insight into how multilingual TTS systems fail when handling challenging inputs such as numbers, dates, named entities, long sentences, code-switched expressions, and punctuation-related structures. This paper proposes a complex-text robustness diagnosis framework for low-resource multilingual TTS. We evaluate robustness from three dimensions: content consistency, language consistency, and generation stability. A multilingual robustness testing scheme is designed for Thai, Vietnamese, Swahili, and Indonesian, covering ordinary sentences and multiple types of complex text inputs. We further introduce automatic diagnostic metrics, including character error rate, language identification accuracy, and duration abnormal rate. To support input-level risk analysis before speech generation, we propose a lightweight Text Risk Score (TRS), which estimates synthesis risk from interpretable text features without manual annotation or model training. Experiments on three representative multilingual TTS systems, including OmniVoice, VoxCPM2, and MMS-TTS, show that complex text inputs expose systematic failure patterns that are not fully reflected by ordinary short-sentence evaluation. Different systems exhibit distinct vulnerabilities in number normalization, named entity handling, long-text generation, and code-switched input processing. Furthermore, TRS shows a positive correlation with content errors and duration abnormalities, demonstrating its usefulness as a low-cost pre-synthesis indicator for complex-text risk diagnosis in low-resource multilingual TTS.</summary>2893    <category term="cs.CL" scheme="http://arxiv.org/schemas/atom"/>2894    <category term="cs.SD" scheme="http://arxiv.org/schemas/atom"/>2895    <published>2026-09-10T13:41:22Z</published>2896    <arxiv:comment>NCMMSC 2026 accepted</arxiv:comment>2897    <arxiv:primary_category term="cs.CL"/>2898    <author>2899      <name>Tianlun Zuo</name>2900    </author>2901    <author>2902      <name>Ziyu Zhang</name>2903    </author>2904    <author>2905      <name>Tingzhi Mao</name>2906    </author>2907    <author>2908      <name>Zhonghua Fu</name>2909    </author>2910    <author>2911      <name>Lei Xie</name>2912    </author>2913  </entry>2914  <entry>2915    <id>http://arxiv.org/abs/2609.11542v1</id>2916    <title>Characterizing Job Power Elasticity for Power-Flexible AI Training</title>2917    <updated>2026-09-10T13:40:13Z</updated>2918    <link href="https://arxiv.org/abs/2609.11542v1" rel="alternate" type="text/html"/>2919    <link href="https://arxiv.org/pdf/2609.11542v1" rel="related" type="application/pdf" title="pdf"/>2920    <summary>Large language model (LLM) training is among the fastest-growing sources of electricity demand in modern data centers, and power availability is a primary bottleneck to continued AI infrastructure growth. Making the power consumption of these workloads flexible could unlock additional power for AI growth, limit increases in electricity prices, and improve the utilization of existing grid infrastructure. However, to realize this flexibility, we must first understand how the performance of training workloads changes when GPU power is reduced.2921  This paper presents the first systematic characterization of \emph{job power elasticity} (the sensitivity of throughput to power reductions) in LLM training. To quantify elasticity, we introduce the \emph{Power Flexibility Index (PFI)}, a normalized metric that quantifies the performance cost of power reductions and provides a control primitive for SLA-aware power flexibility.2922  We collect data from 131 LLM training runs on H200 (plus 24 H200 validation runs and 34 matched H100 runs), including both dense and mixture-of-experts models, pretraining and fine-tuning tasks, and up to 32 GPUs. We find that LLM training jobs exhibit substantial but variable power elasticity, and we identify telemetry signals that predict PFI at runtime. Finally, we demonstrate that PFI-aware power allocation maximizes total tokens/second throughput under power constraints. Under a 30\% power reduction, PFI-aware power allocation recovers ~1.5k tokens/s per job, 63\% of the performance gap between an equal-weight allocation and an oracle with perfect information. Our results establish power elasticity as a measurable property of training jobs and provide a foundation for power-aware, grid-responsive AI infrastructure.</summary>2923    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>2924    <published>2026-09-10T13:40:13Z</published>2925    <arxiv:primary_category term="cs.AI"/>2926    <author>2927      <name>Philip Colangelo</name>2928    </author>2929    <author>2930      <name>Charles Dawson</name>2931    </author>2932    <author>2933      <name>Shayan Sengupta</name>2934    </author>2935    <author>2936      <name>Ayse Coskun</name>2937    </author>2938    <author>2939      <name>Varun Sivaram</name>2940    </author>2941  </entry>2942  <entry>2943    <id>http://arxiv.org/abs/2609.11538v1</id>2944    <title>Particle GFlowNets: Rethinking Generative Marginalization Models</title>2945    <updated>2026-09-10T13:37:28Z</updated>2946    <link href="https://arxiv.org/abs/2609.11538v1" rel="alternate" type="text/html"/>2947    <link href="https://arxiv.org/pdf/2609.11538v1" rel="related" type="application/pdf" title="pdf"/>2948    <summary>Generative Marginalization Models (MaMs) have been recently introduced as efficient neural sampling models for any-order autoregressive modelling of discrete distributions. By learning both the marginal and conditional probabilities of a persistent-block Gibbs sampler, MaMs enable fast posterior evaluation with a single neural network forward pass. While prior work has considered MaMs to be distinct from Generative Flow Networks (GFlowNets), a well-established paradigm for inference in discrete stochastic models, we show that they are equivalent. Then, we also extend MaMs' sampling strategy to non-autoregressive generative processes. In particular, we describe an automatic criterion for full-state rejuvenation of the Gibbs sampler, derived from the Gelman-Rubin statistic, which plays a key role in speeding up learning convergence. Our experiments show that our method, called Particle GFlowNets, markedly accelerates training in large combinatorial spaces.</summary>2949    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>2950    <published>2026-09-10T13:37:28Z</published>2951    <arxiv:comment>Accepted at UAI 2026</arxiv:comment>2952    <arxiv:primary_category term="cs.LG"/>2953    <author>2954      <name>Tiago da Silva</name>2955    </author>2956    <author>2957      <name>Diego Mesquita</name>2958    </author>2959    <author>2960      <name>Salem Lahlou</name>2961    </author>2962  </entry>2963  <entry>2964    <id>http://arxiv.org/abs/2609.11532v1</id>2965    <title>Prompt Revision as a Source of Cultural Bias in Text-to-Image Systems</title>2966    <updated>2026-09-10T13:32:40Z</updated>2967    <link href="https://arxiv.org/abs/2609.11532v1" rel="alternate" type="text/html"/>2968    <link href="https://arxiv.org/pdf/2609.11532v1" rel="related" type="application/pdf" title="pdf"/>2969    <summary>Commercial text-to-image systems silently revise user prompts before generating images, a step users typically cannot disable or even see. Yet, existing audits of cultural bias examine only the final images and treat generation as a single pipeline, so they cannot tell where the bias originates. We introduce WORLDVIEW, a multilingual benchmark of 8,960 prompts across 15 languages and 31 language-context pairings. Using it, we audit the revision layer in three systems (DALL-E-3, Imagen-4, GPT-Image-1.5) through a three-step analysis of how heavily it marks each cultural context, whether it flattens that context into a narrow vocabulary, and whether that vocabulary is stereotypical. Relative to a no-context English baseline, the US is the least-marked context, while non-Western and non-Anglophone contexts are marked far more heavily, flattened into narrow vocabularies applied across topically diverse prompts, and reduced to recognizable cultural stereotypes. Comparing images from original versus revised prompts on models without a revision layer, we identify the layer itself as a previously undocumented, causal source of this stereotyping. To locate cultural bias, and fix it, we must audit the system as deployed, not the model alone.</summary>2970    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>2971    <published>2026-09-10T13:32:40Z</published>2972    <arxiv:comment>Accepted to EMNLP 2026</arxiv:comment>2973    <arxiv:primary_category term="cs.AI"/>2974    <author>2975      <name>Aleksandra Urman</name>2976    </author>2977    <author>2978      <name>Elsa Lichtenegger</name>2979    </author>2980    <author>2981      <name>Salima Jaoua</name>2982    </author>2983    <author>2984      <name>Azza Bouleimen</name>2985    </author>2986    <author>2987      <name>Robin Forsberg</name>2988    </author>2989    <author>2990      <name>Corinna Hertweck</name>2991    </author>2992    <author>2993      <name>Stefania Ionescu</name>2994    </author>2995    <author>2996      <name>Nicolò Pagan</name>2997    </author>2998    <author>2999      <name>Ancsa Hannak</name>3000    </author>3001    <author>3002      <name>Joachim Baumann</name>3003    </author>3004  </entry>3005  <entry>3006    <id>http://arxiv.org/abs/2609.11527v1</id>3007    <title>Lightweight LiDAR-Based Cone Detection Framework Using Random Forest for Formula Student Driverless</title>3008    <updated>2026-09-10T13:28:53Z</updated>3009    <link href="https://arxiv.org/abs/2609.11527v1" rel="alternate" type="text/html"/>3010    <link href="https://arxiv.org/pdf/2609.11527v1" rel="related" type="application/pdf" title="pdf"/>3011    <summary>Reliable, low-latency perception is crucial for Formula Student Driverless vehicles, yet many existing pipelines rely on deep learning and multi-sensor fusion, often requiring GPU acceleration. This paper presents a lightweight LiDAR-only perception pipeline tailored for CPU execution, combining ground removal, IMU-based motion compensation, DBSCAN clustering, and geometric feature-based Random Forest classification. Feature importance analysis reduced the model input from 12 to 7 features while preserving performance. Evaluated on 2,371 labeled clusters collected from real FSD events, the pipeline achieves an F1-score of 98.33% and an end-to-end runtime of 3.13 ms on CPU-only hardware. The released dataset, labeling tool, and trained models provide a practical and reproducible baseline for other resource-constrained autonomous racing teams.</summary>3012    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>3013    <published>2026-09-10T13:28:53Z</published>3014    <arxiv:comment>9 pages, 2 figures, 3 tables. Accepted at the 5th International Conference on Cognitive Mobility (CogMob 2026)</arxiv:comment>3015    <arxiv:primary_category term="cs.AI"/>3016    <author>3017      <name>Márk Mező-Kerekes</name>3018    </author>3019    <author>3020      <name>Péter Praksz</name>3021    </author>3022    <author>3023      <name>Chang Liu</name>3024    </author>3025  </entry>3026  <entry>3027    <id>http://arxiv.org/abs/2609.11524v1</id>3028    <title>Risk-Averse Decision Making with Multi-Level Reliability Guarantees</title>3029    <updated>2026-09-10T13:26:37Z</updated>3030    <link href="https://arxiv.org/abs/2609.11524v1" rel="alternate" type="text/html"/>3031    <link href="https://arxiv.org/pdf/2609.11524v1" rel="related" type="application/pdf" title="pdf"/>3032    <summary>Many applications in engineering, including wireless broadcasting, require designs that provide performance certificates at different target outage levels. This paper studies the problem of maximizing the weighted average of such certificates in the presence of uncertainty about the true system state. The problem is shown to be equivalent to an optimization over nested prediction sets, connecting to the literature on conformal prediction and extending prior art on single-level risk-averse decision making. Furthermore, we derive a dual formulation that decouples optimization across input values. Numerical experiments on a diversity-based wireless transmission system illustrate the cost of enforcing multi-level certificates with a single shared policy and trace the Pareto trade-off between multiple reliability levels.</summary>3033    <category term="stat.ML" scheme="http://arxiv.org/schemas/atom"/>3034    <category term="cs.IT" scheme="http://arxiv.org/schemas/atom"/>3035    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>3036    <published>2026-09-10T13:26:37Z</published>3037    <arxiv:primary_category term="stat.ML"/>3038    <author>3039      <name>Amirmohammad Farzaneh</name>3040    </author>3041    <author>3042      <name>Osvaldo Simeone</name>3043    </author>3044  </entry>3045  <entry>3046    <id>http://arxiv.org/abs/2609.11521v1</id>3047    <title>Generalized Score Matching for Parameter Estimation on Convex Domains</title>3048    <updated>2026-09-10T13:23:41Z</updated>3049    <link href="https://arxiv.org/abs/2609.11521v1" rel="alternate" type="text/html"/>3050    <link href="https://arxiv.org/pdf/2609.11521v1" rel="related" type="application/pdf" title="pdf"/>3051    <summary>Maximum likelihood (ML) estimation is a principled and statistically efficient approach for learning probabilistic models. However, for unnormalized models, ML estimation requires evaluating the partition function and differentiating through it, which may not always be tractable. Score matching provides a practically viable alternative that circumvents this obstacle by fitting the score in a way that eliminates dependence on the normalizing constant. We derive the generalized score matching objective on a convex subset of $\mathbb{R}^{d}$ constructively starting from Minimum Probability Flow (MPF) learning, and show how classical score matching as well as domain-adapted variants for non-negative data arise naturally within the proposed framework. We show that the resulting objective is a {\it proper local scoring rule} of second-order, which provides the theoretical guarantee that the true density is recovered when the objective is minimized. Furthermore, for a model belonging to the exponential family, we establish convexity of the objective together with consistency of the finite-sample estimator under standard regularity conditions. Our derivation sheds new light on the scope and applicability of generalized score matching in various problem settings. We compare generalized score matching-based estimators on constrained domains, where the partition function is analytically intractable. We provide experimental results on parameter estimation for model densities belonging to the exponential family defined over convex subsets of $\mathbb{R}^{d}$, and a generative modeling use-case to demonstrate broader applicability of the proposed generalized score matching framework.</summary>3052    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>3053    <category term="stat.ML" scheme="http://arxiv.org/schemas/atom"/>3054    <published>2026-09-10T13:23:41Z</published>3055    <arxiv:primary_category term="cs.LG"/>3056    <author>3057      <name>Nishanth Shetty</name>3058    </author>3059    <author>3060      <name>Saisuchith Mahajan</name>3061    </author>3062    <author>3063      <name>Chandra Sekhar Seelamantula</name>3064    </author>3065  </entry>3066  <entry>3067    <id>http://arxiv.org/abs/2609.11519v1</id>3068    <title>Learning Interaction between Image and Layout Priors for Joint Image-Layout Generation in Design Templates</title>3069    <updated>2026-09-10T13:22:23Z</updated>3070    <link href="https://arxiv.org/abs/2609.11519v1" rel="alternate" type="text/html"/>3071    <link href="https://arxiv.org/pdf/2609.11519v1" rel="related" type="application/pdf" title="pdf"/>3072    <summary>In this paper, we address the problem of graphic design template creation, which generates a background image and a layout of foreground elements over the background to form a harmonious composition from an input text.3073  Prior work on graphic design generation mostly adopts a sequential paradigm, where design elements are generated sequentially. We argue that such a sequential scheme falls short of faithfully capturing the dependency between the background and layout (and thus the joint image-layout distribution), which limits the quality of generated design templates.3074  To overcome this limitation, we propose a model, InterIL, which jointly generates the two modalities, background image and layout, in a single generative process. The novel design of our joint model connects the backbones of pretrained image and layout diffusion models with a learnable communication module to explicitly model bidirectional image-layout interaction. During training, the image and layout backbones are frozen to maintain and leverage the vast pretrained single-modality prior knowledge, while only the communication module is updated, so that the model can focus on learning image-layout interaction and thereby better capture the joint image-layout distribution for improved composition harmony.3075  Our model has no design-specific inductive bias, which allows it to better preserve the original characteristics of realistic designs. We further introduce a test-time guidance strategy to enable users to impose their specific preferences on generated results.3076  Our experiments show that, compared with prior approaches, our model can generate significantly better results in terms of image, layout and image-layout harmonization, producing outputs closer to real samples. We also demonstrate the flexibility of our model in enforcing user preferences at inference without retraining.</summary>3077    <category term="cs.CV" scheme="http://arxiv.org/schemas/atom"/>3078    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>3079    <category term="cs.GR" scheme="http://arxiv.org/schemas/atom"/>3080    <published>2026-09-10T13:22:23Z</published>3081    <arxiv:comment>Main paper with supplementary material. Submitted to IEEE Transactions on Visualization and Computer Graphics</arxiv:comment>3082    <arxiv:primary_category term="cs.CV"/>3083    <author>3084      <name>Shirong Yang</name>3085    </author>3086    <author>3087      <name>Bo Yang</name>3088    </author>3089    <author>3090      <name>Ying Cao</name>3091    </author>3092  </entry>3093  <entry>3094    <id>http://arxiv.org/abs/2609.11518v1</id>3095    <title>Breaking the Central Bias: Spatially Partitioned Experts for Coordinate-Based Neuroevolution</title>3096    <updated>2026-09-10T13:22:15Z</updated>3097    <link href="https://arxiv.org/abs/2609.11518v1" rel="alternate" type="text/html"/>3098    <link href="https://arxiv.org/pdf/2609.11518v1" rel="related" type="application/pdf" title="pdf"/>3099    <summary>Evolvable-Substrate HyperNEAT (ES-HyperNEAT), a bio-inspired indirect encoding that determines neuron placement and connection weights from spatial coordinates, exhibits a failure mode on MNIST as a diagnostic benchmark. Because input pixels map to a coordinate space centered at the origin, evolved networks converge on a small central cluster of input pixels, a spatial-concentration bias; prior work observed only 21% mean accuracy in this regime. Is this bias an optimization artifact or an architectural ceiling? Inspired by Mixture-of-Experts (MoE) principles, we partition the input into non-overlapping spatial segments, each assigned to a separately evolved specialist network. With 13 such experts, this design reaches 43% mean accuracy, a 106% relative improvement over the baseline. The architectural gain does not depend on data-driven aggregation: equal-weighted averaging, which uses no validation data, already yields a 70% improvement; the gain comes from partitioning, not the weighting. Receptive-field analysis shows the mechanism: partitioning forces evolution to discover features across the entire image, expanding active pixel coverage from 4% to 79%. Absolute accuracy stays below gradient-trained baselines, but the relative gain points to central bias, not the evolutionary search. Two tools are designed to generalize beyond MNIST: a receptive-field diagnostic for silent input-coverage collapse, and a spatial-partitioning remedy that restores coverage.</summary>3100    <category term="cs.NE" scheme="http://arxiv.org/schemas/atom"/>3101    <category term="cs.CV" scheme="http://arxiv.org/schemas/atom"/>3102    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>3103    <published>2026-09-10T13:22:15Z</published>3104    <arxiv:comment>15 pages, 4 figures, 1 table. Author's accepted manuscript, accepted at the BIOMAP workshop (BIO-inspired Methods for Pattern Recognition) of ICPR 2026, Lyon, France</arxiv:comment>3105    <arxiv:primary_category term="cs.NE"/>3106    <author>3107      <name>Romain Claret</name>3108    </author>3109    <author>3110      <name>Arthur Gygax</name>3111    </author>3112    <author>3113      <name>Michael O'Neill</name>3114    </author>3115    <author>3116      <name>Paul Cotofrei</name>3117    </author>3118    <author>3119      <name>Michael Palma Mendes</name>3120    </author>3121    <author>3122      <name>Pascal Felber</name>3123    </author>3124  </entry>3125  <entry>3126    <id>http://arxiv.org/abs/2609.11516v1</id>3127    <title>LoopVAE: Recurrent Depth Across Scales for Visual Tokenization</title>3128    <updated>2026-09-10T13:20:03Z</updated>3129    <link href="https://arxiv.org/abs/2609.11516v1" rel="alternate" type="text/html"/>3130    <link href="https://arxiv.org/pdf/2609.11516v1" rel="related" type="application/pdf" title="pdf"/>3131    <summary>Hierarchical visual tokenizers typically allocate different processing blocks to different spatial scales. We ask how much of this computation can use the same parameters. LoopVAE reuses a scale- and loop-conditioned core within and across scales, while keeping resolution-changing transitions independent. A four-block core executes 28 block applications per encoder or decoder. On ImageNet-256, the 29M-parameter convolutional model reaches 0.28 rFID and 32.54 dB PSNR under an approximately 30-epoch two-stage training budget, using approximately 65% fewer parameters than the 84M reference VAEs. A non-adversarial Transformer ablation with the same execution graph finds competitive PSNR and SSIM under global sharing, although unshared blocks improve LPIPS. Targeted loop interventions show that completing the trained recurrence improves reconstruction and that even small feature updates can have substantial downstream effects. Truncation also exposes output-range errors, distinguishing useful recurrent computation from reliable early exit. Runtime profiling reveals the execution tradeoff: fewer stored weights require more arithmetic and longer runtime in the tested configurations. With convolutional and Transformer operators and single- or multi-resolution latent interfaces, LoopVAE establishes recurrent depth across scales as a parameter-sharing design axis for visual tokenization.</summary>3132    <category term="cs.CV" scheme="http://arxiv.org/schemas/atom"/>3133    <published>2026-09-10T13:20:03Z</published>3134    <arxiv:comment>19 pages, 6 figures, 7 tables</arxiv:comment>3135    <arxiv:primary_category term="cs.CV"/>3136    <author>3137      <name>Zhiying Lu</name>3138    </author>3139  </entry>3140  <entry>3141    <id>http://arxiv.org/abs/2609.11514v1</id>3142    <title>Prototype Matters: Modality-unified Prototype Self-distillation for Unsupervised Visible-infrared Person Re-identification</title>3143    <updated>2026-09-10T13:19:00Z</updated>3144    <link href="https://arxiv.org/abs/2609.11514v1" rel="alternate" type="text/html"/>3145    <link href="https://arxiv.org/pdf/2609.11514v1" rel="related" type="application/pdf" title="pdf"/>3146    <summary>Estimating reliable cross-modality association is crucial to unsupervised visible-infrared person re-ID. While optimal transport is shown to be a practical solution for cross-modality association, it suffers from the rigidness of hard label assignment without considering the impact of cluster noise. Moreover, enforcing only cross-modality contrast is also suboptimal, as it fails to jointly optimize the similarity relation within and across modality. In this paper, we propose a novel framework for cross-modality learning by well exploitation of prototypes: First, instead of contrasting with cross-modality prototypes, we show that modality-unified prototypical contrast facilitates better modality invariance by jointly and simultaneously optimizing similarity relation within and across-modality. Taking self-prototype as a steady teacher, we further refine the instance-prototype online relation through prototype-guided self-distillation. The two components are optimized in a unified framework, leading to a simple yet effective model. On standard VI-ReID benchmarks, we perform extensive comparison and analysis, validating the effectiveness of our proposed method. Code is available at: https://github.com/Terminator8758/PoSeD.</summary>3147    <category term="cs.CV" scheme="http://arxiv.org/schemas/atom"/>3148    <published>2026-09-10T13:19:00Z</published>3149    <arxiv:comment>ACM Multimedia 2026</arxiv:comment>3150    <arxiv:primary_category term="cs.CV"/>3151    <author>3152      <name>Menglin Wang</name>3153    </author>3154    <author>3155      <name>Xiaojin Gong</name>3156    </author>3157  </entry>3158  <entry>3159    <id>http://arxiv.org/abs/2609.11509v1</id>3160    <title>Extending SMT Solving with Non-Ground Clause Learning</title>3161    <updated>2026-09-10T13:17:14Z</updated>3162    <link href="https://arxiv.org/abs/2609.11509v1" rel="alternate" type="text/html"/>3163    <link href="https://arxiv.org/pdf/2609.11509v1" rel="related" type="application/pdf" title="pdf"/>3164    <summary>Quantifier instantiation is currently the main approach to non-ground SMT solving: solvers generate ground instances and solve the resulting ground SMT problems with CDCL(T)-style reasoning. When a conflict is found, conflict analysis learns only a ground clause, even though the conflict comes from instances of non-ground clauses. Yet non-ground reasoning can give exponentially shorter proofs than purely ground reasoning. We propose a calculus that consists of ground instantiations, CDCL(T)-style rules, and non-ground conflict analysis. The solver reasons on ground instances, but the resolution steps of conflict analysis are performed on their original non-ground clauses. This produces learned clauses that are typically more general than the ground conflict. With a suitable strategy, the learned clauses are even non-redundant. We also show how chronological backtracking can be included in SMT solving. Our calculus gives a common setting for CDCL(T)-style SMT solving, a range of instantiation-based procedures, and non-ground clause learning, and we prove that it simulates CDCL, SCL(FOL), SCL(T), and even Resolution.</summary>3165    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>3166    <category term="cs.LO" scheme="http://arxiv.org/schemas/atom"/>3167    <published>2026-09-10T13:17:14Z</published>3168    <arxiv:comment>Extended version of LPAR 2026 paper</arxiv:comment>3169    <arxiv:primary_category term="cs.AI"/>3170    <author>3171      <name>Yasmine Briefs</name>3172    </author>3173    <author>3174      <name>Christoph Weidenbach</name>3175    </author>3176  </entry>3177  <entry>3178    <id>http://arxiv.org/abs/2609.11507v1</id>3179    <title>Harnessing Intrinsic Subject-Aware Attention for Controllable Multi-Subject Video Generation</title>3180    <updated>2026-09-10T13:15:03Z</updated>3181    <link href="https://arxiv.org/abs/2609.11507v1" rel="alternate" type="text/html"/>3182    <link href="https://arxiv.org/pdf/2609.11507v1" rel="related" type="application/pdf" title="pdf"/>3183    <summary>Multi-subject video generation faces two key challenges: uncontrollable fidelity strength and potential semantic drift. We address these by analyzing the internal mechanisms of Diffusion Transformers (DiTs). We found that certain attention blocks naturally form an Intrinsic Spatial Grounding Map (ISGM) that precisely locates reference subjects. Building on this insight, we propose Dual-phase Intrinsic Attention Leveraging (DIAL), a framework that uses these internal signals for both training and inference. In low-noise stages, we use ISGM to guide the attention mechanism, allowing precise control over fidelity strength during inference without retraining. In high-noise stages, we use these same maps to automatically build preference pairs at no additional cost for Reinforcement Learning (RL). This RL procedure effectively anchors the model's attention to reference subjects and mitigates semantic drift. Extensive experiments show that DIAL significantly outperforms baseline models on the OpenS2V-Eval benchmark, consistently improving identity consistency and enabling controllable fidelity strength.</summary>3184    <category term="cs.CV" scheme="http://arxiv.org/schemas/atom"/>3185    <published>2026-09-10T13:15:03Z</published>3186    <arxiv:comment>23 pages, 11 figures, 4 tables</arxiv:comment>3187    <arxiv:primary_category term="cs.CV"/>3188    <author>3189      <name>Niange Yu</name>3190    </author>3191    <author>3192      <name>Ye Tian</name>3193    </author>3194    <author>3195      <name>Biaolong Chen</name>3196    </author>3197    <author>3198      <name>Miao Lu</name>3199    </author>3200    <author>3201      <name>Aixi Zhang</name>3202    </author>3203    <author>3204      <name>Hao Jiang</name>3205    </author>3206    <author>3207      <name>Yunhai Tong</name>3208    </author>3209    <author>3210      <name>Pipei Huang</name>3211    </author>3212  </entry>3213  <entry>3214    <id>http://arxiv.org/abs/2609.11506v1</id>3215    <title>UBone3D: Physics-Rectified Conditional Flow Matching for Anatomical 3D Shape Completion from Ultrasound</title>3216    <updated>2026-09-10T13:13:29Z</updated>3217    <link href="https://arxiv.org/abs/2609.11506v1" rel="alternate" type="text/html"/>3218    <link href="https://arxiv.org/pdf/2609.11506v1" rel="related" type="application/pdf" title="pdf"/>3219    <summary>Three-dimensional ultrasound (US) is a safe, radiation-free complementary modality to CT and X-rays for longitudinal monitoring, yet its segmentation-derived partial point clouds are extremely artifact-laden. Consequently, it is challenging to recover a clean and complete anatomical structure from such US point clouds. In this paper, we present UBone3D, a novel framework based on physics-rectified conditional flow matching (CFM) that performs point cloud completion directly from partial US observations. UBone3D models deterministic physics artifacts (e.g., surface thickening, streaking, dropouts) via a simulated physics proxy, and introduces test-time physics rectification to steer the shape completion. At inference, the completion is jointly steered by two decoupled forces: (1) anatomical plausibility enforced by a CT-trained generative shape prior, BoneFM, and (2) physics consistency enforced by USimNet in the ultrasound formation space. Extensive experiments on simulated and in-vivo data demonstrate significant improvements in reconstruction accuracy and anatomical fidelity over existing baselines.</summary>3220    <category term="cs.CV" scheme="http://arxiv.org/schemas/atom"/>3221    <category term="eess.IV" scheme="http://arxiv.org/schemas/atom"/>3222    <published>2026-09-10T13:13:29Z</published>3223    <arxiv:comment>Accepted to ECCV 2026. Camera-ready Author Version</arxiv:comment>3224    <arxiv:primary_category term="cs.CV"/>3225    <author>3226      <name>Weiying Chen</name>3227    </author>3228    <author>3229      <name>Yuchong Gao</name>3230    </author>3231    <author>3232      <name>Siyuan Li</name>3233    </author>3234    <author>3235      <name>Marek Reformat</name>3236    </author>3237    <author>3238      <name>Rui Zheng</name>3239    </author>3240    <author>3241      <name>Edmond Lou</name>3242    </author>3243  </entry>3244  <entry>3245    <id>http://arxiv.org/abs/2609.11505v1</id>3246    <title>Structural priors for data-efficient language learning</title>3247    <updated>2026-09-10T13:12:48Z</updated>3248    <link href="https://arxiv.org/abs/2609.11505v1" rel="alternate" type="text/html"/>3249    <link href="https://arxiv.org/pdf/2609.11505v1" rel="related" type="application/pdf" title="pdf"/>3250    <summary>Efficient language learning requires methods to reduce the reliance on large data and computational resources. We investigate structural transfer: First training models on non-language data to induce useful priors for natural language. This approach is a form of weight initialization for multilingual language modeling. We evaluate transfer via next-token-prediction loss, weight shifts in the model, and downstream linguistic benchmarks. Several symbolic data types - notably music, probabilistic grammars, and cellular automata - yield lower language-modeling loss than random initialization. These gains coincide with smaller weight shifts during subsequent language training, suggesting that structural transfer positions models in a more favorable region of the parameter space. However, a lower loss does not translate consistently into better downstream linguistic performance, and transfer from non-language data is less efficient than additional language data. We conclude that non-language data can serve as a partial substitute for language data for the training objective of next-token prediction but does not reliably support broader linguistic generalization.</summary>3251    <category term="cs.CL" scheme="http://arxiv.org/schemas/atom"/>3252    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>3253    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>3254    <published>2026-09-10T13:12:48Z</published>3255    <arxiv:comment>EMNLP 2026, BabyLM Challenge; 18 pages, 11 figures</arxiv:comment>3256    <arxiv:primary_category term="cs.CL"/>3257    <author>3258      <name>Yana Veitsman</name>3259    </author>3260    <author>3261      <name>Jonas Mayer Martins</name>3262    </author>3263    <author>3264      <name>Jonathan Lautenschlager</name>3265    </author>3266    <author>3267      <name>Lisa Beinborn</name>3268    </author>3269  </entry>3270  <entry>3271    <id>http://arxiv.org/abs/2609.11504v1</id>3272    <title>DeFiFlowBench: Benchmarking and Improving Safe Executability in Natural-Language DeFi Workflow Synthesis</title>3273    <updated>2026-09-10T13:10:32Z</updated>3274    <link href="https://arxiv.org/abs/2609.11504v1" rel="alternate" type="text/html"/>3275    <link href="https://arxiv.org/pdf/2609.11504v1" rel="related" type="application/pdf" title="pdf"/>3276    <summary>A structurally valid DeFi workflow can still authorize a costly trade. We introduce DeFiFlowBench, a benchmark of 207 team-authored prompts for natural-language DeFi workflow synthesis. It measures graph coverage, configuration completeness, and declared safety predicates, then tests supported trade configurations on a local EVM. Direct, constrained, and few-shot prompting produce 14-19 unsafe held-out executions per configuration under a fixed 5% price-impact cap. A slippage bound derived from a quote does not prevent the price impact of the order itself. We propose Koan-Safe, which combines a prompt-only intent parser, a replaceable generator, and structural repair with default safety parameters. On 75 held-out workflow prompts, its hybrid variant scores 0.67 on the static safety proxy, compared with 0.33 for the best baseline. Koan-Safe records no unsafe executions on the saved benchmark outputs. A matched-candidate ablation produces 14-17 unsafe executions when enforcement is disabled. Additional tests expose the limits of default injection: permissive existing thresholds can still authorize unsafe trades. A separately evaluated policy cap addresses this failure on a 36-case diagnostic grid. These results support explicit trade protections and execution-based evaluation, while distinguishing declared safety from a general guarantee.</summary>3277    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>3278    <category term="cs.SE" scheme="http://arxiv.org/schemas/atom"/>3279    <published>2026-09-10T13:10:32Z</published>3280    <arxiv:comment>Code and benchmark: https://github.com/Varun-2538/Koan</arxiv:comment>3281    <arxiv:primary_category term="cs.LG"/>3282    <author>3283      <name>Abhinav Rajeev Kumar</name>3284    </author>3285    <author>3286      <name>Harshit Arora</name>3287    </author>3288    <author>3289      <name>Varun Singh</name>3290    </author>3291    <author>3292      <name>Manikandan Nanjappan</name>3293    </author>3294  </entry>3295  <entry>3296    <id>http://arxiv.org/abs/2609.11499v1</id>3297    <title>Recursive Code World Models: Building Complex Worlds through Recursive Scene Programs</title>3298    <updated>2026-09-10T13:05:06Z</updated>3299    <link href="https://arxiv.org/abs/2609.11499v1" rel="alternate" type="text/html"/>3300    <link href="https://arxiv.org/pdf/2609.11499v1" rel="related" type="application/pdf" title="pdf"/>3301    <summary>Code world models represent worlds as executable programs, but this representation alone does not determine how to construct a complex world. We introduce Recursive Code World Models (RCWM), a framework for reconstructing complex 3D worlds in code from a single reference image. RCWM couples a Recursive Scene Program (RSP) representation with a construction solver that recursively calls itself. An RSP represents the executable world as compositional scene code, while each solver call follows the same complete process: establish the whole, recursively reconstruct unresolved parts, and revisit the whole to refine their composition. This global-local-global recursion gives fine-scale structures their own perception-and-editing loops while preserving scene-wide geometry and relationships. Reference-aligned views propagate a shared camera projection across levels, while parent revisitation addresses boundaries, spatial relations, and shared errors that emerge after local refinement. A vision-language coding agent directly compares reference images with scene renders to guide refinement, recursive descent, and return. Across complex scenes, RCWM outperforms prior code-based image-to-scene reconstruction methods. Ablation studies further support the benefits of recursive construction and suggest that deeper calls can improve finer-scale reconstruction. RCWM provides a recursive construction principle for building complex executable worlds from visual evidence.</summary>3302    <category term="cs.CV" scheme="http://arxiv.org/schemas/atom"/>3303    <published>2026-09-10T13:05:06Z</published>3304    <arxiv:comment>21 pages, 11 figures</arxiv:comment>3305    <arxiv:primary_category term="cs.CV"/>3306    <author>3307      <name>Zhiqi Li</name>3308    </author>3309    <author>3310      <name>Yuxuan Liao</name>3311    </author>3312    <author>3313      <name>Bo Zhu</name>3314    </author>3315  </entry>3316  <entry>3317    <id>http://arxiv.org/abs/2609.11498v1</id>3318    <title>ActMap: Single-Pass Uncertainty Quantification from Generation-Time Activation Maps</title>3319    <updated>2026-09-10T13:04:24Z</updated>3320    <link href="https://arxiv.org/abs/2609.11498v1" rel="alternate" type="text/html"/>3321    <link href="https://arxiv.org/pdf/2609.11498v1" rel="related" type="application/pdf" title="pdf"/>3322    <summary>Practical uncertainty quantification (UQ) for large language models must decide,3323  from a single generation, whether a specific answer should be trusted. Existing3324  methods either sample multiple generations, read only output-token probabilities,3325  or reduce the model's internal computation to a single hidden state. We introduce3326  ActMap, a white-box representation that compresses the generation-time hidden-3327  state trajectory (every layer, every generated token) into a fixed $12 \times 323328  \times 128$ tensor of temporal-statistic channels that preserves structure across3329  transformer depth and pooled hidden coordinates. The map is captured during the3330  generation pass with no measurable overhead, has a fixed shape across model3331  depths and hidden sizes, and occupies 96 KiB: a compact artifact that can be3332  retained for audit-relevant generations and probed directly, with occlusion3333  analysis localizing the classifier's signal to mid-depth regions of the map. A3334  lightweight classifier, instantiated as a compact Vision Transformer, reads an3335  estimated correctness probability from each map in a fraction of a millisecond;3336  capacity-matched MLPs perform comparably, indicating the representation itself3337  carries the result. Trained and evaluated in-domain on short-answer QA, direct-3338  answer math, and summarization factuality with three instruction-tuned 7-8B3339  models, ActMap consistently outperforms sampling, token-probability, attention,3340  and embedding baselines, and matches ACT-ViT, a detector trained on dense3341  activation tensors $67 \times$ larger, at essentially the same mean AUROC with3342  lower calibration error on ten of twelve pairs. The resulting score supports3343  abstention, routing, and selective verification from a single generation, making3344  it a practical primitive for scalable oversight of deployed models.</summary>3345    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>3346    <published>2026-09-10T13:04:24Z</published>3347    <arxiv:comment>13 pages, 4 figures, 10 tables. Includes technical appendix</arxiv:comment>3348    <arxiv:primary_category term="cs.AI"/>3349    <author>3350      <name>Jacopo Dardini</name>3351      <arxiv:affiliation>University of Bologna</arxiv:affiliation>3352    </author>3353    <author>3354      <name>Roberta Calegari</name>3355      <arxiv:affiliation>University of Bologna</arxiv:affiliation>3356    </author>3357  </entry>3358  <entry>3359    <id>http://arxiv.org/abs/2609.11495v1</id>3360    <title>Combining Synthetic and Real Data for Low-Resource Historical OCR: A Manchu Case Study</title>3361    <updated>2026-09-10T13:02:11Z</updated>3362    <link href="https://arxiv.org/abs/2609.11495v1" rel="alternate" type="text/html"/>3363    <link href="https://arxiv.org/pdf/2609.11495v1" rel="related" type="application/pdf" title="pdf"/>3364    <summary>Manchu, now critically endangered, was one of the principal languages of the Qing empire (1636-1912), and its extensive archival record is increasingly digitized but remains difficult to search and analyze at scale. Previous work showed that vision-language models (VLMs) trained only on synthetic Manchu word images can reach 87.4% word accuracy on real Qing manuscripts and prints, leaving a substantial synthetic-to-real gap. This study examines how synthetic and real historical training data should be combined for low-resource OCR. Using 60,000 synthetic and 20,306 real historical word images, we evaluate three pretrained VLMs and a compact convolutional recurrent neural network (CRNN) under four regimes: synthetic-only, real-only, joint synthetic-real, and sequential synthetic-to-real training, following a common checkpoint-selection and archival evaluation protocol. Introducing real training images raises the leading configurations to between 95.09% and 96.28% word accuracy, while no synthetic-only configuration exceeds 87.92%. Synthetic supplementation substantially improves all three VLMs, whereas its marginal effect for the CRNN is sensitive to the training objective. Joint and sequential training yield broadly similar archival accuracy under the tested practical pipelines. A compact CRNN also reaches the leading performance range once real images are available, showing that model scale alone does not determine recognition accuracy. Finally, complementary errors among strong recognizers allow voting to raise accuracy to 98.27% without additional training, while an eighteenth-century Manchu dictionary provides a principled rule for adjudicating disagreements.</summary>3365    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>3366    <published>2026-09-10T13:02:11Z</published>3367    <arxiv:primary_category term="cs.LG"/>3368    <author>3369      <name>Yan Hon Michael Chung</name>3370    </author>3371    <author>3372      <name>Hanlin Wang</name>3373    </author>3374  </entry>3375  <entry>3376    <id>http://arxiv.org/abs/2609.11493v1</id>3377    <title>From Document Silos to Process Intelligence: A Multi-Layer Knowledge Graph for CMC Process Development</title>3378    <updated>2026-09-10T12:59:04Z</updated>3379    <link href="https://arxiv.org/abs/2609.11493v1" rel="alternate" type="text/html"/>3380    <link href="https://arxiv.org/pdf/2609.11493v1" rel="related" type="application/pdf" title="pdf"/>3381    <summary>Chemistry, Manufacturing and Controls (CMC) process development generates an enormous body of technical information across a multi-stage, knowledge-intensive continuum from drug discovery to commercial manufacturing. This knowledge is traditionally fragmented across functions and heterogeneous formats, causing traceability gaps and significant knowledge-management costs during technology transfer and regulatory filing. We present a modular agentic-AI platform that converts a heterogeneous corpus of process-development documents into a queryable, dual-layer knowledge graph. A base knowledge layer builds a lexical graph with a Document-Section-Chunk hierarchy through lossless ingestion of digital, scanned, handwritten, and multilingual documents, while an intelligence layer extracts ontology-aligned entities and bridges cross-document concepts through a provenance-anchored domain graph. LLM agents operate across both layers, selecting the retrieval path best suited to each question. We evaluate the lexical layer with a novel three-tier protocol measuring the deployment-fidelity of a retrieval-augmented generation (RAG) system on proprietary data, demonstrated on 505 questions curated from 38 development reports of a Sanofi small-molecule program. Tier-1 multiple-choice accuracy of 95% signals strong platform reliability; the stricter Tier-2 LLM-judge pass rate of 85%, which degrades on comparative and corpus-wide questions, reveals a failure taxonomy that Tier-1 accuracy alone fails to capture. A router agent selects between layers according to question type. We anticipate this protocol will enable future designers of agentic platforms to assess their systems against nonpublic databases, and that graph-based architectures will see broader adoption in pharma as a means of transforming fragmented document repositories into structured process intelligence.</summary>3382    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>3383    <category term="cs.MA" scheme="http://arxiv.org/schemas/atom"/>3384    <published>2026-09-10T12:59:04Z</published>3385    <arxiv:primary_category term="cs.AI"/>3386    <author>3387      <name>Reza Amirmoshiri</name>3388    </author>3389    <author>3390      <name>Faryad Sahneh</name>3391    </author>3392    <author>3393      <name>Yasser Jangjou</name>3394    </author>3395  </entry>3396  <entry>3397    <id>http://arxiv.org/abs/2609.11490v1</id>3398    <title>Published Unlearning Numbers Move Per Checkpoint, and Not Because the Removed Data Survives: An Audit of 263 Released Batch-Normalized Checkpoints</title>3399    <updated>2026-09-10T12:57:14Z</updated>3400    <link href="https://arxiv.org/abs/2609.11490v1" rel="alternate" type="text/html"/>3401    <link href="https://arxiv.org/pdf/2609.11490v1" rel="related" type="application/pdf" title="pdf"/>3402    <summary>An unlearning audit reads its verdict off numbers that an unlearned model and its retrained reference each publish, and both also ship batch-normalization statistics that no gradient step wrote and no release records. Refitting them on kept data at bit-identical weights moves 47 of 221 released checkpoints past the spread their own release's seeds show, several inside a method whose average does not move: what moves is the checkpoint's property, not its method's. What does the moving is not the removed data surviving in the state: exchanging kept records for removed ones inside a fixed fitting pool moves a published cell by almost nothing, while how far a checkpoint's shipped state has drifted from any refit does track it. The consequence for a published decision is real but narrow: twelve verdicts cross, four clear a measured recalibration budget, two clear it on every replicate, and a population we trained and sited near its own criterion yields none. A release should therefore name the fitting convention beside the number, on the batch-normalized vision models where this channel exists.</summary>3403    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>3404    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>3405    <published>2026-09-10T12:57:14Z</published>3406    <arxiv:comment>38 pages, 4 figures, 26 tables. Independent of and concurrent with arXiv:2609.08901 (posted 8 Sep 2026): the instrument and protocol here were pre-registered on 29 Aug 2026; dated provenance in Appendix S</arxiv:comment>3407    <arxiv:primary_category term="cs.AI"/>3408    <author>3409      <name>Junlong Shen Xingyu Li</name>3410    </author>3411  </entry>3412  <entry>3413    <id>http://arxiv.org/abs/2609.11489v1</id>3414    <title>The Convention Gap: Towards Measuring Implicit Communication in Cooperative AI Evaluation</title>3415    <updated>2026-09-10T12:57:03Z</updated>3416    <link href="https://arxiv.org/abs/2609.11489v1" rel="alternate" type="text/html"/>3417    <link href="https://arxiv.org/pdf/2609.11489v1" rel="related" type="application/pdf" title="pdf"/>3418    <summary>Cooperative AI agents are evaluated against other AIs, yet human cooperation relies on implicit conventions---shared protocols for reading meaning beyond the literal message---which AI-AI benchmarks may not capture. We propose the \emph{convention gap}, the difference between the failure probability predicted from the literal content of communication and the observed failure rate, as a metric of implicit communication. In the card game Hanabi, the finite deck and deterministic hint constraints make this posterior exactly computable. We replayed about 101,000 play actions from three public datasets of human-human (hanab.live), AI-AI (HOAD), and human-AI (HanabiData) games. The gap was +26.2 percentage points (pp) in human pairs, $-$0.7~pp in AI pairs, and +16.4~pp in human-AI pairs, and was concentrated on plays of cards that had received no hints (+46~pp in human pairs). Within human-AI play, the literal information available to humans was similar across the three AI partners (mean predicted failure 38--41\%), but human failure rates ranged from 14.4\% to 34.4\% and the gap from +24.1 to +6.2~pp; the partner eliciting the largest gap produced the fewest human failures. Game score carried different information: it depended on each corpus's roster composition, whereas the gap separated human from AI play at the agent level. As a known-answer check, Off-Belief Learning agents, whose convention content is controlled by construction, gave a gap of +1.6~pp at the convention-free level, rising monotonically to +21.7~pp. These results suggest that convention compatibility, rather than AI-AI performance, may predict an AI's effectiveness with human partners.</summary>3419    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>3420    <category term="cs.HC" scheme="http://arxiv.org/schemas/atom"/>3421    <published>2026-09-10T12:57:03Z</published>3422    <arxiv:primary_category term="cs.AI"/>3423    <author>3424      <name>Makoto Fukushima</name>3425    </author>3426    <author>3427      <name>Hua-Dong Xiong</name>3428    </author>3429    <author>3430      <name>Ehsan Moradi Pari</name>3431    </author>3432  </entry>3433  <entry>3434    <id>http://arxiv.org/abs/2609.11486v1</id>3435    <title>FreeFlow: A Bias-free Hierarchical Transformer for Optical Flow Estimation</title>3436    <updated>2026-09-10T12:55:04Z</updated>3437    <link href="https://arxiv.org/abs/2609.11486v1" rel="alternate" type="text/html"/>3438    <link href="https://arxiv.org/pdf/2609.11486v1" rel="related" type="application/pdf" title="pdf"/>3439    <summary>Optical flow methods typically rely on task-specific inductive biases, such as correlation volumes, feature warping, and iterative refinement, among others, to reach high accuracy. While effective, such biases constrain the model to predefined heuristics, which can limit its expressivity and lead to more complex pipelines and additional computational cost. We present FreeFlow, a hierarchical transformer built without any flow-specific components, using instead a single feed-forward encoder--decoder. FreeFlow combines three attention variants: window attention for local processing, shifted-window attention for cross-window information exchange, and a global attention operating at a reduced resolution. The resulting architecture scales naturally with model capacity, enabling a consistent accuracy gain from small to large variants. Despite the absence of standard inductive biases, FreeFlow achieves state-of-the-art results on major benchmarks, including Sintel (0.68/1.48 EPE on Clean/Final), KITTI-2015 (3.23 Fl-all), and Spring (3.192 1px), while remaining memory efficient at 1080p inference.</summary>3440    <category term="cs.CV" scheme="http://arxiv.org/schemas/atom"/>3441    <published>2026-09-10T12:55:04Z</published>3442    <arxiv:comment>Accepted at ECCV 2026. Project page: https://github.com/msu-video-group/freeflow</arxiv:comment>3443    <arxiv:primary_category term="cs.CV"/>3444    <author>3445      <name>Vladislav Bargatin</name>3446    </author>3447    <author>3448      <name>Alexander Yakovenko</name>3449    </author>3450    <author>3451      <name>Khaled Abud</name>3452    </author>3453    <author>3454      <name>Dmitriy Vatolin</name>3455    </author>3456    <arxiv:doi>10.1007/978-3-032-37132-4_11</arxiv:doi>3457    <link rel="related" href="https://doi.org/10.1007/978-3-032-37132-4_11" title="doi"/>3458  </entry>3459  <entry>3460    <id>http://arxiv.org/abs/2609.11477v1</id>3461    <title>Pre- and Post-Treatment Brain Metastases Segmentation Using nnU-Net with Post-Processing for BraTS 2026</title>3462    <updated>2026-09-10T12:43:04Z</updated>3463    <link href="https://arxiv.org/abs/2609.11477v1" rel="alternate" type="text/html"/>3464    <link href="https://arxiv.org/pdf/2609.11477v1" rel="related" type="application/pdf" title="pdf"/>3465    <summary>Brain metastases exhibit high inter-lesion variability in size, enhancement pattern, and post-treatment appearance, making volumetric segmentation of both pre- and post-treatment cases the central challenge of the BraTS 2026 Task 1 (Brain Metastases). We build a pragmatic pipeline on a 5-fold nnU-Net ResEnc-L ensemble, in which each fold is trained independently for 1,000 epochs with the standard Dice + cross-entropy loss on 1,296 four-modality training cases. This ensemble is followed by a rule-based post-processing cascade tuned for the lesion-wise Dice similarity coefficient (LW-DSC), a detection-oriented metric that behaves very differently from the traditional global Dice. The final pipeline reaches an LW-DSC of 0.733 / 0.751 / 0.713 / 0.549 on the enhancing tumour (ET), tumour core (TC), whole tumour (WT), and resection cavity (RC) sub-regions on the official validation leaderboard. Rather than trusting these leaderboard gains, we audit every post-processing stage with a five-fold out-of-fold (OOF) analysis with no model-training leakage over all 1,296 training cases, scored with the official BraTS evaluation code (BraTS_evaluation): it confirms two stages as robust, per-fold-consistent improvements while the third improves only the leaderboard and does not reproduce out-of-fold. We further provide a mechanistic analysis of the LW-DSC metric that explains why recall-recovering post-processing carries low risk whereas component deletion does not, and we report thirteen negative results spanning loss engineering, alternative backbones, and inference-time settings, several of which run counter to widely held intuitions. Source code is released under Apache-2.0 at https://github.com/hornbeamliu/brats2026-met.</summary>3466    <category term="cs.CV" scheme="http://arxiv.org/schemas/atom"/>3467    <published>2026-09-10T12:43:04Z</published>3468    <arxiv:comment>Accepted to MICCAI 2026 Challenge BraTS-METS</arxiv:comment>3469    <arxiv:primary_category term="cs.CV"/>3470    <author>3471      <name>Haobin Liu</name>3472    </author>3473    <author>3474      <name>Xin Wang</name>3475    </author>3476  </entry>3477  <entry>3478    <id>http://arxiv.org/abs/2609.11472v1</id>3479    <title>BridgeMatch: Conditional Transport Bridges in Matching Matrix Space for 3D Deformable Registration</title>3480    <updated>2026-09-10T12:39:58Z</updated>3481    <link href="https://arxiv.org/abs/2609.11472v1" rel="alternate" type="text/html"/>3482    <link href="https://arxiv.org/pdf/2609.11472v1" rel="related" type="application/pdf" title="pdf"/>3483    <summary>Reliable non-rigid point cloud correspondences are important for deformable anatomical registration, embodied perception and manipulation, and dynamic 3D reconstruction. Coarse-to-fine methods reduce computational cost by selecting the top-\(K\) coarse regions. However, this pruning may remove weak but correct hypotheses and restrict fine matching to an incomplete search space. We present \paper, a two-stage generative solver that maintains the complete soft matching matrix at both coarse and high resolutions. Stage~I uses denoising diffusion to estimate a global matching matrix in the compact coarse-resolution space. We then lift this matrix to high resolution while preserving its hierarchy. The lifted matrix is rank-bounded and block-constant. Stage~II refines it through a conditional transport bridge. We implement the bridge with two types of dynamics: a deterministic endpoint-parameterized conditional Flow Matching (CFM) ODE and a stochastic Brownian-bridge SDE inspired by Schrödinger bridges. Both variants share the lifted source, a time-conditioned transformer, and a matching-matrix endpoint predictor. Experiments on 4DMatch and 4DLoMatch show that both variants produce more accurate correspondences than the compared methods and improve downstream registration, with larger gains in low-overlap cases. They also improve cross-dataset generalization on CAPE and DeepDeform without target-domain adaptation while using the same deformation solver.</summary>3484    <category term="cs.CV" scheme="http://arxiv.org/schemas/atom"/>3485    <published>2026-09-10T12:39:58Z</published>3486    <arxiv:primary_category term="cs.CV"/>3487    <author>3488      <name>Qianliang Wu</name>3489    </author>3490    <author>3491      <name>Haobo Jiang</name>3492    </author>3493    <author>3494      <name>Guangwei Gao</name>3495    </author>3496    <author>3497      <name>Shuo Chen</name>3498    </author>3499    <author>3500      <name>Jin Xie</name>3501    </author>3502    <author>3503      <name>Jian Yang</name>3504    </author>3505    <author>3506      <name>Yaqing Ding</name>3507    </author>3508  </entry>3509  <entry>3510    <id>http://arxiv.org/abs/2609.11463v1</id>3511    <title>BruNet: A Cross-Domain Transfer Framework for Bruise Segmentation</title>3512    <updated>2026-09-10T12:35:20Z</updated>3513    <link href="https://arxiv.org/abs/2609.11463v1" rel="alternate" type="text/html"/>3514    <link href="https://arxiv.org/pdf/2609.11463v1" rel="related" type="application/pdf" title="pdf"/>3515    <summary>Segmenting bruises is a challenging task in medical imaging due to limited data and annotations, diffuse boundaries, and highly variable appearance. In this work, we propose BruNet, a segmentation framework that combines a ViT-based visual encoder (a self-supervised DINOv3 or a pretrained LingBot-Vision backbone) with a SAM-based mask decoder. BruNet is trained on the HAM10000 skin lesion dataset and evaluated on a separate bruise dataset without additional fine-tuning. Although a small number of prior studies have explored machine learning and computer vision for bruise analysis, existing work has primarily focused on detection, classification, or colour analysis rather than pixel-level localisation. To the best of our knowledge, this is the first study to address automatic bruise segmentation. Our results show that BruNet outperforms CNN-based models, state-of-the-art segmentation models, ChatGPT-4o/5-assisted SAM2 zero-shot baselines, and the medical-oriented MedSAM model, demonstrating strong cross-domain generalisation to bruise segmentation.</summary>3516    <category term="cs.CV" scheme="http://arxiv.org/schemas/atom"/>3517    <published>2026-09-10T12:35:20Z</published>3518    <arxiv:primary_category term="cs.CV"/>3519    <author>3520      <name>Qiming Wang</name>3521    </author>3522    <author>3523      <name>Richard J. Motley</name>3524    </author>3525    <author>3526      <name>Ebube E. Obi</name>3527    </author>3528    <author>3529      <name>Xianfang Sun</name>3530    </author>3531    <author>3532      <name>Paul L. Rosin</name>3533    </author>3534  </entry>3535  <entry>3536    <id>http://arxiv.org/abs/2609.11460v1</id>3537    <title>ReGround: Grounding Reviewer Comments in Multimodal Evidence</title>3538    <updated>2026-09-10T12:32:58Z</updated>3539    <link href="https://arxiv.org/abs/2609.11460v1" rel="alternate" type="text/html"/>3540    <link href="https://arxiv.org/pdf/2609.11460v1" rel="related" type="application/pdf" title="pdf"/>3541    <summary>Reviewer comments naturally relate to specific parts of the reviewed paper, yet grounding these comments to the underlying evidence is difficult due to long multimodal documents. Existing benchmarks do not capture this setting and largely focus on explicit, information-seeking queries. We introduce ReGround, a large-scale dataset for reviewer comment grounding that links 10,267 reviewer comments to 16,274 evidence in the original anonymous submission of 3,656 papers. We build on a simple observation: author rebuttals often include explicit references to content of the submission used to address reviewer comments, providing a high-precision annotation source. We cast grounding as a retrieval task and evaluate a wide range of retrieval methods. Results show that retrieval over the entire paper content performs poorly, evidence-type inference is a major bottleneck, and multimodal evidence provides complementary signals that text alone misses. Our dataset exposes grounding reviewer comments as a difficult and practically important problem for scientific document understanding.</summary>3542    <category term="cs.CL" scheme="http://arxiv.org/schemas/atom"/>3543    <category term="cs.IR" scheme="http://arxiv.org/schemas/atom"/>3544    <published>2026-09-10T12:32:58Z</published>3545    <arxiv:comment>Accepted at EMNLP 2026</arxiv:comment>3546    <arxiv:primary_category term="cs.CL"/>3547    <author>3548      <name>Serwar Basch</name>3549    </author>3550    <author>3551      <name>Lizhen Qu</name>3552    </author>3553    <author>3554      <name>Iryna Gurevych</name>3555    </author>3556  </entry>3557  <entry>3558    <id>http://arxiv.org/abs/2609.11458v1</id>3559    <title>Flexible and Interpretable Accent Distance Measurements</title>3560    <updated>2026-09-10T12:32:13Z</updated>3561    <link href="https://arxiv.org/abs/2609.11458v1" rel="alternate" type="text/html"/>3562    <link href="https://arxiv.org/pdf/2609.11458v1" rel="related" type="application/pdf" title="pdf"/>3563    <summary>Determining the differences between two speakers' accents is a fundamental task in linguistics and speech technology research. The methodology used to measure these differences depends on the specific research area. A phonetics researcher may demonstrate accent variation by comparing vowel formants in paired recordings of individual words. These results will be interpretable, but the recordings will be time-consuming to collect and may not be representative of connected speech. Accented Text-to-Speech (TTS) research has pushed towards using accent embeddings derived from accent classification tasks. These embeddings can be produced from any speech recording, but are not readily interpretable. In this paper, we demonstrate that articulatory representations created through articulatory inversion can be used as an interpretable basis for accent comparison and that optimal transport provides a framework for accent comparison across arbitrary recording types.</summary>3564    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>3565    <published>2026-09-10T12:32:13Z</published>3566    <arxiv:primary_category term="cs.AI"/>3567    <author>3568      <name>Charles McGhee</name>3569    </author>3570    <author>3571      <name>Mark J. F. Gales</name>3572    </author>3573    <author>3574      <name>Kate M. Knill</name>3575    </author>3576  </entry>3577  <entry>3578    <id>http://arxiv.org/abs/2609.11452v1</id>3579    <title>RouteRepair: Instance-Level Failure Diagnosis and Targeted Repair in LLM-Based Automated Heuristic Design for Routing Optimization</title>3580    <updated>2026-09-10T12:21:18Z</updated>3581    <link href="https://arxiv.org/abs/2609.11452v1" rel="alternate" type="text/html"/>3582    <link href="https://arxiv.org/pdf/2609.11452v1" rel="related" type="application/pdf" title="pdf"/>3583    <summary>Efficient routing optimization is essential to freight transportation, urban logistics, and shared mobility, where high-quality heuristics are often required under limited computational budgets. Recent large language model (LLM)-based automated heuristic design methods can generate effective routing rules, but aggregate evaluation may mask recurrent failures on particular instance structures. To address this limitation, this study develops RouteRepair, which diagnoses parent-specific weaknesses from instance-level performance and applies targeted modifications to the corresponding heuristic components while protecting behavior that already performs well. Routing evidence, solver behavior, and program context are combined to define bounded repair objectives, and each intervention is validated through matched parent-child evaluation of failure recovery and collateral degradation. Experiments on the traveling salesman problem (TSP) and capacitated vehicle routing problem (CVRP) span constructive search, guided local search, and ant colony optimization. RouteRepair-GLS reduces the mean TSP optimality gap from 1.7476% to 0.7587%, while the constructive CVRP heuristic lowers average route cost by 1.91% relative to the savings heuristic; the generated ACO priors also outperform matched hand-designed priors. These results show that failure-aware, evidence-constrained refinement can improve routing heuristics on difficult instances while preserving performance on cases they already solve well.</summary>3584    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>3585    <published>2026-09-10T12:21:18Z</published>3586    <arxiv:comment>22 pages, 13 figures, 11 tables</arxiv:comment>3587    <arxiv:primary_category term="cs.AI"/>3588    <author>3589      <name>Binghao Ji</name>3590    </author>3591    <author>3592      <name>Di Huang</name>3593    </author>3594    <author>3595      <name>Jiahui Fang</name>3596    </author>3597    <author>3598      <name>Zhiyuan Liu</name>3599    </author>3600  </entry>3601  <entry>3602    <id>http://arxiv.org/abs/2609.11450v1</id>3603    <title>Cross-Lingual Clinical Annotation Projection as Constrained Text Generation: A Six-Language Study</title>3604    <updated>2026-09-10T12:17:52Z</updated>3605    <link href="https://arxiv.org/abs/2609.11450v1" rel="alternate" type="text/html"/>3606    <link href="https://arxiv.org/pdf/2609.11450v1" rel="related" type="application/pdf" title="pdf"/>3607    <summary>Background: To determine whether cross-lingual clinical annotation projection can be formulated as a text-preserving, document-level generative task that produces verifiable character-level annotations for multilingual clinical corpus construction, and to characterize its robustness and computational trade-offs relative to candidate-based projection pipelines. Methods: We developed a constrained LLM projection workflow that inserts entity tags directly into immutable target-language text, followed by deterministic validation and character-offset reconstruction. We evaluated it alongside supervised candidate-span projection and hybrid ML-LLM refinement for transferring Spanish Disease, Symptom, and Procedure annotations into six languages. Evaluation used MultiClinAI gold standard with strict span matching and character-overlap F1 Results: Direct LLM projection achieved the strongest and most consistent performance. GLM 5.2 obtained a mean Strict F1 of 0.9201 across 18 language-entity combinations, while locally deployable Gemma4:31B achieved 0.9133. The best LLM configuration improved Strict F1 over the previous state of the art in all 18 settings, by 0.0564-0.1512, yielding 55,416 grounded mentions with reconstructed offsets. Conclusions: Direct LLM-based projection enables high-quality multilingual clinical annotation transfer and provides a practical approach for extending clinical NLP resources to languages with fewer annotated datasets and language-specific tools. Combined with local inference and deterministic validation, it can substantially reduce expert time and cost for multilingual clinical corpus construction.</summary>3608    <category term="cs.CL" scheme="http://arxiv.org/schemas/atom"/>3609    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>3610    <published>2026-09-10T12:17:52Z</published>3611    <arxiv:comment>14 pages, 4 figures, 4 tables, submitted to journal</arxiv:comment>3612    <arxiv:primary_category term="cs.CL"/>3613    <author>3614      <name>Álvaro Rey-Blanes</name>3615    </author>3616    <author>3617      <name>Francisco J. Moreno-Barea</name>3618    </author>3619    <author>3620      <name>Francisco J. Veredas</name>3621    </author>3622  </entry>3623  <entry>3624    <id>http://arxiv.org/abs/2609.11449v1</id>3625    <title>Prevalence Determines Precision:Silent Contamination in Detector-Defined Datasets</title>3626    <updated>2026-09-10T12:17:44Z</updated>3627    <link href="https://arxiv.org/abs/2609.11449v1" rel="alternate" type="text/html"/>3628    <link href="https://arxiv.org/pdf/2609.11449v1" rel="related" type="application/pdf" title="pdf"/>3629    <summary>Many ML datasets are constructed by running a detector, heuristic, or model over candidate pools; accepted items become labels. Dataset precision is then governed by true-positive prevalence in each pool via Bayes, not solely by detector quality. Using one instrument and period, we hold a detector-defined event dataset plus an independent official index labeling every detected item as real or phantom. One detector, three pools yield phantom rates 81.7%, 9.0%, and 0.0%. Transferring precision from the two high-rate pools to the low-rate pool predicts 0.955 versus measured 0.183, a +422% error; the Bayes expression predicts all three within 3.3%. The detected response curve is an exact convex combination of a true-event and a phantom component (residual 1.1e-16), with phantoms outnumbering true events 473 to 308, so contamination is a second signal with detector-inherited shape, not additive noise. Contamination direction depends on the estimator: on identical windows one statistic is diluted and another inflated because its denominator is also contaminated. A common normalization turns the estimator into a mean of ratios whose expectation need not exist; on the same 335 events it returns 0.40 where the well-defined estimator returns 0.10.</summary>3630    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>3631    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>3632    <published>2026-09-10T12:17:44Z</published>3633    <arxiv:comment>9 pages</arxiv:comment>3634    <arxiv:primary_category term="cs.LG"/>3635    <author>3636      <name>Jia Huang</name>3637    </author>3638    <author>3639      <name>Yankai Wan</name>3640    </author>3641    <author>3642      <name>Yangjun Ou</name>3643    </author>3644  </entry>3645  <entry>3646    <id>http://arxiv.org/abs/2609.11447v1</id>3647    <title>Investigating catastrophic forgetting in sound event classification</title>3648    <updated>2026-09-10T12:16:56Z</updated>3649    <link href="https://arxiv.org/abs/2609.11447v1" rel="alternate" type="text/html"/>3650    <link href="https://arxiv.org/pdf/2609.11447v1" rel="related" type="application/pdf" title="pdf"/>3651    <summary>This work investigates a number of approaches to prevent catastrophic forgetting in class incremental learning scenarios for sound event classification tasks. We analyze the problem using architectural and regularization approaches, using FSD50K and AudioSet datasets. We design incremental stages and solutions that selectively protect the kernels of the network from weight updates to prevent catastrophic forgetting, and a dynamic head solution that expands itself each time a new task is learned. The findings show that catastrophic forgetting mainly happens in deeper layers, in particular in the classifier head. For the studied in-domain sound classification problem, the solution that seems to alleviate catastrophic forgetting and is the most efficient is a full freezing of the feature extractor with a fine-tuning of the dynamic head classifier, showing little to no forgetting and great training stability, and a good balance between memory-stability and learning plasticity.</summary>3652    <category term="eess.AS" scheme="http://arxiv.org/schemas/atom"/>3653    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>3654    <category term="cs.SD" scheme="http://arxiv.org/schemas/atom"/>3655    <published>2026-09-10T12:16:56Z</published>3656    <arxiv:comment>Accepted in MMSP2026</arxiv:comment>3657    <arxiv:primary_category term="eess.AS"/>3658    <author>3659      <name>Riccardo Casciotti</name>3660    </author>3661    <author>3662      <name>Annamaria Mesaros</name>3663    </author>3664  </entry>3665  <entry>3666    <id>http://arxiv.org/abs/2609.11446v1</id>3667    <title>Calibration-Aware Uncertainty Cascades for Efficient Heterogeneous Model Collaboration</title>3668    <updated>2026-09-10T12:14:54Z</updated>3669    <link href="https://arxiv.org/abs/2609.11446v1" rel="alternate" type="text/html"/>3670    <link href="https://arxiv.org/pdf/2609.11446v1" rel="related" type="application/pdf" title="pdf"/>3671    <summary>Heterogeneous model collaboration seeks to exploit the complementary strengths of different models to balance predictive performance and inference cost. Existing approaches typically rely either on trained routers, which tie routing decisions to a fixed task and model pool, or on raw-confidence cascades, whose thresholds lack consistent reliability semantics across heterogeneous models. Consequently, these approaches adapt poorly to changing model pools and deployment budgets. We propose Calibration-Aware Uncertainty Cascades (CAUC), a simple post-hoc framework that independently calibrates each model's confidence and selects deployment policies using validation data. The resulting calibrated confidence scores establish a common reliability scale for accepting an early prediction, invoking a stronger model, or selectively combining model outputs. This unified decision criterion decouples deployment policies from any particular model pool or operating budget. We further show theoretically that calibration gives confidence thresholds an explicit selective-risk interpretation, whereas uncalibrated scores offer no comparable reliability guarantee. Extensive experiments demonstrate that, across six language benchmarks, CAUC achieves an average relative accuracy improvement of 1.9% over strong-model-only inference while avoiding approximately 47% of strong-model calls. On image classification benchmarks, it maintains or improves predictive performance while reducing measured GFLOPs by up to 57%.</summary>3672    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>3673    <published>2026-09-10T12:14:54Z</published>3674    <arxiv:comment>13 pages, 6 figures, 6 tables, including appendix. Under review</arxiv:comment>3675    <arxiv:primary_category term="cs.AI"/>3676    <author>3677      <name>Yilin Zhang</name>3678    </author>3679    <author>3680      <name>Han Jiang</name>3681    </author>3682    <author>3683      <name>Cai Xu</name>3684    </author>3685    <author>3686      <name>Ying Liu</name>3687    </author>3688    <author>3689      <name>Wei Zhao</name>3690    </author>3691  </entry>3692  <entry>3693    <id>http://arxiv.org/abs/2609.11439v1</id>3694    <title>Multi-Modal Controlled Coherent Motion Generation</title>3695    <updated>2026-09-10T12:10:00Z</updated>3696    <link href="https://arxiv.org/abs/2609.11439v1" rel="alternate" type="text/html"/>3697    <link href="https://arxiv.org/pdf/2609.11439v1" rel="related" type="application/pdf" title="pdf"/>3698    <summary>It is natural for humans to walk and talk simultaneously. This paper tackles the challenge of replicating such natural behaviors in 3D avatar motion generation driven by concurrent multimodal inputs, such as a text description of a man walking alongside speech audio. Existing methods, constrained by the scarcity of aligned multimodal data, typically combine motions from individual modalities sequentially or through weighted sums. However, they often result in mismatched or unrealistic movements. To overcome these limitations, we propose MOCO, a novel diffusion-based framework capable of processing multiple simultaneous inputs, including speech audio, text descriptions, and trajectory data, to generate coherent and lifelike motions without requiring aligned multimodal data. Our key innovation lies in decoupling the motion generation process. During each denoising step, the diffusion model independently generates motions for each modality from the input noise and assembles the body parts according to predefined spatial rules. The resulting combined motion is then diffused and serves as the input noise for the subsequent denoising step. This iterative approach enables each modality to refine its contribution within the context of the overall motion, progressively harmonizing movements across modalities. Consequently, the generated motions become increasingly natural and fluid with each iteration, achieving coherent and synchronized behaviors. We evaluate our approach using a purpose-built multimodal benchmark. Experimental results demonstrate that MOCO outperforms existing baselines, advancing the field of multimodal motion generation for 3D avatars.</summary>3699    <category term="cs.CV" scheme="http://arxiv.org/schemas/atom"/>3700    <published>2026-09-10T12:10:00Z</published>3701    <arxiv:comment>ECCV 2026</arxiv:comment>3702    <arxiv:primary_category term="cs.CV"/>3703    <author>3704      <name>Yifei Liu</name>3705    </author>3706    <author>3707      <name>Qiong Cao</name>3708    </author>3709    <author>3710      <name>Hongwei Yi</name>3711    </author>3712    <author>3713      <name>Huaiguang Jiang</name>3714    </author>3715    <author>3716      <name>Changxing Ding</name>3717    </author>3718  </entry>3719  <entry>3720    <id>http://arxiv.org/abs/2609.11434v1</id>3721    <title>Hologram Representation via Quadratic Phase Gaussian Splatting</title>3722    <updated>2026-09-10T12:06:11Z</updated>3723    <link href="https://arxiv.org/abs/2609.11434v1" rel="alternate" type="text/html"/>3724    <link href="https://arxiv.org/pdf/2609.11434v1" rel="related" type="application/pdf" title="pdf"/>3725    <summary>We introduce Complex-Valued Quadratic Phase Gaussian (CVQPG), a novel hologram representation method that replaces standard 2D Gaussian representations used in 2D Gaussian Splatting with 2D quadratic phase functions. CVQPG incorporates additional learnable parameters to control the curvature of these bases. We evaluate our approach against state-of-the-art methods, exceeding the visual quality by +0.19 dB (RGB) and +0.33 dB (grayscale) on average in holographic reconstructions. Specifically, our equal parameter count evaluations show that modulating the primitive's wavefront is an effective and lightweight enhancement for hologram representations. In addition, our frequency domain analysis illustrates that CVQPG has successfully preserved the mid-to-high frequency band of natural images.</summary>3726    <category term="cs.GR" scheme="http://arxiv.org/schemas/atom"/>3727    <category term="cs.CV" scheme="http://arxiv.org/schemas/atom"/>3728    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>3729    <published>2026-09-10T12:06:11Z</published>3730    <arxiv:comment>SIGGRAPH Asia 2026 Technical Communications</arxiv:comment>3731    <arxiv:primary_category term="cs.GR"/>3732    <author>3733      <name>Haolong Wang</name>3734    </author>3735    <author>3736      <name>Yicheng Zhan</name>3737    </author>3738    <author>3739      <name>Kaan Akşit</name>3740    </author>3741    <author>3742      <name>Simeng Qiu</name>3743    </author>3744  </entry>3745  <entry>3746    <id>http://arxiv.org/abs/2609.11431v1</id>3747    <title>LLMs as Post-hoc Auditors of Physiological Plausibility in Symbolic Regression: A Clinician-Evaluated Case Study</title>3748    <updated>2026-09-10T12:03:33Z</updated>3749    <link href="https://arxiv.org/abs/2609.11431v1" rel="alternate" type="text/html"/>3750    <link href="https://arxiv.org/pdf/2609.11431v1" rel="related" type="application/pdf" title="pdf"/>3751    <summary>Genetic Programming and its variants, such as grammatical evolution, are widely used in Symbolic Regression to derive mathematical expressions from multivariate data. In addition to predictive accuracy, models are appreciated for their potential to provide interpretability, offering explicit equations that relate input variables to outcomes. However, achieving interpretability and plausibility remains challenging, as evolved models may be complex or scientifically inconsistent. In this study, we explore whether Large Language Models, can assist in improving the explainability of Symbolic Regression models generated by evolutionary computation methods. Building upon our previous work on estimating body fat percentage using grammar-based Genetic Programming , we investigate the use of LLMs as post-processing tools to analyze and rank evolved expressions according to their interpretability and medical plausibility. Four symbolic expressions are analysed by three LLMs over three repeated runs, and the resulting interpretations and rankings are assessed by a panel of three clinicians. Across the three LLMs, comparative model-ranking outputs received more favorable clinician assessments than isolated term-level interpretations. However, the LLMs also produced physiologically and mathematically questionable explanations, indicating that they are better suited to comparative auditing under expert oversight than to autonomous validation.\blfootnote{The present work is an extended version of a paper submitted into a journal.</summary>3752    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>3753    <published>2026-09-10T12:03:33Z</published>3754    <arxiv:primary_category term="cs.AI"/>3755    <author>3756      <name>Jorge López-Varela</name>3757    </author>3758    <author>3759      <name>J. Ignacio Hidalgo</name>3760    </author>3761    <author>3762      <name>José-Manuel Muñoz</name>3763    </author>3764    <author>3765      <name>Omar Costilla-Reyes</name>3766    </author>3767    <author>3768      <name>Esther Maqueda</name>3769    </author>3770    <author>3771      <name>Jesus Moreno-Fernandez</name>3772    </author>3773    <author>3774      <name>Tomás González-Vidal</name>3775    </author>3776    <author>3777      <name>J. Manuel Velasco</name>3778    </author>3779    <author>3780      <name>Oscar Garnica</name>3781    </author>3782  </entry>3783  <entry>3784    <id>http://arxiv.org/abs/2609.11414v1</id>3785    <title>SWRouter: Similarity-Contractive Window Routing for Multi-Turn Large Language Model Conversations</title>3786    <updated>2026-09-10T11:45:55Z</updated>3787    <link href="https://arxiv.org/abs/2609.11414v1" rel="alternate" type="text/html"/>3788    <link href="https://arxiv.org/pdf/2609.11414v1" rel="related" type="application/pdf" title="pdf"/>3789    <summary>Large language models exhibit complementary strengths, motivating routing methods that dispatch each query to the most suitable model. Although existing routers are effective in single-turn settings, they do not directly transfer to multi-turn dialogue, where routing performance critically depends on how historical context is segmented, retained, and incorporated into the current prompt. This introduces two fundamental challenges: preventing information loss and information confusion during context construction, and evaluating routing quality without conflating model selection with prompt construction quality. In this paper, we propose SWRouter, a Similarity-Contractive Window Router for multi-turn large language model routing. SWRouter combines a similarity-based context segmentation mechanism for prompt construction with a dual-metric evaluation framework that decouples construction accuracy from router performance. Experiments on multi-turn dialogue benchmarks demonstrate that SWRouter consistently surpasses strong baselines, achieving a 16.26% improvement in evaluation accuracy over the best individual large language model and an additional 8.22% gain over the Conv-ID Context baseline. Our results highlight that multi-turn large language model routing requires a joint design of context construction and evaluation, rather than a direct extension of single-turn routing methods.</summary>3790    <category term="cs.CL" scheme="http://arxiv.org/schemas/atom"/>3791    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>3792    <category term="cs.IR" scheme="http://arxiv.org/schemas/atom"/>3793    <published>2026-09-10T11:45:55Z</published>3794    <arxiv:primary_category term="cs.CL"/>3795    <author>3796      <name>Yu Wang</name>3797    </author>3798    <author>3799      <name>Yuchen Li</name>3800    </author>3801    <author>3802      <name>Rui Kong</name>3803    </author>3804    <author>3805      <name>Xinran Chen</name>3806    </author>3807    <author>3808      <name>Jiamin Chen</name>3809    </author>3810    <author>3811      <name>Hengyi Cai</name>3812    </author>3813    <author>3814      <name>Shuaiqiang Wang</name>3815    </author>3816    <author>3817      <name>Jiashu Zhao</name>3818    </author>3819    <author>3820      <name>Yulun Zhang</name>3821    </author>3822    <author>3823      <name>Zhonghao Lyu</name>3824    </author>3825    <author>3826      <name>Haoyi Xiong</name>3827    </author>3828    <author>3829      <name>Linghe Kong</name>3830    </author>3831    <author>3832      <name>Jimmy Xiangji Huang</name>3833    </author>3834    <author>3835      <name>Dawei Yin</name>3836    </author>3837  </entry>3838  <entry>3839    <id>http://arxiv.org/abs/2609.11412v1</id>3840    <title>X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation</title>3841    <updated>2026-09-10T11:43:17Z</updated>3842    <link href="https://arxiv.org/abs/2609.11412v1" rel="alternate" type="text/html"/>3843    <link href="https://arxiv.org/pdf/2609.11412v1" rel="related" type="application/pdf" title="pdf"/>3844    <summary>Reducing audio-encoder depth lowers the inference cost of speech large language models, but removing complete blocks perturbs the embeddings consumed by the decoder and can cause deletion and premature end-of-sequence errors. We introduce X-AuT, a progressive framework that selects layer combinations through short behavioral probes and restores the pruned model through representation alignment, cross-scale distillation, scheduled student-policy supervision, and LoRA finetuning. The language-model backbone remains frozen, while attention LoRA adapters and the tied output embedding adapt during distillation. Training uses the highest-agreement tier from a transcript-consistency pipeline, followed by source reweighting during finetuning. On ten public Chinese--English benchmarks, compressing Qwen3-ASR-0.6B from 18 to 16 audio-encoder layers reduces macro-average error from 5.61% to 5.27%. The 14-layer model reaches 5.75% with 20.7% fewer audio-tower parameters. Under the matched recipe, the 1.7B teacher yields 5.55% mean error, compared with 8.45% for self-distillation, and progressive 18$\rightarrow$14 pruning outperforms direct pruning (5.75% vs. 6.73%). These single-run results establish two practical operating points and show that the accuracy effects vary across benchmarks. Project website: https://xpeng-ai.github.io/x-aut</summary>3845    <category term="cs.SD" scheme="http://arxiv.org/schemas/atom"/>3846    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>3847    <published>2026-09-10T11:43:17Z</published>3848    <arxiv:primary_category term="cs.SD"/>3849    <author>3850      <name>Haojun Zhang</name>3851    </author>3852    <author>3853      <name>Yi Zou</name>3854    </author>3855    <author>3856      <name>Min Chen</name>3857    </author>3858    <author>3859      <name>Qize Yu</name>3860    </author>3861    <author>3862      <name>Lianrui Fan</name>3863    </author>3864    <author>3865      <name>Xini Ding</name>3866    </author>3867    <author>3868      <name>Hao Li</name>3869    </author>3870    <author>3871      <name>Shuchang Zhou</name>3872    </author>3873    <author>3874      <name>Xianming Liu</name>3875    </author>3876    <author>3877      <name>Shiyu Huang</name>3878    </author>3879  </entry>3880  <entry>3881    <id>http://arxiv.org/abs/2609.11404v1</id>3882    <title>Deep-Fake CAPTCHA: Mitigating Next-Generation Social Engineering Attacks</title>3883    <updated>2026-09-10T11:35:11Z</updated>3884    <link href="https://arxiv.org/abs/2609.11404v1" rel="alternate" type="text/html"/>3885    <link href="https://arxiv.org/pdf/2609.11404v1" rel="related" type="application/pdf" title="pdf"/>3886    <summary>This paper presents DF-CAPTCHA, an active defense against real-time deepfake impersonation in voice and video calls. Instead of passively searching for artifacts, DF-CAPTCHA prompts the caller to perform simple challenge-response tasks that are easy for humans but difficult for current real-time deepfake systems to generate convincingly. The framework verifies the response using four criteria: realism, identity consistency, task completion, and response time. We evaluate the approach across both audio and video modalities using user studies and experiments with real-time deepfake models. Results show that people often struggle to distinguish real-time deepfakes from authentic media, while DF-CAPTCHA substantially improves detection performance over passive methods, reaching high accuracy in both modalities. These findings suggest that active challenge-based verification is a practical and robust defense against next-generation social engineering attacks based on real-time deepfakes.</summary>3887    <category term="cs.CR" scheme="http://arxiv.org/schemas/atom"/>3888    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>3889    <published>2026-09-10T11:35:11Z</published>3890    <arxiv:comment>Expanded work from the original ASIA CCS paper on DF-CAPTCHA (now evaluates video deepfakes too)</arxiv:comment>3891    <arxiv:primary_category term="cs.CR"/>3892    <author>3893      <name>Guy Frankovits</name>3894    </author>3895    <author>3896      <name>Lior Yasur</name>3897    </author>3898    <author>3899      <name>Fred M. Grabovski</name>3900    </author>3901    <author>3902      <name>Yisroel Mirsky</name>3903    </author>3904  </entry>3905  <entry>3906    <id>http://arxiv.org/abs/2609.11403v1</id>3907    <title>From Queries to Narratives: Cultural Heritage Data Stories for Knowledge Graph Exploration and Quality Assessment</title>3908    <updated>2026-09-10T11:34:36Z</updated>3909    <link href="https://arxiv.org/abs/2609.11403v1" rel="alternate" type="text/html"/>3910    <link href="https://arxiv.org/pdf/2609.11403v1" rel="related" type="application/pdf" title="pdf"/>3911    <summary>Cultural-heritage KGs such as the NFDI4Culture-KG contain millions of triples about artworks, music, inscriptions, historical events, and the people and places connected to them. For many users, however, discovering this knowledge can be difficult. While SPARQL can be learned, writing meaningful queries first requires an in-depth understanding of the graph's data model, an investment many domain researchers and practitioners are unwilling to make. Even with existing user interfaces, a starting point and some guidance are usually needed, because the data contained in the graph is highly specialized, heterogeneous, and constantly growing, making it challenging to know what it contains or which questions it can answer. In this paper, we present data stories as a way not only to lower this barrier, but also to turn exploration into data-quality assessment, and thus combine accessible querying with the discovery of issues that remain hidden in aggregate statistics. In this contribution, a data story is understood as a narrative document that integrates explanatory text and images with executable SPARQL queries and their visualized results. It is described how they are authored against the graph and how they serve several purposes: guiding users through an unfamiliar graph, creating reproducible narratives, and surfacing data-quality issues previously hidden in aggregate statistics. The authoring platform LODEON including its Sparnatural and AI-supported authoring assistants is introduced as a proof-of-concept. Within the authoring environment, every claim made about the data can be backed by an explicit query, making these narratives transparent and reproducible. This paper also reflects on lessons learned from hands-on seminars and workshops. Early experience suggests that such data stories make cultural-heritage knowledge graphs more accessible for both exploration and quality assessment.</summary>3912    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>3913    <category term="cs.DL" scheme="http://arxiv.org/schemas/atom"/>3914    <published>2026-09-10T11:34:36Z</published>3915    <arxiv:primary_category term="cs.AI"/>3916    <author>3917      <name>Tabea Tietz</name>3918    </author>3919    <author>3920      <name>Torsten Schrade</name>3921    </author>3922    <author>3923      <name>Etienne Posthumus</name>3924    </author>3925    <author>3926      <name>Linnaea Söhn</name>3927    </author>3928    <author>3929      <name>Jonatan Jalle Steller</name>3930    </author>3931    <author>3932      <name>Jörg Waitelonis</name>3933    </author>3934    <author>3935      <name>Harald Sack</name>3936    </author>3937  </entry>3938  <entry>3939    <id>http://arxiv.org/abs/2609.11401v1</id>3940    <title>Improving the Sensitivity of Gravitational Wave Detection with Weighted Conformal Prediction</title>3941    <updated>2026-09-10T11:33:49Z</updated>3942    <link href="https://arxiv.org/abs/2609.11401v1" rel="alternate" type="text/html"/>3943    <link href="https://arxiv.org/pdf/2609.11401v1" rel="related" type="application/pdf" title="pdf"/>3944    <summary>In the last decade, kilometre-scale interferometric gravitational-wave detectors have observed hundreds of compact binary mergers, the majority of which are binary black holes. However, the data are noise-dominated, and multiple independent search algorithms (pipelines) are used to enhance sensitivity and improve robustness. Rather than the standard approach of selecting the most significant pipeline output, we combine the outputs from all pipelines using a conformal prediction-based framework to provide statistically rigorous confidence estimates for candidate events. While combining pipelines improves sensitivity and ranking robustness, it requires a principled statistical framework that remains valid as data properties evolve across observing runs. A key challenge is distribution shifts between simulated datasets used for training and calibration and the real, unlabelled, observations used for testing, which can invalidate coverage guarantees and bias confidence estimates. In this work, we address this challenge by incorporating likelihood-ratio reweighting into our conformal prediction framework to account for covariate shift. Using mock datasets containing simulated signals, we demonstrate that weighted conformal prediction restores well-calibrated coverage under covariate shift and increases the confidence of events near the detection threshold, recovering true signals that would otherwise be missed.</summary>3945    <category term="gr-qc" scheme="http://arxiv.org/schemas/atom"/>3946    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>3947    <category term="stat.ML" scheme="http://arxiv.org/schemas/atom"/>3948    <published>2026-09-10T11:33:49Z</published>3949    <arxiv:primary_category term="gr-qc"/>3950    <arxiv:journal_ref>Proceedings of the Fifteenth Symposium on Conformal and Probabilistic Prediction with Applications, PMLR 329:937-957, 2026</arxiv:journal_ref>3951    <author>3952      <name>Ann-Kristin Malz</name>3953    </author>3954    <author>3955      <name>Gregory Ashton</name>3956    </author>3957    <author>3958      <name>Nicolo Colombo</name>3959    </author>3960  </entry>3961  <entry>3962    <id>http://arxiv.org/abs/2609.11399v1</id>3963    <title>TransClean: A Benchmark for Detecting and Extracting Clean Translations from Large Language Model Outputs</title>3964    <updated>2026-09-10T11:32:18Z</updated>3965    <link href="https://arxiv.org/abs/2609.11399v1" rel="alternate" type="text/html"/>3966    <link href="https://arxiv.org/pdf/2609.11399v1" rel="related" type="application/pdf" title="pdf"/>3967    <summary>Large language models (LLMs) are increasingly used for machine translation, yet their outputs often contain additional text beyond the translation itself, such as language labels, explanations or bilingual repetitions, which we term translation noise. Despite its prevalence, this problem lacks dedicated benchmarks and systematic study. We analyze over 790,000 translation outputs from 12 LLMs across 22 language pairs (LPs) and identify 12 recurring noise patterns, which we group into formatting and content noise. Building on the observed patterns, we construct TransClean, a controlled benchmark of 9,900 pairs of noisy and clean translation outputs, comprising 8,800 synthetically generated instances and 1,100 manually curated authentic instances. We evaluate two extraction approaches on the TransClean benchmark: 1) a span-based extraction method leveraging translation quality estimation models for span detection, and 2) an LLM-based extraction method that prompts an LLM to isolate the translation. Our benchmark and analysis provide the first systematic framework to evaluate and improve the cleanliness of LLM translation outputs.</summary>3968    <category term="cs.CL" scheme="http://arxiv.org/schemas/atom"/>3969    <published>2026-09-10T11:32:18Z</published>3970    <arxiv:comment>Accepted to the Eleventh Conference on Machine Translation (WMT26)</arxiv:comment>3971    <arxiv:primary_category term="cs.CL"/>3972    <author>3973      <name>Shenbin Qian</name>3974    </author>3975    <author>3976      <name>Yves Scherrer</name>3977    </author>3978  </entry>3979  <entry>3980    <id>http://arxiv.org/abs/2609.11393v1</id>3981    <title>Beyond Confidence: Stability-Aware Test-Time Adaptation for LLM Reasoning</title>3982    <updated>2026-09-10T11:29:18Z</updated>3983    <link href="https://arxiv.org/abs/2609.11393v1" rel="alternate" type="text/html"/>3984    <link href="https://arxiv.org/pdf/2609.11393v1" rel="related" type="application/pdf" title="pdf"/>3985    <summary>Test-time adaptation has emerged as a lightweight alternative to costly post-training for improving the reasoning capabilities of Large Language Models (LLMs) on downstream tasks. Predictive entropy provides a model-derived signal for such adaptation, guiding models toward higher-confidence reasoning states without external verifiers or reward models. However, higher confidence does not necessarily imply correctness, as LLMs may remain highly confident along incorrect reasoning trajectories. We observe that high-confidence reasoning is more likely to be correct when confidence remains stable under local perturbations. Based on this observation, we propose Test-Time Adaptation via Stability-Aware Confidence Optimization (TASCO), a framework that incorporates local stability into confidence-based test-time adaptation while keeping the LLM frozen. TASCO operationalizes local stability by optimizing a lightweight task-level prefix under two alternative perturbation strategies: Random Perturbation promotes distributional stability across trajectories induced by nearby perturbed prefixes, whereas Sharpness-Aware Perturbation targets worst-case local sensitivity. Experiments demonstrate that TASCO improves reasoning accuracy and token efficiency across diverse LLMs and reasoning benchmarks, while behavioral analyses show that it maintains stable confidence under local perturbations without prematurely concentrating the model's predictive distribution.</summary>3986    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>3987    <published>2026-09-10T11:29:18Z</published>3988    <arxiv:primary_category term="cs.AI"/>3989    <author>3990      <name>Bincheng Gu</name>3991    </author>3992    <author>3993      <name>Min Gao</name>3994    </author>3995    <author>3996      <name>Zongwei Wang</name>3997    </author>3998    <author>3999      <name>Yibing Bai</name>4000    </author>4001    <author>4002      <name>Yulan He</name>4003    </author>4004    <author>4005      <name>Junliang Yu</name>4006    </author>4007  </entry>4008  <entry>4009    <id>http://arxiv.org/abs/2609.11391v1</id>4010    <title>Buyer Artificial Intelligence-Enabled Environmental Governance and Supplier Environmental Controversies: An Organizational Information Processing and Signaling</title>4011    <updated>2026-09-10T11:27:27Z</updated>4012    <link href="https://arxiv.org/abs/2609.11391v1" rel="alternate" type="text/html"/>4013    <link href="https://arxiv.org/pdf/2609.11391v1" rel="related" type="application/pdf" title="pdf"/>4014    <summary>Environmental controversies in global supply chains pose significant risks for global buyers. This study examines whether overseas suppliers' exposure to buyers' artificial intelligence (AI)-enabled environmental governance reduces supplier environmental controversies. Drawing on organizational information processing theory and signaling theory, we investigate how suppliers' exposure to AI-enabled governance influences their environmental controversies and the institutional contingencies under which this effect varies. Using text analysis to measure buyer AI-enabled environmental governance, we analyze panel data on 2,505 suppliers of U.S.-listed firms across 41 countries from 2020 to 2024 with multidimensional fixed-effects models. We find that suppliers' exposure to buyer AI-enabled environmental governance is negatively associated with supplier environmental controversies in the following year. This negative relationship is stronger in supplier countries with higher AI readiness and regulatory quality. The study contributes to research on AI-enabled sustainability governance and sustainable supply chain risk management.</summary>4015    <category term="cs.CY" scheme="http://arxiv.org/schemas/atom"/>4016    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>4017    <category term="stat.AP" scheme="http://arxiv.org/schemas/atom"/>4018    <published>2026-09-10T11:27:27Z</published>4019    <arxiv:primary_category term="cs.CY"/>4020    <author>4021      <name>Yongchao Martin Ma</name>4022    </author>4023    <author>4024      <name>Xinya Guan</name>4025    </author>4026  </entry>4027  <entry>4028    <id>http://arxiv.org/abs/2609.11390v1</id>4029    <title>VikingRAG: Accurate and Token-efficient Retrieval-augmented Generation over Structured Documents</title>4030    <updated>2026-09-10T11:25:59Z</updated>4031    <link href="https://arxiv.org/abs/2609.11390v1" rel="alternate" type="text/html"/>4032    <link href="https://arxiv.org/pdf/2609.11390v1" rel="related" type="application/pdf" title="pdf"/>4033    <summary>State-of-the-art retrieval-augmented generation (RAG) methods exploit document structures to acquire sufficient evidence, but often incur substantial token costs. To reduce structural-context tokens without compromising high RAG accuracy, we present {\sf VikingRAG}, a directory-aware semantic data management system that tightly integrates semantic and structural access to support structural-context-efficient, evidence-gap-driven multi-round retrieval. To further reduce token overhead of multi-round interaction, we materialize agentic multi-round retrieval traces as experience edges, and reuse these edges for similar queries, avoiding repeated multi-round exploration. To additionally reduce token costs when agentic multi-round retrieval is unnecessary, we introduce an adaptive escalation strategy that answers from one-round experience-augmented retrieval when the evidence is sufficient, and invokes agentic multi-round retrieval only otherwise. Experiments on real datasets show that the base system {\sf VikingRAG} matches high accuracy of state-of-the-art methods while consuming only 11.6\%--51.9\% of their tokens. With retrieval-trace reuse and adaptive escalation, token costs drop to 5.1\%--32.5\% while maintaining competitive accuracy and practical document-storage performance, showing the utility of this work for emerging AI knowledge bases.</summary>4034    <category term="cs.IR" scheme="http://arxiv.org/schemas/atom"/>4035    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>4036    <category term="cs.CL" scheme="http://arxiv.org/schemas/atom"/>4037    <category term="cs.DB" scheme="http://arxiv.org/schemas/atom"/>4038    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>4039    <published>2026-09-10T11:25:59Z</published>4040    <arxiv:primary_category term="cs.IR"/>4041    <author>4042      <name>Peiyuan Gao</name>4043    </author>4044    <author>4045      <name>Gaoyuan Zhang</name>4046    </author>4047    <author>4048      <name>Haojie Qin</name>4049    </author>4050    <author>4051      <name>Yahui Sun</name>4052    </author>4053    <author>4054      <name>Qianyi Zhang</name>4055    </author>4056    <author>4057      <name>Yunhao Zhang</name>4058    </author>4059    <author>4060      <name>Zeyu Wang</name>4061    </author>4062    <author>4063      <name>Wei Lu</name>4064    </author>4065  </entry>4066  <entry>4067    <id>http://arxiv.org/abs/2609.11381v1</id>4068    <title>Agent-Integrated Software: Interaction Contracts and Continuous Assurance</title>4069    <updated>2026-09-10T11:17:19Z</updated>4070    <link href="https://arxiv.org/abs/2609.11381v1" rel="alternate" type="text/html"/>4071    <link href="https://arxiv.org/pdf/2609.11381v1" rel="related" type="application/pdf" title="pdf"/>4072    <summary>Embedding an intelligent agent in an existing application creates a persistent coordination problem: users can revise goals and manipulate shared objects while delegated execution continues. We argue that dependable integration requires an explicit correspondence between task-level interaction and application behavior. We introduce Agent-Integrated Software (AIS) as a software pattern combining a conventional core, direct interaction, and a built-in agent, and Intent-Level Interaction Abstraction (IIA) as the task semantics through which users inspect and control delegated work. An open transition-system model relates AIS execution to IIA states and events. Interaction contracts constrain this relation through task bindings, role-specific authority, control transitions, and outcome evidence; continuous assurance maintains scoped claims as their dependencies change. A compact disclosure contract and conditional propositions illustrate why local component validity is insufficient and how selected admission invariants can be separated from planning. Contrasting software domains expose the framework's assumptions and limits. This perspective develops a research agenda spanning application abstraction, development support, controlled execution, quality assessment, and human supervision, with the aim of making agent integration a maintainable software engineering discipline.</summary>4073    <category term="cs.SE" scheme="http://arxiv.org/schemas/atom"/>4074    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>4075    <published>2026-09-10T11:17:19Z</published>4076    <arxiv:primary_category term="cs.SE"/>4077    <author>4078      <name>Shengcheng Yu</name>4079    </author>4080    <author>4081      <name>Chunrong Fang</name>4082    </author>4083    <author>4084      <name>Zhenyu Chen</name>4085    </author>4086  </entry>4087  <entry>4088    <id>http://arxiv.org/abs/2609.11380v1</id>4089    <title>DINO-Med: A Unified Patch-Based Adaptation Framework for Multi-Modal Medical Image Analysis Applied to Liver Fibrosis Staging</title>4090    <updated>2026-09-10T11:16:26Z</updated>4091    <link href="https://arxiv.org/abs/2609.11380v1" rel="alternate" type="text/html"/>4092    <link href="https://arxiv.org/pdf/2609.11380v1" rel="related" type="application/pdf" title="pdf"/>4093    <summary>Adapting natural-image foundation models like DINOv3 to multi-modal medical imaging is challenging due to the significant domain gap between natural color images and multi-channel medical scans. We present a unified, patch-based framework that processes raw multimodal imaging through training-free registration, automated localization, and mask-filtered patch extraction. This architecture culminates in a hierarchical strategy that aggregates patch-level insights into subject-level diagnostics. Using liver fibrosis staging as a case study, we evaluate four patch-level feature representations: handcrafted Radiomics features, learned ResNet features, pre-trained foundation model SAM-Med2D features, and frozen DINOv3 features. To ensure a controlled comparison, all models utilize the same lightweight MLP head and are evaluated across both rigid and deformable registration settings. Our training protocol focuses on mild fibrosis (S1) and cirrhosis (S4) classes only, enabling a single classifier to address both substantial fibrosis detection and cirrhosis staging. Evaluated via 10 random train (90%)/ test (10%) splits on 360 subjects from the CARE 2025 Liver Track 4 cohort, our DINOv3-based framework significantly outperforms all baselines, achieving the best classification accuracy of 78.4% for S1 and 75.8% for S4.</summary>4094    <category term="cs.CV" scheme="http://arxiv.org/schemas/atom"/>4095    <published>2026-09-10T11:16:26Z</published>4096    <arxiv:comment>12 pages, 2 figures. Accepted by AIiH</arxiv:comment>4097    <arxiv:primary_category term="cs.CV"/>4098    <author>4099      <name>Boya Wang</name>4100    </author>4101    <author>4102      <name>Ruizhe Li</name>4103    </author>4104    <author>4105      <name>Chao Chen</name>4106    </author>4107    <author>4108      <name>Xin Chen</name>4109    </author>4110  </entry>4111  <entry>4112    <id>http://arxiv.org/abs/2609.11378v1</id>4113    <title>Brain-PACE: A Deep Siamese MRI Framework for Modelling Longitudinal Brain Acceleration</title>4114    <updated>2026-09-10T11:14:32Z</updated>4115    <link href="https://arxiv.org/abs/2609.11378v1" rel="alternate" type="text/html"/>4116    <link href="https://arxiv.org/pdf/2609.11378v1" rel="related" type="application/pdf" title="pdf"/>4117    <summary>Brain age estimation has become a popular research proxy for assessing brain health and disease, yet longitudinal trajectories of brain ageing are still poorly defined, and clinical use is limited. Building on existing Siamese longitudinal frameworks, we develop Brain-Predicted Age Acceleration (Brain-PACE) to directly estimate the pace of structural brain ageing from paired T1-weighted MRI. Brain-PACE identified accelerated ageing in $42.6$% of participants with mild cognitive impairment. Faster Brain-PACE was associated with greater functional and cognitive impairment (FAQ; $r=0.35$, ADAS13; $r=0.30$, CDR-SB; $r=0.32$) and greater regional tau burden in the posterior cingulate ($r=0.59$), precuneus ($r=0.47$), and entorhinal cortex ($r=0.37$). These associations were stronger than those observed when pace was calculated indirectly from repeated cross-sectional brain age estimates, suggesting that direct longitudinal modelling captures complementary information relevant to ongoing pathological change. Methodologically, Brain-PACE extends the LILAC framework by combining spatial attention with soft label distribution learning and a Cramér distance objective, improving probabilistic performance and reducing prediction bias while providing measures of predictive uncertainty. Together, these findings support Brain-PACE as a complementary longitudinal imaging phenotype with sensitivity to relevant clinical and biological changes in early neurodegeneration.</summary>4118    <category term="cs.CV" scheme="http://arxiv.org/schemas/atom"/>4119    <published>2026-09-10T11:14:32Z</published>4120    <arxiv:comment>20 pages, 6 figures</arxiv:comment>4121    <arxiv:primary_category term="cs.CV"/>4122    <author>4123      <name>Samuel Maddox</name>4124      <arxiv:affiliation>School of Computing Sciences, University of East Anglia</arxiv:affiliation>4125    </author>4126    <author>4127      <name>Jacob Newman</name>4128      <arxiv:affiliation>School of Computing Sciences, University of East Anglia</arxiv:affiliation>4129    </author>4130    <author>4131      <name>Saber Sami</name>4132      <arxiv:affiliation>Norwich Medical School, University of East Anglia</arxiv:affiliation>4133    </author>4134    <author>4135      <name>Michal Mackiewicz</name>4136      <arxiv:affiliation>School of Computing Sciences, University of East Anglia</arxiv:affiliation>4137    </author>4138    <author>4139      <name>for the Alzheimer's Disease Neuroimaging Initiative</name>4140    </author>4141    <author>4142      <name>the Australian Imaging Biomarkers</name>4143    </author>4144    <author>4145      <name>Lifestyle flagship study of ageing</name>4146    </author>4147  </entry>4148  <entry>4149    <id>http://arxiv.org/abs/2609.11376v1</id>4150    <title>Deep operator learning for efficient sampling from invariant measures of stochastic differential equations</title>4151    <updated>2026-09-10T11:12:18Z</updated>4152    <link href="https://arxiv.org/abs/2609.11376v1" rel="alternate" type="text/html"/>4153    <link href="https://arxiv.org/pdf/2609.11376v1" rel="related" type="application/pdf" title="pdf"/>4154    <summary>We introduce an amortized neural sampler that combines operator learning with flow methods for sampling. It maps SDE coefficient functions to pushforwards from a reference measure to the invariant measures, enabling efficient sampling across families of stochastic differential equations. Our framework shifts traditional sampling cost to an initial training phase, after which new SDE instances require only one encoder pass and a few ODE solver steps, independent of mixing time. To handle problems in high dimensions, we use Lagrangian trajectory sensors for the coefficient functions and cross attention in the architecture. We also theoretically establish the expressivity and resolution invariance of our framework. Experiments on 1D and 2D SDE families show competitive accuracy with substantial speedups over MCMC in regimes with slow mixing, transfer across sensor counts, and demonstration results on a 64D interacting particle SDE where traditional grid approaches are infeasible.</summary>4155    <category term="math.NA" scheme="http://arxiv.org/schemas/atom"/>4156    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>4157    <published>2026-09-10T11:12:18Z</published>4158    <arxiv:primary_category term="math.NA"/>4159    <author>4160      <name>Lin Guo</name>4161    </author>4162    <author>4163      <name>Lei Li</name>4164    </author>4165    <author>4166      <name>Jingtong Zhang</name>4167    </author>4168  </entry>4169  <entry>4170    <id>http://arxiv.org/abs/2609.11375v1</id>4171    <title>Vision Transformer-Based Multi-Level Feature Fusion for Multi-Label Sewer Defect Classification</title>4172    <updated>2026-09-10T11:11:03Z</updated>4173    <link href="https://arxiv.org/abs/2609.11375v1" rel="alternate" type="text/html"/>4174    <link href="https://arxiv.org/pdf/2609.11375v1" rel="related" type="application/pdf" title="pdf"/>4175    <summary>Automated classification of sewer defects is essential for infrastructure condition assessment and maintenance decision-making, but existing deep learning methods struggle to balance classification accuracy and computational complexity in large-scale multi-label scenarios. This study develops Sewer-Transformer-ML, a hierarchical vision Transformer with multi-level feature fusion, together with two lightweight architectures, Sewer-MobileNet-ML and Sewer-Mobile-TransNet, for resource-constrained inspection scenarios. On the Sewer-ML test set, Sewer-Transformer-ML-Base achieved an $F2_{\text{CIW}}$ of 65.68% and an $F1_{\text{Normal}}$ of 92.68%, ranking first on the public leaderboard and exceeding the second-ranked method by 7.6 percentage points in $F2_{\text{CIW}}$. Sewer-MobileNet-ML achieved an $F2_{\text{CIW}}$ of 65.73% with only 17 M parameters, representing an approximately 95% parameter reduction relative to the base model. Under the standard Sewer-Capsule data split, Sewer-Mobile-TransNet achieved 96.43% classification accuracy. When the training set was reduced to 1,177 images, pretraining on Sewer-ML consistently improved model performance. Ablation experiments further showed that direct concatenation was more effective for Transformer features, whereas attention-based fusion better supported multiscale CNN features. These findings provide a computational basis for automated sewer inspection, lightweight model design, and adaptation across civil infrastructure inspection platforms.</summary>4176    <category term="cs.CV" scheme="http://arxiv.org/schemas/atom"/>4177    <published>2026-09-10T11:11:03Z</published>4178    <arxiv:primary_category term="cs.CV"/>4179    <author>4180      <name>Xu Fang</name>4181    </author>4182    <author>4183      <name>Zhuoran Wang</name>4184    </author>4185    <author>4186      <name>Qing Li</name>4187    </author>4188    <author>4189      <name>Shengyu Zhang</name>4190    </author>4191    <author>4192      <name>Guanzhi Deng</name>4193    </author>4194    <author>4195      <name>Jianbiao He</name>4196    </author>4197    <author>4198      <name>Qingquan Li</name>4199    </author>4200  </entry>4201  <entry>4202    <id>http://arxiv.org/abs/2609.11373v1</id>4203    <title>Characterizing Bluesky Content Moderation Service: From Automation of Service to Landscape of Harms</title>4204    <updated>2026-09-10T11:08:48Z</updated>4205    <link href="https://arxiv.org/abs/2609.11373v1" rel="alternate" type="text/html"/>4206    <link href="https://arxiv.org/pdf/2609.11373v1" rel="related" type="application/pdf" title="pdf"/>4207    <summary>Empirical research on content moderation is fundamentally constrained by the opaque deployment of moderation systems on major social media platforms. To this end, the recent emergence of decentralized platforms with transparent, public moderation logs presents an unprecedented opportunity for independent audits. In this work, we leverage this architectural transparency to conduct the first large-scale audit of the default moderation system on Bluesky, the Bluesky Moderation Service (BMS). Analyzing its 10.6M moderation labels from 2025, we investigate three foundational aspects: (i) its mechanism (the degree of automation versus human oversight), (ii) its efficacy (accuracy in detecting harms), and (iii) its purpose (the landscape of harms it identifies). Our findings reveal a human-AI collaborative system where labels for sexual and graphic content are applied automatically in seconds, while nuanced and high stakes labels require more human oversight, taking hours or days. Through a manual annotation study, we find the BMS operates with high precision (0.837), but struggles with low recall (0.222), with our annotators identifying 4.5$\times$ more harmful content than the moderation system in a random sample. Finally, unsupervised clustering of the most frequently applied labeled posts uncovers detected harms ranging from hostility in discourse toward protected groups to the spread of sexually explicit and other graphic content. Our work offers a look into the operational realities of a deployed moderation system, providing a concrete data-driven foundation for designing more effective and transparent moderation systems.</summary>4208    <category term="cs.CY" scheme="http://arxiv.org/schemas/atom"/>4209    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>4210    <published>2026-09-10T11:08:48Z</published>4211    <arxiv:comment>18 pages, 7 figures, 12 tables. Accepted for publication at ICWSM 2027</arxiv:comment>4212    <arxiv:primary_category term="cs.CY"/>4213    <author>4214      <name>Pushpdeep Singh</name>4215    </author>4216    <author>4217      <name>Sayeh Jarollahi</name>4218    </author>4219    <author>4220      <name>Ayan Majumdar</name>4221    </author>4222    <author>4223      <name>Vabuk Pahari</name>4224    </author>4225    <author>4226      <name>Abhijnan Chakraborty</name>4227    </author>4228    <author>4229      <name>Krishna P. Gummadi</name>4230    </author>4231    <author>4232      <name>Ingmar Weber</name>4233    </author>4234    <author>4235      <name>Abhisek Dash</name>4236    </author>4237  </entry>4238  <entry>4239    <id>http://arxiv.org/abs/2609.11372v1</id>4240    <title>RAMamba-Net: A Reliability-Aware and Mamba-Based Multimodal Fusion Network for Auditory Attention Detection</title>4241    <updated>2026-09-10T11:08:15Z</updated>4242    <link href="https://arxiv.org/abs/2609.11372v1" rel="alternate" type="text/html"/>4243    <link href="https://arxiv.org/pdf/2609.11372v1" rel="related" type="application/pdf" title="pdf"/>4244    <summary>Auditory attention decoding (AAD) identifies the attended speaker from physiological signals, supporting neuro-steered hearing devices and natural human-machine interaction. Electroencephalography (EEG) is the dominant modality for AAD but provides incomplete evidence in naturalistic audio-visual scenes, motivating EEG and electrooculography (EOG) fusion. Existing approaches remain limited by weak cross-modal interaction, inefficient temporal modeling, and low robustness to sample variations. To address the limitations, we propose RAMamba-Net, a reliability-aware Mamba-based multimodal fusion network for AAD. RAMamba-Net employs a Mamba-enhanced band-aware convolutional Transformer to capture band-specific EEG patterns and long-range temporal dynamics. A dual-branch temporal-spatial encoder models EOG temporal and inter-channel dependencies. Cross-modal attention enables explicit modality interaction. Then, a reliability-aware module is introduced to estimate sample-wise modality weights for feature and prediction consistency, thereby enhancing multimodal fusion. Experiments on two AAD benchmarks demonstrate that RAMamba-Net effectively exploits complementary EEG-EOG information, yielding accuracy gains of 5.76% over unimodal baselines, together with more robust decoding and discriminative representations. Further analyses show that explicit cross-modal interaction improves multimodal alignment, while the reliability-aware module suppresses unreliable modality evidence and is robust to signal perturbation and parameter variation.</summary>4245    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>4246    <published>2026-09-10T11:08:15Z</published>4247    <arxiv:comment>RAMamba-Net, a reliability-aware Mamba-based multimodal fusion network for auditory attention decoding</arxiv:comment>4248    <arxiv:primary_category term="cs.AI"/>4249    <author>4250      <name>Xingyi He</name>4251    </author>4252    <author>4253      <name>Ziwei Wang</name>4254    </author>4255    <author>4256      <name>Dongrui Wu</name>4257    </author>4258  </entry>4259  <entry>4260    <id>http://arxiv.org/abs/2609.11366v1</id>4261    <title>Local Robustness Quantification for Naive Bayes Classifiers and Generative Forests: a General Approach</title>4262    <updated>2026-09-10T10:55:57Z</updated>4263    <link href="https://arxiv.org/abs/2609.11366v1" rel="alternate" type="text/html"/>4264    <link href="https://arxiv.org/pdf/2609.11366v1" rel="related" type="application/pdf" title="pdf"/>4265    <summary>We provide methods for calculating the robustness of the predictions of two types of generative classifiers whose underlying distribution is a Probabilistic Graphical Model (PGM): naive Bayes classifiers and generative forests (a probabilistic extension of random forests). Following the paradigm of robustness quantification, we define the robustness of a prediction as the extent to which the distribution of the classifier can be perturbed without changing this prediction. We consider perturbations obtained by varying the local models of the PGMs within general neighborhoods and focus in particular on epsilon-contamination, total variation distance and chi-squared divergence balls. We test our methods on benchmark datasets, demonstrate that the robustness value of a prediction serves as an indicator for its trustworthiness and compare our approach with other such indicators.</summary>4266    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>4267    <published>2026-09-10T10:55:57Z</published>4268    <arxiv:primary_category term="cs.LG"/>4269    <author>4270      <name>Adrián Detavernier</name>4271    </author>4272    <author>4273      <name>Jasper De Bock</name>4274    </author>4275  </entry>4276  <entry>4277    <id>http://arxiv.org/abs/2609.11365v1</id>4278    <title>Portable Semantics, Private Dialects: Reuse and Negative Transfer in Latent Communication Between Language-Model Cells</title>4279    <updated>2026-09-10T10:53:47Z</updated>4280    <link href="https://arxiv.org/abs/2609.11365v1" rel="alternate" type="text/html"/>4281    <link href="https://arxiv.org/pdf/2609.11365v1" rel="related" type="application/pdf" title="pdf"/>4282    <summary>In shared-genome language-model societies, restricted evidence visibility favors reusable, value-indexed latent packet interfaces, whereas the sole high-performing globally visible model in the parent study learned an episode-entangled code. This companion study asks whether independently trained societies share one packet language, where strict zero-shot transfer fails, and whether inherited interface state helps or harms later learning. First, a leakage-controlled causal interoperability audit over all 30 ordered pairs of six independently trained restricted societies -- under sealed held-out structure and a preregistered raw/orthogonal/linear/nonlinear alignment ladder -- shows the six semantically similar interfaces do not form one raw language: one same-initialization pair is exactly interoperable in both directions, a second shows asymmetric partial compatibility, and all 26 cross-initialization directions fail every frozen alignment rung. Second, within the tested decomposition and a single sealed source formulation, a source-span control localizes strict zero-shot failure to interpretation and execution of the new operator instructions. Third, in a matched adaptation factorial, the globally trained communication interface acts as a severe negative-transfer prior: reinitializing only the packet reader, writer, and mouth raises final depth-three accuracy from 0.169 to 0.857. Fourth, across two restricted checkpoints and two independently frozen target streams each, inherited interfaces never exceeded fresh-interface controls by the preregistered 0.10 margin. All primary conclusions are bounded to a near-transfer 17-state setting; the negative-transfer factorial concerns one globally visible parent-cohort checkpoint, while an appendix adds a post hoc tagged-global twin case study.</summary>4283    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>4284    <published>2026-09-10T10:53:47Z</published>4285    <arxiv:comment>17 pages, 1 figure, 5 tables. Companion to arXiv:2608.20054. Code and evaluation records: https://github.com/tokenosopher/populus-evidence-partitioning ; checkpoints and fitted alignment maps: https://huggingface.co/tokenosopher/populus-evidence-partitioning-checkpoints</arxiv:comment>4286    <arxiv:primary_category term="cs.AI"/>4287    <author>4288      <name>Narcis Marincat</name>4289    </author>4290  </entry>4291  <entry>4292    <id>http://arxiv.org/abs/2609.11360v1</id>4293    <title>R4Tun: LLM-guided adaptive segmental tunnel lining segmentation in point clouds</title>4294    <updated>2026-09-10T10:46:18Z</updated>4295    <link href="https://arxiv.org/abs/2609.11360v1" rel="alternate" type="text/html"/>4296    <link href="https://arxiv.org/pdf/2609.11360v1" rel="related" type="application/pdf" title="pdf"/>4297    <summary>Automated inspection of segmental tunnel linings requires adaptive segmentation from 3D point clouds, yet expert-tuned pipelines often degrade when tunnel conditions vary. This paper presents R4Tun, a large language model (LLM)-driven adaptation framework that extends an expert-designed pipeline (SAM4Tun) with bounded parameter tuning informed by structured context: memory ($m$), state ($s$), and knowledge ($k$). Evaluated on 30 selected Seg2Tunnel subsets (13 regular, 17 complex) across three LLMs, the full $m+s+k$ design raised mean Intersection-over-Union (mIoU) from 0.18 to 0.43--0.48 and overall accuracy (OA) from 0.42 to 0.59--0.65 relative to the static SAM4Tun baseline, with the near-reference regular (staggered) subsets reaching mIoU 0.784--0.796 across LLMs. Across 270 (30 tunnels $\times$ 3 different LLMs $\times$ 3 context settings) runs, the LLMs showed similar parameter-adjustment trends (with overlapping 95\% CIs on mean gains) and consistently adjusted a shared set of critical parameters. These results support R4Tun as a controlled, label-free, cross-LLM adaptation mechanism in the tested SAM4Tun--Seg2Tunnel setting, demonstrating consistent accuracy gains; we position R4Tun as a mechanism contribution rather than a deployable final-inspection system, in which each bounded parameter change is auditable via logged rationales.</summary>4298    <category term="cs.CV" scheme="http://arxiv.org/schemas/atom"/>4299    <published>2026-09-10T10:46:18Z</published>4300    <arxiv:primary_category term="cs.CV"/>4301    <author>4302      <name>Xinghui Tao</name>4303    </author>4304    <author>4305      <name>Zehao Ye</name>4306    </author>4307    <author>4308      <name>Guangming Wang</name>4309    </author>4310    <author>4311      <name>Jelena Ninić</name>4312    </author>4313    <author>4314      <name>Brian Sheil</name>4315    </author>4316  </entry>4317  <entry>4318    <id>http://arxiv.org/abs/2609.11355v1</id>4319    <title>SEAR: Segment-Evidence-Aware Routing for Weak-to-Strong Multilingual Speech MCQ</title>4320    <updated>2026-09-10T10:38:19Z</updated>4321    <link href="https://arxiv.org/abs/2609.11355v1" rel="alternate" type="text/html"/>4322    <link href="https://arxiv.org/pdf/2609.11355v1" rel="related" type="application/pdf" title="pdf"/>4323    <summary>This paper describes our system for Task~2 of the second Multilingual Conversational Speech Language Model (MLC-SLM) Challenge. We adapt Qwen3-Omni-30B-A3B-Instruct with a segment-evidence-aware data and post-training pipeline. A language model converts timestamped ASR into coherent event spans, which are expanded by a boundary margin and cropped from the original recording. We then synthesize complementary semantic MCQs with Qwen3.6-27B and acoustic MCQs with Gemini~3.1 Flash-Lite, followed by structural, grounding, answer-consistency, and target-model trainability checks, yielding 359,825 verified MCQs across 21 language and accent variants. A text-only probe partitions the data into weak, text-answerable items used for supervised fine-tuning and strong, audio-dependent items used for reinforcement learning with Group Sequence Policy Optimization (GSPO), stabilized by debiased advantages, sequence-level importance correction, and dynamic filtering. Our system obtains 90.92% accuracy on the final official evaluation set.</summary>4324    <category term="cs.CL" scheme="http://arxiv.org/schemas/atom"/>4325    <category term="cs.SD" scheme="http://arxiv.org/schemas/atom"/>4326    <published>2026-09-10T10:38:19Z</published>4327    <arxiv:primary_category term="cs.CL"/>4328    <author>4329      <name>Huy Hoang Le</name>4330    </author>4331    <author>4332      <name>Long-Bao Nguyen</name>4333    </author>4334    <author>4335      <name>Minh Tri Dao</name>4336    </author>4337  </entry>4338  <entry>4339    <id>http://arxiv.org/abs/2609.11347v1</id>4340    <title>Reification as a Transferable Vocabulary: Zero-Shot Link Prediction with Vanilla GNNs</title>4341    <updated>2026-09-10T10:26:09Z</updated>4342    <link href="https://arxiv.org/abs/2609.11347v1" rel="alternate" type="text/html"/>4343    <link href="https://arxiv.org/pdf/2609.11347v1" rel="related" type="application/pdf" title="pdf"/>4344    <summary>Knowledge graph foundation models such as ULTRA achieve zero-shot link prediction on unseen graphs through dedicated architectures that hard-code a transfer mechanism. In this work we move that mechanism out of the architecture and into the representation, by \emph{reifying} the input graph: every fact becomes a node, connected to its subject, object, and relation type through a fixed vocabulary of six meta-relations, with relation types as anonymous shared nodes rather than model parameters. On this representation, five textbook GNNs (GAT, GINE with sum and with mean+max aggregation, GraphSAGE, R-GCN), each trained on a single knowledge graph of 4,245 triples for 30 minutes on one NVIDIA A100, transfer zero-shot to 40 inductive link-prediction benchmarks. The best of them, an off-the-shelf GAT, matches ULTRA, a dedicated foundation model pretrained on three graphs, across ULTRA's own evaluation suite. The same fixed vocabulary extends to relational databases, a row becoming an entity and a foreign-key column a relation type; a preliminary probe on two unseen databases, with no cell values, schema text or in-context labels, shows a model of this family pretrained on three knowledge graphs ranking foreign-key targets far above random-initialization and degree controls. We release the code, the checkpoints, and the evaluation pipeline for all 40 benchmarks.</summary>4345    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>4346    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>4347    <published>2026-09-10T10:26:09Z</published>4348    <arxiv:primary_category term="cs.LG"/>4349    <author>4350      <name>Camille Pradel</name>4351    </author>4352  </entry>4353  <entry>4354    <id>http://arxiv.org/abs/2609.11341v1</id>4355    <title>Exploring Diffusion Transformers for Cross-Modal Augmentation in Multimodal Brain State Decoding</title>4356    <updated>2026-09-10T10:20:10Z</updated>4357    <link href="https://arxiv.org/abs/2609.11341v1" rel="alternate" type="text/html"/>4358    <link href="https://arxiv.org/pdf/2609.11341v1" rel="related" type="application/pdf" title="pdf"/>4359    <summary>Multimodal brain state decoding has largely focused on fusing paired modalities for prediction, but has rarely explored how their correspondence can be further exploited to enrich training data and improve multimodal representation learning. To address this gap, we propose CoMA-DiT, a bidirectional cross-modal Diffusion Transformer for latent augmentation that treats paired modalities as sources of mutual generative supervision rather than merely as inputs to be fused. CoMA-DiT conditions velocity prediction on the paired modality through cross-modal attention and adaptively injects the resulting variation via a reliability-gated residual mechanism. Experiments on multimodal auditory attention decoding and emotion recognition showed that CoMA-DiT consistently outperformed 20 representative baselines, achieving absolute gains of 4.28% and 6.70% in accuracy and macro-F1 over the no-augmentation baseline, respectively. Extensive ablation, sensitivity, visualization, and interpretability analyses further demonstrated its robustness, generalizability, and ability to capture functionally relevant cross-modal interactions. These findings support a broader view of multimodal learning: Paired modalities can serve not only as inputs for fusion but also as supervision sources that augment one another.</summary>4360    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>4361    <published>2026-09-10T10:20:10Z</published>4362    <arxiv:comment>CoMA-DiT, a cross-modal augmentation framework built on Diffusion Transformer, extends multimodal learning beyond fusion by leveraging paired modalities as mutual generative supervision to enrich training data and improve brain state decoding</arxiv:comment>4363    <arxiv:primary_category term="cs.AI"/>4364    <author>4365      <name>Ziwei Wang</name>4366    </author>4367    <author>4368      <name>Xingyi He</name>4369    </author>4370    <author>4371      <name>Hongbin Wang</name>4372    </author>4373    <author>4374      <name>Tianwang Jia</name>4375    </author>4376    <author>4377      <name>Bohan Fang</name>4378    </author>4379    <author>4380      <name>Dongrui Wu</name>4381    </author>4382  </entry>4383  <entry>4384    <id>http://arxiv.org/abs/2609.11335v1</id>4385    <title>On the Impact of Anonymization on the Performance of Large Language Models</title>4386    <updated>2026-09-10T10:12:36Z</updated>4387    <link href="https://arxiv.org/abs/2609.11335v1" rel="alternate" type="text/html"/>4388    <link href="https://arxiv.org/pdf/2609.11335v1" rel="related" type="application/pdf" title="pdf"/>4389    <summary>As large language models are increasingly deployed in sensitive domains, anonymizing input data to protect personally identifiable information has become a critical practice. However, the impact of this anonymization on model utility is not well understood. This paper presents a systematic empirical study of the trade-off between privacy and performance. We evaluate five prominent language models across eleven diverse benchmarks, comparing their performance on original versus pseudonymized inputs. Our results reveal that while anonymization generally degrades performance, the effect is highly nuanced. We find that more capable models, such as Qwen2.5-72B and GPT-4o mini, suffer the largest performance drops, suggesting a stronger reliance on specific entity information. The impact is also task-dependent: performance on TruthfulQA improves with anonymization, while retrieval-focused tasks like RGB experience a catastrophic decline. Further experiments show that reversible anonymization techniques that preserve entity uniqueness significantly outperform irreversible ones like redaction, and that explicitly prompting models about anonymization offers no discernible benefit. We conclude that anonymization is not a one-size-fits-all solution and must be co-designed with the model and task in mind to balance privacy and utility effectively. Our findings provide a crucial baseline for developing more robust, privacy-aware AI systems.</summary>4390    <category term="cs.CL" scheme="http://arxiv.org/schemas/atom"/>4391    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>4392    <published>2026-09-10T10:12:36Z</published>4393    <arxiv:primary_category term="cs.CL"/>4394    <author>4395      <name>Tobias Deußer</name>4396    </author>4397    <author>4398      <name>Max Hahnbück</name>4399    </author>4400    <author>4401      <name>Lorenz Sparrenberg</name>4402    </author>4403    <author>4404      <name>Tobias Uelwer</name>4405    </author>4406    <author>4407      <name>Christian Bauckhage</name>4408    </author>4409    <author>4410      <name>Rafet Sifa</name>4411    </author>4412  </entry>4413  <entry>4414    <id>http://arxiv.org/abs/2609.11334v1</id>4415    <title>E-CONAN (Entailment, CONtradition And Neutral) Benchmarks: Arabic Textual Entailment and Natural Inference Datasets</title>4416    <updated>2026-09-10T10:11:36Z</updated>4417    <link href="https://arxiv.org/abs/2609.11334v1" rel="alternate" type="text/html"/>4418    <link href="https://arxiv.org/pdf/2609.11334v1" rel="related" type="application/pdf" title="pdf"/>4419    <summary>Natural Language Inference processes pairs of sentences to extract their semantic relations. NLI has been a hot research topic, integrated as a main component in other NLP applications. Despite significant advancements in textual inference across various languages all around the world, Arabic language still suffers from limited resources in this domain. To address this gap, this paper introduces E-CONAN benchmarks that are composed of sentences pairs from various sources: (1) automatically-translated pairs, (2) human-validated machine-translated pairs, (3) hand-crafted pairs from teaching Arabic as foreign language books, and (4) headlines pairs from different news channels containing rumors. E-CONAN contains two benchmark datasets, E-CONAN-2, a 2-way dataset (RTE) and E-CONAN-3, a 3-way dataset (NLI). Additionally, we have used E-CONAN benchmarks to evaluate 9 state-of-the-art multilingual pretrained models using zero-shot classification. Models were evaluated across the ArNLI, XNLI, and E-CONAN datasets. Results show that E-CONAN is a potentially valuable resource for evaluating model generalization and even for fine-tuning pre-trained models. Its diverse composition, derived from a combination of sources, offers a broader and more robust assessment compared to XNLI and ArNLI. In addition, we have evaluated 5 LLMs on E-CONAN-3 dataset. Moreover, we incorporated MARBERT as a representative Arabic-specific baseline and conducted performance evaluation comparison to demonstrate how Arabic-specific models scale against cross-lingual and LLM-based approaches on the E-CONAN benchmarks. Furthermore, we conducted detailed qualitative and quantitative error analysis to analyze frequent error patterns. E-CONAN benchmarks will be publicly available, we hope that it will enrich research community in Arabic textual entailment and natural language inference.</summary>4420    <category term="cs.CL" scheme="http://arxiv.org/schemas/atom"/>4421    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>4422    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>4423    <published>2026-09-10T10:11:36Z</published>4424    <arxiv:primary_category term="cs.CL"/>4425    <author>4426      <name>Khloud AL Jallad</name>4427    </author>4428    <author>4429      <name>Nada Ghneim</name>4430    </author>4431    <author>4432      <name>Ghaida Rebdawi</name>4433    </author>4434    <arxiv:doi>10.1109/ACCESS.2026.3732060</arxiv:doi>4435    <link rel="related" href="https://doi.org/10.1109/ACCESS.2026.3732060" title="doi"/>4436  </entry>4437  <entry>4438    <id>http://arxiv.org/abs/2609.11331v1</id>4439    <title>Estimating Inconsistency Response Surfaces under Uncertainty in Cyber-Physical System Development</title>4440    <updated>2026-09-10T10:03:26Z</updated>4441    <link href="https://arxiv.org/abs/2609.11331v1" rel="alternate" type="text/html"/>4442    <link href="https://arxiv.org/pdf/2609.11331v1" rel="related" type="application/pdf" title="pdf"/>4443    <summary>Cyber-Physical Systems (CPS) are commonly represented through multiple interconnected models. During development, CPS consistency requires that shared model elements remain compatible across these models. Uncertainty, for example, due to sensor noise or model abstraction, changes the admissible values of model elements and can introduce inconsistencies, i.e., situations in which models can no longer be jointly satisfied. While existing approaches can determine consistency for a given uncertainty configuration, they provide limited support for systematically exploring, analyzing, and explaining inconsistency across large uncertainty spaces. We address this challenge by reformulating inconsistency as an intervention response modeling problem. Using Saltelli sampling and multi-fidelity Monte Carlo estimation, we generate intervention-response datasets and train a surrogate model that directly predicts inconsistency from the propagated uncertainty geometry. Experiments on 48 scenarios and 10 CPS domains show that the surrogate matches Monte Carlo estimates while reducing evaluation time from milliseconds to microseconds, enabling orders-of-magnitude more response-surface evaluations within fixed computational budgets. Building on the learned response surfaces, we perform sensitivity analysis to identify dominant uncertainty drivers and introduce a gradient-based consistency recourse method to determine minimal uncertainty interventions that restore consistency. The results show that inconsistency under uncertainty can be effectively learned, analyzed, and repaired through response-surface modeling, providing a scalable foundation for uncertainty-aware consistency management in CPS development.</summary>4444    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>4445    <category term="cs.SE" scheme="http://arxiv.org/schemas/atom"/>4446    <category term="eess.SY" scheme="http://arxiv.org/schemas/atom"/>4447    <published>2026-09-10T10:03:26Z</published>4448    <arxiv:comment>Extended version of the paper accepted at IEEE ICDM 2026; 10 pages + appendix, 11 figures, 3 tables</arxiv:comment>4449    <arxiv:primary_category term="cs.LG"/>4450    <author>4451      <name>Johannes Mäkelburg</name>4452    </author>4453    <author>4454      <name>Tim Schwabe</name>4455    </author>4456    <author>4457      <name>Maribel Acosta</name>4458    </author>4459  </entry>4460  <entry>4461    <id>http://arxiv.org/abs/2609.11330v1</id>4462    <title>Predictive Multi-Landmark OCT Tracking for Increased Motion Robustness</title>4463    <updated>2026-09-10T10:01:04Z</updated>4464    <link href="https://arxiv.org/abs/2609.11330v1" rel="alternate" type="text/html"/>4465    <link href="https://arxiv.org/pdf/2609.11330v1" rel="related" type="application/pdf" title="pdf"/>4466    <summary>Optical coherence tomography is a promising modality for markerless motion tracking due to its high spatial resolution and inherent depth perception. However, existing OCT-based tracking approaches are limited in terms of trackable velocity, particularly when multiple landmarks are tracked sequentially for 6D pose estimation. In this work, we present a predictive tracking approach that propagates positional updates between multiple tracked landmarks to obtain a global pose prediction. This enables more robust tracking under high velocities. Our results demonstrate RMSEs below 1 mm for velocities up to 100 mm/s and up to nine consecutively tracked landmarks, highlighting the potential of global motion propagation and prediction for improving the robustness of OCT-based tracking.</summary>4467    <category term="cs.CV" scheme="http://arxiv.org/schemas/atom"/>4468    <published>2026-09-10T10:01:04Z</published>4469    <arxiv:comment>Accecpted at CURAC conference 2026</arxiv:comment>4470    <arxiv:primary_category term="cs.CV"/>4471    <author>4472      <name>Konrad Reuter</name>4473    </author>4474    <author>4475      <name>Suresh Guttikonda</name>4476    </author>4477    <author>4478      <name>Chaitali Uday Karekar</name>4479    </author>4480    <author>4481      <name>Christian Betz</name>4482    </author>4483    <author>4484      <name>Alexander Schlaefer</name>4485    </author>4486  </entry>4487  <entry>4488    <id>http://arxiv.org/abs/2609.11326v1</id>4489    <title>The Semantic Elevation Operator and the Closure of the Undecidable Class under Preservation</title>4490    <updated>2026-09-10T09:56:17Z</updated>4491    <link href="https://arxiv.org/abs/2609.11326v1" rel="alternate" type="text/html"/>4492    <link href="https://arxiv.org/pdf/2609.11326v1" rel="related" type="application/pdf" title="pdf"/>4493    <summary>The undecidability of a program's static semantic properties is governed by Rice's theorem. Self-modifying systems, however, require analysing not whether a property holds now, but whether it is preserved when the system rewrites itself. We formalise this transition through a semantic elevation operator ΛΦ, which turns the static question "does x satisfy P?" into the dynamic question "is P preserved after x is transformed by Φ?". We prove that when Φ is intensional (depending on the source code, not only on the computed function), the elevated property remains undecidable even though it breaks the extensionality that Rice's theorem requires; the proof rests on Kleene's recursion theorem, not on Rice. Consequently the class U of non-verifiable properties is closed under the elevation operator. Unbounded iteration of the operator climbs the arithmetical hierarchy -to Π02-completeness- consolidating non-verifiability as a structural fact. We further show that the supervisory regress does not terminate: no fnite tower of increasingly capable verifiers yields an unconditional certificate. A categorical reading of these results in the efective topos, in which elevation appears as an instance of Lawvere's fxed-point theorem, is left as a direction for future work.</summary>4494    <category term="cs.LO" scheme="http://arxiv.org/schemas/atom"/>4495    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>4496    <category term="cs.CL" scheme="http://arxiv.org/schemas/atom"/>4497    <category term="math.LO" scheme="http://arxiv.org/schemas/atom"/>4498    <published>2026-09-10T09:56:17Z</published>4499    <arxiv:primary_category term="cs.LO"/>4500    <author>4501      <name>Jose Pascual Gumbau Mezquita</name>4502    </author>4503  </entry>4504  <entry>4505    <id>http://arxiv.org/abs/2609.11322v1</id>4506    <title>MultiHuSE: A Multimodal Dataset for Humour Styles and Emotions</title>4507    <updated>2026-09-10T09:50:00Z</updated>4508    <link href="https://arxiv.org/abs/2609.11322v1" rel="alternate" type="text/html"/>4509    <link href="https://arxiv.org/pdf/2609.11322v1" rel="related" type="application/pdf" title="pdf"/>4510    <summary>Computational recognition of verbal humour remains a challenging task, requiring an understanding of language, delivery style, emotions, and cultural context. Most existing approaches focus on binary classification and lack datasets that capture psychological dimensions of humour alongside variations in expression. We introduce MultiHuSE, a multimodal dataset comprising 2,407 high-definition videos of 50 demographically diverse actors performing 1,463 text samples across four psychological humour styles (affiliative, aggressive, self-enhancing, and self-deprecating), as well as neutral content. A subset is additionally annotated for underlying emotions. The dataset uniquely captures multiple actor interpretations of the same texts, enabling systematic analysis of expressive diversity. Baseline experiments show that multimodal fusion outperforms unimodal approaches (80.1% vs. 77.4% accuracy) in humour style classification, with particularly strong gains for affiliative humour (66% to 74%). While text provides the strongest individual signal, fusion models deliver meaningful improvements. We hope that MultiHuSE provides empirical support for psychological theories linking humour and emotion, while also opening new avenues for research in human communication, well-being, and AI-driven interaction. The dataset is available for academic use under an End-User Licence Agreement.</summary>4511    <category term="cs.CL" scheme="http://arxiv.org/schemas/atom"/>4512    <category term="cs.CV" scheme="http://arxiv.org/schemas/atom"/>4513    <category term="cs.MM" scheme="http://arxiv.org/schemas/atom"/>4514    <published>2026-09-10T09:50:00Z</published>4515    <arxiv:comment>7 pages, 3 figures, 5 tables. Accepted at IEEE CBMI 2025 (International Conference on Content-Based Multimedia Indexing), Dublin, Ireland</arxiv:comment>4516    <arxiv:primary_category term="cs.CL"/>4517    <arxiv:journal_ref>2025 International Conference on Content-Based Multimedia Indexing (CBMI), Dublin, Ireland, 2025, pp. 1-7</arxiv:journal_ref>4518    <author>4519      <name>Mary Ogbuka Kenneth</name>4520    </author>4521    <author>4522      <name>Foaad Khosmood</name>4523    </author>4524    <author>4525      <name>Abbas Edalat</name>4526    </author>4527    <arxiv:doi>10.1109/CBMI66578.2025.11339313</arxiv:doi>4528    <link rel="related" href="https://doi.org/10.1109/CBMI66578.2025.11339313" title="doi"/>4529  </entry>4530  <entry>4531    <id>http://arxiv.org/abs/2609.11321v1</id>4532    <title>AI Exposure and AI Resilience: A Two-Dimensional Assessment Framework for Software and Software-Based Business Model</title>4533    <updated>2026-09-10T09:49:02Z</updated>4534    <link href="https://arxiv.org/abs/2609.11321v1" rel="alternate" type="text/html"/>4535    <link href="https://arxiv.org/pdf/2609.11321v1" rel="related" type="application/pdf" title="pdf"/>4536    <summary>Artificial intelligence is changing both software production and the economics of software-based business models. Classical technology due diligence mainly examines technical properties such as architecture, scalability, and technical debt. These criteria do not fully capture how AI can affect a company's value proposition, competitive position, margins, or access to customers. This paper develops Artificial Intelligence Exposure and Resilience (AI-ER) as a two-dimensional assessment framework. AI exposure describes the pressure for change that AI creates for a business model. AI resilience describes the company's ability to absorb that pressure, adapt to changed conditions, and use AI in an economically viable way. Metrics for both dimensions are derived from current AI capabilities, their deployment conditions, and relevant research on business models and organizational adaptability. The model keeps exposure and resilience separate and adds an explicit assessment of evidence quality and confidence. It can be applied first with public information and later refined with internal evidence. The result is a traceable company profile that supports comparison without concealing uncertainty in the underlying evidence. The paper also specifies an initial score logic and a procedure for empirical validation.</summary>4537    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>4538    <published>2026-09-10T09:49:02Z</published>4539    <arxiv:comment>14 pages, 3 figures, 6 tables. Preprint</arxiv:comment>4540    <arxiv:primary_category term="cs.AI"/>4541    <author>4542      <name>Paul Darius Mandl</name>4543      <arxiv:affiliation>Findustrial GmbH</arxiv:affiliation>4544    </author>4545    <author>4546      <name>Peter Mandl</name>4547      <arxiv:affiliation>Munich University of Applied Sciences</arxiv:affiliation>4548    </author>4549    <author>4550      <name>Martin Häusl</name>4551      <arxiv:affiliation>Munich University of Applied Sciences</arxiv:affiliation>4552    </author>4553  </entry>4554  <entry>4555    <id>http://arxiv.org/abs/2609.11319v1</id>4556    <title>Magenta: Closing the Loop Between Mathematical Reasoning and Lean Verification</title>4557    <updated>2026-09-10T09:48:07Z</updated>4558    <link href="https://arxiv.org/abs/2609.11319v1" rel="alternate" type="text/html"/>4559    <link href="https://arxiv.org/pdf/2609.11319v1" rel="related" type="application/pdf" title="pdf"/>4560    <summary>Most of mathematical knowledge has been communicated through so-called informal use of mathematics and natural language. With large language models (LLMs) being highly adept in using natural language, they achieve strong performance, yet not perfect, in informal mathematical reasoning. Restraining LLMs to informal reasoning misses out on the opportunity to use the discrete verification abilities that machines offer through machine-checkable proofs. In this paper, we bridge the gap between informal and formal reasoning by integrating Lean signals into the informal reasoning process. We introduce Magenta, a training-free agentic pipeline that, given only a natural-language problem, produces an answer, expresses it as a Lean 4 statement, and constructs a machine-checked proof. A statement judge verifies whether the formalisation preserves the original problem, while an error-attribution judge routes failed attempts either to mathematical re-derivation or local Lean repair. Magenta achieves 100% accuracy across all evaluated olympiad benchmarks, including AIME 2025, AIME 2026, and HMMT February 2026. When paired with the open-weight K2-Horizon-7B reasoner, it solves all six IMO 2026 problems. Our analysis shows that statement adjudication is essential for preventing false certificates and that feedback-guided correction outperforms independent resampling on difficult problems.</summary>4561    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>4562    <published>2026-09-10T09:48:07Z</published>4563    <arxiv:comment>9 pages, preprint</arxiv:comment>4564    <arxiv:primary_category term="cs.AI"/>4565    <author>4566      <name>Joshua Ong Jun Leang</name>4567    </author>4568    <author>4569      <name>Haonan Li</name>4570    </author>4571    <author>4572      <name>Zheng Zhao</name>4573    </author>4574    <author>4575      <name>Xinyi Shang</name>4576    </author>4577    <author>4578      <name>Wenda Li</name>4579    </author>4580    <author>4581      <name>Zhengzhong Liu</name>4582    </author>4583    <author>4584      <name>Erix Xing</name>4585    </author>4586    <author>4587      <name>Shay Cohen</name>4588    </author>4589    <author>4590      <name>Eleonora Giunchiglia</name>4591    </author>4592  </entry>4593  <entry>4594    <id>http://arxiv.org/abs/2609.11318v1</id>4595    <title>Mr.LHDR: A Benchmark for Multimodal Real-World Long-Horizon Deep Research Agents</title>4596    <updated>2026-09-10T09:47:44Z</updated>4597    <link href="https://arxiv.org/abs/2609.11318v1" rel="alternate" type="text/html"/>4598    <link href="https://arxiv.org/pdf/2609.11318v1" rel="related" type="application/pdf" title="pdf"/>4599    <summary>Deep research agents are increasingly capable of web search, tool use, multimodal evidence analysis, and information synthesis. However, existing benchmarks mainly evaluate medium-horizon exploration and rarely test whether agents can sustain long, dependency-heavy research processes. We introduce Mr.LHDR (Multimodal real-world Long-Horizon Deep Research), a benchmark for evaluating real-world deep research over long, irreducible chains of interdependent evidence across eight categories. Each question is constructed from a hidden Node-Relation graph and requires an average of 12.1 necessary intermediate conclusions with a mean dependency depth of 10.4 before reaching a short, unique, and verifiable answer. Questions incorporate multimodal evidence, including images, maps, PDFs, logos, charts, tables, and video frames, with at least one non-text element that changes the reasoning state. Mr.LHDR evaluates both final answers and the correctness of intermediate conclusions under annotated dependencies. We evaluate general models, deep research systems, and agent frameworks using Overall Accuracy (OA), Strict Accuracy (SA), Checklist Score (CS), and Dependency-Aware Checklist Score (DACS). Results show that even the strongest system achieves only 43.1% OA and 34.3% SA, indicating that final-answer accuracy substantially overestimates complete research success. Removing images reduces DACS by 12.6 points, demonstrating the importance of multimodal evidence, while SA consistently declines as reasoning chains become longer. These findings reveal sustained, dependency-consistent evidence integration, rather than isolated fact retrieval, as a key bottleneck for current deep research agents.</summary>4600    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>4601    <published>2026-09-10T09:47:44Z</published>4602    <arxiv:comment>Code and data are available at https://github.com/minghaoguo20/Mr-LHDR-eval</arxiv:comment>4603    <arxiv:primary_category term="cs.AI"/>4604    <author>4605      <name>Minghao Guo</name>4606    </author>4607    <author>4608      <name>Meng Cao</name>4609    </author>4610    <author>4611      <name>Sui Zhao</name>4612    </author>4613    <author>4614      <name>Siyu Ning</name>4615    </author>4616    <author>4617      <name>Xin Wang</name>4618    </author>4619    <author>4620      <name>Haoze Zhao</name>4621    </author>4622    <author>4623      <name>Jiaxuan Yang</name>4624    </author>4625    <author>4626      <name>Haihong Hao</name>4627    </author>4628    <author>4629      <name>Mingfei Han</name>4630    </author>4631    <author>4632      <name>Shunlin Rong</name>4633    </author>4634    <author>4635      <name>Haijun Wu</name>4636    </author>4637    <author>4638      <name>Xiaodan Liang</name>4639    </author>4640    <author>4641      <name>Xiaojun Chang</name>4642    </author>4643  </entry>4644  <entry>4645    <id>http://arxiv.org/abs/2609.11317v1</id>4646    <title>Mi-Ripple: Restoring Images Degraded by Iterative AI Editing</title>4647    <updated>2026-09-10T09:46:42Z</updated>4648    <link href="https://arxiv.org/abs/2609.11317v1" rel="alternate" type="text/html"/>4649    <link href="https://arxiv.org/pdf/2609.11317v1" rel="related" type="application/pdf" title="pdf"/>4650    <summary>Iterative reference-conditioned image editing can introduce grid-like and granular textures, commonly described as digital ripple. We present Mi-Ripple, a diagnosis-guided restoration workflow that suppresses this digital ripple while protecting image structure. Mi-Ripple separates periodic lattice artifacts from content-entangled granular texture, then combines selective spectral notching, structure-aware smoothing, and cleaned-reference regeneration. This separation enables low-distortion filtering when artifacts are spectrally isolated and visual reconstruction when filtering would erase legitimate detail. Across fourteen notch-only executions, whole-image residual standard deviation is 0.08--0.44 in CIELAB lightness units. In a paired regeneration example, reference cleaning reduces output debris density by 45\%. Mi-Ripple links measurable artifact reduction to visibly cleaner generated images, rather than optimizing a spectral score alone.</summary>4651    <category term="cs.CV" scheme="http://arxiv.org/schemas/atom"/>4652    <published>2026-09-10T09:46:42Z</published>4653    <arxiv:primary_category term="cs.CV"/>4654    <author>4655      <name>Jiayin Chen</name>4656    </author>4657    <author>4658      <name>Yicheng Xu</name>4659    </author>4660    <author>4661      <name>Muting Wang</name>4662    </author>4663  </entry>4664  <entry>4665    <id>http://arxiv.org/abs/2609.11315v1</id>4666    <title>Routing by Reasoning Need: Trajectory-Aware Decoding Control for Diffusion Vision-Language Models</title>4667    <updated>2026-09-10T09:41:19Z</updated>4668    <link href="https://arxiv.org/abs/2609.11315v1" rel="alternate" type="text/html"/>4669    <link href="https://arxiv.org/pdf/2609.11315v1" rel="related" type="application/pdf" title="pdf"/>4670    <summary>Diffusion vision-language models generate answers through iterative refinement, exposing intermediate answer trajectories that can be inspected and controlled at inference time. However, this controllability creates a reasoning-need mismatch, where a universal generation length is applied to questions with different reasoning demands. Visually closed questions may be harmed by continued refinement after a stable answer has formed, whereas reasoning-sensitive questions may be harmed by premature commitment. We formulate this problem as reasoning-budget mismatch and study it in LLaDA-V. Rather than choosing a universal generation length, our training-free controller routes each example to early commitment, baseline preservation, or reasoning-supportive decoding using trajectory signals from answer closure, commitment evidence, and representation revision pressure, without using ground-truth answers. Across answer-focused, mixed-reasoning, and CoT-sensitive benchmarks, routed control improves robustness over fixed long decoding, pure short decoding, and single-rule interventions. The gains are not explained by shorter outputs alone. Answer-closed examples often benefit from commitment, whereas CoT-sensitive examples require preserving or supporting intermediate reasoning. Taken together, these results suggest diffusion VLM decoding should route inference-time control by the state suggested by the observed trajectory instead of relying on a universal decoding length.</summary>4671    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>4672    <published>2026-09-10T09:41:19Z</published>4673    <arxiv:comment>17 pages, 9 figures. Accepted to Findings of EMNLP 2026</arxiv:comment>4674    <arxiv:primary_category term="cs.AI"/>4675    <author>4676      <name>Yixiang Liu</name>4677    </author>4678    <author>4679      <name>Zhongxing Xu</name>4680    </author>4681    <author>4682      <name>Zhonghua Wang</name>4683    </author>4684    <author>4685      <name>Xiaoying Tang</name>4686    </author>4687  </entry>4688  <entry>4689    <id>http://arxiv.org/abs/2609.11314v1</id>4690    <title>A Dynamic Fusion Large Language Model for Traffic Flow Prediction</title>4691    <updated>2026-09-10T09:40:21Z</updated>4692    <link href="https://arxiv.org/abs/2609.11314v1" rel="alternate" type="text/html"/>4693    <link href="https://arxiv.org/pdf/2609.11314v1" rel="related" type="application/pdf" title="pdf"/>4694    <summary>Traffic flow prediction is a core supporting technology for intelligent transportation systems. It uses historical data to infer future traffic dynamics in specific areas, thereby helping to alleviate congestion and improve resource allocation efficiency. Traditional neural networks struggle to break through accuracy limits due to their reliance on singular feature modeling, while large language models (LLMs) suffer from insufficient capture of spatial topological information and mining spatiotemporal correlation. This study proposes a Dynamic Fusion Large Language Model (DF-LLM) for traffic flow prediction. The model incorporates three core components: spatiotemporal embedding module, spatiotemporal fusion module, and LLM backbone. The spatiotemporal embedding module enables synergistic representation of multi-scale spatiotemporal features. The spatiotemporal fusion module integrates spatial topology and dynamic dependencies via graph convolution. The LLM backbone adopts a differentiated parameter adaptation strategy to balance training efficiency and traffic data adaptability. Additionally, it introduces a context aggregation attention module to strengthens global dependencies. More importantly, the LLM backbone takes the residual connections to mitigate the gradient vanishing in deep networks. Experiments show that DF-LLM has achieved better performance by comparing the metrics on all the four datasets.</summary>4695    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>4696    <published>2026-09-10T09:40:21Z</published>4697    <arxiv:comment>Accepted by WISA 2026</arxiv:comment>4698    <arxiv:primary_category term="cs.LG"/>4699    <author>4700      <name>Xue Qiu</name>4701    </author>4702    <author>4703      <name>Jianli Xiao</name>4704    </author>4705  </entry>4706  <entry>4707    <id>http://arxiv.org/abs/2609.11312v1</id>4708    <title>GRIPNet: Gaussian Radial Intensity Prior Guided Architecture for Pulmonary Nodule Detection in CT</title>4709    <updated>2026-09-10T09:38:44Z</updated>4710    <link href="https://arxiv.org/abs/2609.11312v1" rel="alternate" type="text/html"/>4711    <link href="https://arxiv.org/pdf/2609.11312v1" rel="related" type="application/pdf" title="pdf"/>4712    <summary>Lung cancer causes more deaths than any other malignancy, and low-dose CT screening is the main pathway to early diagnosis. That pathway hinges on the smallest lesions, yet nodules below six millimeters remain hard to detect, because most methods treat a nodule as a generic object and ignore the imaging physics behind its appearance. We show that this appearance is highly regular. Intensity peaks at the geometric center of a nodule and decays radially in a Gaussian pattern, and a fit to 18,218 annotated lesions from three public benchmarks yields a mean radial coefficient of determination above 0.86 in every dataset and size stratum. A square convolution samples both axes uniformly and is mismatched to this radial signal, most severely for small nodules. Guided by this evidence, we propose GRIPNet (Gaussian Radial Intensity Prior Network), a detector in which every module maps to a measurable property of the intensity distribution. Pinwheel convolutions decompose radial gradients, a dual-frequency module separates boundary detail from structural context, dilated masked attention matches the decay extent, and an adaptive loss reweights samples by conspicuity. GRIPNet raises mAP@0.5 to 95.3, 91.6 and 97.9 percent on KanserSet, LUNA16 and Lung-PET-CT-Dx while sharpening high-IoU localization at real-time speed.</summary>4713    <category term="cs.CV" scheme="http://arxiv.org/schemas/atom"/>4714    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>4715    <published>2026-09-10T09:38:44Z</published>4716    <arxiv:primary_category term="cs.CV"/>4717    <author>4718      <name>Haojie Yang</name>4719    </author>4720    <author>4721      <name>Ran Su</name>4722    </author>4723  </entry>4724  <entry>4725    <id>http://arxiv.org/abs/2609.11310v1</id>4726    <title>Your Model Already Knows Don't Teach It, Learn to Ask It: Soft Prompting for Few-Shot Adaptation of Vision-Language Models</title>4727    <updated>2026-09-10T09:38:11Z</updated>4728    <link href="https://arxiv.org/abs/2609.11310v1" rel="alternate" type="text/html"/>4729    <link href="https://arxiv.org/pdf/2609.11310v1" rel="related" type="application/pdf" title="pdf"/>4730    <summary>We address few-shot object detection with vision-language models (VLMs) in out-of-domain settings such as aerial, industrial, and medical imagery, using only ten annotated images for supervision. Existing adaptation methods are discrete prompt optimization and LoRA fine-tuning. We revisit a third option: soft prompting, where a small number of continuous prompt tokens are optimized while the pretrained backbone remains frozen.4731  We identify two key design choices. First, placing prompt tokens at the cross-modal boundary between visual and text tokens outperforms other placements (10.0 vs. 8.4 mAP). Second, initializing prompts from the empty space token outperforms semantic and random initialization.4732  With these choices, one to three learned tokens (7,168 parameters on average) match the best LoRA configuration on Roboflow20-VL (14.2 mAP, 10-shot) while training over 20,000x fewer parameters. Soft prompting remains harder to optimize, exhibiting higher variance across random seeds. Unlike LoRA, however, it causes no forgetting: the LoRA rank matching our accuracy reduces NaturalBench VQA accuracy by 35% relative, rising to 56% at the largest rank, whereas soft prompting leaves pretrained performance unchanged.4733  The learned tokens behave like prompts rather than weights. They transfer to a newer model without retraining (+0.8 mAP on Qwen3.5-9B) and can be verbalized into readable prompts competitive with prompt-search methods (matching DetPO and outperforming GEPA).4734  The approach also extends beyond detection. On RoboCasa manipulation tasks, the frozen $π_{0.5}$ vision-language-action policy benefits from soft prompting, matching the LoRA baseline on two of three tasks when tokens are placed at the gradient bottleneck. These results suggest modern VLMs already encode much of what is needed for specialized domains; the challenge is learning how to ask.</summary>4735    <category term="cs.CV" scheme="http://arxiv.org/schemas/atom"/>4736    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>4737    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>4738    <category term="eess.IV" scheme="http://arxiv.org/schemas/atom"/>4739    <category term="stat.ML" scheme="http://arxiv.org/schemas/atom"/>4740    <published>2026-09-10T09:38:11Z</published>4741    <arxiv:primary_category term="cs.CV"/>4742    <author>4743      <name>Gautam Rajendrakumar Gare</name>4744    </author>4745    <author>4746      <name>Siyi Li</name>4747    </author>4748    <author>4749      <name>Hewei Wang</name>4750    </author>4751    <author>4752      <name>Cesar Daniel Hernandez</name>4753    </author>4754    <author>4755      <name>Wei Zhao</name>4756    </author>4757    <author>4758      <name>Wolfgang M. Pauli</name>4759    </author>4760    <author>4761      <name>John Galeotti</name>4762    </author>4763    <author>4764      <name>Deva Ramanan</name>4765    </author>4766  </entry>4767  <entry>4768    <id>http://arxiv.org/abs/2609.11308v1</id>4769    <title>2AM: Grounding Agent-Side Memory as Guidance for Steerable Action Models in Long-Horizon Manipulation</title>4770    <updated>2026-09-10T09:35:56Z</updated>4771    <link href="https://arxiv.org/abs/2609.11308v1" rel="alternate" type="text/html"/>4772    <link href="https://arxiv.org/pdf/2609.11308v1" rel="related" type="application/pdf" title="pdf"/>4773    <summary>Long-horizon robot manipulation requires memory, but not necessarily inside the action policy. To address such tasks, current agentic systems often combine VLAs with planners and geometric tools, sometimes using additional depth or calibrated geometry. These systems confound attribution: gains may come from richer observations or alternative motor tools, while failures may stem from either the policy or an under-specified language interface. We isolate this question through a deliberately constrained design: less tool breadth, but greater interface bandwidth. 2AM makes a multimodal Agent the sole holder of task memory and a single RGB-based, episodically stateless Action Model the sole executor of task-relevant motion. The Agent compiles interaction history into subtask language and optional 2D grasp, place, and move hints that bind its physical intention at different time scales. To teach this steerability to the VLA, we augment demonstrations with structured hint labels and train under condition dropout, spatial noise, and temporal jitter to tolerate imperfect Agent outputs. On LIBERO-Mem, without depth, online geometry, or planner-based object motion, 2AM reaches 76.3% average completion, a 61.5-point improvement over the strongest reported baseline of 14.8%, together with 63.0% relaxed and 11.8% strict success. These results show that task memory can remain Agent-side. They further show that Action Model capability depends not only on what the policy has learned, but on how precisely the Agent can steer it.</summary>4774    <category term="cs.RO" scheme="http://arxiv.org/schemas/atom"/>4775    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>4776    <published>2026-09-10T09:35:56Z</published>4777    <arxiv:primary_category term="cs.RO"/>4778    <author>4779      <name>Yutong Hu</name>4780    </author>4781    <author>4782      <name>Fengjiao Chen</name>4783    </author>4784    <author>4785      <name>Xuezhi Cao</name>4786    </author>4787    <author>4788      <name>Renaud Detry</name>4789    </author>4790  </entry>4791  <entry>4792    <id>http://arxiv.org/abs/2609.11302v1</id>4793    <title>Automatic Lyric Transcription for Greek Songs: Scaling and Task Composition Effects in Whisper Adaptation</title>4794    <updated>2026-09-10T09:30:59Z</updated>4795    <link href="https://arxiv.org/abs/2609.11302v1" rel="alternate" type="text/html"/>4796    <link href="https://arxiv.org/pdf/2609.11302v1" rel="related" type="application/pdf" title="pdf"/>4797    <summary>Automatic Lyric Transcription (ALT) remains substantially more challenging than speech recognition due to melodic variability, rhythmic irregularity, and accompaniment interference. This is heightened in low-resource languages like Greek, where no prior benchmark for ALT exists. We present the first controlled study of Whisper adaptation for Greek ALT, investigating model scaling effects, task composition via multitask training in transcribe-translate ratios, and two-stage speech-to-singing adaptation. We also curate a segment-level aligned singing dataset based on the Greek Audio Dataset (GAD) using source separation and CTC forced alignment. Results show that scaling consistently improves performance, while multitask learning acts as a beneficial regularizer primarily for smaller-capacity models. The 2-stage adaptation in Whisper Large-v3 achieves a Word Error Rate (WER) of 27.2%, a significant improvement over zero-shot baselines, establishing the first Greek ALT benchmark.</summary>4798    <category term="cs.CL" scheme="http://arxiv.org/schemas/atom"/>4799    <category term="cs.SD" scheme="http://arxiv.org/schemas/atom"/>4800    <published>2026-09-10T09:30:59Z</published>4801    <arxiv:comment>Accepted at Interspeech 2026</arxiv:comment>4802    <arxiv:primary_category term="cs.CL"/>4803    <author>4804      <name>Maria Frangiadaki</name>4805    </author>4806    <author>4807      <name>Dimitrios Damianos</name>4808    </author>4809    <author>4810      <name>Kosmas Kritsis</name>4811    </author>4812    <author>4813      <name>Vassilis Katsouros</name>4814    </author>4815  </entry>4816  <entry>4817    <id>http://arxiv.org/abs/2609.11299v1</id>4818    <title>A Two-Mirror Faceted Projection System for EUV Lithography</title>4819    <updated>2026-09-10T09:28:02Z</updated>4820    <link href="https://arxiv.org/abs/2609.11299v1" rel="alternate" type="text/html"/>4821    <link href="https://arxiv.org/pdf/2609.11299v1" rel="related" type="application/pdf" title="pdf"/>4822    <summary>We propose an all-reflective two-mirror projection system for extreme ultraviolet (EUV) lithography operating at exposure wavelengths of $13.5$~nm (Mo/Si) and $11.2$~nm (Ru/Be), delivering a fourfold ($4\times$) demagnification of the periodic mask pattern at a numerical aperture approaching unity ($\mathrm{NA}_{\max} \approx 0.993$). In contrast to conventional EUV projection objectives that incorporate 6--10 aspheric mirrors with an overall optical throughput of less than $15\%$, the proposed design redirects each accepted discrete spatial diffraction order scattered by the mask onto the wafer via a dedicated pair of planar mirror facets. The number of reflections is strictly fixed at two for all accepted orders, retaining $50$--$60\%$ of the power leaving the mask in each accepted order. We derive a spatial geometry providing rigorous optical path length equalization across all diffraction orders, thereby removing order-dependent propagation phase shifts. Individually optimized 30-bilayer Bragg multilayer coatings are designed for each facet using the transfer matrix method combined with global evolutionary optimization algorithms. The architecture is generalized to a three-dimensional vector formulation with a two-dimensionally periodic mask. Utilizing inverse lithography technology, Fourier parameterization, and a differentiable electromagnetic modal waveguide solver, we solve the synthesis problem for binary absorber masks (La absorber on a Ru/Be/Sr multilayer mirror). We demonstrate simulated aerial images of sub-10-nm features on the wafer (isolated peaks with a full width at half maximum (FWHM) of approximately $5.4$~nm and line pairs with a critical dimension of $6$~nm) and find that the two peaks remain resolved for the tested wafer defocus values from $0$ to $5$~nm along the $z$-axis.</summary>4823    <category term="physics.optics" scheme="http://arxiv.org/schemas/atom"/>4824    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>4825    <category term="physics.app-ph" scheme="http://arxiv.org/schemas/atom"/>4826    <category term="physics.class-ph" scheme="http://arxiv.org/schemas/atom"/>4827    <category term="physics.comp-ph" scheme="http://arxiv.org/schemas/atom"/>4828    <published>2026-09-10T09:28:02Z</published>4829    <arxiv:primary_category term="physics.optics"/>4830    <author>4831      <name>Vasiliy A. Es'kin</name>4832    </author>4833    <author>4834      <name>Egor V. Ivanov</name>4835    </author>4836    <author>4837      <name>Olga V. Martynova</name>4838    </author>4839  </entry>4840  <entry>4841    <id>http://arxiv.org/abs/2609.11295v1</id>4842    <title>A Hilbert-Valued Functional Decomposition Framework for Explaining Time-Dependent Outputs</title>4843    <updated>2026-09-10T09:25:43Z</updated>4844    <link href="https://arxiv.org/abs/2609.11295v1" rel="alternate" type="text/html"/>4845    <link href="https://arxiv.org/pdf/2609.11295v1" rel="related" type="application/pdf" title="pdf"/>4846    <summary>Feature-based explanations quantify features' influence on model predictions, but are primarily designed for scalar outputs. In many applications, however, outputs are functional or multivariate, such as time-dependent trajectories in demand forecasting. Consequently, existing approaches typically explain each output location independently, ignoring dependencies across the output components. We address this limitation by developing a unified framework for feature-based explanations of time-dependent outputs. Specifically, we generalize functional decomposition to Hilbert-valued prediction functions and extend an existing feature-based explanation framework to this setting. Our framework introduces kernel-based output representations that enable time-dependency-aware explanations at multiple levels of temporal granularity, including time-specific, time-resolved, and time-aggregated, while providing a unified view in which existing methods arise as special cases. We validate our framework on synthetic and real-world data, including intraday financial market volatility prediction and energy demand forecasting.</summary>4847    <category term="stat.ML" scheme="http://arxiv.org/schemas/atom"/>4848    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>4849    <published>2026-09-10T09:25:43Z</published>4850    <arxiv:primary_category term="stat.ML"/>4851    <author>4852      <name>Sophie Hanna Langbein</name>4853    </author>4854    <author>4855      <name>Niklas Koenen</name>4856    </author>4857    <author>4858      <name>Marvin N. Wright</name>4859    </author>4860    <author>4861      <name>Julia Herbinger</name>4862    </author>4863  </entry>4864  <entry>4865    <id>http://arxiv.org/abs/2609.11294v1</id>4866    <title>Memory Compression for High-Fanout Agent Sandboxes</title>4867    <updated>2026-09-10T09:25:43Z</updated>4868    <link href="https://arxiv.org/abs/2609.11294v1" rel="alternate" type="text/html"/>4869    <link href="https://arxiv.org/pdf/2609.11294v1" rel="related" type="application/pdf" title="pdf"/>4870    <summary>High-fanout agent workloads create a growing memory bottleneck because a single task may spawn many concurrent sandbox sessions. Yet these sandboxes are far from independent: they originate from a shared template and execute related trajectories, exposing substantial template-relative and cross-sandbox memory redundancy. Conventional memory compression is poorly matched to this setting in three fundamental dimensions: how to compress, because they fail to exploit similarity across non-identical sandbox pages; what to compress, because they control page-fault overhead through conservative page selection; and when to compress, because compression is either triggered by memory pressure or performed without awareness of agent execution phases.4871  We present AgentZip, the first memory compression system designed specifically for AI-agent sandboxes. AgentZip introduces compression mechanisms that exploit both the template-relative and cross-sandbox redundancy. It broadens the compression scope to any page with a profitable representation and shifts overhead control from compression-time page selection to restore-time prefetching. It further aligns expensive compression with LLM waiting periods to avoid interfering with foreground tool execution. Across LLM training and inference workloads, AgentZip reduces sandbox-owned memory by up to 8.7x, compared with 2.1x for the Linux configuration. Restore prefetching and agent-execution-aware scheduling reduce the slowdown of aggressive compression from as high as 3.1x to 1.40x while retaining nearly all of its memory-saving benefit.</summary>4872    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>4873    <category term="cs.OS" scheme="http://arxiv.org/schemas/atom"/>4874    <published>2026-09-10T09:25:43Z</published>4875    <arxiv:primary_category term="cs.AI"/>4876    <author>4877      <name>Mengming Li</name>4878    </author>4879    <author>4880      <name>Ceyu XU</name>4881    </author>4882    <author>4883      <name>Qijun Zhang</name>4884    </author>4885    <author>4886      <name>Jiangnan Yu</name>4887    </author>4888    <author>4889      <name>Xiangfeng Sun</name>4890    </author>4891    <author>4892      <name>Haohui Mai</name>4893    </author>4894    <author>4895      <name>Zhiyao Xie</name>4896    </author>4897  </entry>4898  <entry>4899    <id>http://arxiv.org/abs/2609.11291v1</id>4900    <title>Off-Target Effects of Response-Style Alignment in a Korean 27B Language Model</title>4901    <updated>2026-09-10T09:23:04Z</updated>4902    <link href="https://arxiv.org/abs/2609.11291v1" rel="alternate" type="text/html"/>4903    <link href="https://arxiv.org/pdf/2609.11291v1" rel="related" type="application/pdf" title="pdf"/>4904    <summary>We post-train Qwen3.8-27B for Korean response style -- verbosity, list and markdown usage, discourse structure and register -- and measure two behaviours the objective never targets: abstention on ambiguous social questions in KoBBQ, where the benchmark-correct answer is UNKNOWN, and unprompted disclosure in securities guidance. Both move, and the changes are expressed primarily through the model's emission policy: how often it answers and how much it says.4905  Matched target-form controls show that answer propensity depends on the training target, not the prompt set or recipe alone. Holding prompts, recipe, data volume and serving fixed and changing only the target text, three style seeds give positive answer-rate point estimates (mean +0.82 pp) and three neutral seeds negative ones (mean -1.53 pp); the observed seed ranges do not overlap and the means differ by 2.34 pp. A length-matched arm lies between them, and a fourth arm that stays short while preserving hedging is unstable across seeds, so which feature of the form is responsible is unresolved.4906  For absolute stereotyped exposure the decomposition into an answer-propensity term and a conditional-composition term is an algebraic identity, not a finding; its empirical content is where the movement went. Across the trained checkpoints the changes are dominated by answer propensity while the composition term stays small, and because that term is evaluated on treatment-dependent answered subsets we do not read it as evidence about latent preference.4907  Two measurement results follow. A between-arm contrast in conditional stereotyped share does not identify a change in conditional content preference when answer status is treatment-dependent. And agreement between two rule detectors for the same construct runs from 0.44 to 0.99 depending on which checkpoint produced the text -- observable without any reference labels.</summary>4908    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>4909    <published>2026-09-10T09:23:04Z</published>4910    <arxiv:comment>19 pages. Korean-language evaluation (KoBBQ); all uncertainty estimates over KoBBQ items are clustered on the benchmark template</arxiv:comment>4911    <arxiv:primary_category term="cs.AI"/>4912    <author>4913      <name>Hyojung Han</name>4914    </author>4915  </entry>4916  <entry>4917    <id>http://arxiv.org/abs/2609.11286v1</id>4918    <title>Generating a Consistent Enterprise: Synthesis and Reference-Free Evaluation of Multi-System Business Data</title>4919    <updated>2026-09-10T09:17:21Z</updated>4920    <link href="https://arxiv.org/abs/2609.11286v1" rel="alternate" type="text/html"/>4921    <link href="https://arxiv.org/pdf/2609.11286v1" rel="related" type="application/pdf" title="pdf"/>4922    <summary>Synthetic relational data is normally produced by a model trained on a real dataset, and its quality is measured as the distance to that dataset. This paper describes a generator that has no real dataset at either end. Given an industry, a company size, a business model, a set of business applications, and a random seed, it produces a complete fictional enterprise: a workforce, a customer base, sales deals, support tickets, recorded calls, chat messages, and documents, all consistent with one another. One entity graph is projected into the native formats of 66 business products, so the same customer appears in the CRM, the support desk, and the call system under one identity. Because no real counterpart exists, realism is built in from cited reference statistics and verified by reference-free measurement: a five-axis scorecard of 28 statistical checks, an adversarial detector that hunts for the marks of synthetic generation, and a set of soundness checks that include a classifier test against an independently shuffled copy of the data. Because these instruments existed before the generator was tuned, progress is measured under a fixed yardstick: over 23 generated companies, mean realism climbed from 60.3 to 99.1, the weakest company from 41.1 to 94.9, and the detector, which initially flagged 55.2% of all records, now flags none. The scores hold on a seed never used during development. A second generator builds relational databases from a list of business questions. It forces qualifying rows for each answerable question, adds controlled near misses, and computes exact labels from the finished tables. The generator runs as a hosted service at https://console.era.eon.io. A company built there to a specification is served through its simulators over MCP and REST, and the simulators are also published as container images for offline use</summary>4923    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>4924    <published>2026-09-10T09:17:21Z</published>4925    <arxiv:comment>10 pages</arxiv:comment>4926    <arxiv:primary_category term="cs.AI"/>4927    <author>4928      <name>Benjamin Gruenbaum</name>4929    </author>4930    <author>4931      <name>Doron Porat</name>4932    </author>4933    <author>4934      <name>Assaf Natanzon</name>4935    </author>4936    <author>4937      <name>Roy Zavida</name>4938    </author>4939    <author>4940      <name>Chen Dinachi</name>4941    </author>4942    <author>4943      <name>Or Itzahary</name>4944    </author>4945    <author>4946      <name>Omer Niv</name>4947    </author>4948  </entry>4949  <entry>4950    <id>http://arxiv.org/abs/2609.11282v1</id>4951    <title>When Does Text Inform? Benchmarking Information-Theoretic Metrics for Multimodal Time-Series Forecasting</title>4952    <updated>2026-09-10T09:15:47Z</updated>4953    <link href="https://arxiv.org/abs/2609.11282v1" rel="alternate" type="text/html"/>4954    <link href="https://arxiv.org/pdf/2609.11282v1" rel="related" type="application/pdf" title="pdf"/>4955    <summary>Multimodal forecasting models that combine time series with text annotations promise richer prediction through textual context, but how do we know whether a text annotation meaningfully contributes to the forecasters prediction? This is an information-theoretic question, but to evaluate whether information-theoretic metrics can reliably measure the predictive value an annotation provides, a ground truth benchmark is needed, and none currently exist. We create a synthetic time series signal with annotations in three categories: semantically correct, incorrect, and irrelevant. Because the data generation process is fully controlled, ground-truth information content is known exactly, enabling principled evaluation of six complementary mutual information estimators (KSG, MINE, InfoNCE, CCA, PID and V-information). We show that all six estimators identify correct annotations as most informative, and are able to audit the quality of mixed text corpora, choosing the annotations that result in the best downstream forecasting results without the need for model training. Our benchmark identifies limitations of each estimator, and these are validated on seven real-world datasets, which show how estimator performance differs on weak signals. Finally, we establish practical rules for implementing these metrics for annotation auditing and fusion selection.</summary>4956    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>4957    <category term="cs.IT" scheme="http://arxiv.org/schemas/atom"/>4958    <published>2026-09-10T09:15:47Z</published>4959    <arxiv:primary_category term="cs.AI"/>4960    <author>4961      <name>Emma Andrews</name>4962    </author>4963    <author>4964      <name>Gianmarco Mengaldo</name>4965    </author>4966  </entry>4967  <entry>4968    <id>http://arxiv.org/abs/2609.11281v1</id>4969    <title>Bio-inspired Learning and Decision-Making with Probabilistic In-Memory Computing Hardware: Part 1</title>4970    <updated>2026-09-10T09:13:38Z</updated>4971    <link href="https://arxiv.org/abs/2609.11281v1" rel="alternate" type="text/html"/>4972    <link href="https://arxiv.org/pdf/2609.11281v1" rel="related" type="application/pdf" title="pdf"/>4973    <summary>Learning and decision-making in animals are often modeled as Bayesian processes, where sensory evidence is integrated with prior beliefs to guide behavior in the face of uncertainty. But what are the inherent neural dynamics that give rise to this ability, and how could they be replicated in computing systems? This abstract discusses a biologically grounded framework in which noisy neural and synaptic dynamics perform inference and learning via stochastic sampling from an internal energy function, capturing uncertainty over latent states and model parameters through neural and synaptic variability, respectively. This enables approaches such as predictive coding networks to account for epistemic uncertainty via Markov chain Monte Carlo sampling. Drawing a parallel between intrinsic noise in biological systems and electrical noise in emerging probabilistic analogue memory technologies, we highlight how analogue in-memory computing hardware naturally emerges as the solution for massively scalable and energy-efficient probabilistic inference.</summary>4974    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>4975    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>4976    <published>2026-09-10T09:13:38Z</published>4977    <arxiv:primary_category term="cs.AI"/>4978    <author>4979      <name>Thomas Dalgaty</name>4980    </author>4981    <author>4982      <name>Eiji Kawasaki</name>4983    </author>4984    <author>4985      <name>Miguel de Prado</name>4986    </author>4987    <author>4988      <name>Devendra Vyas</name>4989    </author>4990    <author>4991      <name>Tommaso Salvatori</name>4992    </author>4993  </entry>4994  <entry>4995    <id>http://arxiv.org/abs/2609.11279v1</id>4996    <title>SAMV-DUSt3R: Instance-Centric 3D Scene Decoupling from Sparse Multi-Views</title>4997    <updated>2026-09-10T09:12:48Z</updated>4998    <link href="https://arxiv.org/abs/2609.11279v1" rel="alternate" type="text/html"/>4999    <link href="https://arxiv.org/pdf/2609.11279v1" rel="related" type="application/pdf" title="pdf"/>5000    <summary>With the rising demand to decouple objects from 3D scenes, we propose SAMV-DUSt3R, an end-to-end model that injects SAM2 2D masks into MV-DUSt3R reconstruction. A Cross Flow Mask Block uses these masks to steer the network toward the target instance, jointly improving shape accuracy and achieving object-level disentanglement without multi-stage pipelines. To ensure reconstruction stability, a lightweight Spatial RankGNN selects the optimal reference view with a selection accuracy of 73.5\%. Extensive experiments demonstrate that our method boosts average reconstruction precision by 11\% across various metrics compared to state-of-the-art baselines. These results reveal a strong instance-disentanglement capability and clear benefits for driving, robotics, AR/VR, and heritage digitisation.</summary>5001    <category term="cs.CV" scheme="http://arxiv.org/schemas/atom"/>5002    <published>2026-09-10T09:12:48Z</published>5003    <arxiv:primary_category term="cs.CV"/>5004    <author>5005      <name>Langxu Zhao</name>5006    </author>5007    <author>5008      <name>Zuan Gu</name>5009    </author>5010    <author>5011      <name>Yingdan Zhang</name>5012    </author>5013    <author>5014      <name>Pengfei Zhao</name>5015    </author>5016    <author>5017      <name>Tianhan Gao</name>5018    </author>5019  </entry>5020  <entry>5021    <id>http://arxiv.org/abs/2609.11277v1</id>5022    <title>Predicting Train Delays in Finland Using Machine Learning and Weather Data</title>5023    <updated>2026-09-10T09:08:43Z</updated>5024    <link href="https://arxiv.org/abs/2609.11277v1" rel="alternate" type="text/html"/>5025    <link href="https://arxiv.org/pdf/2609.11277v1" rel="related" type="application/pdf" title="pdf"/>5026    <summary>Reliable railway operations depend increasingly on real-time environmental intelligence delivered through wireless sensor infrastructures, a capability that 6G networks will substantially enhance through integrated sensing and edge computing. Adverse weather, particularly in Arctic regions with extreme temperatures and heavy precipitation, remains a leading cause of train delays, yet most prediction approaches rely on raw meteorological inputs without exploiting domain-informed feature engineering. This paper investigates machine learning for train delay prediction using the Finland Integrated Train-Weather (FI-TW) dataset, which fuses railway operational records with observations from the Finnish Meteorological Institute's nationwide sensor network of approximately 200 stations communicating over wireless links. We evaluate three feature configurations using XGBoost at Oulu central station (101,146 observations): full weather features, instant weather observations only, and derived weather category scenarios. The category-based approach, employing hierarchical classifications such as Blizzard, Heavy Snow, and Extreme Cold, achieved an R^2 of 0.78, root mean squared error of 8.5 minutes, and mean absolute error of 3.7 minutes, representing an 11% R^2 improvement and 10% error reduction over alternative configurations. These results demonstrate that compact, domain-informed features derived from sensor streams outperform raw meteorological observations, offering bandwidth-efficient representations suitable for edge deployment over current and emerging wireless infrastructures.</summary>5027    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>5028    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>5029    <published>2026-09-10T09:08:43Z</published>5030    <arxiv:comment>6 pages, 3 Figures, 4 tables, presented at Wireless Europe 2026, Rimini, Italy, June 2026</arxiv:comment>5031    <arxiv:primary_category term="cs.AI"/>5032    <author>5033      <name>Vinicius Pozzobon Borin</name>5034    </author>5035    <author>5036      <name>Jean Michel de Souza Sant'Ana</name>5037    </author>5038    <author>5039      <name>Nurul Huda Mahmood</name>5040    </author>5041  </entry>5042  <entry>5043    <id>http://arxiv.org/abs/2609.11274v1</id>5044    <title>Xiaomi-CocktailASR-1 Technical Report</title>5045    <updated>2026-09-10T09:06:00Z</updated>5046    <link href="https://arxiv.org/abs/2609.11274v1" rel="alternate" type="text/html"/>5047    <link href="https://arxiv.org/pdf/2609.11274v1" rel="related" type="application/pdf" title="pdf"/>5048    <summary>Recently, large language model (LLM) based ASR models have achieved significant progress, yet they generally lack support for multi-speaker scenarios, where the cocktail party problem remains a critical bottleneck for further advancing ASR. Existing TS-ASR methods, including end-to-end architectures with speaker embeddings and latest LLM-based explorations suffer from degraded single-speaker performance and the inability to reject when the target speaker is absent. In this paper, we propose Xiaomi-CocktailASR-1, an LLM-based end-to-end TS-ASR architecture. By utilizing reference speech as voiceprint prompts, it directly transcribes the target speaker's speech without requiring speech separation. Xiaomi-CocktailASR-1 maintains competitive performance in single-speaker scenarios, comparable to mainstream ASR models. It also features a negative sample rejection capability, outputting empty text when the target speaker is absent from the mixed speech. Additionally, Xiaomi-CocktailASR-1 supports a Chain-of-Thought (CoT) reasoning mode to provide explicit reasoning steps. Extensive experiments on various synthetic and real-world multispeaker benchmarks demonstrate that Xiaomi-CocktailASR-1 achieves state-of-the-art performance, effectively addressing the cocktail party problem through a unified architecture that balances multispeaker and single-speaker recognition accuracy, along with rejection capability.</summary>5049    <category term="cs.SD" scheme="http://arxiv.org/schemas/atom"/>5050    <category term="cs.CL" scheme="http://arxiv.org/schemas/atom"/>5051    <category term="eess.AS" scheme="http://arxiv.org/schemas/atom"/>5052    <published>2026-09-10T09:06:00Z</published>5053    <arxiv:primary_category term="cs.SD"/>5054    <author>5055      <name>Yiru Zhang</name>5056    </author>5057    <author>5058      <name>Hang Su</name>5059    </author>5060    <author>5061      <name>Lichun Fan</name>5062    </author>5063    <author>5064      <name>Ying Zeng</name>5065    </author>5066    <author>5067      <name>Chang Liu</name>5068    </author>5069    <author>5070      <name>Yifeng Wang</name>5071    </author>5072    <author>5073      <name>Yuquan Liang</name>5074    </author>5075    <author>5076      <name>Tao Li</name>5077    </author>5078    <author>5079      <name>Lian Li</name>5080    </author>5081    <author>5082      <name>Wenhao Yang</name>5083    </author>5084    <author>5085      <name>Jian Luan</name>5086    </author>5087    <author>5088      <name>Cong Zou</name>5089    </author>5090    <author>5091      <name>Heng Qu</name>5092    </author>5093  </entry>5094  <entry>5095    <id>http://arxiv.org/abs/2609.11271v1</id>5096    <title>Order-Aware 2.5D Multiple Instance Learning for Preoperative MRI-Based Perineural Invasion Risk Assessment in Intrahepatic Cholangiocarcinoma</title>5097    <updated>2026-09-10T09:04:52Z</updated>5098    <link href="https://arxiv.org/abs/2609.11271v1" rel="alternate" type="text/html"/>5099    <link href="https://arxiv.org/pdf/2609.11271v1" rel="related" type="application/pdf" title="pdf"/>5100    <summary>Perineural invasion (PNI) is an adverse histopathologic marker in intrahepatic cholangiocarcinoma (ICC), but it is usually confirmed only after resection. Preoperative T2-weighted MRI may provide noninvasive imaging cues predictive of PNI, although labels are available only at the patient level without slice- or voxel-level annotations. We propose Order-Aware Slab Multiple Instance Learning (OAS-MIL), a weakly supervised framework for patient-level PNI prediction. Each tumor-centered MRI crop is represented as an ordered sequence of overlapping 2.5D slabs formed from contiguous axial slices. A shared encoder extracts slab-level features, which are aggregated by a permutation-invariant set-attention branch and a bidirectional sequence-attention branch. Using five-fold label-stratified cross-validation at the patient level, OAS-MIL achieved a mean AUROC of 0.770, outperforming the evaluated volumetric and MIL baselines. These results suggest that axial order provides a useful inductive bias for weakly supervised PNI prediction from MRI.</summary>5101    <category term="cs.CV" scheme="http://arxiv.org/schemas/atom"/>5102    <published>2026-09-10T09:04:52Z</published>5103    <arxiv:primary_category term="cs.CV"/>5104    <author>5105      <name>Hyunsu Go</name>5106    </author>5107    <author>5108      <name>Youngung Han</name>5109    </author>5110    <author>5111      <name>Kyeonghun Kim</name>5112    </author>5113    <author>5114      <name>Jinyong Jun</name>5115    </author>5116    <author>5117      <name>Junbeom Lee</name>5118    </author>5119    <author>5120      <name>Dohyun Kweon</name>5121    </author>5122    <author>5123      <name>Yului Jeong</name>5124    </author>5125    <author>5126      <name>Suah Park</name>5127    </author>5128    <author>5129      <name>Sungha Park</name>5130    </author>5131    <author>5132      <name>Anna Jung</name>5133    </author>5134    <author>5135      <name>Woo Kyoung Jeong</name>5136    </author>5137    <author>5138      <name>Ken Ying-Kai Liao</name>5139    </author>5140    <author>5141      <name>Hyuk-Jae Lee</name>5142    </author>5143    <author>5144      <name>Nam-Joon Kim</name>5145    </author>5146  </entry>5147  <entry>5148    <id>http://arxiv.org/abs/2609.11269v1</id>5149    <title>Improving Faint Object Detection for Space Situational Awareness with Variational Autoencoders</title>5150    <updated>2026-09-10T09:02:13Z</updated>5151    <link href="https://arxiv.org/abs/2609.11269v1" rel="alternate" type="text/html"/>5152    <link href="https://arxiv.org/pdf/2609.11269v1" rel="related" type="application/pdf" title="pdf"/>5153    <summary>We present a deep-learning pipeline for enhancing the detection of faint moving objects in optical space situational awareness (SSA) imagery through automated star removal and background reconstruction. Detecting low signal-to-noise ratio (SNR) objects remains extremely challenging in optical observations, particularly in the cislunar (X-GEO) environment, where structured sky backgrounds, dense stellar fields, and scattered moonlight significantly degrade the performance of classical detection algorithms. To address this problem, the proposed pipeline combines a lightweight segmentation network (Tiny-U-Net) to generate stellar masks with a partial-convolution variational autoencoder (astro-VAE), designed to learn the statistical distribution of astronomical backgrounds and perform context-aware inpainting of masked regions. The reconstructed background maps can then be used as a preprocessing step to suppress fixed sources and background inhomogeneities prior to detection. As a proof of concept, the approach is integrated with a shift-and-stack scheme and evaluated on real ground-based telescope observations targeting the X-GEO region. Results demonstrate that the method reconstructs star-free backgrounds with high fidelity, while preserving moving targets and significantly enhancing detectability, thereby providing an effective data-driven preprocessing strategy for faint moving-object detection in optical SSA scenarios.</summary>5154    <category term="cs.CV" scheme="http://arxiv.org/schemas/atom"/>5155    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>5156    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>5157    <published>2026-09-10T09:02:13Z</published>5158    <arxiv:comment>Accepted at SPAICE 2026: the 3rd European Space Agency Conference on AI in and for Space</arxiv:comment>5159    <arxiv:primary_category term="cs.CV"/>5160    <author>5161      <name>Angela Cratere</name>5162    </author>5163    <author>5164      <name>Luca Ghilardi</name>5165    </author>5166    <author>5167      <name>Vishnu Reddy</name>5168    </author>5169    <author>5170      <name>Francesco Dell'Olio</name>5171    </author>5172    <author>5173      <name>Charalampos S. Kouzinopoulos</name>5174    </author>5175    <author>5176      <name>Roberto Furfaro</name>5177    </author>5178  </entry>5179  <entry>5180    <id>http://arxiv.org/abs/2609.11265v1</id>5181    <title>Uncertainty DMD: Restoring Diversity in Few-Step Autoregressive Video Distillation</title>5182    <updated>2026-09-10T08:59:47Z</updated>5183    <link href="https://arxiv.org/abs/2609.11265v1" rel="alternate" type="text/html"/>5184    <link href="https://arxiv.org/pdf/2609.11265v1" rel="related" type="application/pdf" title="pdf"/>5185    <summary>Few-step distillation improves the efficiency of autoregressive (AR) video generation, but often causes diversity collapse: under the same prompt, different noise samples tend to produce highly similar videos with weakened motion dynamics. We analyze this degradation in Distribution Matching Distillation (DMD)-distilled AR video generators and find that, in the autoregressive setting, it takes the form of a structured uncertainty collapse: the mode-seeking bias of DMD maps different noise samples to nearly identical first chunks, and the deterministic AR cache then propagates this collapsed state to all subsequent chunks, turning a local loss of stochasticity at the rollout root into a global suppression of temporal variation. Based on this analysis, we propose Uncertainty DMD, a simple uncertainty-injection framework that restores stochasticity at two key stages of AR generation: a timestep perturbation for the first chunk to increase first-chunk diversity, and a stochastic cache-writing mechanism for later chunks to preserve uncertainty in autoregressive conditioning. The method requires no architectural changes and introduces only lightweight perturbation operations. The same perturbation mechanisms are used during both training and inference. Experiments show that Uncertainty DMD consistently improves diversity and motion dynamics while maintaining comparable per-sample visual quality.</summary>5186    <category term="cs.CV" scheme="http://arxiv.org/schemas/atom"/>5187    <published>2026-09-10T08:59:47Z</published>5188    <arxiv:comment>Project page: https://scdzx.github.io/Uncertainty-DMD</arxiv:comment>5189    <arxiv:primary_category term="cs.CV"/>5190    <author>5191      <name>Zixuan Duan</name>5192    </author>5193    <author>5194      <name>Xunzhi Xiang</name>5195    </author>5196    <author>5197      <name>Yabo Chen</name>5198    </author>5199    <author>5200      <name>Xin Zhang</name>5201    </author>5202    <author>5203      <name>Changhan Liu</name>5204    </author>5205    <author>5206      <name>Haibin Huang</name>5207    </author>5208    <author>5209      <name>Chi Zhang</name>5210    </author>5211    <author>5212      <name>Qi Fan</name>5213    </author>5214    <author>5215      <name>Xuelong Li</name>5216    </author>5217  </entry>5218  <entry>5219    <id>http://arxiv.org/abs/2609.11262v1</id>5220    <title>AI-Powered Flare Combustion Efficiency Estimation</title>5221    <updated>2026-09-10T08:56:57Z</updated>5222    <link href="https://arxiv.org/abs/2609.11262v1" rel="alternate" type="text/html"/>5223    <link href="https://arxiv.org/pdf/2609.11262v1" rel="related" type="application/pdf" title="pdf"/>5224    <summary>Achieving high combustion efficiency in flare stacks is crucial for adhering to regulatory standards and controlling the release of hydrocarbons into the environment. Traditional instruments like gas analyzers and hyperspectral cameras are expensive, fragile, and require frequent calibration, which makes them impractical for remote or budget constrained industrial sites. We propose an innovative solution that combines a lightweight vision-language encoder with a compact multi-layer perceptron to predict combustion efficiency directly from low-cost thermal video footage. The fully trained model is integrated into an easy-to-deploy graphical user interface. This interface overlays predicted combustion efficiency values on each video frame, displays real-time trends in combustion efficiency, shows the distribution of combustion efficiency across all frames in the video, and allows users to export CSV reports. Over a six-month period, the system achieved 99% uptime and required less than 15 minutes of maintenance per week.</summary>5225    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>5226    <category term="cs.CV" scheme="http://arxiv.org/schemas/atom"/>5227    <published>2026-09-10T08:56:57Z</published>5228    <arxiv:comment>Accepted at the 4th International Conference on Machine Learning and Data Engineering (ICMLDE 2025). 5 pages</arxiv:comment>5229    <arxiv:primary_category term="cs.AI"/>5230    <author>5231      <name>Afeefa Azam</name>5232    </author>5233    <author>5234      <name>Iyyakutti Iyappan Ganapathi</name>5235    </author>5236    <author>5237      <name>Fares Ossama Abdelhafez</name>5238    </author>5239    <author>5240      <name>Divya Velayudhan</name>5241    </author>5242    <author>5243      <name>Maregu Assefa Habtie</name>5244    </author>5245    <author>5246      <name>Hamad Karki</name>5247    </author>5248    <author>5249      <name>Khalid Yousef Al Awadhi</name>5250    </author>5251    <author>5252      <name>Naoufel Werghi</name>5253    </author>5254  </entry>5255  <entry>5256    <id>http://arxiv.org/abs/2609.11261v1</id>5257    <title>INDRA: A New AI Tool for Exploring Tobacco, Fossil Fuel, and Chemical Industry Archives</title>5258    <updated>2026-09-10T08:55:00Z</updated>5259    <link href="https://arxiv.org/abs/2609.11261v1" rel="alternate" type="text/html"/>5260    <link href="https://arxiv.org/pdf/2609.11261v1" rel="related" type="application/pdf" title="pdf"/>5261    <summary>Five decades of litigation have disgorged hundreds of millions of pages of formerly secret business records from the tobacco industry, along with documents from the makers of drugs, chemicals, food, firearms, and fossil fuels. Yet these archives have been effectively inaccessible to general-purpose large language models (LLMs) because they have never been compiled into an LLM-readable corpus. Chatbots may be familiar with some of the materials contained in such archives but, with no direct access to the documents, they are vulnerable to hallucination and other defects. Here we introduce INDRA, a research platform designed to remedy such failures by embedding the conventions of archival historiography into a system-level protocol governing every output. The platform federates UCSF's Industry Documents Library, Columbia and CUNY's ToxicDocs, Stanford's SRITA, and other heretofore siloed collections, and provides three interlinked safeguards: (1) a closed evidentiary sandbox confines the model to a user-selected corpus, blocking retrieval from external sources that could introduce bias; (2) real-time provenance tagging marks the boundary between archival evidence and parametric inference; and (3) a system-level protocol enforced by deterministic scripts guides the structure of every output. Together these safeguards prevent the model from conflating "the documents say X" with "I think X" or "I learned X from prior training." The result is an LLM-powered research partner enabling massive multi-archival investigations, a tool whose outputs are designed to be checked rather than trusted, and whose architecture makes the conditions of knowledge production visible and auditable. Three case studies demonstrate the method's analytical value and limitations, including what we call the Heraclitus effect, the steppingstone dilemma, and the gullibility (or mafia) problem.</summary>5262    <category term="cs.DL" scheme="http://arxiv.org/schemas/atom"/>5263    <category term="cs.CL" scheme="http://arxiv.org/schemas/atom"/>5264    <category term="cs.CY" scheme="http://arxiv.org/schemas/atom"/>5265    <published>2026-09-10T08:55:00Z</published>5266    <arxiv:comment>35 pages, 6 figures, Appendices available at https://indra.stanford.edu/methods/appendices</arxiv:comment>5267    <arxiv:primary_category term="cs.DL"/>5268    <author>5269      <name>Daniel Akselrad</name>5270    </author>5271    <author>5272      <name>Robert N. Proctor</name>5273    </author>5274  </entry>5275  <entry>5276    <id>http://arxiv.org/abs/2609.11255v1</id>5277    <title>Rethinking Radiomap Blind Prediction with Limited Environment and Configuration Representations</title>5278    <updated>2026-09-10T08:52:01Z</updated>5279    <link href="https://arxiv.org/abs/2609.11255v1" rel="alternate" type="text/html"/>5280    <link href="https://arxiv.org/pdf/2609.11255v1" rel="related" type="application/pdf" title="pdf"/>5281    <summary>Radiomap blind prediction infers radiomaps from observable representations of the propagation environment and base station (BS) configuration without field measurements. These representations are inherently incomplete and cannot uniquely determine the target radiomap. Under squared loss, we identify the conditional-mean radiomap as the population-optimal deterministic target and decompose domain risk into target-approximation error and irreducible uncertainty. The train-test risk gap motivates propagation priors as cross-domain guidance, although their partial or simplified forms may bias the attainable predictor. We therefore propose RadioDecomp, which treats a prior-guided predictor as a correctable base and uses deterministic residual refinement to learn its remaining predictable discrepancy. We instantiate RadioDecomp as RadioLSR (LoS-Shadow-Residual). Experiments under cross-configuration and cross-environment settings show that RadioLSR is especially effective for cross-configuration generalization and provides overall gains over a controlled monolithic counterpart under cross-environment generalization.</summary>5282    <category term="eess.SP" scheme="http://arxiv.org/schemas/atom"/>5283    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>5284    <published>2026-09-10T08:52:01Z</published>5285    <arxiv:comment>This paper has been accepted for presentation at IEEE Globecom 2026</arxiv:comment>5286    <arxiv:primary_category term="eess.SP"/>5287    <author>5288      <name>Xiaojie Li</name>5289    </author>5290    <author>5291      <name>Yu Han</name>5292    </author>5293    <author>5294      <name>Han Fang</name>5295    </author>5296    <author>5297      <name>Shangqing Liu</name>5298    </author>5299    <author>5300      <name>Shi Jin</name>5301    </author>5302    <author>5303      <name>Chao-Kai Wen</name>5304    </author>5305  </entry>5306  <entry>5307    <id>http://arxiv.org/abs/2609.11253v1</id>5308    <title>MUtE: A Dual Framework for Concept Erasure and Counterfactual Interventions</title>5309    <updated>2026-09-10T08:47:11Z</updated>5310    <link href="https://arxiv.org/abs/2609.11253v1" rel="alternate" type="text/html"/>5311    <link href="https://arxiv.org/pdf/2609.11253v1" rel="related" type="application/pdf" title="pdf"/>5312    <summary>Erasing concept-specific information from representations has been proven useful for mitigating bias or interpreting model decisions. The joint objective is to transform the original representations such that the target concept becomes unpredictable, while maximally preserving concept-unrelated information. In this work, we revisit the optimal bounds of concept erasure to derive a novel class of erasure functions that naturally induce a deterministic, dual counterfactual mapping. Bridging the gap between theoretical optimality and practical representation learning, we design an implementation that imposes a translational bias on counterfactual trajectories - a constraint that aligns with how many concepts geometrically manifest in modern language models. Our framework enables seamless navigation between concept erasure and counterfactual generation. We empirically demonstrate its efficacy in improving downstream algorithmic fairness and generating counterfactual texts.</summary>5313    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>5314    <category term="cs.CL" scheme="http://arxiv.org/schemas/atom"/>5315    <published>2026-09-10T08:47:11Z</published>5316    <arxiv:comment>21 pages, 3 figures, 6 tables</arxiv:comment>5317    <arxiv:primary_category term="cs.LG"/>5318    <author>5319      <name>Antoine Saillenfest</name>5320    </author>5321  </entry>5322  <entry>5323    <id>http://arxiv.org/abs/2609.11248v1</id>5324    <title>Generative Replay Mitigates Sample Starvation in Quantum Architecture Search</title>5325    <updated>2026-09-10T08:45:37Z</updated>5326    <link href="https://arxiv.org/abs/2609.11248v1" rel="alternate" type="text/html"/>5327    <link href="https://arxiv.org/pdf/2609.11248v1" rel="related" type="application/pdf" title="pdf"/>5328    <summary>Reinforcement learning (RL) can automate quantum architecture search, but its scalability is limited when useful circuit trajectories become rare in the rapidly expanding search space. Existing replay mechanisms reuse observed transitions; the proposed learned model produces additional predicted one step transitions from real state-action seeds. Here we introduce GenQAS, a tensor network-guided RL framework that combines a fixed matrix product state warm-start with prioritized generative replay. A learned local transition model generates synthetic circuit transitions on demand and mixes them with real experience during Double Deep Q-Network updates. Under a random exploration analysis, near ground state circuits occupy a rapidly shrinking region of the accessible state space. We investigate whether real data anchored synthetic replay can improve the effective training signal in this regime. Across chemical Hamiltonian benchmarks from 6 to 12 qubits, GenQAS improves fixed-budget success probability and identifies compact circuits at competitive energy error. At 12 qubits, it improves final success probability by up to $7.0\times$ over passive replay. On a 15-qubit transverse field Ising model, GenQAS increases success probability from $12\%$ to $21\%$. In a noisy 6-qubit BeH$_2$ transfer experiment, generative replay reduces the steps to chemical accuracy by $92.7\%$. These results show that generative replay can mitigate sample starvation in quantum architecture search and support more resource efficient circuit discovery.</summary>5329    <category term="quant-ph" scheme="http://arxiv.org/schemas/atom"/>5330    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>5331    <category term="cs.ET" scheme="http://arxiv.org/schemas/atom"/>5332    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>5333    <published>2026-09-10T08:45:37Z</published>5334    <arxiv:comment>GenQAS: 38 pages, 7 figures, 2 tables and 1 algorithm in main text</arxiv:comment>5335    <arxiv:primary_category term="quant-ph"/>5336    <author>5337      <name>Akash Kundu</name>5338    </author>5339    <author>5340      <name>Amit Kumar Jaiswal</name>5341    </author>5342    <author>5343      <name>Sebastian Feld</name>5344    </author>5345    <author>5346      <name>Prayag Tiwari</name>5347    </author>5348  </entry>5349  <entry>5350    <id>http://arxiv.org/abs/2609.11247v1</id>5351    <title>The Illusion of Balanced Multimodal Sentiment Analysis: Beyond the Limits of Optimization-Based Methods</title>5352    <updated>2026-09-10T08:44:51Z</updated>5353    <link href="https://arxiv.org/abs/2609.11247v1" rel="alternate" type="text/html"/>5354    <link href="https://arxiv.org/pdf/2609.11247v1" rel="related" type="application/pdf" title="pdf"/>5355    <summary>Multimodal Sentiment Analysis (MSA) remains constrained by modality imbalance, yet the field continues to rely on optimization-based balancing methods that promise more than they deliver. We provide three contributions: 1) a unified evaluation framework testing gradient and loss-based balancing strategies under controlled settings; 2) a theoretical diagnosis explaining why these methods fail, as they conflate fitting speed with discriminative contribution; and 3) a research agenda toward held-out discriminative modality valuation. Experiments on CMU-MOSI and CMU-MOSEI reveal three shortcomings: no strategy reliably outperforms Late Concatenation; performance is sensitive to hyperparameters; and even ratio calibration fails to yield consistent gains. The core issue is fundamental: loss is not utility, and gradients are not importance. Modality imbalance remains unresolved, motivating utility estimation from held-out performance.</summary>5356    <category term="cs.CL" scheme="http://arxiv.org/schemas/atom"/>5357    <published>2026-09-10T08:44:51Z</published>5358    <arxiv:comment>Accepted at Interspeech 2026</arxiv:comment>5359    <arxiv:primary_category term="cs.CL"/>5360    <author>5361      <name>Ioanna Kaffeza</name>5362    </author>5363    <author>5364      <name>Efthymios Georgiou</name>5365    </author>5366    <author>5367      <name>Alexandros Potamianos</name>5368    </author>5369  </entry>5370  <entry>5371    <id>http://arxiv.org/abs/2609.11246v1</id>5372    <title>Assessing the Reusability of Public Speech Resources for Low-Resource Languages: A Central Kurdish Case Study</title>5373    <updated>2026-09-10T08:43:42Z</updated>5374    <link href="https://arxiv.org/abs/2609.11246v1" rel="alternate" type="text/html"/>5375    <link href="https://arxiv.org/pdf/2609.11246v1" rel="related" type="application/pdf" title="pdf"/>5376    <summary>Kurdish is spoken by millions of people, but little technology can read it aloud. A recent study released three Kurdish voices, 35 hours of recorded speech, and a paper describing the work, all free to download. This review checks how well those public files match the paper. The research is careful about its limits, but the files contain several problems: a settings file lists equipment that was never used, test recordings are left unlabeled among training data, and a coding fault mishandles long numbers. The download page also claims a stronger result than the paper reports and recommends one voice for general use. That recommendation matters because Kurdish has major regional and written variation, while these voices were built from three people reading prepared texts. The process therefore removes much everyday and regional speech. English and German benefit from long traditions of dictionaries and linguistic description that help identify wrong pronunciations; Kurdish has far less such support, so software choices can go unchecked. The voices sound fluent, but they represent the reading styles of their speakers rather than Kurdish as a whole. Most of these issues can be fixed using information the team already has, without changing the reported results. Better records would mainly make the work easier for others, especially community linguists, to check and reuse. The license is the main exception: whether audiobook owners allow corrected versions to be shared will affect whether future Kurdish voices can build on this work or must start again.</summary>5377    <category term="cs.CL" scheme="http://arxiv.org/schemas/atom"/>5378    <published>2026-09-10T08:43:42Z</published>5379    <arxiv:primary_category term="cs.CL"/>5380    <author>5381      <name>Hiwa Asadpour</name>5382    </author>5383  </entry>5384  <entry>5385    <id>http://arxiv.org/abs/2609.11244v1</id>5386    <title>OmniHallu: Unified Hallucination Detection for Cross-Modal Comprehension and Generation in Multimodal Large Language Models</title>5387    <updated>2026-09-10T08:41:12Z</updated>5388    <link href="https://arxiv.org/abs/2609.11244v1" rel="alternate" type="text/html"/>5389    <link href="https://arxiv.org/pdf/2609.11244v1" rel="related" type="application/pdf" title="pdf"/>5390    <summary>While Multimodal Large Language Models (MLLMs) have achieved remarkable progress across diverse tasks, they suffer from hallucinations where generated outputs contradict or misrepresent input semantics. Existing research typically addresses hallucination detection within a single modality or task type, limiting generalizability. We introduce OmniHallu, a unified hallucination detection framework spanning both comprehension and generation tasks across image, video, and audio modalities. We contribute OmniHallu-Bench, a 10,000-sample benchmark with claim-level human annotations covering six cross-modal tasks: image-to-text (I2T), video-to-text (V2T), audio-to-text (A2T), text-to-image (T2I), text-to-video (T2V), and text-to-audio (T2A). Our multi-agent architecture decomposes model outputs into atomic claims, verifies them through modality-specific experts, and aggregates evidence via structured reasoning. We further propose a preference-optimized trainable verifier that approximates the multi-agent decision boundary, reducing expert calls by 66% with minimal performance loss. Extensive experiments reveal a consistent modality-dependent performance gradient and provide fine-grained insights into cross-modal hallucination patterns.</summary>5391    <category term="cs.CL" scheme="http://arxiv.org/schemas/atom"/>5392    <category term="cs.CV" scheme="http://arxiv.org/schemas/atom"/>5393    <published>2026-09-10T08:41:12Z</published>5394    <arxiv:comment>Accepted to Findings of EMNLP 2026. 12 pages, 4 figures</arxiv:comment>5395    <arxiv:primary_category term="cs.CL"/>5396    <author>5397      <name>Jianjiang Yang</name>5398    </author>5399    <author>5400      <name>Peihang Li</name>5401    </author>5402    <author>5403      <name>Shanqing Xu</name>5404    </author>5405    <author>5406      <name>Mengchen Qian</name>5407    </author>5408    <author>5409      <name>Lu Zhang</name>5410    </author>5411    <author>5412      <name>Meng Luo</name>5413    </author>5414  </entry>5415  <entry>5416    <id>http://arxiv.org/abs/2609.11243v1</id>5417    <title>Sci-MMR: Benchmarking Multi-Step Evidence-Grounded Scientific Reasoning in Multimodal Agents</title>5418    <updated>2026-09-10T08:41:07Z</updated>5419    <link href="https://arxiv.org/abs/2609.11243v1" rel="alternate" type="text/html"/>5420    <link href="https://arxiv.org/pdf/2609.11243v1" rel="related" type="application/pdf" title="pdf"/>5421    <summary>Autonomous research agents are increasingly expected to search the literature, analyze experimental evidence, and generate scientific hypotheses. These capabilities require multi-step evidence grounded reasoning that progressively acquires, integrates, and verifies evidence before reaching a conclusion. Existing multimodal benchmarks, however, largely evaluate final-answer accuracy, leaving open whether predictions are actually supported by traceable scientific evidence. We introduce Sci-MMR, a benchmark for multi-step evidence-grounded scientific reasoning built on structured argument graphs linking scientific claims, citation-grounded knowledge, visual evidence, and supporting regions. Sci-MMR comprises 235 multi-hop reasoning tasks spanning four scientific disciplines, with an average of nine figure panels per task. Evaluating eight frontier multimodal models, we find that answer accuracy consistently exceeds complete-evidence recovery rate by more than 20%, revealing a substantial gap that answer-only evaluation is structurally unable to capture. Through controlled interventions, we identify two fundamental bottlenecks. First, evidence acquisition: models struggle to extract complete structured evidence from scientific figures, accounting for 57.2% of failures. While cropping tools yield modest gains (+4.5 points), providing gold evidence improves accuracy by up to 37.0 points, indicating difficulty in assembling complete multi-region evidence. Second, evidence integration: models struggle to translate available evidence into correct conclusions, accounting for 31.8% of failures, while even with gold evidence the strongest model achieves only 69.1% accuracy on the hardest tasks. These findings indicate that current answer-centric benchmarks substantially overestimate the evidence-grounded reasoning capabilities of multimodal research agents</summary>5422    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>5423    <published>2026-09-10T08:41:07Z</published>5424    <arxiv:primary_category term="cs.AI"/>5425    <author>5426      <name>Jiaqiang Li</name>5427    </author>5428    <author>5429      <name>Yajie Yang</name>5430    </author>5431    <author>5432      <name>Zhiheng Xi</name>5433    </author>5434    <author>5435      <name>Jiadong Chen</name>5436    </author>5437    <author>5438      <name>Enyu Zhou</name>5439    </author>5440    <author>5441      <name>Senjie Jin</name>5442    </author>5443    <author>5444      <name>Yang Nan</name>5445    </author>5446    <author>5447      <name>Jiazheng Zhang</name>5448    </author>5449    <author>5450      <name>Han Wang</name>5451    </author>5452    <author>5453      <name>Yanxin Li</name>5454    </author>5455    <author>5456      <name>Dingwei Zhu</name>5457    </author>5458    <author>5459      <name>Bicheng Deng</name>5460    </author>5461    <author>5462      <name>Yuhui Wang</name>5463    </author>5464    <author>5465      <name>Xiang Zheng</name>5466    </author>5467    <author>5468      <name>Qi Zhang</name>5469    </author>5470    <author>5471      <name>Lei Bai</name>5472    </author>5473    <author>5474      <name>Xingjun Ma</name>5475    </author>5476    <author>5477      <name>Tao Gui</name>5478    </author>5479  </entry>5480  <entry>5481    <id>http://arxiv.org/abs/2609.11242v1</id>5482    <title>From Evaluation to Enhancement: Benchmarking and Improving Think-with-Video Reasoning for Video Generative Models</title>5483    <updated>2026-09-10T08:40:58Z</updated>5484    <link href="https://arxiv.org/abs/2609.11242v1" rel="alternate" type="text/html"/>5485    <link href="https://arxiv.org/pdf/2609.11242v1" rel="related" type="application/pdf" title="pdf"/>5486    <summary>Video generation has advanced to produce visually compelling and temporally coherent results. Yet, whether these models can genuinely think with video--executing symbolic rules, respecting physical laws, and pursuing intentional goals--remains an open question. Existing benchmarks only partially address this, often conflating visual quality with cognitive correctness. We introduce VWG-Bench (Video World Generalist Benchmark), a comprehensive benchmark spanning 9 reasoning dimensions and 38 fine-grained tasks. To enable precise diagnosis, we design a three-level VLM-as-Judge protocol that independently assesses video-level fluency, task-level rule adherence, and sample-level goal realization. Evaluations of leading models reveal a striking gap: while models achieve strong rendering scores, they consistently fail on logic-heavy and rule-constrained tasks. To address this, we propose Vid-PRE (Video Prompt Reasoner and Enhancer), a model-agnostic prompt rewriter that offloads the cognitive burden of reasoning to a dedicated VLM. Trained via reinforcement learning with purely text-based rewards, Vid-PRE produces concise, constraint-aware prompts without the instability of video-level reward signals. Experiments show that Vid-PRE yields substantial reasoning improvements across multiple generators without architectural modifications. Together, VWG-Bench and Vid-PRE offer a rigorous diagnostic lens and a scalable path toward true think-with-video capabilities. All data and code are publicly available at https://huggingface.co/datasets/KlingTeam/VWG-Bench.</summary>5487    <category term="cs.CV" scheme="http://arxiv.org/schemas/atom"/>5488    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>5489    <published>2026-09-10T08:40:58Z</published>5490    <arxiv:comment>Accepted to ECCV 2026. 46 pages, 41 figures</arxiv:comment>5491    <arxiv:primary_category term="cs.CV"/>5492    <author>5493      <name>Meng Luo</name>5494    </author>5495    <author>5496      <name>Yicheng Liu</name>5497    </author>5498    <author>5499      <name>Jiahao Wang</name>5500    </author>5501    <author>5502      <name>Yuanxing Zhang</name>5503    </author>5504    <author>5505      <name>Xin Tao</name>5506    </author>5507    <author>5508      <name>Pengfei Wan</name>5509    </author>5510    <author>5511      <name>Kun Gai</name>5512    </author>5513    <author>5514      <name>Hao Fei</name>5515    </author>5516  </entry>5517  <entry>5518    <id>http://arxiv.org/abs/2609.11240v1</id>5519    <title>Fast and Accurate Monomodal 3D High Resolution Deep Registration of Drosophila Larval Brain Volumes</title>5520    <updated>2026-09-10T08:38:02Z</updated>5521    <link href="https://arxiv.org/abs/2609.11240v1" rel="alternate" type="text/html"/>5522    <link href="https://arxiv.org/pdf/2609.11240v1" rel="related" type="application/pdf" title="pdf"/>5523    <summary>The larval stage of Drosophila melanogaster is a compact model system for neuroscience whose genetic toolkit allows fluorescent markers to be expressed in defined neural populations, and comparing the resulting expression patterns across animals requires every brain to be registered into a shared anatomical reference space. Existing pipelines for this task are predominantly based on classical registration methods, which perform a new optimization for each volume, often require per-case parameter tuning, and can take minutes per brain, limiting their use as a routine preprocessing step. We present a trained deep registration pipeline that deformably aligns a larval brain to a reference template in a single forward pass at high spatial resolution, on volumes that hold several times more voxels than those learned 3D registration is normally reported on, together with the preprocessing and anatomy-anchored evaluation pipeline required to apply it. Against eleven classical and seven further learned baselines on a held-out collection acquired with different acquisition and quality strata, the proposed pipeline is the most accurate, improving on the strongest classical baseline by 23 percentage points of anatomical landmark-local mutual information. It registers a volume one to two orders of magnitude faster than the classical deformable pipelines, and it retains more of its accuracy than any other method as acquisition quality degrades. The network, its trained weights and the full pipeline are released as the open-source deep larval brain registration framework: https://github.com/agentdr1/deep-larval-brain-reg</summary>5524    <category term="cs.CV" scheme="http://arxiv.org/schemas/atom"/>5525    <published>2026-09-10T08:38:02Z</published>5526    <arxiv:primary_category term="cs.CV"/>5527    <author>5528      <name>Daniel Reisenbüchler</name>5529    </author>5530    <author>5531      <name>Yousef Sadegheih</name>5532    </author>5533    <author>5534      <name>Michael Dittrich</name>5535    </author>5536    <author>5537      <name>Pratibha Kumari</name>5538    </author>5539    <author>5540      <name>Muhammad Usman</name>5541    </author>5542    <author>5543      <name>Dorit Merhof</name>5544    </author>5545  </entry>5546  <entry>5547    <id>http://arxiv.org/abs/2609.11237v1</id>5548    <title>SCINTILLA-SNN: A Spiking Multi-Scale Selective Aggregation Network for Perineural Invasion Prediction</title>5549    <updated>2026-09-10T08:36:34Z</updated>5550    <link href="https://arxiv.org/abs/2609.11237v1" rel="alternate" type="text/html"/>5551    <link href="https://arxiv.org/pdf/2609.11237v1" rel="related" type="application/pdf" title="pdf"/>5552    <summary>Preoperative prediction of perineural invasion (PNI) in cholangiocarcinoma (CCA) is clinically valuable but remains challenging because PNI-related cues on magnetic resonance imaging (MRI) are subtle, sparse, and spatially localized around the tumor boundary. Standard 3D CNN and transformer architectures process volumetric data in a dense or spatially uniform manner, which can dilute subtle PNI-related evidence while requiring a large number of multiply-accumulate operations over 3D feature grids. To address these limitations, we propose SCINTILLA-SNN, a 3D spiking network composed of a four-stage hierarchical backbone and a Multi-Scale Spike Aggregation (MSSA) module for PNI prediction. The backbone extracts hierarchical volumetric representations through spiking convolutional stages and local spike window modulation stages. Given the resulting stage-wise representations, MSSA maps each spatial token to a learnable content value and modulates it with a spike-dynamics gate derived from firing rate and timestep-wise membrane-potential variability. The resulting score, referred to as the diagnostic token score, is used to selectively aggregate sparse PNI-related evidence. Experiments on a 10-year retrospective cohort of 182 CCA patients show that SCINTILLA-SNN achieves an AUROC of 0.748 under 5-fold cross-validation, while reducing the estimated inference energy by 23.18$\times$ compared with dense MAC-only computation of the same network.</summary>5553    <category term="cs.CV" scheme="http://arxiv.org/schemas/atom"/>5554    <published>2026-09-10T08:36:34Z</published>5555    <arxiv:primary_category term="cs.CV"/>5556    <author>5557      <name>Youngung Han</name>5558    </author>5559    <author>5560      <name>Yului Jeong</name>5561    </author>5562    <author>5563      <name>Kyeonghun Kim</name>5564    </author>5565    <author>5566      <name>Dohyun Kweon</name>5567    </author>5568    <author>5569      <name>Suah Park</name>5570    </author>5571    <author>5572      <name>Hyunsu Go</name>5573    </author>5574    <author>5575      <name>Sungha Park</name>5576    </author>5577    <author>5578      <name>Anna Jung</name>5579    </author>5580    <author>5581      <name>Jinyong Jun</name>5582    </author>5583    <author>5584      <name>Yunho Choe</name>5585    </author>5586    <author>5587      <name>Yunjin Seo</name>5588    </author>5589    <author>5590      <name>Ken Ying-Kai Liao</name>5591    </author>5592    <author>5593      <name>Hyuk-Jae Lee</name>5594    </author>5595    <author>5596      <name>Nam-Joon Kim</name>5597    </author>5598  </entry>5599  <entry>5600    <id>http://arxiv.org/abs/2609.11236v1</id>5601    <title>HALDETECT at ImageEval 2026 Shared Tasks: Answer-First Contrastive Grounding with QLoRA</title>5602    <updated>2026-09-10T08:35:29Z</updated>5603    <link href="https://arxiv.org/abs/2609.11236v1" rel="alternate" type="text/html"/>5604    <link href="https://arxiv.org/pdf/2609.11236v1" rel="related" type="application/pdf" title="pdf"/>5605    <summary>Large multimodal models tend to hallucinate visual detail fluently, which limits their deployment for fine-grained interpretation. We present HALDETECT, our system for the English hallucination-detection track (Task 1b) of ImageEval 2026, in which a system must identify, from an image and three culturally plausible statements, the single visually grounded one. We frame the item as one contrastive decision, emit the answer before its explanation, and structure reasoning around colour/texture, shape/form, and context. Our best submitted adapter fine-tunes Qwen2.5-VL-7B-Instruct with 4-bit QLoRA while freezing the vision encoder and reaches Contrastive Instability (CI) 0.035 on the 1,000-item test set; we placed third of eight teams. Development experiments show that answer order can matter more than model scale and that adaptation beats prompting alone. Retrospective paired analysis of the released gold labels confirms the QLoRA gain over the best prompt but not the small gap between the devtest-selected and best-test adapters, and reseeding all four training sizes shows that the apparent data-scaling curve does not survive a seed change. The 35 residual errors are culturally plausible function, material, and recognition distinctions; naive adapter voting does not help.</summary>5606    <category term="cs.CV" scheme="http://arxiv.org/schemas/atom"/>5607    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>5608    <published>2026-09-10T08:35:29Z</published>5609    <arxiv:comment>10 pages, 3 figures, 10 tables (including appendices). System description paper for Task 1b (English) of ImageEval 2026 Shared Tasks (Fourth Arabic Natural Language Processing Conference), to appear in the Shared Tasks proceedings</arxiv:comment>5610    <arxiv:primary_category term="cs.CV"/>5611    <author>5612      <name>Syed Mohaiminul Hoque</name>5613    </author>5614    <author>5615      <name>Md Sakhawat Hossain</name>5616    </author>5617  </entry>5618  <entry>5619    <id>http://arxiv.org/abs/2609.11235v1</id>5620    <title>When is Test-Time Adaptation Identifiable From Unlabeled Evidence?</title>5621    <updated>2026-09-10T08:35:06Z</updated>5622    <link href="https://arxiv.org/abs/2609.11235v1" rel="alternate" type="text/html"/>5623    <link href="https://arxiv.org/pdf/2609.11235v1" rel="related" type="application/pdf" title="pdf"/>5624    <summary>Test-time adaptation (TTA) offers many ways to update a deployed model without labels, but choosing the wrong update can make a strong source model worse. Recent methods therefore try to predict which adaptation will work from unlabeled test data. We ask a prior question: does the evidence given to the selector contain enough information to determine the best action at all? We show that this is not guaranteed, even with a perfect selector. If an observation channel makes two deployments look the same while their TTA rankings differ, reliable selection is impossible from that channel; richer evidence can restore the decision only when it resolves the relevant ambiguity. We make this boundary exact in a finite-batch Gaussian TTA model, where doing nothing beats mean recentering for small shifts, recentering wins beyond a unique critical shift, and the boundary shrinks as $1/\sqrt n$. Public benchmark studies on CIFAR-100-C and DomainNet-126 show the same failure mode with modern TTA methods: changing only deployment structure can reverse the oracle action while global order-blind evidence remains unchanged. The result is a practical way to separate two failure modes that are usually mixed together: a weak selector versus an information channel that cannot support the desired decision in the first place.</summary>5625    <category term="cs.CV" scheme="http://arxiv.org/schemas/atom"/>5626    <published>2026-09-10T08:35:06Z</published>5627    <arxiv:primary_category term="cs.CV"/>5628    <author>5629      <name>Kartik Jhawar</name>5630    </author>5631    <author>5632      <name>Lipo Wang</name>5633    </author>5634  </entry>5635  <entry>5636    <id>http://arxiv.org/abs/2609.11234v1</id>5637    <title>NovGauge: A Fine-Grained Benchmark for Diagnosing LLMs' Capability in Paper Novelty Assessment</title>5638    <updated>2026-09-10T08:34:56Z</updated>5639    <link href="https://arxiv.org/abs/2609.11234v1" rel="alternate" type="text/html"/>5640    <link href="https://arxiv.org/pdf/2609.11234v1" rel="related" type="application/pdf" title="pdf"/>5641    <summary>Large language models (LLMs) are increasingly used in peer review at major AI conferences, yet novelty remains a persistent weak point. Existing benchmarks assess novelty as a single holistic score, making it difficult to diagnose which dimension a model misjudges or whether its evidence is faithful. We present NovGauge, a human-anchored benchmark for fine-grained novelty assessment diagnosis. The benchmark contains 619 paper pairs and 50 multi-paper sets, drawn from two expert sources: ICLR reviewer overlap claims and survey co-citations. Instances are independently labeled along three dimensions: task, problem, and method, capturing application goals, technical challenges, and solution approaches. We propose a cascading diagnostic pipeline that verifies per-dimension correctness, evidence grounding, and logical support. Evaluation of 18 LLMs shows hallucination rates ranging from 0% to 39% across dimensions, and among non-hallucinated correct-positive judgments, over 70% cite evidence fails to logically support the stated reason. The best-performing model, GPT-5.5, achieves 43-72% Verified F1 across dimensions, while most models retain less than half of their raw F1 after faithfulness verification. These results suggest that current LLMs remain far from reliable scientific novelty assessment, particularly when correctness is conditioned on faithful evidence grounding.</summary>5642    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>5643    <published>2026-09-10T08:34:56Z</published>5644    <arxiv:primary_category term="cs.AI"/>5645    <author>5646      <name>Guoqiang Zhang</name>5647    </author>5648    <author>5649      <name>Kexin Tan</name>5650    </author>5651    <author>5652      <name>Ming Zhang</name>5653    </author>5654    <author>5655      <name>Li Ju</name>5656    </author>5657    <author>5658      <name>Wenqing Jing</name>5659    </author>5660    <author>5661      <name>Zhonghan Yue</name>5662    </author>5663    <author>5664      <name>Jiayi Chen</name>5665    </author>5666    <author>5667      <name>Shiqiang Wu</name>5668    </author>5669    <author>5670      <name>Shaofan Liu</name>5671    </author>5672    <author>5673      <name>Yue Zhang</name>5674    </author>5675    <author>5676      <name>Yuankai Ying</name>5677    </author>5678    <author>5679      <name>Yang Shi</name>5680    </author>5681    <author>5682      <name>Tao Gui</name>5683    </author>5684    <author>5685      <name>Qi Zhang</name>5686    </author>5687    <author>5688      <name>Xuanjing Huang</name>5689    </author>5690  </entry>5691  <entry>5692    <id>http://arxiv.org/abs/2609.11231v1</id>5693    <title>A Voice-Interactive Multi-Agent System for Smart Operating Rooms: Architecture Design and Key Technologies</title>5694    <updated>2026-09-10T08:29:27Z</updated>5695    <link href="https://arxiv.org/abs/2609.11231v1" rel="alternate" type="text/html"/>5696    <link href="https://arxiv.org/pdf/2609.11231v1" rel="related" type="application/pdf" title="pdf"/>5697    <summary>This paper presents SurgicalRoomAgent, a voice-interactive multi-agent system for smart operating rooms based on large language models (LLMs). The system achieves natural language understanding, device control, intraoperative recording, and surgical report generation through a layered architecture comprising a voice interaction pipeline (wake, ASR, turn detection, agent reasoning, TTS) and an agent core (skill registry, task planner, device manager). Three key technologies are investigated: (1) KV Cache prefix warming for low-latency inference, reducing recomputation overhead from approximately 500 ms to tens of milliseconds via byte-level Longest Common Prefix reuse; (2) streaming partial JSON parsing with early parallel task execution, reducing end-to-end latency by approximately 30%; and (3) progressive skill prompt disclosure, which dynamically filters system prompts based on user role, connected devices, and surgical phase to maximize information density within limited context windows. The system is implemented using the Qwen3-27B model with llama.cpp/sglang inference engines. Experimental analysis demonstrates effective operation within a 16,384-token context limit and multi-device parallel control response times meeting OR real-time requirements.</summary>5698    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>5699    <category term="cs.CL" scheme="http://arxiv.org/schemas/atom"/>5700    <category term="cs.HC" scheme="http://arxiv.org/schemas/atom"/>5701    <published>2026-09-10T08:29:27Z</published>5702    <arxiv:primary_category term="cs.AI"/>5703    <author>5704      <name>Tianxiang Zhou</name>5705    </author>5706  </entry>5707  <entry>5708    <id>http://arxiv.org/abs/2609.11228v1</id>5709    <title>Solving Few-Shot Multiobjective Multitask Optimization via Iterative Sequential Transfer</title>5710    <updated>2026-09-10T08:28:28Z</updated>5711    <link href="https://arxiv.org/abs/2609.11228v1" rel="alternate" type="text/html"/>5712    <link href="https://arxiv.org/pdf/2609.11228v1" rel="related" type="application/pdf" title="pdf"/>5713    <summary>Applying knowledge transfer across multiple optimization tasks, multitask optimization (MTO) emerges as a promising approach to solving synergistic optimization tasks simultaneously. However, the development of effective knowledge transfer mechanisms in MTO fundamentally relies on aligning elite solution distributions across tasks. This dependency creates a critical bottleneck in few-shot optimization regimes, as restricted evaluation budgets impede the identification of elite solution distributions required for beneficial transfer. This challenge is exacerbated in multiobjective multitask problems, where each optimizer must approximate a continuous Pareto manifold rather than a single optimal point. This paper introduces Iterative Sequential Transfer (IST) to circumvent this bottleneck. We model MTO as a sequence of sequential transfer optimization problems, concentrating evaluations on a single target per iteration. We propose a likelihood-informed task prioritization mechanism to maximize transfer utility by identifying the task most likely ready for knowledge integration. Empirical results on benchmark and real-world problems verify the effectiveness of the proposed method under tight budgets.</summary>5714    <category term="cs.LG" scheme="http://arxiv.org/schemas/atom"/>5715    <category term="cs.AI" scheme="http://arxiv.org/schemas/atom"/>5716    <category term="cs.NE" scheme="http://arxiv.org/schemas/atom"/>5717    <published>2026-09-10T08:28:28Z</published>5718    <arxiv:comment>Accepted paper in WCCI/CEC 2026</arxiv:comment>5719    <arxiv:primary_category term="cs.LG"/>5720    <author>5721      <name>Tingyang Wei</name>5722    </author>5723    <author>5724      <name>Haofeng Wu</name>5725    </author>5726    <author>5727      <name>Ananda Phan Iman</name>5728    </author>5729    <author>5730      <name>Zhao Wei</name>5731    </author>5732    <author>5733      <name>Jiao Liu</name>5734    </author>5735    <author>5736      <name>Yew-Soon Ong</name>5737    </author>5738  </entry>5739</feed>5740