arXiv Papers with Code in Computer Science (August 2026)

PaperId: 1, https://arxiv.org/pdf/2608.31075.pdf   GitHub GitHub
Authors:Zhiqin Yang, Jingwen Fu, Yuhan Liu, Hengyu Liu, Yonggang Zhang, Kainan Cao, Zizhuo Zhang, Chenxin Li, Ruibin Yuan, Jiahao Pan, Jiankai Sun, Zhenyuan Zhang, Yibo Li, Yunlong Lin, Jing Xiong, Sida Lin, Bo Han, Wei Xue, Yike Guo
Title: Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence
Abstract:
Recent advances in large reasoning models (LRMs) have shown that reinforcement learning with verifiable rewards (RLVR) can substantially improve reasoning in mathematics and code, where outcomes can be checked automatically. Extending this progress to open‑ended and agentic tasks remains difficult because reliable rewards are harder to obtain and direct human supervision cannot keep pace with the scale and complexity of model‑generated experience. This paper studies how LRMs can continue to improve as human supervision gradually recedes from the learning loop. We examine two connected dimensions of this problem. The reward axis traces the development from per‑instance human judgments to reusable verifiers and rewards that operate even without human feedback. The experience axis examines how learning can progress from human‑curated tasks and environments toward self‑generated curricula, constructed environments, and autonomous co‑evolution. We connect these dimensions through a five‑level ladder from L0 to L4 that identifies which parts of the learning process remain under continued human control. Our analysis further highlights the risks introduced by increasingly autonomous rewards and experience generation, including reward hacking, feedback drift, curriculum collapse, and environment errors. Consequently, we also provide the evaluation around three complementary objects: policy capability, feedback fidelity, and experience quality. This analysis provides a structured account of current approaches to scaling LRMs beyond human supervision and the open problems involved in developing self‑sustaining learning systems toward superintelligence. Furthermore, we maintain a continuously updated \hrefhttps://github.com/visitworld123/Awesome‑Scaling‑LRM‑Beyond‑Human‑SupervisionGitHub repository to track the latest advances.

Authors:Shangqing Tu, Daniel Zhang-Li, Yucheng Wang, Shiyu Gan, Yanpeng Wang, Huiqiang Rong, Mofei Chen, Shen Yang, Yini Chen, Yinuo Duan, Binglin Liu, Ye He, Danqi Zheng, Zhanxin Hao, Yuxuan Wu, Mengting Tao, Yuqiu Liu, Jifan Yu, Juanzi Li, Bin Xu, Lei Hou, Huiqin Liu, Yu Zhang
Title: CogEvol: Towards Efficient and Reliable Learning Environment Generation
Abstract:
We present CogEvol, a family of models trained specifically for Learning Environment Generation: turning a course brief into a finished learning artifact (structured‑JSON slides or self‑contained interactive HTML pages) in a single pass. Across 220k production requests, CogEvol completes a slide in a median of 17 seconds and an interactive page in 59, replacing minutes‑long multi‑turn agent scaffolding. Reliability is enforced rather than hoped for: a production‑grounded data pipeline turns real failures into 53,687 verified SFT samples, and a hybrid rule‑plus‑VLM reward drives GRPO‑based RL, hardened after we caught and fixed a reward‑hacking episode that produced visually convincing but unplayable games. CogEvol‑27B scores 83.7 on slide quality and 63.7 on a 500‑case interactive‑HTML benchmark with 26.9x fewer parameters than flagship coding models, and, in collaboration with the OpenMAIC team, serves their live production traffic. CogEvol‑4B is released openly under the Apache 2.0 license at https://github.com/CogEvol/CogEvol‑4B; external flagships are measured on the same suites under the identical harness. Scaffold editing cuts interactive‑page generation cost by a further ~76%, and the full stack runs on domestic Ascend accelerators at application‑level parity with A800 GPUs, lowering the unit cost of AI‑native education at scale.

Authors:Mathias Zinnen, Alisha Mund, Sabine Lang, Lukas Hüttner, Thomas Gorges, Vincent Christlein
Title: Lot Machine: Multimodal Lot Extraction from Auction Catalogs
Abstract:
For provenance research and art market studies, auction catalogs are an essential resource to trace specific objects over time and space. While historical auction catalogs follow established domain conventions, their internal formatting remains highly variable, and their large‑scale analysis is currently restricted by the lack of machine‑readable representations of the auction lots. We propose a pipeline to automatically extract structured lot‑level metadata from German Sales, a large database of historical auction and sales catalogs from the 19th and 20th centuries. Using a manually annotated test set of representative catalog pages, we evaluate Vision‑Language Models (VLMs) under varying prompt strategies and constrained decoding frameworks. To reflect the practical constraints faced by cultural heritage institutions, including budget, compute resources, and data privacy requirements, we benchmark the methods across different deployment modes ranging from commercial providers to locally hosted, quantized models. We find that commercial endpoints establish the performance ceiling, while institutional gateways offer a viable, privacy‑preserving alternative. Local deployments remain feasible, but strictly require enforcing the output structure during generation to guarantee a valid JSON format. While varying degrees of human‑in‑the‑loop correction are still necessary, this work demonstrates that a VLM‑based pipeline can successfully unlock historical auction catalogs for large‑scale automated analysis.

Authors:Yulin Zhang, Yukun Huang, Sanxing Chen, Tianyi Lin, Ziang Yang, Xunjian Yin, Bhuwan Dhingra
Title: Lazy Grounding: Attacking Search Agents with Factual Evidence
Abstract:
Search agents mitigate hallucination by grounding their answers in retrieved web results. However, retrieval‑based approaches also introduce an attack surface: agents may cite misinformation from poisoned search corpora containing false or malicious documents. We demonstrate that, in some cases, search agents' reasoning and responses may be steered by completely factual but distracting information. We refer to this failure as lazy grounding. We expose lazy grounding by injecting nearby evidence from answer‑changing rewrites of benchmark questions into the search corpora. Each document contains factual evidence that supports a neighboring rewritten question but is retrieved for the original question. Across 12 model‑benchmark pairs, the attack causes the accuracy of search agents' responses to drop by 5.9 points on average and by up to 17.3 points, while inducing nearby‑answer adoption in every setting. The effect is even stronger when nearby evidence appears later or is more answer‑shaped. Our results show that robust search agents must defend against not only misinformation but also the misapplication of factual evidence. The code is publicly available at https://github.com/frankyzha/lazy‑grounding.

Authors:Yuyang Hong, Jinhui Guo, Jiaqi Gu, Lubin Fan, Ruixiang Wang, Kun Ding, Yue Wu, Shiming Xiang, Jieping Ye
Title: DICS: Exploring Data Intrinsic Consistency for Visual Instruction Selection
Abstract:
Visual instruction tuning is crucial for advancing the vision‑language alignment and instruction‑following capabilities of Vision‑Language Models (VLMs). However, identifying optimal subsets under a fixed ratio constraint from rapidly expanding datasets remains a significant bottleneck. While existing methods largely depend on distribution diversity or heuristic filtering, they often overlook the internal coherence within individual samples. To bridge this gap, we propose Data Intrinsic Consistency (DIC), a self‑scoring metric designed to quantify the sample‑level inter‑component consistency. DIC consists of two modules: Visual Information Consistency (VIC), evaluating the alignment between visual content and instructions, and Response Information Consistency (RIC), assessing response coherence relative to the instruction. Building upon DIC, we introduce Data Intrinsic Consistency Selection (DICS), an adaptive data selection method that optimizes the trade‑off between high intra‑sample consistency and global distributional diversity under varying data budgets. Extensive experiments demonstrate that DICS consistently outperforms state‑of‑the‑art methods across diverse dataset scales and model architectures, surpassing full‑dataset fine‑tuning while using only 25% of the LLaVA‑1.5‑665K data. We further curate DICS‑6M, a 6M‑sample multi‑modal instruction corpus that enables the largest‑scale visual instruction selection study to date; remarkably, DICS reaches 94.52% of the official InternVL3‑8B‑Instruct performance using less than 25% of its reported training data. Code can be seen at https://github.com/cqu‑student/DICS

Authors:Kaishuu Shinozaki-Conefrey, Olivier Pascaud, Robin Courant, Xi Wang, Dimitris Samaras, Vicky Kalogeiton
Title: TAKE 85: Testing Audiovisual filmmaKer's intEnt across 85 Hours of Film
Abstract:
Films communicate through deliberate creative choices, including lighting, color, composition, editing, dialogue, music, and sound. Humans naturally interpret these signals as directorial intent, yet current multimodal large language models (MLLMs) are evaluated almost exclusively on understanding what happens rather than why it is presented that way. We introduce TAKE 85, the first benchmark for directorial‑intent understanding, comprising 398 short films (85 hours) with expert‑verified question‑answer pairs spanning global and fine‑grained visual and audio intent. Through controlled modality ablations, TAKE 85 enables systematic evaluation of multimodal reasoning. Experiments on state‑of‑the‑art MLLMs reveal a substantial gap between perceptual recognition and intentional understanding: while models accurately describe events and narratives, they consistently fail to infer the communicative role of filmmaking decisions. Our results establish directorial intent as a previously overlooked dimension of multimodal understanding: even the strongest model reaches only 58 out of 100, and our ablations show that no input modality is sufficient on its own. All code, Q&As, and models are publicly available from https://github.com/KaiShinozakiConefrey/Take‑85

Authors:Lucas A. Dias, Henrique A. Schulz, Rafaela de Miranda, Guilherme L. Peres, Pedro L. Bittencourt, Rayson Laroca
Title: Confidence-Aware Ensemble and Long-Word Refinement for Artistic Text Recognition
Abstract:
Artistic Text Recognition (ATR) remains challenging because word images often combine decorative fonts, curved layouts, object‑like characters, clutter, and severe distortions. This paper studies WordArt‑V1.5 as a standardized benchmark for this setting and evaluates recent scene and artistic text recognizers under a common protocol. We propose a confidence‑aware ensemble that combines SVTRv2, PARSeq, and MAERec after fine‑tuning on the official training split. The ensemble selects predictions using the minimum confidence over disagreement positions, emphasizing characters that separate competing hypotheses. For long words, where a single character error can invalidate the whole prediction, we add a targeted refinement stage based on Needleman‑Wunsch alignment and lexicon‑guided correction. On the WordArt‑V1.5 Test B split, the proposed system reaches 89.90% Word Recognition Accuracy, improving the best individual fine‑tuned model by 1.77 percentage points. The long‑word refinement produces a modest global gain, but improves the targeted long‑word subset by 2.72 percentage points. Finally, an error analysis of all remaining mistakes shows that 48.8% are associated with labeling issues, visual ambiguity, or illegible samples, highlighting the value of diagnostic reporting for future ATR benchmarks and models. Our source code is available at https://github.com/lucas‑azdias/Artistic‑Text‑Recognition/.

Authors:Kun Efimov-Zhang, Yifei Song, Claire Gardent
Title: XQDT: eXplainable and Quantitative Data-Text Alignment Metric with Feedback Signals
Abstract:
Evaluating data‑text alignment remains challenging: existing metrics often provide limited explanations for the scores, while prompt‑based LLM‑as‑Judge methods can be expensive and unreliable. We present an end‑to‑end explainable evaluation metric that fine‑tunes a language model to identify omitted, extra, incorrect, and correct data units in a data‑text pair. These local judgements are aggregated into precision, recall, and F1 scores, providing both fine‑grained diagnostic feedback and an interpretable measure of alignment quality. Across benchmarks, our fine‑tuned models outperform LLM‑as‑Judge methods in error prediction and achieve competitive precision, recall, and F1 scores, while maintaining strong correlation with human judgements. Beyond evaluation, our verifier outputs also provide useful feedback signals for downstream correction and refinement, supporting alignment‑oriented improvement of data‑to‑text and text‑to‑data. Code and resources are available at https://github.com/guihuzhang/xqdt.

Authors:Henrique Zan Grande, João G. Pitol, Lucas B. Schuck, Rafael V. Serenato, Rayson Laroca, Andre Gustavo Hochuli
Title: On the Role of MRI Sequences in Cross-Dataset Generalization for Brain Tumor Segmentation
Abstract:
Brain tumor segmentation in magnetic resonance imaging (MRI) is a critical task for diagnosis and treatment planning. Despite the success of deep learning architectures such as U‑Net and its variants, performance degradation across datasets remains a major challenge, particularly under domain shift and limited annotated data. To address this issue, this study systematically evaluates how individual MRI sequences influence model robustness across two well‑known datasets. A ResUNet‑based framework is employed, where each modality is trained independently to isolate its effect under a controlled cross‑dataset evaluation protocol with tumor size stratification, without target‑domain training, or with limited domain adaptation. Results show that the T2f/FLAIR sequence achieves the best cross‑dataset performance, with Dice scores exceeding 75%. It consistently outperforms other modalities across most tumor size ranges, while multi‑sequence training further improves performance. Additionally, even limited target‑domain adaptation yields rapid initial gains, reducing the need for extensive annotations and costly retraining. Our source code is publicly available at https://github.com/henrique‑zan/brain_tumor_segmentation/.

Authors:Mohammadali Khodabandehlou, Bhaskar Krishnamachari
Title: Compression-Aware Abstention: Teaching LLMs to Refuse When KV-Compression Masks Remove Answer Evidence
Abstract:
KV‑cache compression reduces LLM inference memory by evicting context tokens, but when the evicted tokens contain answer‑bearing evidence, the model may hallucinate instead of recognizing that the compressed context is insufficient. We address this failure from a behavioral perspective: to our knowledge, this is the first work to formulate compression‑aware abstention as a learning problem, in which a model learns to answer when supporting evidence survives compression and abstain when it does not. We construct supervision from compressor survival masks and tight answer‑bearing spans, labeling examples as Confident when evidence survives and Abstain when it is removed. A 10.1M‑parameter LoRA adapter trained on ~2.6K MuSiQue 2‑hop QA examples reduces base‑model hallucinations by 97% under prompt‑style truncation while preserving correct answering on evidence‑retaining examples. Unlike prompt‑only abstention baselines, which over‑abstain on many answerable high‑retention examples, the trained adapter learns a conditional policy. We also evaluate the method under actual compressed‑cache decoding, where multi‑compressor training yields a 6‑22x relative lift over the unaided base on evidence‑retaining examples. Controlled‑deletion experiments show that the learned behavior is driven by evidence content rather than input length alone.

Authors:Alexandre V. Delazeri, Gabriel E. Lima, Eduil Nascimento, Rayson Laroca, David Menotti
Title: Evaluating 2D and 3D-Aware Vision Foundation Models for Vehicle Attribute Recognition
Abstract:
Vehicle attribute recognition is an important task in intelligent transportation systems, particularly when Automatic License Plate Recognition (ALPR) is unavailable or unreliable. Although vision foundation models have shown strong transferability across domains, their effectiveness for fine‑grained vehicle classification remains underexplored. Moreover, given the inherently three‑dimensional structure of vehicles, it is unclear whether emerging 3D‑aware foundation models offer advantages over standard 2D architectures. This paper presents an empirical benchmark of 14 state‑of‑the‑art 2D and 3D‑aware vision foundation models. Using the challenging real‑world UFPR‑VeSV dataset, we evaluate these models as frozen feature extractors via linear probing for vehicle type, make, and model recognition. We further stress‑test the best‑performing models under few‑shot learning and Out‑of‑Distribution (OOD) domain shifts. Our results show that standard 2D self‑supervised models, particularly DINOv3, substantially outperform 3D‑aware models in fine‑grained tasks, achieving over 93% Macro‑Accuracy for make and model recognition. However, the 3D‑aware Depth Anything v2 exhibits stronger invariance to viewing angles in vehicle type classification. These findings motivate hybrid approaches that combine 2D and 3D priors for robust vehicle recognition. Our code is publicly available at https://github.com/UFPR‑IPASPPR/3D‑Vision‑Benchmark/.

Authors:Chandler Timm C. Doloriel, Yunbei Zhang, Sarthak Kumar Maharana, Muhammad Salman Siddiqui, Tor Kristian Stevik, Fadi Al Machot, Kristian Hovde Liland, Habib Ullah
Title: Towards Continual Test-Time Adaptation of Vision-Language Models in Open-Vocabulary Semantic Segmentation
Abstract:
Open‑vocabulary semantic segmentation (OVSS) relies on vision‑language alignment to recognize arbitrary text‑defined categories, yet this alignment is fragile under continual test‑time distribution shift. Our diagnostic analysis reveals that entropy minimization drives patch‑level class collapse, continual updates erode vision‑language alignment, and redundant gradients from low‑shift samples waste computation. We propose Diversify, Anchor, and Filter (DAF), a stabilization framework that augments entropy‑based adaptation with a marginal diversity loss that resists collapse, a cross‑modal anchor consistency loss that constrains feature drift relative to a frozen source model, and feature salience filtering that skips low‑value backward passes to offset part of the source‑anchor overhead. We evaluate on five datasets spanning natural scenes, autonomous driving, underwater imagery, and remote sensing with their corrupted variants. Across the evaluated continual shifts, DAF remains stable where entropy minimization collapses, improving mIoU by over 8 points on Pascal VOC20‑C, over 9 points on LoveDA, and over 3 points on Foggy Cityscapes compared to the source model, and is robust to aggressive adaptation and learning rate choices.

Authors:Chandler Timm C. Doloriel, Yunbei Zhang, Muhammad Salman Siddiqui, Tor Kristian Stevik, Fadi Al Machot, Kristian Hovde Liland, Habib Ullah
Title: Continual Test-Time Adaptation via Entropy Sensitivity-Guidance in Strict Online Setting
Abstract:
Test‑time adaptation (TTA) promises robustness under distribution shift by updating a pretrained model on unlabeled test data, but strict online TTA with batch size one and no access to source data is especially prone to drift or collapse. We introduce Sensitivity‑Guided Erasing Adaptation (SEGA), a method for strict online continual TTA (CTTA) on corruption‑style streams. SEGA uses a small number of structured erasures to probe how predictive entropy changes as information is removed, and uses the resulting per‑sample sensitivity trajectories to coordinate recovery and sample selection rather than relying on raw entropy or batch statistics. This yields a practical feedback signal for long‑horizon batch‑size‑one adaptation without periodic resets or model reservoirs. In experiments on ImageNet‑C, CIFAR10/100‑C, and corruption‑generated aquaculture streams treated as controlled corruption‑style proxies, SEGA yields consistent robustness and stability gains over strong CTTA baselines while reducing backward passes through sensitivity‑based gating.

Authors:Gabriele Trivigno, Marcos Alfaro, Claudia Cuttano, Gabriele Berton, Luis Payá, Carlo Masone
Title: FoundYou: A Unified Model for Personalized Segmentation and Retrieval
Abstract:
Personalized segmentation and personalized retrieval both aim to identify the same physical object across different images. While the former localizes the object within a target image, the latter retrieves images where it appears. Despite this shared instance‑level objective, the two tasks have largely evolved separately and are addressed with distinct solutions. In this work, we introduce FoundYou, a unified framework built on the observation that Segment Anything 2 (SAM 2), trained to preserve object identity across video frames, inherently captures instance‑level cues. We leverage this property to match objects across independent images, enabling segmentation and retrieval to emerge as two outcomes of the same instance alignment process. This unified view unlocks new capabilities beyond traditional benchmarks, including few‑shot personalized retrieval and promptable personalized segmentation with flexible prompts. Extensive experiments show consistent gains over unified and task‑specific methods, including +18.4 mIoU on PerMIS and +17.8 mAP on ILIAS. Performance scales with additional references and remains robust to weaker prompts. Beyond personalization, FoundYou achieves state‑of‑the‑art results on category‑level retrieval benchmarks. Notably, our approach keeps the SAM 2‑small model entirely frozen and adds only 5.9 M trainable parameters, yielding a 52 M‑parameter model that is over 75x faster and 20x smaller than the only prior unified solution. Code is available at https://github.com/ga1i13o/FoundYou .

Authors:Tomohiro Aizawa, Shigeru Kuriyama, Chunzhi Gu
Title: OrnaStyler: Ornament-Aware Latent Editing for Content-Preserving 3D Stylization
Abstract:
Text‑guided style editing of 3D assets is essential for adapting existing objects to diverse visual aesthetics in digital content creation. Despite rapid progress in 3D shape modeling, faithfully stylizing an existing asset remains challenging when the desired stylization involves fine‑grained structural ornamentation, which requires the model to preserve the source geometry and object identity, while coherently integrating new style‑specific details. We propose OrnaStyler, a zero‑shot framework for text‑guided ornament‑aware 3D stylization. Built upon rectified flow‑based generative modeling, OrnaStyler introduces an inversion‑guided editing strategy that recovers content‑aware latent representations at both geometry and appearance levels in a staged manner to facilitate faithful editing. Our core idea is to explicitly model the spatial configuration of stylistic elements, thereby mitigating the fundamental tension between content preservation and style expression in the voxel space. Specifically, at the geometry level, we manipulate voxel representations through flow inversion to synthesize ornament‑enhanced structures while preserving the spatial identity of the source asset. Then, at the appearance level, we introduce an adjacency‑aware feature inpainting mechanism to harmonize newly generated ornaments with the original content, yielding coherent geometry‑appearance integration. Our approach operates solely in the inference phase and enables selective editing over geometric augmentation or appearance stylization. Extensive experiments on both generated and real‑world 3D assets against prior methods demonstrate that OrnaStyler achieves state‑of‑the‑art editing performance in terms of content preservation, style fidelity, and overall visual realism. Code is available at: https://github.com/tomohiro0427/OrnaStyler

Authors:Hai Nguyen-Truong, Tuan-Anh Vu, Dang Huynh
Title: Off-Manifold Refinement: Guiding Video Generators with a Frozen World Model
Abstract:
Modern video generators routinely fail at physical dynamics: objects float, trajectories violate gravity, contacts vanish. Standard denoising and flow‑matching objectives fit visual data distributions but do not explicitly penalize such physical violations. Existing remedies can improve physical consistency, but typically add substantial inference or training cost. Candidate‑selection methods generate and score multiple videos, while gradient‑based world‑model guidance repeatedly decodes and re‑encodes intermediate estimates. Generator‑internal refinement adds perturbation and re‑denoising loops, whereas post‑training requires curated data and additional optimization. We propose Off‑Manifold Refinement (OMR), an inference‑time method that instead injects world‑model feedback directly into a single sampling trajectory. During scheduled middle ODE steps, we augment the generator velocity with the gradient of an adapter‑space V‑JEPA 2.1 surprise energy. This external correction can move the latent away from the uncorrected sampling trajectory and toward regions ranked as more physically plausible by the frozen predictor, after which the generator continues rendering from the corrected state. A small trained latent‑to‑embedding adapter keeps the gradient tractable at inference, and both the video generator and the world model remain frozen. On our fixed 400‑prompt VideoPhy‑2 detailed subset, OMR lifts the joint Semantic‑Adherence‑and‑Physical‑Commonsense metric from 47.0% to 52.0% (+5.0pp absolute, +10.6% relative) over the base Wan2.2‑T2V‑A14B sampler. On a separate fixed 50‑prompt efficiency subset, it requires 1.71 × the base runtime rather than the multiplicative cost of reward/search alternatives. Project page: https://itruonghai.github.io/omr.

Authors:Devrim Çavuşoğlu, Emre Akbaş
Title: REIGN: Refurbished Embeddings with Integrated Guidance Networks for Efficient Context-Length Scaling
Abstract:
Dense retrieval over long documents is expensive. Token‑level encoders scale quadratically in sequence length, and most long‑context embedding models reach 32K tokens only through architectural workarounds or by stretching billion‑parameter LLMs. We propose REIGN (Refurbished Embeddings with Integrated Guidance Networks), a contrastively trained bi‑encoder that operates on sequences of contextualised chunk embeddings from a frozen Guidance Network (GN) rather than on raw tokens. REIGN targets multi‑chunk inputs, primarily for document‑to‑document retrieval; single‑chunk inputs stay with the GN. Decoupling token‑level processing from document‑level reasoning, and caching the GN embeddings to disk, cuts per‑document training cost by roughly four orders of magnitude relative to chunked Transformer fine‑tuning. We also release a synthetic long‑document retrieval benchmark for contrastive training and evaluation at long context lengths. Across an in‑distribution Wikipedia benchmark, the LoCo out‑of‑distribution suite, and a real‑world patent retrieval case study, REIGN matches dense long‑context retrievers at smaller parameter budgets in each regime. A paired significance test puts it on par with models 1.6‑4.3x larger on the patent task, and it stays within 0.65 nDCG@10 of a 20x‑larger model on LoCo.

Authors:Muxin Liu, Tianbo Liu, Jing Xia, Xiaoyang Lyu, Xiaoshan Wu, Bo Wang, Peng Dai, Zhongrui Wang, Shaoshuai Shi, Xiaojuan Qi
Title: OptiGeo: Efficient Monocular Geometry for Embodied Perception in Optically Challenging Scenes
Abstract:
Monocular depth estimation has achieved strong open‑domain generalization, yet reliable robotic deployment remains difficult in transparent, reflective, and specular environments, where depth sensors often produce missing or biased depth. Existing methods often handle such optical failures with scene‑specific preprocessing, auxiliary modules, or post‑hoc fine‑tuning. While effective in constrained settings, these designs increase architectural redundancy and can over‑specialize general geometry models to narrow optical scenarios. We revisit this problem as a localized failure mode within base‑model training and identify sensor‑induced supervision bias as a key bottleneck: models inherit sensor failure patterns from biased real‑depth supervision in optically challenging regions. We then introduce OptiGeo, a bias‑aware training framework that rehabilitates biased real supervision using a clean‑geometry teacher and residual‑trimmed alignment. We redefine transparency‑targeted rendering as a compact source of clean optical geometry, rather than a large domain‑specific fine‑tuning set. With only a small targeted rendering set, OptiGeo learns the geometric structure of transparent objects and regions, correcting local geometry distortions that real sensors cannot reliably supervise. Despite only 30M parameters, OptiGeo outperforms substantially larger 300M‑scale monocular models and billion‑scale multi‑view baselines on transparent‑scene benchmarks, while remaining competitive on general zero‑shot depth and boundary sharpness. Real‑world navigation cases further validate its practicality as an efficient perception module in optically challenging scenes.

Authors:Yutian Jiang, Ruijie Li, Sisuo Lyu, Xixuan Hao, Qingxiang Liu, Yongzi Yu, Yuxuan Liang
Title: Perceive to Hypothesize, Verify to Ground: An Agentic Reasoning Framework for Open-World Geo-Localization
Abstract:
Open‑world geo‑localization requires models to reason over ambiguous visual cues through multi‑step reasoning and external knowledge grounding. While recent large vision‑language models exhibit strong multimodal reasoning capabilities, existing approaches still suffer from perceptual hallucination and context drift due to the lack of explicit evidence‑grounded verification. In this work, we reformulate geo‑localization as a human‑like perceive‑then‑verify reasoning problem and propose GeoPAVE (Geo‑localization Perception‑and‑Verification‑Engine), a bi‑level agentic framework that contains perception‑based hypothesis generation via single‑pass rollouts and verification‑based evidence grounding for decision actions: support, refute, and refine. To support rigorous evaluation, we further introduce PAVED, a novel dataset derived from real‑world user check‑in data, equipped with comprehensive reasoning trajectories featuring multi‑hop queries, multi‑round tool invocations, and structured perception‑verification traces. The dataset and code are available at https://github.com/Arandinglv/GeoPAVE.

Authors:Yifan Chen, Haitao Li, Qingyao Ai, Fengbin Zhu, Tat-Seng Chua, Min Zhang, Yiqun Liu
Title: GenRubric: Self-Evolving Rubric Generation for Scalable LLM Evaluation
Abstract:
Large language models are increasingly used as scalable evaluators for open‑ended tasks. However, many LLM judges derive query‑specific criteria during scoring, leaving the evaluation requirements insufficiently specified and their coverage difficult to audit. Query‑specific rubrics make these requirements explicit, but expert‑written rubrics are costly to construct, while existing automatic methods typically rely on inference‑time refinement or external supervision. We introduce GenRubric, a self‑evolving framework that improves rubric generation from unlabeled queries without requiring additional human annotations during self‑evolution. Our approach is based on rubric‑induced self‑consistency: independently sampled rubrics for the same query provide partial views of its latent evaluation requirements, and a comprehensive rubric should induce a response that generalizes across these complementary evaluation views. We implement this principle through reinforcement learning, combining a cross‑rubric comprehensiveness signal with group‑level and criterion‑level rewards for rubric quality. We train GenRubric models at 4B, 8B, and 14B scales across multiple domains. Experiments on human‑annotated rubric benchmarks show that self‑evolution improves the agreement between evaluations induced by generated rubrics and those induced by expert‑written rubrics. The improvements further generalize to held‑out domains, demonstrating the potential of self‑evolving rubric generation for scalable and query‑specific LLM evaluation. Code and models are publicly available at https://github.com/foggpoy/GenRubric.

Authors:Amir Abbes, Ines Harrabi, Lucas Justin Yirepoa Kinda, Rim Trabelsi, Adnane Cabani, Fatma Abdelkefi
Title: MariSat: A Maritime Dataset for Instance Segmentation of Objects in Satellite and Aerial Images
Abstract:
Automated maritime surveillance from satellite and aerial imagery requires large, precisely annotated datasets, which remain scarce for the instance‑segmentation task, particularly for small vessels in cluttered port environments. We present MariSat, a new benchmark dataset of 1260 aerial and satellite images covering diverse port and coastal scenes, annotated at the pixel level for eight maritime object classes (sailboat, yacht, jet‑ski, fishing boat, cruise ship, military vessel, tugboat and cargo ship). The dataset was produced through a semi‑automatic annotation pipeline combining the textpromptable segmentation model SAM 3 with a cascade of geometric and colorimetric post‑processing filters, followed by a manual correction and quality‑control pass performed with the CVAT annotation platform. We describe the image‑collection methodology, the annotation and correction process, and the resulting data organization. We also report class‑wise statistics for the training, validation, and test splits. MariSat has already been used to fine‑tune and benchmark segmentation and detection models (SAM 3 and YOLO11) for real‑time maritime monitoring. We report detailed quantitative and per‑class results for both tasks. The MariSat dataset is publicly available on GitHub : https://github.com/amirabbes/P2M‑Maritime‑Segmentation

Authors:Shaghayegh Kolli, Sina Emami, Moreno D'Incà, Pouyan Nejadi, Nicu Sebe, Massimiliano Mancini, Jana Diesner
Title: ContextBias: Controlled Evaluation of Bias Persistence Under Context Shift in Text-to-Image Models
Abstract:
Text‑to‑image models learn associations between concepts ‑ in the case of this paper, people's professions, which we refer to as roles ‑ and visual attributes. These associations can underpin many observed forms of stereotypical bias. A key open question in this area is whether these associations are stable or change when visual representations of people in professional roles are placed in different prompted contexts. We introduce ContextBias, a controlled evaluation framework, and ContextBench, a benchmark spanning 92 roles and 1,656 semantically controlled prompts, designed to isolate the effect of contextual variation on role‑linked visual representations. Evaluating four state‑of‑the‑art models on 66,240 generated images, we find that placing a role in a semantically unrelated context does not suppress role‑linked attributes; instead, cross‑role attribute concentration increases (pooled BI +0.047). Demographic cues, characteristic garments, and role‑specific tools remain highly prevalent across context‑free, related, and unrelated conditions, and are robust to semantic prompt reformulation. Scene composition and camera framing show the greatest context‑sensitivity. These findings reveal a form of stereotypical persistence that remains largely invisible to context‑free evaluations, highlighting the need for controlled contextual variation in bias benchmarking. Code and dataset: https://huggingface.co/datasets/shaghayegh/ContextBias , https://github.com/Sina‑Emami/ContextBias

Authors:Doyeon Kim, Suyoung Bae, Yumin Lee, Jee-Hyong Lee
Title: A^2Agent: Action-Aware Reinforcement Learning for Repository-Level Code Localization Agents
Abstract:
Localizing issue‑relevant code regions is a critical step in automated software engineering. However, due to their reliance on sparse trajectory‑level signals, existing methods cannot identify which per‑turn actions are effective and often discover correct code regions during exploration but fail to commit them. To address these limitations, we propose an action‑aware reinforcement learning method that combines a per‑turn reward sequence rewarding both the discovery and commitment of gold code regions with an action‑level advantage estimation scheme that isolates each action's credit by grouping turns sharing the same exploration context. Extensive evaluations show that our method improves the average F1 over the state‑of‑the‑art (SOTA) by 1.58% on SWE‑Bench Verified and 8.55% on SWE‑Bench Pro, with our 4B model outperforming baselines up to 8x larger. Our code is available at https://github.com/donian00/A2Agent.

Authors:Torsten Keßler, Bas W. T. Gieling, René R. Hiemstra, Michael R. A. Abdelmalik
Title: Wigner-Eckart Factorization of the Polyatomic Boltzmann Collision Operator
Abstract:
We extend the Wigner‑Eckart factorization of the spectral Boltzmann collision operator to polyatomic gases with continuous internal energy. Because internal energies are invariant under spatial rotations, the SO(3) reduction survives the Borgnakke‑Larsen energy exchange, and the twelve‑dimensional collision integral collapses onto a nine‑dimensional kinematic core. The core splits into a sparse geometric tensor, evaluated exactly, and a dense physical tensor, integrated by singularity‑resolving Gauss rules with an auxiliary Laplace representation of the fractional energy couplings. The quadrature attains near machine precision at the fractional exponents of real gases. The collision invariants are embedded exactly, preserving the translational‑internal energy exchange. The factorization compresses the operator by three to nearly four orders of magnitude and accelerates its evaluation 40‑fold over dense formulations. The method is validated against the exact monatomic limit, Landau‑Teller relaxation, and an analytic frozen‑channel Prandtl number, and it matches a published calibration of the same kernel for N2, CO, and H2.

Authors:Juneyong Lee, Jaeyoung Choi
Title: Null-Space Diffusion Restoration with Adaptive Uncertainty-Guided Fusion for Ultrasound Speckle Reduction
Abstract:
Ultrasound B‑mode imaging commonly suffers from speckle noise and artifacts, requiring a delicate balance between contrast, resolution, and preservation of anatomical structures. Although recently developed despeckling methods have achieved some progress, supervised learning approaches remain fundamentally limited by the ground truth paradox, which arises from the absence of noise‑free, ground truth reference images in in vivo scenarios. Existing unsupervised diffusion‑based methods typically enforce data consistency directly in the nonlinear log‑compressed domain, which can disproportionately amplify background artifacts when mapped back to the envelope domain. To overcome these limitations, we propose an uncertainty‑guided null‑space diffusion (UGNS) framework, a novel label‑free solution that enforces consistency correction on a stabilized positive‑envelope proxy obtained via inverse log compression. The proposed UGNS introduces several technical novelties: (a) extraction of a structural prior in the stabilized envelope domain to produce a robust signal envelope that preserves anatomical structure, (b) development of an adaptive range‑null reconstruction mechanism that uses an adaptive weight mask to preserve tissue regions via range‑space projection, and (c) introduction of uncertainty‑guided fusion in an adaptive way to mitigate sampling variability. Extensive and comparative experiments were conducted using the PICMUS benchmark and in vivo datasets. The results demonstrate that UGNS achieves competitive generalized contrast‑to‑noise ratio (gCNR) values across diverse datasets. In addition, it is successfully validated that UGNS effectively suppresses speckle noise while preserving fine spatial resolution. Code is available at https://github.com/yousirong/UGNS.git.

Authors:Peizheng Li, Xin Ai, Hanyuan Liu, Qiange Wang, Yanfeng Zhang
Title: RegionCache: Semantic-Aware Region Reuse for Efficient Multi-Turn Image Generation
Abstract:
Real‑world image generation often involves multi‑turn editing, where users iteratively modify small regions while most image content remains unchanged. However, existing diffusion transformer (DiT)‑based editing pipelines recompute the entire image at every turn, causing substantial redundant computation. Existing DiT acceleration methods further ignore semantic correspondence across prompts, leading to unnecessary recomputation or unsafe reuse that harms editing quality. To address this, we propose RegionCache, a semantic‑aware reuse framework for multi‑turn image editing that selectively reuses diffusion states from unchanged regions. RegionCache detects reusable regions through semantic overlap between consecutive prompts and cross‑attention localization, and adopts an adaptive reuse schedule based on prompt similarity and contextual consistency. Experiments on PixArt‑alpha demonstrate that RegionCache achieves 1.43x‑‑2.55x end‑to‑end speedup while maintaining comparable image quality. Code is available at https://github.com/hebutBryant/RegionCache.

Authors:Lin Chen, Yitong Chen, Yong Li
Title: Do LLMs Change Their Minds Like Humans? Diagnosing Human--LLM Divergence in Single-Turn Persuasion Judgments
Abstract:
Large language models (LLMs) are increasingly deployed as proxies for human participants in social simulations, yet whether they update their beliefs in response to persuasive arguments, as humans do, remains poorly understood. We conduct a systematic comparison using a naturally occurring online persuasion corpus in which original posters explicitly verify whether a reply changed their view. Our results show that LLMs achieve only slight agreement with humans (Cohen's kappa ranging from 0.079 to 0.178). Content‑level analyses show that humans and LLMs agree on the strongest persuasion cues but diverge on finer ones: humans are more swayed by novel content and assertive language, whereas LLMs favor topical similarity and surface‑level formatting. At the level of persuasion strategy, LLMs underweight emotional appeals and overweight credibility signals relative to humans, while the type of proposition under debate exerts no measurable effect on the degree of divergence. Furthermore, switching from first‑person role‑playing to third‑person observation shifts all models toward greater resistance to persuasion, with the effect varying across persuasion strategies and textual features. These findings highlight the risk of treating LLM judgments as faithful proxies for human belief updating and point to structural differences in how LLMs and humans process persuasive discourse. Our code is available at https://github.com/tsinghua‑fib‑lab/LLM‑belief‑update‑cmv.

Authors:Gonglin Chen, Ben Southall, Hanyuan Xiao, Wenbin Teng, Haolin Xiong, Tianwen Fu, Junyi Ouyang, Kshitij Singh Minhas, Supun Samarasekera, Rakesh Kumar, Yajie Zhao
Title: XDG: Accelerated Visual Disambiguation
Abstract:
Visual aliasing, also known as the doppelganger problem, remains a key challenge for structure‑from‑motion (SfM): visually similar but physically distinct surfaces can produce incorrect image matches and degrade reconstruction quality. Previous work mitigates this issue with geometry‑aware foundation‑model features, but places a heavy transformer classifier on top of the backbone, making large‑scale disambiguation expensive. We introduce XDG, an efficient visual disambiguation model designed for scalable SfM. Our key observation is that a 3D foundation model already performs the cross‑view geometric reasoning necessary for visual disambiguation, so doppelganger classification should adapt the backbone representation directly rather than relearn pair reasoning in a separate heavy decoder. XDG fine‑tunes Depth Anything 3 with lightweight LoRA adapters and repurposes its camera tokens as compact pair‑level classification tokens. A compact MLP head predicts whether a candidate image pair observes the same 3D surface. Extensive experiments show that XDG provides a favorable accuracy‑efficiency tradeoff: it remains competitive with the state‑of‑the‑art disambiguation method across pairwise and reconstruction benchmarks and delivers more than a 3x inference speedup. On individual LaMAR scenes containing thousands of images, XDG saves more than 10 hours of visual disambiguation processing. Code is available at https://github.com/xtcpete/xdg.

Authors:Bomiao Wang, Zekai Shao, Jiexiang Lan, Xiaoliang Fu, Xingchen Zeng, Siming Chen
Title: DVBench: Benchmarking MLLMs for Understanding Dynamic Charts and Narratives in Data Videos
Abstract:
While MLLMs have made significant strides in chart comprehension and video understanding, current evaluations largely isolate these capabilities, leaving a critical gap in understanding temporally evolving structured visual information. To address this gap, we introduce DVBench, a benchmark for evaluating MLLMs on data videos, a storytelling medium that integrates dynamic charts with structured narratives. We decompose data video understanding into five dimensions. DVBench comprises 300 real‑world data videos and 1,000 human‑verified QA pairs curated through a rigorous semi‑automated pipeline. Extensive evaluations of nine MLLMs show that Gemini‑3.1‑Pro achieves the best overall performance, while Kimi‑k2.5 is the strongest open‑source model. We further identify two notable phenomena: open‑source model performance does not scale strictly with parameter size, and narrative proficiency does not guarantee visual capability. Fine‑grained analyses and ablation studies further reveal dimension‑specific weaknesses and the effects of frame configurations and subtitle inputs, informing future MLLM development. DVBench is publicly available at https://bomiaowang.github.io/DVBench/.

Authors:Zhiyu Chen, Keyu Zhao, Jigao Fu, Dong Liang, Yanbiao Wu, Jiaoyang Li, Haidong Xue, Xinhua Zeng, Yuanyi Zhen, Fengli Xu, Yong Li
Title: Ideation Arena: Evaluating LLM Generated Research Ideas with Battle-style Human Expert Assessment
Abstract:
Evaluating research ideas generated by LLMs is difficult because their scientific value cannot be fully determined by objective criteria, and no single reference answer specifies what counts as a good idea. To address this challenge, we introduce Ideation Arena, a battle style platform that evaluates research ideas through pairwise human assessment. Ideation Arena evaluates ideas generated by 14 frontier LLMs and 5 research agent architectures built on 2 base models. To ensure a common starting point, Ideation Arena builds shared literature contexts from papers familiar to the participating researchers and provides the same contexts to all LLMs and agents. We collect over 6,000 double blind pairwise comparisons from 105 active computer science researchers and construct an Elo rating leaderboard of proposal‑stage expert preferences in computer science under a shared closed‑context protocol. We validate the rankings through interrater agreement and robustness analyses, showing that the leaderboard remains stable under changes in annotator composition and domain coverage. Our results show substantial variation in agent effectiveness, with some frameworks improving ideation quality over their backbones and others offering little benefit or even underperforming their base models. We further construct Ideation Arena Eval, a benchmark for assessing whether automated evaluators align with human preferences in research ideation. Experiments with current LLM judges show that they still cannot reliably reproduce expert preferences, with the best judge reaching 72.56% Soft Accuracy on Overall Quality. Our code, data, and leaderboards are available at https://github.com/foss12138/Research‑Ideation‑Arena.

Authors:Zhe Dong, Wanqing Wu, Yuzhe Sun, Haochen Jiang, Yuchen Ma, Lecheng Ren, Tianzhu Liu, Yanfeng Gu
Title: GeoRay: Gauge-Aware Feed-Forward Satellite 3D Reconstruction in the Geodetic Frame
Abstract:
Feed‑forward 3D foundation models reconstruct perspective scenes in one pass. Satellite photogrammetry needs a different product, one that domain adaptation alone does not deliver: dense surface height in an absolute geodetic frame under non‑central rational polynomial cameras (RPCs). Perspective‑pretrained features are not reliably observable along RPC height rays, absolute elevation carries a low‑order height‑‑datum gauge exchangeable with sensor bias to first order, and monocular and multi‑view cues fail in different regions. \method treats all three. Lightweight ray‑consistent adapters make a frozen backbone matchable along native RPC rays. An explicit datum mechanism separates relief from absolute level and is equivariant to the vertical origin by construction, so one trained model serves zero‑, one‑, and sparse‑control inference. Calibrated inverse‑variance fusion combines the two relief streams. \bench, our absolute‑frame benchmark of eighteen systems across in‑domain, cross‑dataset, and cross‑city tiers, scores absolute placement without registration or test‑reference leakage. On 26 held‑out US3D tiles, \method attains 2.99\,m absolute MAE at 91.9% coverage, improves completeness‑aware accuracy by 46.4 points over the strongest compliant feed‑forward baseline, remains the most accurate such system under both transfer shifts, and runs in 24\,s model‑forward time per tile. Code and models will be released at https://github.com/HIT‑SIRS/GeoRay

Authors:Xinke Jiang, Zhixin Zhang, Zhibang Yang, Jiaran Gao, Rihong Qiu, Shijin Chen, Xu Chu, Junfeng Zhao, Yasha Wang
Title: Harness-RL: Black-Box Reinforcement Learning with Action-Args Decoupling for Central-Agent Multi-Agent Harnesses
Abstract:
Large language model agents increasingly solve long‑horizon tasks through multi‑agent harnesses in which a central agent coordinates specialized sub‑agents, tools, and environments. Training the central policy in such a harness raises two challenges. First, an action label is a low‑cardinality decision, whereas its args form a high‑dimensional conditional sequence; optimizing both with a shared sequence‑level signal can produce conflicting gradients. Second, dynamic scheduling creates interdependent sessions with branches, parallel calls, and rewritten contexts, which cannot be faithfully reduced to one flat token sequence. We introduce Harness‑RL, a structured reinforcement learning framework that combines Conflict‑Aware Policy Optimization (CAPO) with interface‑level black‑box trajectory construction. The black‑box component captures Interface Call Records, builds per‑session prefix trees, and aligns outcome and process rewards with trainable tokens. CAPO uses forward activations to identify parameter partitions associated with action and args tokens, then routes their policy gradients to the corresponding subspaces. Harness‑RL supports both central‑only and joint multi‑agent training. Across seven multi‑hop question answering and agentic retrieval benchmarks, it reaches average F1 scores of 42.93 and 47.79 with Qwen2.5‑1.5B and Qwen2.5‑3B, respectively, while ablations validate the contribution of CAPO and favor central‑only optimization in the evaluated setting. Our code is available at https://github.com/jiangxinke/Harness‑RL.

Authors:Jiaze Li, Aocheng Shen, Bing Liu, Boyu Zhang, Xiaoxuan Fan, Qiankun Zhang, Xianjun Deng
Title: InteractBench: Benchmarking LLMs on Competitive Programming under Unrevealed Information
Abstract:
Competitive programming is increasingly being used to evaluate the algorithmic reasoning capabilities of large language models (LLMs). However, existing benchmarks primarily focus on full‑information tasks where all problem inputs are provided upfront. This overlooks a critical dimension of algorithmic reasoning: the ability of generated programs to operate when key information is not revealed upfront. Interactive problems, a distinctive component of competitive programming, embody this challenge. These problems require programs to engage in multi‑round interaction with an interactor (a judge program) under strict protocol constraints and limited query budgets, with new information revealed only in response to queries. To address this gap, we introduce InteractBench, a benchmark comprising 322 high‑quality interactive problems curated from Codeforces, AtCoder, IOI, and ICPC. Each problem is packaged with executable local interactors, enabling fully offline evaluation. Unlike existing benchmarks, InteractBench assesses whether model‑generated code can acquire information and track state dynamically. Our evaluation reveals a significant interaction gap: even the most advanced reasoning models achieve limited success on interactive problems. Beyond success rates, we propose a fine‑grained failure taxonomy to diagnose the root causes of these deficiencies. Although algorithmic logic errors remain dominant, protocol violations and query‑budget overruns are frequent. Code is available at https://github.com/kmsgk0/InteractBench.

Authors:Yi Xu, Ruichao Hou, Tongwei Ren, Gangshan Wu
Title: SPLG-Mamba: Structure-Preserving Local-Global Mamba Network for Salient Object Detection in Optical Remote Sensing Images
Abstract:
Salient object detection in optical remote sensing images (ORSI‑SOD) requires dense predictions that preserve object completeness and structural continuity under complex backgrounds, scale variation, and irregular object shapes. Existing methods often localize salient regions, but their predictions may still suffer from structural degradation, including fragmented, incomplete, or locally missing foreground responses. This degradation is closely related to hierarchical feature propagation, where shallow details can introduce texture‑induced background responses, deep semantics may over‑smooth weak structures, and uncontrolled cross‑scale fusion can disturb coherent regions. To address this issue, we propose a novel Structure‑Preserving Local‑Global Mamba Network, SPLG‑Mamba, for ORSI‑SOD. Specifically, SPLG‑Mamba integrates Smooth‑Detail Recalibration (SDR), hierarchy‑aware Local‑Global Mamba, and Gated Cross‑Scale Fusion (GCSF). SDR recalibrates smoothed responses and detail residuals before state‑space modeling, Local‑Global Mamba assigns local modeling to shallow feature levels and global modeling to deep feature levels, and GCSF controls cross‑scale detail injection during decoding. Experiments on ORSSD, EORSSD, and ORSI‑4199 demonstrate state‑of‑the‑art results and improved structural completeness and continuity. The code is available at https://github.com/yxu9910/SPLG‑Mamba

Authors:Xinke Jiang, Yue Fang, Zhibang Yang, Jiaran Gao, Zhixin Zhang, Tao Feng, Rihong Qiu, Wentao Zhang, Hongxin Ding, Ruizhe Zhang, Yongxin Xu, Yuheng Huang, Xu Chu, Junfeng Zhao, Yasha Wang
Title: AgenticRag-R1: Agentic Reinforcement Learning with Stack Memory for Multi-Step Reasoning, Retrieval and Memorizing
Abstract:
Retrieval‑Augmented Generation (RAG) improves the factuality of large language models (LLMs), yet existing RAG systems often struggle with complex, multi‑step reasoning that requires adaptive retrieval and continuous revision of intermediate contexts. Recent reinforcement learning (RL)‑based agentic RAG methods partially alleviate this issue, but typically rely on coarse‑grained action spaces and trajectory‑level rewards, resulting in weak reward assignment and a bias toward short‑horizon, stereotyped reasoning template. To address, we propose AgenticRag‑R1, a RL framework that deeply integrates reasoning, retrieval, and memory via a memory stack and fine‑grained action space, supported by hierarchical action‑aware rewards and an information‑aware trajectory rejection strategy to enable effective long‑horizon learning. Experiments across a diverse set of multi‑hop, open‑domain, and agentic reasoning benchmarks, spanning multiple backbone model sizes, demonstrate that AgenticRag‑R1 consistently outperforms strong baselines. Moreover, AgenticRag‑R1 learns more robust, interpretable, and memory‑aware reasoning behaviors, highlighting the effect of fine‑grained action modeling and information‑aware optimization for long‑horizon reasoning. Our code is anonymous available at https://github.com/jiangxinke/Harness‑RL/tree/AgenticRAG‑R1‑Whitebox.

Authors:Jieying Xue, Phuong Minh Nguyen, Minh Le Nguyen, Shogo Okada
Title: Cross-lingual Functional Vectors for Emotion Detection in Large Language Models
Abstract:
Function vectors (FVs) have recently emerged as a promising mechanism for steering the behavior of large language models (LLMs) by injecting task‑specific latent direction representations derived from in‑context demonstrations. While prior studies have shown that FVs can recover task behavior in structured in‑context learning settings, their effectiveness on semantically complex tasks and their ability to generalize across languages remain underexplored. We investigate the cross‑lingual transferability of FVs using multilingual multi‑label emotion recognition as a challenging semantic classification benchmark. Specifically, we examine whether FVs extracted from a source language can steer task behavior in another language under both standard clean and perturbed zero‑shot settings without providing demonstrations during inference. Across diverse cross‑lingual settings, applying FVs substantially improves performance, suggesting that FVs capture language‑agnostic, task‑relevant signals rather than purely language‑specific lexical patterns, and highlighting their potential as a lightweight and transferable mechanism for multilingual task adaptation. We observe that each LLM exhibits a relatively stable optimal range of attention heads for constructing effective FVs, and the pattern remains consistent across languages. In addition, FVs can partially replicate the task‑steering effects of standard few‑shot in‑context learning while avoiding the computational overhead of processing multiple demonstrations, making them effective for large‑scale practical applications. Our code is available at https://github.com/yingjie7/cross_lingual_fvs.

Authors:Ming-Han Lee, Chi-Yeh Chen
Title: nnMNet: Baseline for Martian Terrain Semantic Segmentation
Abstract:
Semantic segmentation is a crucial task for understanding Mars, the most Earth‑like planet in our solar system. However, it is challenging because the Martian surface is highly unstructured and complex, making accurate pixel‑level prediction and fine‑grained annotation difficult. Recent advancements in deep learning have introduced numerous methods and datasets to address these challenges. Nevertheless, the field lacks a robust, publicly available, and reproducible baseline, as well as a unified benchmark to facilitate fair evaluations. In this work, we present nnMNet, a new baseline model designed for Martian terrain semantic segmentation. Building upon nnWNet, we integrate linear attention to better capture global context and employ lightweight convolutions to reduce computational overhead. To bridge the gap between local and global representations, we introduce the Spatially‑Aware Fusion Block (SAFB), which augments and combines features with diverse characteristics. Furthermore, we establish a new benchmark by curating and standardizing three high‑quality datasets for thorough evaluation. nnMNet achieves new state‑of‑the‑art 86.61%, 83.25%, and 88.24% mIoU on SynMars‑TW, SynMars‑Air, and MarsScapes, respectively. Our code, models, and datasets are publicly available at https://github.com/dereklee0310/nnMNet.

Authors:Zachary Coalson, A M Aahad, Stella Doehring, Zane Ma, Sanghyun Hong
Title: On the Resilience of Text-to-Video Diffusion Models to Hardware Faults
Abstract:
We present the first systematic study of the resilience of text‑to‑video (T2V) diffusion models under random hardware‑level faults. While T2V models are widely used for automated video generation due to their ability to produce high‑quality, temporally coherent, and realistic videos, their iterative denoising process and spatiotemporal dependencies introduce unique failure modes. We perform an extensive fault‑injection study covering both computational and memory faults across three T2V models and a representative benchmark. Our results show that (1) a single fault can degrade overall performance by up to 3.7%, with semantic correctness more affected than perceptual quality; (2) memory faults are more damaging than computational faults, high‑order exponent bits are particularly vulnerable, and the widely‑used bfloat16 is more susceptible than alternative formats; and (3) 7‑28% of faults cause visible artifacts, including semantic changes such as added objects, suggesting that single faults are sufficient to alter output semantics. Our findings reveal reliability risks in deployed T2V systems and motivate further research on improving fault resilience. Code: \hrefhttps://github.com/ztcoalson/T2V‑Resiliencehttps://github.com/ztcoalson/T2V‑Resilience.

Authors:Minkyu Kim, Juhwan Choi, YoungBin Kim
Title: Not Safe for All: Auditing the Dialect Penalty in Text-to-Image Safety Pipelines
Abstract:
Text‑to‑image (T2I) safety guardrails fail to generalize equitably to non‑standard dialects. Evaluating 23,080 paired prompts across five English dialects, we formalize this failure as the dialect penalty, where filters trigger based on linguistic surface features rather than semantic intent. Text‑level filters fail in opposing directions: NSFW‑T over‑flags benign dialect prompts and LatentGuard over‑flags toxic ones (bias gaps up to +28.29 pp), while the OpenAI Moderation API under‑detects them. A controlled typo ablation confirms this penalty originates from flagging dialectal features, not generic out‑of‑distribution sensitivity. The pixel‑level generator is largely dialect‑agnostic; the penalty enters at text processing and cascades unevenly to post‑hoc guardrails. We show this bias tracks training data imbalance and is mitigable via group‑balanced retraining, with an ablation attributing the gain to balanced exposure rather than to the worst‑group objective of GroupDRO (group distributionally robust optimization). Current pipelines systematically fail dialect speakers, an equity failure masked by mean accuracy benchmarks. Our official code and dataset are publicly available at https://github.com/minguinho26/dialect‑penalty‑t2i. Content Warning: This paper contains offensive, toxic, or disturbing text prompts and generated images.

Authors:Yilun Liu, Boyu Luo, Yanran Tang, Ruihong Qiu, Zi Huang
Title: Call Neighbours Yourself: Graph Walks with Destination-Conditioned On-Policy Self-Distillation
Abstract:
Reasoning over text‑attributed graphs (TAGs) requires large language models (LLMs) to combine a node's text with evidence distributed across its neighbourhood. Existing methods fix the set of accessible neighbours before generation, forcing reasoning to operate over a static context and preventing the model from acquiring missing evidence during inference. We argue that neighbour selection should itself be part of the reasoning process. To this end, we propose Call Neighbours Yourself (CNY), a framework that enables LLMs to proactively explore graph neighbourhoods through topology‑constrained graph‑walk actions. Instead of reasoning over a pre‑selected neighbour set, CNY exposes lightweight neighbour previews and learns when to expand candidate neighbours for additional evidence. To address the delayed‑credit challenge of neighbour exploration, we introduce destination‑conditioned on‑policy self‑distillation, which retrospectively evaluates a selected neighbour after its content is revealed and converts the resulting change in action preference into an action‑level training signal. Experiments on standard TAG reasoning benchmarks under a unified raw‑text setting show that CNY consistently outperforms fixed‑context post‑training baselines. Furthermore, the learned exploration policy transfers to unseen graphs and to a graph‑level task not encountered during training. Code is available at https://github.com/superallen13/CNY.

Authors:Tian Yu, Lu Feng, Sebastian Elbaum
Title: Drive the Thoughts: Runtime Monitoring of VLA Reasoning-Trajectory Consistency
Abstract:
Autonomous vehicles (AVs) operate in complex environments where failures are consequential. Sophisticated machine learning models for perception and planning are key to overcoming at least part of that complexity, but their black‑box nature complicates validation and verification (V&V). The recent integration of Vision‑Language‑Action (VLA) models into AVs introduces a unique opportunity: besides generating trajectories, these models produce an explicit Chain‑of‑Thought (CoT) explaining their underlying rationale. This CoT provides a rich specification to cross‑check model outputs and detect inconsistencies that may expose unsafe or unintended behavior. This paper assesses whether CoTs from a recent open driving VLA can support such monitoring. We curate DriveAlignBench, a specialized dataset from NVIDIA's Alpamayo 1.5 VLA for AVs containing 150 CoT‑trajectory pairs, which we manually annotate for reliability, trajectory consistency, and safety. Our analysis reveals that 33.3% of CoTs are unreliable. Among reliable CoTs, the generated trajectory is consistent with the CoT in 74% of cases. Leveraging this potential, we propose integrating a CoT‑trajectory consistency check into a runtime monitor. The check is nontrivial: CoTs express open‑vocabulary, scene‑relative driving commitments, while trajectories are low‑level ego‑motion sequences whose semantics depend on road geometry and motion context. To bridge this gap, we develop a family of automated consistency monitors. Our best monitor, lane‑relative F‑LLM with GPT‑5.5, achieves F1 = 0.75, improving over the strongest raw‑waypoint LLM baseline by +0.13 absolute F1 and over a rule‑based monitor by +0.38. We release DriveAlignBench, the monitor implementations, and annotation tools at https://github.com/776styjsu/drive‑the‑thoughts.

Authors:Keigo Sakurai, Takahiro Ogawa, Miki Haseyama
Title: The Edge Spectrum of Choice-Derived Item Graphs: Strong and Weak Edges Encode Different Relations in Collaborative Filtering
Abstract:
Graph collaborative filtering relies on item‑‑item graphs whose edges are used for positive smoothing, under the implicit assumption that stronger edges encode more of the same relation as weaker ones. We show that this assumption fails for a practically important class of graphs: those whose edge weights come from a choice model. On such graphs, strong and weak edges encode qualitatively different relations, which we call an edge spectrum. Specifically, strong edges concentrate on the in‑slate competitors of clicked items, exactly the pairs that the within‑slate ranking gradient pushes apart, while weak edges do not. We formalize this as a sign mismatch between the smoothing operator and the ranking gradient, and prove that co‑click graphs cannot exhibit the same misalignment by construction. This diagnosis explains three empirical observations on MIND and EB‑NeRD: (i) drop‑in choice‑derived operators do not beat co‑click, despite indexing structurally distinct neighborhoods; (ii) uniform scalar fixes (sign flip, in‑slate margin loss) fail predictably, because the misalignment lives in the graph, not in the loss; (iii) only edge‑magnitude‑aware operators, with the regime boundary located by the diagnosis rather than by tuning, recover the predicted ordering. The neighbor cutoff k is therefore a semantic switch, not a sparsification hyperparameter. Our claim concerns which interventions fail or succeed and why, not absolute headline gains, which the diagnosis itself predicts to be small under the attenuated propagation channel we observe. We turn the diagnosis into a reusable protocol practitioners can run before deploying any choice‑derived item‑side operator. Code: https://github.com/kyomusso/Edge‑Spectrum‑in‑CF.

Authors:Qianqian Chen, Hyun Bin Kim, Denzel Elden Wijaya, Yang Yi, Bo Liu, Yangkai Ding
Title: TRINITY: A Multi-Perspective Benchmark for Personal-Style Video Highlight Detection
Abstract:
Traditional video highlight detection relies on a narrow, event‑centric definition of saliency, which often fails to generalize to unconstrained personal videos where highlights are heterogeneous and perspective‑dependent. To address this, we introduce TRINITY, a multi‑perspective benchmark that decomposes highlight saliency into three complementary dimensions, Event, Emotion, and Nature, within a unified temporal framework. Leveraging this multi‑faceted view, we propose a shared‑backbone multi‑branch architecture designed for parallel multi‑perspective prediction via view‑specific experts. Comprehensive experiments demonstrate that our method significantly outperforms state‑of‑the‑art baselines, achieving gains of +7.15/+3.62 mAP (rho=15%/50%) on Mr. HiSum and +10.82 mAP on YouTube Highlights. These results validate that multi‑perspective modeling provides a more robust and comprehensive formulation of video saliency, especially for complex real‑world scenarios. The benchmark and relevant codes will be released upon acceptance. The benchmark is available at https://huggingface.co/datasets/vanilladucky/TRINITY and the code is available at https://github.com/vanilladucky/TRINITY.

Authors:Yibo Gong, Cong Guo, Jiacheng Ding
Title: HoopMind: A Real-Time Neural Game-Tree System for Opponent-Aware Possession Planning
Abstract:
School coaches prepare for opponents with game film and intuition. The analytics tools of professional teams stay out of reach. We ask how far public data can close this gap. Professional basketball is our case study, chosen for its data rather than the league. We fuse five public sources into one per‑shot dataset of 4.23M shots over 21 seasons. The sources are shot locations, two play‑by‑play feeds, official matchup tracking, and player biometrics. Alignment across them is 99.5% to 100%. We also report two data pitfalls that are easy to miss. We then model a half‑court possession as a sequential game. Shot values come from ShotNet, an embedding multilayer perceptron (MLP). On a held‑out season it beats a zone‑rate baseline and a logistic baseline, and its probabilities are well calibrated. A depth‑limited expectimax search then solves the offensive decision tree, with branch‑and‑bound pruning to keep it real time. All training runs offline, so the online system stays light. A scouting planner and a playable simulator both run in a single browser page.

Authors:Lingfeng Yao, Chenpei Huang, Xingke Yang, Ziye Geng, Changqing Luo, Hao Wang, Jiang Liu, Miao Pan
Title: PhysWave: Physics-Guided Latent Diffusion Models for Controllable Spatial Audio Generation
Abstract:
Text‑to‑spatial audio generation, such as text‑to‑First‑Order Ambisonics (FOA), provides a convenient way to create spatial audio for billion‑dollar gaming and film industries. However, existing text‑to‑FOA methods are largely data‑driven and may produce audio that violates acoustic relations between source direction and distance. They also separate descriptive and parametric control, forcing users to trade usability for precision. In this paper, we present PhysWave, a physics‑guided latent diffusion model for controllable text‑to‑FOA generation. PhysWave unifies natural‑language and trajectory control through a shared waypoint‑caption representation, and augments diffusion training with two differentiable acoustic priors: spherical‑harmonic direction consistency and inverse‑square distance consistency. To support dynamic spatial generation, we further construct a 300K‑clip FOA dataset with diverse sound categories and source trajectories. Extensive results show that the proposed priors help PhysWave generate spatially consistent FOA audio while maintaining competitive audio quality. Further analyses show that these physics priors improve spatial consistency during training and can also be used as inference‑time guidance for training‑free spatial refinement.

Authors:Yingdan Shi, Xiang Xu, Kaize Ding, Alfred O. Hero, Ren Wang
Title: On the Plasticity Collapse in Continual Machine Unlearning
Abstract:
Machine unlearning enables deep neural networks to selectively remove the influence of specific data in response to privacy and regulatory requirements. While prior work largely studies single‑shot unlearning, real‑world systems must accommodate continual unlearning, where multiple unlearning requests occur sequentially over time. In this work, we identify a fundamental limitation of this setting: plasticity collapse, a progressive breakdown in a model's ability to effectively forget. Through theoretical analysis of continual unlearning dynamics, we show that continual unlearning operations accumulate geometric constraints in parameter space, leading to saturated subspaces that restrict future updates. This structural effect induces two distinct failure modes: (1) Forward failure ‑‑ diminishing forgetting quality for subsequent tasks, and (2) Backward failure ‑‑ spontaneous re‑memorization of previously forgotten information. Extensive experiments across multiple architectures, datasets, and methods in image classification confirm that plasticity collapse is not an artifact of specific implementations, but a pervasive phenomenon inherent to continual unlearning. Our findings reveal a critical barrier to the long‑term reliability of machine unlearning systems and motivate the development of plasticity‑preserving unlearning algorithms. Our code is available at https://github.com/TIML‑Group/Continual‑Machine‑Unlearning‑Plasticity‑Collapse

Authors:Joe Eappen, Zikang Xiong, Shreyash S. Iyengar, Suresh Jagannathan
Title: Generalizable Multi-Agent Planning from Signal Temporal Logic Specifications via Diffusion
Abstract:
Multi‑agent systems in the real‑world (e.g., drone swarms, autonomous cars, warehouse robots) must satisfy rich, temporal tasks while avoiding collisions. Signal Temporal Logic (STL) elegantly encodes such objectives, but current STL planning methods face critical limitations. State‑of‑the‑art optimization‑based approaches can handle arbitrary STL specifications but struggle with scalability, becoming computationally impractical as the number of agents grows. Learning‑based methods efficiently handle a large number of agents with rapid planning times but fare poorly when deployment‑time objectives differ from those used during training, and do not support planning tasks that require different specifications to be ascribed to different agents (i.e., heterogeneity) or team‑level specifications requiring coordination of multiple agents. This fundamental trade‑off between generalizability and scalability presents a challenge for realizing multi‑agent STL planning algorithms in practice. To overcome this challenge, we introduce a new diffusion method for multi‑agent planning with STL specifications. Using a differentiable approximation of STL, we integrate the STL gradient in the denoising process, making our approach generalizable to novel formulas whose predicates are placed anywhere within the goal region covered during training, while achieving the same scalability as existing learning‑based methods. Our method supports heterogeneous specifications, and by using diffusion models, naturally enhances plan diversity, thereby significantly reducing safety‑related violations (e.g., collisions) among agents. A detailed evaluation study justifies the utility of STL‑guided diffusion‑based multi‑agent planners for constructing generalizable, scalable, and diverse plans. Videos and code are available at https://www.jeappen.com/diff‑ma‑stl/ and https://github.com/jeappen/diff‑ma‑stl .

Authors:Zihan Wang, Anita Marie Slominska, Rennie Bimman, Elizabeth Di Flumeri, Amanda Mayappo-Neeposh, Conall Francoeur, Tamara Ellen Carver, Xiao-Wen Chang, Doina Precup, Esin Darici Haritaoglu, Ismail Haritaoglu, Akshatha Arodi, Naomi Goloff
Title: SIC-Agents: Benchmarking and Building an Adaptive Simulator for Pediatric Serious Illness Communication Training
Abstract:
Pediatric serious illness communication (SIC) is critically important, yet scalable communication training for clinicians remains limited. Compared with other dialogue simulation settings, pediatric SIC poses additional challenges, including multi‑party interactions, response to parental distress and strong dependence on feedback dynamics. Existing LLM‑based simulators optimize generic dialogue quality rather than curriculum‑contingent behavior required for effective SIC training. In collaboration with educators and pediatric clinicians, we introduce the first benchmark suite and simulation framework tailored to pediatric SIC training. Our benchmarks, PitfallBench and DialogueBench, evaluate simulators both at the turn‑level and across full dialogues. We further propose SIC‑Agents, a self‑improving framework that generates a clinician‑editable skill document to guide simulator behavior. Our experiments show that SIC‑Agents outperforms static expert prompting. To support future research, we release our benchmarks for parent simulation in pediatric SIC at https://github.com/Beikewzh/sic‑benchmarks

Authors:Jason Luo, Saibilila Abudukelimu, Judy Song, Andrew Feng, Shivank Garg, Vasu Sharma, Kevin Zhu
Title: MUDDLE: Measuring Understanding of Documents under Distractor and Length Effects
Abstract:
Document question‑answering systems increasingly answer questions over collections of retrieved documents rather than one clean source, so robustness to distracting context matters as much as reading ability. When such systems fail, it is often unclear whether the context was too long or the distractors were too close to the topic, because prior work tends to conflate these two effects. We present MUDDLE, a controlled benchmark that separates them. MUDDLE uses 270 human‑annotated questions, each tied to a single source document, and instantiates every question in five conditions: the source alone, the source with two or four topically similar hard negatives, and the source with two or four random distractors. The random distractors are matched to the hard negatives in length and provenance, so an accuracy gap between the two arms reflects topical similarity rather than length. All five conditions are rendered in markdown, page images, and raw PDF, but the distractor sweep reported here is run in markdown, since a source plus its distractors exceeds current image and PDF input limits. We score answers with an LLM judge across three model families. In the complete markdown sweep, hard negatives lower accuracy more than length‑matched random documents at both context sizes for gpt‑5‑mini, while random documents stay near the no‑distractor baseline. The effect is small but directionally consistent, and for gpt‑5‑mini hard negatives significantly underperform length‑matched random distractors when pooled across context sizes. We release the data and evaluation code for a reproducible study of context degradation.

Authors:Sisi Zhu, Changwei Yu, Renshuai Tao, Zhenliang Ni
Title: Co-Evolutionary Prompt Optimization with Cross-Category Transfer for Zero-Shot Anomaly Detection
Abstract:
Zero‑shot anomaly detection (ZSAD) has gained significant attention for its practical value in industrial inspection. Recently, CLIP‑based approaches have been widely adopted in ZSAD due to their strong vision‑language generalization capabilities. However, existing methods commonly employ continuous prompt embeddings for prompt optimization and encode semantics in latent vectors, which lack interpretability and scalability. To this end, we propose CoEvoAD, a co‑evolutionary framework for discrete prompt selection. CoEvoAD performs prompt search in the discrete natural‑language space using an evolutionary algorithm. Candidate prompts are iteratively generated, evaluated, and selected throughout population evolution, thus preserving the interpretability and composability of natural language. Furthermore, we introduce a Cross‑Category Transfer Objective (CCTO), which treats held‑out source categories as proxies for unseen categories and scores prompt rules based on their estimated cross‑category transferability, effectively improving cross‑category generalization. Extensive experiments are conducted to validate the effectiveness of CoEvoAD, and the results show that it achieves state‑of‑the‑art performance across multiple anomaly detection datasets. The code is available at https://github.com/rstao‑bjtu/CoEvoAD.

Authors:Johanna Angulo, Víctor Yeste, Hector Espinos-Morato
Title: Benchmark Contamination: A Taxonomy Organized by Defeated Mitigation
Abstract:
A benchmark score is a joint property of the model, the evaluation harness, the elicitation budget, the sampled population, and contamination status. Leaderboards publish the model and the score, so capability and leakage stay observationally equivalent. Existing taxonomies classify contamination for automated detection, not the question a reporter faces at publication: given the mitigations already applied, which validity threats remain open? We introduce a taxonomy organized by the mitigation each type defeats ‑‑ direct, derivative, temporal, distributional, and acquired ‑‑ spanning training‑time and evaluation‑time leakage. Holding out a private test set closes the first alone. The fifth is acquired during the evaluation itself; because it is a property of one run, it must be recorded with the reported score rather than with the benchmark release. We operationalize it as a four‑field disclosure protocol in which "unknown" is a valid entry, released under CC BY 4.0 with a JSON Schema, a validator, and worked examples. Two coders external to the design team applied a pre‑registered instrument to 41 documents. Per‑variable linear‑weighted κ runs from 0.00 to 0.35 (median 0.21) over 29 main‑pass documents against a single‑coder test‑retest ceiling of 0.84, collapsing under the class skew the registration anticipated; pooling raises it to 0.46 through chance correction rather than better agreement. Two variables fall below the prevalence‑robust threshold registered in advance: strata reporting and the acquired type introduced here. Disagreement concentrates on when a variable applies rather than on what a document states. Elicitation budgets are reported in 13% of documents, and no document addresses all five types. The contribution is the taxonomy, the score‑side artifact that follows from it, and a pre‑registered measurement of instrument reliability and current disclosure.

Authors:Louis Yiven Zhu
Title: One Capability or Many? Testing the Economic Validity of Frontier AI Evaluation
Abstract:
Frontier‑model leaderboards now rank systems based on economic benchmarks, tests of how well models carry out professional tasks from software engineering to banking workflows, and those rankings inform what organisations buy, what regulators scrutinise, and expectations of how work will change. Whether such benchmarks measure a capability distinct from general test‑taking, or re‑express the one axis along which every benchmark rises as models improve, is a question of construct validity that has not yet been studied. We test it on a hash‑pinned leaderboard snapshot of 421 model configurations across twelve benchmarks, four of them economic, treating benchmarks as items and models as respondents in a latent‑variable model with four hypotheses and their thresholds fixed before analysis. A single factor explains 74.5% of common variance and tracks model release date (R^2 = 0.505), so the leading axis of capability is substantially a time trend; where prior work controls for scale, compute adds little once date is removed. Removing the date trend lowers that share by 14.9 points, and by 24.1 with one row per base model. Under the dimensionality rule fixed in advance the economic benchmarks form no distinct factor, yet a leave‑one‑benchmark‑out test with factors re‑estimated inside every fold shows that a multi‑factor representation predicts held‑out economic scores better than a single general index (pooled Delta‑MSE 0.037, 95% bootstrap interval [0.019, 0.055]). Economic benchmarks therefore add incremental predictive information to a largely date‑driven general factor, and the evidence does not support treating them as a distinct latent capability. Leaderboards remain a sound guide to overall progress, but most of the gap between models released months apart is calendar, so a small gap between contemporaneous models should be date‑adjusted before being read as a capability difference.

Authors:Xuanyou Liu, Novel Alam, Karan Ahuja
Title: EITWatch: Smartwatch-Integrated Planar Electrical Impedance Tomography for Hand Gesture Recognition
Abstract:
Wrist Electrical Impedance Tomography (EIT) senses hand gestures from muscle‑ and tendon‑driven impedance changes, but prior wrist‑EIT systems require electrode coverage beyond the watch‑back contact patch and separate analog front ends. We present EITWatch, the first wrist‑EIT system built around smartwatch case‑back geometry, asking whether this contact patch alone can support gesture recognition: eight planar electrodes in a 31 mm ring acquire 35 impedance measurements at 48 Hz. Because a planar array cannot encircle the wrist, EITWatch uses multi‑depth scanning to sample multiple source‑sink distances and current paths; it beat matched adjacent injection by 15.1/10.4 percentage points (macro/micro) across all 12 participants. In a prompted study, within‑session leave‑one‑round‑out accuracy reached 91.4%/92.5% (window/trial) for six macro‑gestures, and 90.1%/91.5% (window/segment) for five micro‑gestures plus relax; window‑level cross‑session and leave‑one‑user‑out transfer reached 73.2%/70.4% and 63.1%/55.3% (macro/micro).

Authors:Sabilashan Ganeshan
Title: Where Induction Runs Out: Description-Length Difficulty and the Memorisation Gap in Integer-Sequence Benchmarks
Abstract:
Integer sequences from the On‑Line Encyclopedia of Integer Sequences (OEIS) are increasingly used to benchmark mathematical reasoning in language models. We ask what such benchmarks actually measure, using an exactly computable reference learner: two‑part minimum description length (MDL) over the class of P‑recursive (holonomic) recurrences, evaluated on every prefix of a sequence as terms arrive. Three findings follow. First, MDL difficulty is a parameter count. The discovery point nd, the first prefix length at which a symbolic hypothesis beats verbatim storage, is predicted almost exactly by a combinatorial identifiability bound on the selected operator's order and degree. It is invariant to term magnitude: scaling Fibonacci over twelve orders of magnitude leaves nd unchanged, because a hypothesis must encode its own initial conditions and the magnitude cancels. Second, at scale the learner exhibits a regime our curated corpus could not produce even once: across 20,000 OEIS sequences, 89.98% of those that fit a recurrence on some prefix fit none at full length. We call this the wilderness ‑‑ induction acquires a theory, loses it, and never recovers. Third, evaluating three language models on sequences stratified by these MDL regimes refuted our pre‑registered hypothesis: models do not confabulate where MDL reports no theory, but hedge appropriately. Confident errors are inverted, concentrating on the easy stratum, where apparent competence tracks recognition of the sequence rather than induction of its rule. OEIS‑derived benchmarks therefore substantially measure memorisation, and MDL supplies a cheap, contamination‑free difficulty signal they currently lack. Code and data are released.

Authors:Zixiang Xu, Jiaan Wang, Fandong Meng
Title: AlgoWorlds: Benchmarking Tool Use for Global Optimization in Algorithmic Worlds
Abstract:
Tool‑use benchmarks generally evaluate whether an agent completes a workflow using appropriate tools and valid arguments. However, feasibility alone is insufficient in real‑world decision settings such as route planning and fleet dispatch. Individual choices interact through shared constraints and costs, so a feasible solution may still be substantially suboptimal. This raises a harder question: can an agent turn information gathered through tools into a globally optimal decision? We introduce AlgoWorlds, a benchmark that transforms formally specified combinatorial optimization problems into partially observed decision environments with verifiable global optima. Each environment contains a hidden instance observed only through task‑specific information tools, after which the agent commits to one structured decision evaluated for feasibility and optimality. AlgoWorlds contains 240 environments covering ten combinatorial optimization families and four workload levels. Family‑specific deterministic programs generate the instances, exact algorithms certify their optima and determine workload levels, and two structurally different tool interfaces present each underlying instance. We evaluate seven leading LLMs, including Claude Opus 4.8 and GPT‑5.6 Sol. Achieving global optimality remains highly challenging: although leading models produce feasible decisions in most cases, the best‑performing model reaches exact optimality in only 38.61% of cases. Even when agents collect sufficient information to reconstruct the hidden instance, most failures end in feasible but suboptimal decisions. The challenge therefore extends beyond information acquisition to information integration, global constraint reasoning, and decision verification. The project homepage is available at https://xzx34.github.io/AlgoWorlds/, and the code is available at https://github.com/xzx34/AlgoWorlds.

Authors:Qiming Guo, Wenbo Sun, Chen Pan, Ye Wang, Wenlu Wang
Title: Unlearning on Spatio-Temporal Graphs through Subgraph Virtual Edge Reconstruction
Abstract:
Spatio‑temporal graphs are widely used in modeling complex dynamic processes such as temporal forecasting, molecular dynamics, and healthcare monitoring. Recently, stringent privacy regulations such as GDPR and CCPA have introduced significant new challenges for existing spatio‑temporal graph models, requiring complete unlearning of unauthorized data. Since each node in a spatio‑temporal graph diffuses information globally across both spatial and temporal dimensions, existing unlearning methods primarily designed for static graphs and localized data removal cannot efficiently erase a single node without incurring costs nearly equivalent to full model retraining. To address this, we propose CallosumNet, a spatio‑temporal graph unlearning framework biologically inspired by the corpus callosum structure. CallosumNet makes two key technical contributions: (1) it reconstructs subgraphs using biologically‑inspired virtual edges; and (2) it restores interlinked spatio‑temporal dependencies among subgraphs via a lightweight meta‑graph integration layer. Empirical results on four diverse real‑world datasets show that CallosumNet achieves complete unlearning while maintaining accuracy very close to the gold model. The code is publicly available at https://github.com/wenlu‑lab/STGraphUnlearning.

Authors:Qiming Guo, Wenbo Sun, Ye Wang, Wenlu Wang
Title: Spatial Entropy based Partitioning for Spatiotemporal Graph Unlearning
Abstract:
Spatiotemporal graphs underpin applications such as traffic forecasting, weather forecasting, and healthcare monitoring. Privacy regulations such as the GDPR and the CCPA require the complete removal of unauthorized data from trained models, but achieving this on a spatiotemporal graph is difficult: because information propagates globally through both spatial and temporal message passing, fully erasing a node's influence forces costly full‑graph retraining. ST‑graph unlearning requires both exactness and efficiency. We propose IsleNet, which uses spatial‑entropy‑guided partitioning to create balanced, locally coherent subgraphs and reconnects them with lightweight virtual edges. Upon an unlearning request, only the affected subgraph encoder and virtual‑edge layer are retrained, ensuring exact removal with low cost. Experiments on four real‑world benchmarks show that IsleNet attains up to 94% of full‑graph accuracy while reducing unlearning time by up to an order of magnitude. Our code is publicly available at https://github.com/wenlu‑lab/STGraphUnlearning.

Authors:Wenhua Huo, Fenglei Han, Wangyuan Zhao, Xiao Peng, Chunhui Wang, Jialin Wu, Jiayi Han
Title: APPSolver: Adaptive Patch Partitioning for Point-Wise Ship Flow Prediction on Unstructured Meshes
Abstract:
Large non‑uniform point sets make direct attention‑based surrogate modeling costly for ship hydrodynamics. We introduce APPSolver, a point‑wise flow‑prediction framework built around Adaptive Patch Partitioning (APP), a deterministic quadtree representation for fixed two‑dimensional horizontal slices extracted from ship CFD simulations. APP assigns finer patches near the hull and coarser patches farther away, downsamples patch contents, and recovers predictions to the full reference point set. Under a corrected protocol that constructs natural (t,t+1) pairs before splitting, reuses training‑set normalization statistics, and reports three model seeds, learned tokenizers are more accurate than APP‑Transformer, and a persistence baseline has lower one‑step MAE on all three ShipBench hulls. The supported benefit of APP is therefore computational rather than universal predictive superiority: on a representative DTC input, APP‑Transformer requires 1.815 GFLOPs and 1.309 ms per model forward, while a matched ablation shows that adaptive partitioning reduces MAE by 16.4‑24.9% relative to a uniform partition augmented with learned slicing. Condition encoders provide setting‑dependent gains in leave‑one‑hull‑out evaluation, but the current absolute next‑state objective does not establish accurate long‑horizon dynamics. These results characterize APP as a compact spatial representation with an explicit accuracy‑‑efficiency trade‑off. Code is available at https://github.com/wenhuahuo/APPSolver .

Authors:Jakob Wasserthal, Joshy Cyriac, Michael Bach, Kimia Mozahheb Yousefi, Minh-Son To, Máté Sik, Cédric Hémon, Thomas Weikert, Martin Segeroth
Title: Extending TotalSegmentator: Predicting Patient and Acquisition Characteristics from CT and MR Images
Abstract:
Background: Patient details and acquisition metadata are important for clinical decisions, image quality control, and automated research pipelines, but may be missing or unreliable in imaging archives. Purpose: To develop and evaluate a fast open‑source model that predicts patient and acquisition characteristics directly from CT and MR images. Materials and Methods: Separate 3D ResNet‑10 ensembles for CT and MR were trained on 57,291 and 43,200 clinical examinations acquired from 2011 to 2025. Both predicted weight, height, age, sex, contrast presence, vertebral coverage, and image noise. The CT model additionally predicted scanner manufacturer, tube voltage, tube current, convolution kernel, and post‑injection time; the MR model predicted sequence class. Performance was evaluated on internal CT (n=501) and MR (n=636) test sets and an external CT dataset (n=54). Results: Internal CT MAEs were 3.90 kg, 3.68 cm, and 4.42 years for weight, height, and age, with sex F1=0.990; corresponding MR results were 4.34 kg, 4.62 cm, 7.13 years, and F1=0.970. The CNN outperformed a segmentation‑derived XGBoost baseline for all four core targets in both modalities (adjusted P<=.042). F1 scores were 0.963 for CT contrast, 0.953 for MR sequence, and 0.823 for MR contrast. External CT MAEs were 4.45 kg, 4.05 cm, and 5.17 years, with sex F1=0.971. CPU inference required 20 seconds for CT and 12 seconds for MR. Conclusion: One 3D multitask model per modality can rapidly recover patient and acquisition characteristics from heterogeneous CT and MR examinations. Models are available in TotalSegmentator: https://github.com/wasserth/TotalSegmentator

Authors:Yunfan Zhou, Qiming Shi, Yizhou Yang, Di Weng, Yingcai Wu
Title: BIRD-History: A Benchmark for History-Driven Text-to-SQL with Fine-Grained Knowledge Annotations
Abstract:
While recent Large Language Model (LLM)‑based text‑to‑SQL systems achieve impressive performance on standard benchmarks, they struggle when user queries implicitly rely on domain‑specific knowledge, such as business logic, data conventions, and analytical practices, that is neither captured by the schema nor explicitly stated in the natural language question. Historical SQL query logs offer a valuable source of such knowledge, yet existing benchmarks do not adequately support evaluation of history‑driven approaches. To address this gap, we introduce BIRD‑History, a benchmark consisting of 1,393 tasks across 11 databases, designed to evaluate text‑to‑SQL systems' ability to ground underspecified natural language questions using historical SQL scripts. Each task is annotated with ground‑truth labels specifying which historical queries contain relevant knowledge and which SQL clauses encode it, enabling systematic evaluation of both retrieval effectiveness and knowledge utilization. Alongside the benchmark, we propose a plug‑in retriever that extracts five types of external knowledge from historical SQL scripts, then retrieves and reranks relevant fragments for query generation. The retriever integrates seamlessly into existing few‑shot text‑to‑SQL pipelines without requiring prompt modifications. Experiments demonstrate consistent improvements across four text‑to‑SQL systems, highlighting the value of leveraging historical query logs for handling underspecified queries. Dataset and code are open‑sourced on https://github.com/zjuidg/BIRD‑History.

Authors:Nikolay Safonov, Nikita Gornostaev, Alexandra Dubonos, Dmitriy Vatolin
Title: Neural video codecs quality assessment dataset and benchmark
Abstract:
Video traffic constitutes a significant share of global web traffic. To reduce its volume, video codecs have been developed and continuously improved. While the industry has achieved substantial progress in traditional video coding, neural video codecs (NVCs) have recently emerged as a new approach that applies deep learning to video compression. This creates new challenges for compression quality assessment, which is essential for the further development and improvement of such codecs. In particular, it is important to evaluate the novel temporal compression paradigms introduced by NVCs. In this work, we present a large‑scale subjective dataset of videos compressed with both neural and traditional video codecs. The subjective scores were collected through crowd‑sourced pairwise comparisons. The proposed dataset provides a valuable resource for the development and benchmarking of video quality metrics tailored to neural video codecs. The dataset is available at the following link: https://videoprocessing.github.io/nvc‑dataset‑benchmark

Authors:Matin Mahmood, Antonio Rueda-Toicen, Mohamed ElBassat, Seifeldin Elkerdany, Weixing Wang, Gerard de Melo
Title: Hyper3-CLIP: Hierarchy-Conditioned Hyperbolic Vision-Language Training
Abstract:
CLIP‑like vision‑language models (VLMs) trained with contrastive objectives learn strong global image‑text representations, but their Euclidean embeddings and global pooling fail to encode relational structure such as part‑whole and parent‑child relations. Hyperbolic VLMs address this gap with entailment‑based objectives, and text‑conditioned variants improve fine‑grained alignment through sentence‑ and phrase‑level queries. However, these two lines of work remain separate: hyperbolic VLMs use static image and region features, while query‑conditioned methods lack hierarchical geometric structure. We present Hyper3‑CLIP, a hierarchy‑conditioned hyperbolic VLM that combines global, local, and global‑local contrastive learning with query‑conditioned visual pooling. To train the model, we construct lightweight query hierarchies from text, comprising full captions, sentence fragments, localized part descriptions, and extracted phrases. Each query conditions the pooling of visual patches, and the resulting representations support image‑text, whole‑part, and parent‑child entailment losses. Query‑conditioned pooling is active only during training. Hyper3‑CLIP improves R@5 and R@10 retrieval on COCO and Flickr, as well as multi‑label classification on VOC and COCO, while remaining competitive on hierarchy metrics. We also audit zero‑shot prompt sensitivity under fixed prompt regimes and study the effect of the localized GRIT part budget used during training. Code is available at https://github.com/Hyper3Labs/hyper3‑clip.

Authors:Daegyu Sung, Yukyeong Lee, Geon Park, Yumin Choi, Sung Ju Hwang
Title: Super Library Agent: Joint Generation and Maintenance of Multiple Applications Beyond the Single Codebase
Abstract:
Organizations often develop and maintain portfolios of related applications: independently deployable codebases that share substantial domain logic, interface patterns, or operational conventions. As LLM coding agents are increasingly used to generate and maintain such software, a naive application‑by‑application workflow duplicates shared logic across codebases and allows prolonged agentic maintenance to accumulate verbosity, dead code, and structural erosion. We introduce the Super Library Agent problem, where an agent sequentially generates a portfolio of N related applications while maintaining a shared Super Library of reusable cross‑application components. A minimal sequential scaffold can in principle extract shared code and migrate applications to the evolving library, but in practice suffers from low extraction recall and fragile dependency migration. We address these failures with candidate‑guided extraction over code chunk summaries, pre‑extraction codebase consolidation, and context‑aware migration using extraction traces and call‑graph information. Across WebGen‑Bench and PaperBench, our method preserves application functionality while significantly reducing redundancy and token footprint (verbosity, token length) over zero‑shot, and avoiding the structural erosion introduced by naive library construction, with additional reductions in LOC and MDL. Our code is available at https://github.com/sbigstar0310/super‑library‑agent.

Authors:Junxuan Li, Zijun Liu, Ziyi Huang, Peng Li, Yuzhou Liu, Ming Yan, Yang Liu
Title: Learning Simple Test-Time Environments for LLM Web Agents
Abstract:
Large language model (LLM) agents have demonstrated remarkable proficiency in manually constructed environments, yet their performance frequently collapses when transitioned to complex real‑world settings. Existing research largely attribute this degradation to the compositional generalization gaps in LLMs on combinations of multiple simple, well‑structured environments. In this work, we propose that LLM web agents can learn simple environment observations at test time. Specifically, we introduce trial steps for agents to decompose a complex environment observation into sub‑modules, and implement a label‑free learning method, Test‑Time Environment Decomposition (TTED), to adapt agent behaviors with experience during inference. Our empirical evaluations demonstrate the framework's efficacy across both synthetic and realistic benchmarks, showing (1) experience gains acquired within simpler sub‑environments can be effectively composed to improve performance in the full one, and (2) test‑time training on sub‑environments can significantly enhance the compositional generalization of agents in real‑world web automation tasks. We also provide key insights in the design of the label‑free learning algorithm. As more complex environments are accessed by LLM agents, we believe learning environment decomposition skills at test time will be critical for robust real‑world deployment.

Authors:Zhanpei Huang, Binbin Sun, Jialiang Chen, Yiou Wang, Taochen Chen, Yuzhu Ji, Yiqun Zhang, Yiu-Ming Cheung
Title: AOI-Net: Structural Face AOI-Guided Eye-Gaze Track Representation Learning for Autism Spectrum Disorder Detection
Abstract:
Eye‑movement tracking has emerged as a promising non‑invasive approach to Autism Spectrum Disorder (ASD) screening, with systematic differences in attentional allocation and revisit behaviors observed during socially interactive tasks. Existing computational methods typically characterize eye‑movements using discrete gaze trajectories and fixation events, yielding representations dominated by short‑range temporal dynamics and limiting models that primarily emphasize long‑range dependencies. Meanwhile, gaze behavior is naturally organized across semantically meaningful Areas of Interest (AOIs), whose attention allocation and transitions provide important structural cues, yet their relationships are rarely modeled explicitly. To address these limitations, we propose a structural face AOI‑guided Eye‑Gaze Track Network (AOI‑Net) that jointly models short‑term temporal dynamics and AOI‑level structural organization. A network gating mechanism adaptively integrates the complementary temporal and structural representations according to their contributions to gaze‑behavior characterization. To mitigate the pronounced class imbalance commonly encountered between individuals with ASD and Typically Developing (TD) participants in clinical datasets, class‑distribution‑aware learning is further employed to facilitate discriminative embedding learning under skewed class distributions. Experiments on a unique and large‑scale clinical eye‑tracking database comprising eight stimulus subsets and more than 1,300 participants show that AOI‑Net consistently outperforms state‑of‑the‑art methods. The proposed framework also enables interpretable gaze‑behavior modeling and provides a practical basis for scalable AI‑driven ASD screening in real‑world healthcare. The code is available at https://github.com/Zhanpei‑ai/CIM‑AOI‑Net/tree/main/Code

Authors:Tingyu Lin, Christian Stippel, Armin Dadras, Jakob Zenzmaier, Florian Kleber, Wolfgang Aigner, Robert Sablatnig
Title: PERSIST: Persistent-State Discrimination for Shot Boundary Detection
Abstract:
Shot boundary detection (SBD) is widely treated as the localisation of local visual discontinuities, yet many false positives such as hand‑held shake, illumination flicker, motion blur, occlusion, and damaged archival material produce equally sharp local change without introducing a new shot. We reformulate SBD as boundary semantic discrimination: a frame is favoured as a boundary only when its local change evidence is accompanied by a persistent update of the video's latent temporal state, rather than a transient excursion that returns to the surrounding trend. This persistence test is operationalised with a continuous latent state from a FiLM‑conditioned sinusoidal representation network and a structured discriminator that combines three semantic cues, local change, transient impulse, and return‑to‑trend, into a single interpretable per‑frame signal over a dual‑rate temporal backbone. The resulting framework, PERSIST, turns every decision into an inspectable one: the persistence criterion is trained into the classifier, its per‑frame effect stays readable from the gate triple, and its learned latent state is measurably boundary‑discriminative. On a 2,727‑video per‑subtype diagnostic it removes 33‑80% of flash, text‑overlay, and archival false positives relative to an identically trained cue detector, and at matched true‑transition recall it roughly halves TransNetV2's pseudo‑event false positives on that diagnostic and cuts its false positives on ClipShots footage by about a quarter, while preserving recall. It does so while reaching parity with the strongest public detector across online, broadcast, short‑form, and historical‑archive transfer evaluations, under markedly stricter training: it learns from ClipShots real transitions only, whereas the anchor draws on additional corpora whose transitions are 85% synthetic. Code is available at https://github.com/linty5/PERSIST.

Authors:Jinzhe Li, Gengxu Li, Jinnan Li, Yuan Wu, Yi Chang
Title: MMPCBench: Benchmarking Multimodal Large Language Models on Proactive Critique of Flawed Inputs
Abstract:
As Multimodal Large Language Models (MLLMs) evolve into sophisticated interactive assistants, their reliability depends not only on following instructions but also on validating them. We define Proactive Critique as the model's autonomous ability to identify, analyze and fix faulty user inputs without extra prompts. However, evaluations mainly test models under ideal circumstances or simple refusal behaviors, largely ignoring active error processing. To fill this gap, we propose MMPCBench, a comprehensive framework for evaluating MLLMs' proactive critique competence. It features a fine‑grained taxonomy of 4 primary error types spanning 12 subcategories, ranging from cross‑modal contradictions to missing visual premises. We adopt a hierarchical evaluation protocol to measure models' error detection, diagnosis and resolution performance, and apply alignment‑aware metrics to assess the coherence between internal reasoning and final responses. Tests on 14 mainstream MLLMs show obvious weaknesses in proactive critique, especially in dealing with subtle visual anomalies. Notably, we identify a pervasive "consistency gap": reasoning models can often correctly identify and analyze errors during internal reasoning yet suppress these valid insights in final outputs to prioritize response compliance. The code and data is available at https://github.com/ALIENS32/MMPCBench.

Authors:Márcus Lobo, Vitor Matias, Jeová Farias, Moacir Ponti
Title: 3D-MRL: Nested Multimodal 3D Representations via Matryoshka Representation Learning
Abstract:
Vision‑Language Models align point clouds with image and text embeddings, enabling zero‑shot recognition, retrieval, and open‑vocabulary understanding of 3D shapes. Existing multimodal 3D pre‑training methods produce fixed‑dimensional embeddings, requiring separate models for different computational budgets. We propose 3D Matryoshka Representation Learning (3D‑MRL), a multimodal 3D pre‑training framework based on Matryoshka Representation Learning. 3D‑MRL learns nested 3D representations by aligning point clouds with frozen CLIP image and text embeddings while applying contrastive supervision across multiple embedding dimensions. The Matryoshka objective is applied only to the 3D encoder, allowing a single model to produce representations at different dimensionalities without retraining. Experiments on the Objaverse‑LVIS, ModelNet40, and ScanNet datasets show that 3D‑MRL achieves competitive performance on zero‑shot and few‑shot 3D recognition tasks. In addition, the learned representations support retrieval across different embedding dimensions within a single model. On Objaverse‑LVIS, 3D‑MRL improves Top‑1 accuracy from 46.8% to 50.9%. Retrieval experiments further show that different embedding dimensions yield varying levels of semantic and geometric specificity.

Authors:Tatul Danielyan, Mariam Avetisyan, Hrant Davtyan
Title: Cloud and On-Premises Deployment of Uzbek Legal RAG via Targeted Retriever Fine-Tuning
Abstract:
Deploying large language models for legal question answering raises challenges that general‑purpose leaderboards do not capture, particularly for low‑resource languages and under hard operational constraints. We report on building and operating a retrieval‑augmented (RAG) legal assistant for Uzbek that must run in two regimes: a managed cloud service that maximizes answer quality within a per‑token cost ceiling, and an on‑premises deployment for clients whose legal data may not leave their infrastructure, restricting us to open‑weight models on limited local hardware under latency constraints. Because no evaluation existed for this setting, we build two domain benchmarks: a retrieval benchmark of 178 expert‑annotated legal queries with gold provision spans, and an end‑to‑end benchmark of 504 expert‑curated question‑‑answer pairs scored by an LLM judge whose ratings we validate against human judgments and against an independent‑family judge. Applying these benchmarks under each regime, we find the open‑versus‑proprietary gap is small and cheaply closed by fine‑tuning. Therefore, we train UTE‑1, which is a state‑of‑the‑art text embedder among open models for Uzbek. We also demonstrate that closing the performance gap via fine‑tuning is both impractical due to the intensive hardware demands of long‑context legal Q\&A and unnecessary, given that legal acts change frequently. We support this by reporting a negative result from a QLoRA experiment. We distill practical guidance for similar deployments, drawn from a system serving real users in production. We release our benchmarks, evaluation code and the fine‑tuned embedder (UTE‑1) \hrefhttps://metric‑ai‑lab.github.io/Uzbek‑Legal‑RAG/at this https URL to support future work on low‑resource legal NLP.

Authors:Eduard Zamfir, Christian Reisswig, Zongwei Wu, Yongqin Xian, Radu Timofte
Title: Elastic Token Compression for Pixel-Space Diffusion Transformers
Abstract:
Natural images concentrate their detail in a small fraction of the frame, yet diffusion models spend a full token on every patch, in every layer and at every timestep. The waste is largest in pixel‑space models, with no autoencoder to absorb low‑level redundancy first. Probing a pretrained pixel text‑to‑image transformer, we find its middle‑block tokens redundant wherever the image is flat. The redundancy occupies connected, content‑shaped regions, and exploiting it requires tokens with the same geometry. Cutting a Hilbert ordering of the patches provides them. Consecutive positions are always image neighbours, so any contiguous run is a connected region whose size and shape follow the content, and grouping in two dimensions becomes a cut in one. Existing reductions each lose part of this. Similarity merging scatters its groups, latent bottlenecks discard position, and skipping deletes what it should summarize. We cut where the model's features change most and pool each run into one region token. Our Region Token Interface (\method) adapts a diffusion model to these tokens, with the region count drawn at random during fine‑tuning so one checkpoint serves every budget. \method leads prior reduction methods at matched budgets, matches dense quality at 2.0× the speed, and stays close at 2.6×. The code and models are open‑sourced at https://eduardzamfir.github.io/rti

Authors:Haonan Zhou, Gaoxiang Linghu, Youlin Jia, Hongyu Cui, Kewei Wei, Kaiyue Zhou, Bruce X. B. Yu, Gaoang Wang
Title: LightFuse: Relightable Interactive Gaussian Scene Reconstruction via Multi-Scan Fusion and 2D Gaussian Ray Tracing
Abstract:
Relightable interactive scene reconstruction aims to build an editable 3D model from scans of different object arrangements and render new layouts under novel illumination. Existing methods either bake lighting into appearance or recover material and illumination only for fixed scenes, leaving edited layouts with inconsistent shadows and indirect lighting. We present LightFuse, a 2D Gaussian framework that extends interactive scene reconstruction with explicit material‑illumination decomposition and physically based relighting. LightFuse first fuses observations across states to reconstruct a shared background and movable objects. It then conducts ray‑tracing‑oriented geometry refinement to produce more complete and consistent surfaces. On the refined geometry, staged training with differentiable one‑bounce ray tracing separates shared metallic‑‑roughness material from state‑specific environment lighting. The resulting scene supports object rearrangement, material editing, and relighting, while ray tracing recomputes appearance after each interaction. Experiments across synthetic scenes demonstrate state‑of‑the‑art relighting quality, outperforming the strongest baseline by +9.74\,dB PSNR and +0.121 SSIM on average. Project page: https://zhn202.github.io/LightFuse/

Authors:Abdelrahman Abdallah, Mohammed Ali, Bhawna Piryani, Mahmoud Abdalla, Adam Jatowt
Title: Large Language Models Systematically Favor Popular Options: Evidence and Mitigation Across MCQs
Abstract:
Multiple‑choice questions (MCQs) are a standard format for evaluating large language models (LLMs), yet the popularity of answer options can confound evaluation. Modern LLMs systematically prefer popular but incorrect options over less popular correct ones, a vulnerability we call popularity bias. This pattern aligns with confidence miscalibration: model confidence remains high even as accuracy collapses for popular options. To systematically isolate this phenomenon, we introduce PopMCQ, a benchmark with six controlled strategies that vary option popularity while keeping the correct answer fixed. In our most adversarial setting, where all distractors are more popular than the correct option, models choose popular but wrong answers 66% of the time. To mitigate this bias, we propose PopDebias, a lightweight inference‑time correction that estimates and removes a popularity prior from model predictions. It requires no fine‑tuning, is label‑free at test time (using only a small calibration split for parameter fitting), and adds negligible computational cost. Experiments on 22 open‑source LLMs (0.5B to 32B parameters) show consistent improvements, with accuracy gains up to 54.1 percentage points under strong popularity pressure. The code and data are available https://github.com/DataScienceUIBK/PopMCQ

Authors:Yaroslav Prytula, Anton Popov, Dmytro Fishman
Title: QCell: Recombining and Aligning Cell Queries for Overlapping Instance Segmentation
Abstract:
Instance segmentation of overlapping cells in microscopy remains challenging due to semi‑transparent structures that produce weak boundaries and mixed visual evidence in overlap regions. Existing methods address this through local regions of interest or shape priors but lack global reasoning across overlapping objects. We present QCell, a novel query‑based model that de‑overlaps cell instances in microscopy scenes. Our approach combines (i) an instance recombination module that decomposes and recombines query representations in latent space, enabling the model to reason about complete object structure under overlap, and (ii) a contrastive query alignment objective that combines distinctive instance feature learning and separation of overlapping cell queries. We additionally introduce a new Organoid dataset benchmark for overlapping cell segmentation. We show that QCell outperforms state‑of‑the‑art methods across multiple benchmarks, achieving +2.2 AP and +2.7 AJI on ISBI2014. Code is available at https://github.com/SlavkoPrytula/QCell

Authors:Haoru Tan, Sitong Wu, Yanfeng Chen, Shizhen Zhao, Yang-Tian Sun, Tianjia Liu, Chirui Chang, Shaofeng Zhang, Samm Sun, Xiuzhe Wu, Ruobing Xie, Xiaojuan Qi
Title: Dynamic Important Example Mining for Reinforcement Finetuning
Abstract:
Reinforcement fine‑tuning (RFT) is increasingly used to strengthen the reasoning abilities of large models, yet its effectiveness is bound by how training data are selected and used. Most data‑centric RFT methods rely on static or heuristic sample selection, implicitly assuming a sample's value is fixed over training. This overlooks the non‑stationary dynamics of policy learning and can lead to suboptimal updates. We propose Dynamic Important Example Mining (DIEM), a principled and fully automated framework that makes data utilization adaptive throughout RFT. DIEM integrates two components into each optimization step: (i) a gradient‑alignment importance estimator that efficiently approximates each sample's marginal contribution to policy improvement; and (ii) a constrained batch reweighting scheme that maximizes aggregate utility while preserving the update's gradient magnitude to stabilize optimization. Across several reasoning benchmarks, DIEM consistently outperforms strong static and dynamic baselines. The code will be released via https://github.com/hrtan/DIEM.

Authors:Cheng Chen, Jerry Bai, Jiacheng Wei, Boyu Chen, Xiaoji Zheng, Fan Wu, Minghao Yang, Tianrun Chen, Ruibo Li, Xiaoyu Yue, Xiaoyang Guo, Yixiao Ge, Guosheng Lin, Fayao Liu
Title: AnyWorld: Factorized Egocentric World Models for Cross-Embodiment Generalization
Abstract:
Collecting contact‑rich robot experiences at scale remains a major bottleneck for generalizable manipulation. Beyond data quantity, robot learning also requires diverse experiences across embodiments, viewpoints, and scenes. Human egocentric videos provide abundant physical interactions, but each video captures only a narrow slice of experience under a single body, camera trajectory, and environment. We propose AnyWorld, a cross‑embodiment world modeling framework that expands a single human interaction into diverse robot‑native rollouts without paired human‑robot demonstrations. Our model factorizes an interaction into action, camera, and embodiment: action controls capture the motion structure, camera controls specify viewpoint evolution, and the target embodiment context defines the acting body and its interaction geometry. This formulation enables independent recomposition of embodiment, viewpoint, and scene factors, allowing a single model to generate many robot‑domain experiences while preserving the underlying dynamics and object interactions. We train the model with large‑scale human interaction pretraining followed by mixed‑embodiment fine‑tuning. Experiments show that our model supports controllable recomposition across embodiments, viewpoints, and scenes, and we further demonstrate that the generated data can improve manipulation performance on the RoboCasa GR1 tabletop benchmark and a real IRON humanoid robot. Beyond aggregate gains, we test whether unpaired human experience can be recomposed into robot‑native video‑action pairs that target a policy gap. Controlled IRON interventions correct a spurious completion prior and establish language‑grounded spatial target selection; an action‑only counterfactual intervention fails to learn the latter reliably, showing that both action calibration and visual recomposition are necessary.

Authors:Dania Batool, Liliana Lo Presti, Marco La Cascia, Filippo Vella
Title: Background-Free Objectness Learning for Class-Agnostic Detection
Abstract:
Object detectors are typically trained under closed‑set supervision, where unlabeled regions are implicitly treated as background. Under incomplete annotations, this assumption introduces objectness bias: visually valid but unlabeled objects are used as negatives, tying objectness to the annotated taxonomy rather than generic object structure. This limitation is particularly problematic for class‑agnostic and open‑world detection. This paper proposes Background‑Free Objectness Learning (B‑FOR), a dense class‑agnostic detection framework that learns objectness without explicit background supervision on unlabeled regions. B‑FOR formulates detection as the prediction of dense multi‑scale object‑center and scale fields, from which object hypotheses emerge as local spatial structures. Supervision is confined to reliable annotated regions through spatially structured soft targets, avoiding foreground‑background discrimination. To support decoding from emergent local maxima, the paper further introduces displacement‑aware scale fields that model object extent as a spatially varying property of the learned objectness field. Experiments on PASCAL VOC, MS‑COCO, and Open Images demonstrate strong generalization to unseen categories and cross‑dataset object distributions. B‑FOR improves recall by more than +10 AR points over prior class‑agnostic baselines. Ablation studies show that both localized objectness supervision and displacement‑aware scale fields are critical for class‑agnostic localization under incomplete annotations. Code available at: https://github.com/Daniaawan/B‑FOR.

Authors:Qiancheng Zhou, Ruizhe Li
Title: Locked at the Entrance, Open Inside: Where RLVR Narrows the Solution Space
Abstract:
Reinforcement learning with verifiable rewards (RLVR) substantially improves single‑sample accuracy (pass@1) but causes the policy's solution space to contract, diminishing the returns of test‑time scaling. In this work, we investigate where inside a reasoning trajectory this breadth is lost: does the policy fail to access a valid solution family, or does it fail to execute computation once initiated? To disentangle access from execution, we analyze the Countdown task, whose solution space can be exhaustively enumerated into discrete entrance families defined by the first operand and operator, across PPO on Qwen2.5‑3B and GRPO on Qwen2.5‑3B‑Instruct. Across both training setups, solution coverage falls by up to 67%, halving even on problems solved across all checkpoints. We show that this contraction is heavily concentrated at the entrance: per‑token likelihood shifts are 11x‑‑16x larger prior to the first arithmetic operation than during downstream reasoning. Supplying only an unselected entrance prefix restores completion rates in low‑access families by over an order of magnitude (0.018 ‑> 0.212 under PPO), demonstrating that alternative solutions remain executable but are no longer initiated. Guided by this localization, we find that while surface prompting fails to recover diversity, entrance‑targeted interventions succeed: late‑layer parameter interpolation with early checkpoints increases solution coverage by 37% at no loss in pass@1. Finally, we show that early‑step entropy collapse recurs across six math benchmarks with 7B and 14B models, but is not an inevitable byproduct of reasoning optimization: an SFT baseline preserves more than double the coverage, and staged SFT‑‑DPO‑‑RLVR pipelines retain early‑step entropy. In summary, reasoning breadth is lost at the door, not inside the room. Code: https://github.com/ershiyidian/early‑branch‑locking.

Authors:Amir Hamza, Davide Boscaini, Fabio Poiesi
Title: Foundational feature fusion for conditional flow matching in 6D pose estimation
Abstract:
Conditional flow matching has enabled a step forward in object 6D pose estimation, achieving state‑of‑the‑art performance by progressively denoising and registering object representations to observed scenes. Existing methods require training task‑specific encoders supervised on object‑scene overlap and rely on trivial feature fusion strategies to resolve pose ambiguities. We present FunFlow6D, a novel flow matching‑based formulation that leverages features from geometric and appearance foundation models for pose estimation, eliminating the need for task‑specific encoder training. We also introduce a cross attention‑based fusion mechanism that dynamically combines geometric and appearance features to provide richer conditioning for the flow matching module. Experiments on four datasets from the BOP benchmark show that FunFlow6D outperforms the previous state of the art while reducing supervision requirements and memory overhead. Extensive ablations validate the contribution of each proposed component. Project website: https://tev‑fbk.github.io/FunFlow6D/.

Authors:Boyu Cai, Li Yang, Yan Xu, Wei Liu, Nian Liu, Sikui Zhang, Yan Wang, Chunfeng Yuan, Weiming Hu
Title: Dynamic-Robust Photometric-Semantic Reconstruction for Open-Vocabulary 3D Scene Understanding
Abstract:
The integration of novel view synthesis (NVS) and open‑vocabulary segmentation (OVS) has recently yielded powerful feed‑forward 3D foundation models. However, their inherent reliance on static‑scene assumptions leads to severe misalignment of spatial features in unconstrained dynamic environments. To bridge this critical gap, we propose SPAR, a novel joint semantic‑geometric encoding architecture that explicitly isolates transient dynamic noise prior to latent space aggregation. Furthermore, we introduce a dynamic‑region‑aware end‑to‑end training paradigm that structurally couples motion estimation with multi‑view visual and semantic learning. This unified approach enables the network to inherently resolve motion conflicts and distill multi‑view consistent, temporally stable scene representations from dynamic inputs. Extensive experiments on the challenging D‑RE10K benchmark demonstrate that SPAR achieves state‑of‑the‑art performance. Our end‑to‑end approach achieves exceptional novel view synthesis quality, yielding a PSNR of 22.15 dB and 23.33 dB given only 3 and 4 input views respectively. Despite being trained in a self‑supervised manner, our model achieves an mIoU of 88.5% for motion mask prediction. Furthermore, our analysis reveals a strong inter‑task synergy between photometric scene reconstruction and semantic understanding, where semantic synthesis learning consistently enhances photometric fidelity in novel view rendering. Code will be available at https://github.com/dmucby/SPAR.

Authors:Kai Geissler, Raphael Schäfer
Title: Subtraction-Based Tumor Segmentation and Lesion-Centered pCR Prediction for the MAMA-MIA Challenge
Abstract:
We describe the submission of team FME to the MAMA‑MIA Challenge, which evaluated primary tumor segmentation and prediction of pathological complete response (pCR) from pretreatment dynamic contrast‑enhanced breast MRI on an external multi‑country cohort. For segmentation, we trained a five‑fold residual‑encoder nnU‑Net ensemble using only the first post‑contrast minus pre‑contrast image, combined with mirroring test‑time augmentation and largest‑connected‑component filtering. For pCR prediction, we ensembled 25 pretrained 3D video classifiers trained on lesion‑centred crops from the pre‑contrast and first two post‑contrast volumes. FME ranked second in both tasks. The segmentation method achieved a combined performance‑fairness score of 0.882, with Dice 0.713 and normalized Hausdorff distance 0.099. The pCR method achieved a combined score of 0.664, balanced accuracy of 0.541, and equalized‑odds disparity of 0.212. The results indicate that subtraction‑based input and ensembling support robust tumor segmentation under cross‑site domain shift, whereas pCR prediction from baseline DCE‑MRI alone remains limited. For the submission repository, see https://github.com/FraunhoferMEVIS/MAMA‑MIA‑Challenge‑FME

Authors:Shingeon Kim, Hyeyoon Lee, Dain Kwon, Kanghyun Choi, Sunjong Park, Mi-Ryang Kim, Jeong-Eun Lee, Jinho Lee
Title: STARLINC: Satellite Trail Artifact Removal using Inter-Frame Correlation
Abstract:
The rapid expansion of low Earth orbit satellites such as Starlink is increasingly contaminating astronomical surveys. In practice, contaminated images are often identified through inspection. However, modern surveys generate terabytes of data each night, making manual screening infeasible and necessitating reliable automated methods for satellite trail removal. Unfortunately, existing general‑domain line detection methods fail to generalize to astronomical images due to domain mismatch, which are mostly grayscale with sparse bright stars and have a low signal‑to‑noise ratio. Moreover, training new models from scratch is impractical due to the lack of large‑scale annotated astronomical datasets. To address these challenges, we introduce STARLINC, the first ML‑based framework for satellite trail removal without requiring tedious pixel‑level annotation of astronomical images. STARLINC combines synthetic satellite trail generation for training, inter‑frame differential maps from temporally adjacent exposures to highlight transient trails, and heatmaps to provide additional localization cues for pixel‑level segmentation. Extensive experiments on real‑world data demonstrate substantial improvements over baselines, establishing STARLINC as a scalable solution for next‑generation astronomical surveys. Code is available at https://github.com/starioKim/STARLINC.

Authors:Jiaxiang Liu, Chenhao Yuan, Shuwen Xu, Boxuan Xing, Xiusheng Huang, Yinhao Xu, Hao Liu, Wenhao Teng, Xiangwen Liao, Pengfei Cao, Jun Zhao, Kang Liu
Title: Quantifying Error Tolerance in Synthetic Data: An Atomic-level Operand vs. Operator Perturbation Study
Abstract:
Synthetic data generation has become a cornerstone for advancing large language models. However, the lack of the quantitative analysis for error tolerance became a critical bottleneck. Consequently, current filtering strategies fluctuate between two extremes: they are either overly aggressive, risking the exclusion of potentially valuable samples, or overly permissive, failing to eliminate erroneous samples effectively. To bridge this gap, this paper introduces Atomic Tree Operation Modeling (ATOM), a framework that decomposes data into functional units (f(x)\rightarrow y). ATOM distinguishes benign Operand x perturbations from fatal Operator f perturbations. The former are needlessly discarded by aggressive filtering, while the latter slip through permissive filtering. Our experiments reveal a double dissociation: models are robust to operand perturbations but collapse under operator perturbations. By prioritizing operator over aggressive operand precision, our ATOM‑synthesized data outperforms rigorous baselines (e.g., +3.1% gain over LIMA), suggesting that operator diversity matters more than operand precision. Our code is available at https://github.com/Lut‑hub/ATOM.

Authors:Chenyi Xiong, Yan Zhang, Jing Hu, Ziyue Qin, Kui Xiao, Xiaopan Lyu, Xiaoju Hou, Zhifei Li
Title: More Perspectives, Stronger Signals: Multi-Perspective Enhancement and Progressive Fusion for Multimodal Entity Representation Learning
Abstract:
Learning effective multimodal entity representations is fundamental for reasoning tasks such as multimodal knowledge graph completion (MMKGC). However, existing methods often suffer from semantic over‑smoothing within modalities and ineffective noise filtration across modalities, particularly under sparse or ambiguous conditions. To overcome these limitations, we propose PrismF, a unified framework that synergizes multi‑perspective enhancement with progressive fusion to extract stronger signals from diverse inputs. PrismF enhances fine‑grained intra‑modal semantics through a multi‑perspective mechanism that decomposes each modality into complementary views and constrains them with a decoupling loss to reduce representation collapse. Furthermore, it improves cross‑modal integration through a progressive fusion strategy that dynamically calibrates inter‑modal interactions, enabling the model to emphasize informative signals while suppressing noisy or unreliable ones. Extensive experiments on three public benchmarks show that PrismF achieves the strongest overall performance, including relative improvements of 4.04% in MRR and 11.17% in Hits@1 on KVC16K. Our code can be found at https://github.com/HubuKG/PrismF.

Authors:Shuangkang Fang, Yufeng Wang, Yi-Hsuan Tsai, Wenrui Ding, Yi Yang, Shuchang Zhou, Ming-Hsuan Yang
Title: Chat-Edit-3D++: Interactive 3D and 4D Scene Editing via Large Language Models
Abstract:
Recent work on image content manipulation based on vision‑language pre‑training models has been effectively extended to text‑driven 3D scene editing. However, existing schemes for 3D scene editing still have certain shortcomings, hindering their further development as interactive design tools. Such schemes typically adhere to fixed input patterns, limiting flexibility in text input. Furthermore, their editing capabilities are constrained by a single or a few 2D visual models and require intricate pipeline design to integrate these models into 3D reconstruction processes. To address the aforementioned issues, we propose the Hash‑Atlas network, which reformulates 3D scene editing as operations on 2D atlas images, thereby achieving a workflow decoupling of the 2D editing and 3D reconstruction processes. Building on this foundation, we introduce a dialogue‑based 3D scene editing approach, termed CE3D++, which is centered on a large language model (LLM) that allows arbitrary textual input from users and interprets their intentions, subsequently facilitating the autonomous invocation of the corresponding visual models. Additionally, we extend CE3D++ to monocular 4D scenes by imposing motion constraints on moving objects and further fine‑tuning the LLM by creating a trajectory dataset related to editing tasks, which enables the smaller LLM to schedule up to 30 different visual tools accurately. Experimental results demonstrate that CE3D++ effectively integrates multiple visual models to achieve diverse visual editing effects, possessing strong scene comprehension and multi‑round dialog capabilities. The source codes and trained models are available at https://github.com/Fangkang515/CE3D.

Authors:Jinwook Kim
Title: Mechanizing Typed Regulatory Actions for Security Tokens: Semantics, Falsification, and Bounded EVM Evidence
Abstract:
Security‑token standards expose privileged controls without identifying the legal effect executed or the evidence and reversal obligations it carries. We formalize in Isabelle/HOL a reference execution semantics for the six ERC‑8319 meanings: FREEZE, SEIZE, CONFISCATE, LIQUIDATE, RESTRICT, and RECOVER. It distinguishes applied, rejected, and operational‑failure outcomes and mechanizes action‑specific reversals, replay and epoch rules, complete frames, case‑local terminality, and final receipts; the session builds without unproved placeholders or additional axioms. An indistinguishability theorem shows that bound kernel inputs cannot establish external facts about title, settlement, or entitlement. Constructive witnesses and direct mutations establish reachability and sensitivity for the declared fault set. For one ERC‑TRUST Solidity/EVM candidate, we report separately scoped Foundry, Certora, Kontrol/KEVM, mutation, deterministic‑build, and runtime‑identity evidence. The publication profile qualifies 7/7 reusable packages, 49/49 Core obligations, and 24/24 mandatory Supporting obligations; six optional ERC‑3643 obligations remain unclaimed. A current‑state abstraction relation is unique and functional under pinned‑runtime premises, while package and row corollaries remain conditional on hash‑bound certificates. These results do not establish complete Isabelle‑to‑Solidity‑to‑EVM refinement, compiler correctness, audit completion, production readiness, deployment verification, or external legal truth. They provide a machine‑checked domain semantics and an explicit map of proved, bounded, assumed, and open results.

Authors:Yifeng Lu, Zijie Yang, Jie Li, Qingkai Min, Yue Zhang
Title: AI Historian: Helping historians organize and verify person-centred temporal clues from dispersed historical narratives
Abstract:
History is not preserved in complete, continuous form. Accounts of a person's activities, relationships and historical contexts are scattered across texts, chapters and narrative perspectives; historians must retrieve, identify and compare these materials to reconstruct temporal sequences and verify them against sources. Here we present AI Historian (AIH), an AI agent system that helps historians organize person‑time evidence from dispersed biographical narratives. It takes source sentences as evidence units, identifies people and temporal cues, verifies candidate cross‑text associations and infers comparable temporal ranges while preserving traceable source‑text evidence. We evaluated AIH on six Shiji cases concerning Liu Bang, Xiang Yu and Xiao He. AIH Agent achieved a temporal‑localization MicroIoU of 86.2%, compared with 81.3% for human‑only annotation and 17.1% for direct large‑language‑model prompting; it required about 14 min, versus 1 h 32 min for human‑only annotation. We further applied AIH to the Twenty‑Four Histories and other ancient Chinese histories, ancient Japanese and Korean histories, and modern and contemporary historical materials, and released the results through Westlake Historian. These results indicate that AIH can reduce the cost of organizing historical materials at scale while turning connections obscured by chapter‑based narration into traceable, revisable research questions for collaborative testing.

Authors:Zelin Wan, Arash Nourian, Xiaoxiao Li, Nihar Nandan, Kamalakannan Nandagopal
Title: APIFlow-Bench: Measuring Whether Agents Survive Long, Dependent API Workflows
Abstract:
Tool‑using agents are commonly evaluated by a single bit: whether an end‑to‑end workflow completed. This metric fails to distinguish failures that matter in production, such as expired credentials, malformed payloads, or correct execution followed by incorrect final delivery. We introduce APIFlow‑Bench, a fully auditable benchmark for long‑horizon, dependent REST‑API workflows that decomposes performance into seven engineering capabilities and requires agents to produce answers supported by the actual call path. We generate synthetic API worlds forward, subtask by subtask; each subtask is admitted only after a zero‑LLM self‑test triad verifies its grader and an oracle establishes solvability, and an adversarial audit identified and fixed six grader exploits. Grading is deterministic and provenance‑sensitive: a state check traces a mock‑minted canary through the API data flow to the response the answer must originate from, and a typed answer card is verified field by field. We release all answer keys and 44,362 unredacted execution transcripts. Across 19 frontier and open‑weight models under one neutral scaffold, we find: (1) longer dependency chains degrade success, from 93% on individual subtasks to 74% on clean 20‑subtask chains and 61% when including the 8% of chain trials that a model‑consensus screen flags as passed by no model; (2) reliability separates models more than best‑case capability, with best‑of‑five spanning seven points but all‑five‑of‑five reliability spanning 44 points; (3) the independent‑error account of compounding failure does not fit the data: pass rates on 20‑subtask chains are 33 percentage points above the product of subtask‑level rates, and on the clean slice 77% of failing runs reached the correct final state and failed only at delivery.

Authors:Tianrui Hui, Shaofei Huang, Qisong Han, Yaxiong Wang, Lechao Cheng, Zhedong Zheng, Zhun Zhong, Richang Hong, Meng Wang
Title: Acoustically Grounded Cost Learning for Open-Vocabulary Audio-Visual Semantic Segmentation
Abstract:
Open‑Vocabulary Audio‑Visual Semantic Segmentation (OV‑AVSS) aims to perform pixel‑level segmentation of sound‑emitting objects from an open set of categories. The previous method relies on a class‑agnostic foreground definition, which groups semantically diverse objects into a heterogeneous positive set, causing the model to learn unstable sounding patterns and produce unreliable proposals. To address this, we reformulate the objective to be category‑specific and propose a novel Acoustically Grounded Cost Learning (AGCL) framework to transform the static, audio‑agnostic visual‑text priors into dynamic, audio‑grounded cost representations. For intra‑category soundingness discovery, we devise Audio‑Modulated Cost Generation (AMCG) and Audio‑Guided Temporal Aggregation (AGTA) modules to enable both frame‑level sounding region highlighting and video‑level temporal refinement with a low‑intrusive audio injection mechanism. For inter‑category distractor discrimination, we introduce a Synergistic Distractor Mining (SDM) strategy, which selectively penalizes acoustically and semantically confusing negative categories to learn more discriminative decision boundaries. Extensive experiments on the AVSBench‑OV dataset demonstrate that our method significantly outperforms previous state‑of‑the‑art approaches, particularly on unseen categories. Code is available at https://github.com/spyflying/AGCL.

Authors:Dongwook Lee, Sangkwon Park, Eunwoo Song, Che Hyun Lee, Youngho Cho, Junho Kim, June Young Yi, Heeseung Kim, Sungroh Yoon
Title: HEAR Who Said What: Unlocking Speaker-Attributed Reasoning via Counterfactual Voice Grounding
Abstract:
Speech Language Models (SLMs) are increasingly deployed in multi‑speaker environments, yet their ability to attribute speech to the correct speaker and reason over speaker identities remains unclear. Hence, we introduce HEAR, a conceptually hierarchical benchmark diagnosing the foundational capabilities of speaker‑attributed reasoning, comprising 2.4K human‑verified samples from 887 diverse multi‑party audio clips. Evaluating 20 leading SLMs on HEAR reveals they struggle with these foundational tasks, often relying on semantic priors rather than actual vocal cues. To address this, we present A2R, a 30B model optimized on Counterfactual Audio with Speaker‑level Hard negatives (CASH), a dataset designed to guide the model to prioritize acoustic vocal cues over linguistic signals. A2R achieves strong performance on HEAR and exhibits zero‑shot generalization to diverse multi‑speaker downstream tasks, demonstrating that learned speaker attribution unlocks the model's latent capacity for speaker‑aware reasoning. All resources are available at https://attributetoreason.github.io/AttributeToReason/

Authors:Yang Chen, Canyu Shen, Xinzhe Rao, Yuanyi Yan, Yunlu Chen, Meng Tang, Teng Long, Vincent Tao Hu
Title: GramLoop: Training-Free Gram-Gated Replay for Robust Dense Prediction
Abstract:
We aim to improve frozen DINOv3 dense‑prediction models under distribution shift by adding inference computation inside the visual backbone, without changing model weights, task adapters, or prediction heads. The challenge is that repeated transformer‑block computation must refine dense features without disrupting the pairwise patch relations that DINOv3 uses to preserve spatial structure. We introduce GramLoop, a training‑free framework that replays a short transformer window and controls each replay through final‑layer cosine‑Gram consistency. Each proposal is propagated through the frozen suffix, measured against the standard DINOv3 trajectory, and accepted through a patchwise gate at the replay‑window endpoint. Across object detection and semantic segmentation under corruptions, perturbations, and natural shifts, GramLoop improves all five shifted benchmarks over the paired DINOv3 baseline. On COCO‑O, it improves mAP by +0.252 and Effective Robustness by +0.250, while preserving clean ADE20K performance. Code will be released at https://github.com/cheyan9/GramLoop.

Authors:Zafar Ali, Asad Khan, Nimbeshaho Thierry, Nabila Amir, Adam A. Q. Mohammed, Pavlos Kefalas
Title: HANIA: Planner-Guided Multimodal Graph Evidence Selection for Grounded Question Answering
Abstract:
Multimodal question answering remains sensitive to noisy, incomplete, and weakly grounded evidence. Long unstructured contexts can introduce redundancy and encourage unsupported generation, while flat retrieval may overlook relations needed for multi‑step reasoning. We present HANIA, a planner‑guided multimodal graph framework for evidence‑grounded question answering. HANIA processes the supplied image and text using a frozen vision‑language model to extract concise question‑relevant visual evidence with explicit abstention. It then constructs an input‑grounded multimodal graph and applies a two‑group finite‑state planner to coordinate descriptive and relational evidence. Coverage‑aware pruning retains a compact evidence set based on relevance, graph confidence, concept coverage, and modality diversity. The selected passages, visual statements, and graph triples are provided to a frozen instruction‑tuned decoder. We evaluate HANIA on ScienceQA using answer accuracy, evidence‑filtering quality, evidence‑budget sensitivity, and efficiency. The results show that structured evidence planning and compact graph‑guided retrieval can support competitive multimodal question answering without target‑dataset fine‑tuning or iterative retrieval. The code is available at https://github.com/Zafar‑southeast/HANIA.

Authors:Yichen Wei, Faisal Zaghloul, Soujanya C Aryal, Aanya K. Agrawal, Chengfan Li, Jason Xinyu Liu, James Tompkin, Stefanie Tellex
Title: GHOST in the Robots: Real-Time Exocentric Dual-Robot VR Teleoperation from Onboard Cameras
Abstract:
Teleoperating multiple robots simultaneously enables additional views and coordinated control. Yet, it poses fundamental challenges: the system must present sensor data cohesively and allow operators to manage multiple robot bases, arms, and cameras while maintaining low latency. Current multi‑robot teleoperation systems require multiple operators, rely on autonomy, or restrict operators to high‑level commands. We present GHOST: an open‑source VR teleoperation system that enables single operator control of two mobile manipulators via direct lowlevel commands using only onboard sensing. GHOST creates an exocentric 3D workspace by aligning real‑time point clouds from the robots' RGB‑D cameras, where scene coverage is improved through learning‑based completion to aid operator spatial awareness. For control, the operator uses a mode‑switching architecture to command either robot individually or both robots simultaneously. Experiments with 15 novice participants demonstrate 1.6‑4x the success rate of an off‑the‑shelf tablet interface. For experts across nine challenging dual‑robot tasks, our system enabled completion of two tasks that were infeasible with the tablet, and was 1.47x faster on average than the tablet. Website and code: https://h2r.github.io/GHOST/.

Authors:Soohyun Choi, Seonvin Cho, Songnam Hong
Title: PathBridger: Subgoal Bridges for Offline Goal-Conditioned Reinforcement Learning
Abstract:
Offline goal‑conditioned reinforcement learning (GCRL) aims to learn policies for reaching diverse goals entirely from fixed trajectory data. Long‑horizon offline GCRL remains challenging because sparse goal‑reaching signals must be propagated over many steps, while execution errors cannot be corrected through additional environment interaction. Existing methods address these challenges by improving long‑range value estimation or reducing the effective decision horizon through subgoals, options, and action chunks. In several hierarchical methods, however, a selected subgoal specifies where to go, while the intervening state‑space path remains implicit in an endpoint‑conditioned low‑level policy. To address this interface, we propose PathBridger, a hierarchical offline GCRL method that explicitly connects subgoal selection to short‑horizon execution. PathBridger constructs a state‑space bridge toward the selected intermediate endpoint and decodes it into a short executable action chunk using an inverse dynamics model. Experiments across the evaluated OGBench tasks demonstrate strong aggregate performance, with particularly large gains on the multi‑object Cube manipulation tasks. Code: https://github.com/SChoish/PathBridger

Authors:Shuomin Xue, Jingyuan Li, Ju Jia, Jingxuan Yu, Xiaojun Jia
Title: Let Prompts Bridge Defense Knowledge: Transferable Graph Purification via Vulnerability-Aware GPL
Abstract:
Graph Neural Networks (GNNs) have emerged as a cornerstone for representing complex relational dependencies in diverse multimedia tasks, particularly in cross‑platform user interest modeling and cross‑modal semantic alignment. In the real world, a practical defense against graph adversarial perturbations is needed. However, we observe that the prevailing adversarial purification methods are essentially domain‑restricted defenses, which leads to the following shortcomings: (1) single‑domain data provides insufficient structural and semantic diversity for learning robust purification criteria; (2) training of domain‑specific defense strategies from scratch consumes substantial computational cost. To address the above limitations, we propose a transferable graph purification scheme, named ProGAP, to bridge adversarial defense knowledge via vulnerability‑aware graph prompt learning. Firstly, to capture universal adversarial patterns, a perturbation‑capture edge detector is pretrained on data‑rich graphs by jointly modeling topological and semantic information. Subsequently, to achieve more knowledge transfer w.r.t. robustness, vulnerability‑aware prompts are designed that inject targeted purification guidance into biased nodes, during which the pretrained detector adapts to distribution shifts in downstream graphs without parameter‑laborious updates. Experimental results demonstrate that compared with state‑of‑the‑art baselines, our ProGAP achieves 1%‑9% improvement, and reduces the time consumption by up to 2.2x. The code for ProGAP is available at https://github.com/Lieyoufffff/ProGAP.

Authors:Zhongliang Liu, Wenjie Liu, Yang Li
Title: DReSG: Diffusion Residuals for Stylized Gaussian Splatting
Abstract:
Reference‑guided stylization of scenes represented by 3D Gaussian Splatting (3DGS) is important for efficient and controllable 3D content creation. Existing VGG‑feature‑based 3D stylization methods provide stable rendered‑view optimization, but often under‑represent expressive reference style cues; diffusion models offer stronger image priors, yet direct per‑view or score‑based diffusion guidance can lead to view drift, local artifacts, and hard‑to‑control appearance updates. We present DReSG, a 3D‑grounded residual‑feedback framework for stylized Gaussian splatting. DReSG represents attention‑guided diffusion proposals as residual targets relative to the current render, and progressively absorbs these residuals into a shared Gaussian scene through multi‑view Gaussian feedback. To make this feedback stable and controllable, DReSG modulates residual strength during target construction and combines coverage‑aware view selection with conflict‑filtered color updates during multi‑view fitting. Extensive experiments demonstrate that DReSG achieves competitive reference‑guided stylization while better preserving scene structure and cross‑view stability. Our project page is available at https://vpx‑ecnu.github.io/DReSG‑website/.

Authors:Hanting Li, Xin Sun, Wei Ye, Jungong Han, Liang-jie Zhang
Title: Di$^2$CycleSB: Towards High-Quality Unsupervised Nighttime Visibility Enhancement via Schrödinger Bridge Transformer
Abstract:
Light‑effect contamination poses a significant challenge to nighttime visibility enhancement. Most methods suppress light effects by estimating and decomposing them through prior‑driven regularization, yet they are often limited by hand‑crafted priors and ill‑posed nature of decomposition. This work proposes Di^2CycleSB, a unsupervised Cycle Schrödinger Bridge Transformer framework guided by dynamic integral image priors, for high‑quality unsupervised nighttime visibility enhancement. Specifically, a novel light‑effect estimator is introduced to parameterize Gaussian‑like adaptive priors by aggregating dynamic integral image representations for non‑uniform glow estimation. Then, we propose a prior‑informed Generator that exploits light‑effect representations to guide long‑range dependency modeling within our specific Transformer blocks. We formulate light‑effect suppression as a Schrödinger bridge problem and construct forward and backward bridges with cycle consistency constraints to achieve visually pleasing enhancement. Extensive experiments on real‑world datasets demonstrate the remarkable effectiveness of our Di^2CycleSB in enhancing nighttime visibility. In particular, it achieves effective end‑to‑end light‑effect suppression without any regularization constraints and image decomposition. The code and models are available at https://github.com/LHTcode/Di2CycleSB.

Authors:Ke Zhao
Title: Disentangling Representation using Attributes-based Gaussian Estimation for Medical Sound Diagnosis
Abstract:
Deep learning has a powerful capability of feature extraction. However, the lack of fairness and interpretability in deep neural networks poses limitations to their adoption in the medical domain. This paper proposes a disentangled representation learning (DisenRL) framework, named the Attributes‑based Gaussian Estimation for Disentangled Representation (AGEDR), which incorporates Attribute Mapping Embedding (AME) modules designed to map attributes into vectors and align them with a subset of the latent vectors in a Variational AutoEncoder (VAE). This part of the latent vector will be disentangled from the remaining latent vectors by minimizing mutual information. A classifier is then trained using the mean parameters of the latent vectors from the VAE. Extensive experiments demonstrate that AGEDR outperforms both conventional classification models and existing disentangled representation learning methods. The ablation experiments also indicate the disentangling capability and fairness of AGEDR. The source code is publicly available at https://github.com/ZhaoKe1024/DisentangledRepr.

Authors:Hui Huang, Ye Sun, Shiyan Hu
Title: Frequency Selective Neural Networks as a Foundation Architecture for Time Series Learning
Abstract:
Time‑series data across physical and biological domains are fundamentally driven by complex, non‑stationary oscillatory modes. While deep learning models, such as Convolutional Neural Networks (CNNs), Recurrent Neural Networks, and Transformers, have dominated sequential analysis, they remain fundamentally "spectral‑blind". By mapping continuous physical waves into unconstrained spatial or discrete token spaces, these architectures suffer from severe spectral entanglement, acting as opaque black boxes that decouple predictive accuracy from physical reality. In this paper, we introduce the Frequency Selective Neural Network (FSNN), pioneering a foundation architecture guaranteeing physical interpretability without sacrificing expressive power of deep learning. FSNN addresses spectral entanglement by explicitly embedding the rigorous mathematics of advanced signal processing into its neural topology. Through a fully differentiable Wiener‑like filter bank optimized via complex‑domain backpropagation, FSNN autonomously discovers and isolates the precise physical modes of a given task. Extensive evaluations demonstrate that FSNN establishes state‑of‑the‑art predictive performance, achieving 77.0% average accuracy on the standard 10 multivariate UEA datasets and leading across all major metrics on the highly imbalanced PTB‑XL clinical ECG benchmark. Crucially, in contrast to yielding abstract feature maps, FSNN converges directly on physically meaningful frequency bands, such as isolating the cardiac QRS complex, providing a highly scalable, interpretable paradigm for robust pattern recognition in complex temporal domains. Our code is available at: https://github.com/ad6174hhhh/FSNN.

Authors:Mohammad Nazeri, Alexandyr Card, Samira Huber, Anuj Pokhrel, Yujun Wang, Ruben Hammele, Daeun Song, Sören Pirk, Xuesu Xiao
Title: Hydra: A Navigation World Action Model with Discrete Latent Planning and Continuous Flow-Matching Execution
Abstract:
World models let robots imagine possible futures, but exploiting this capability for real‑time control is bottlenecked by a representation misalignment: the generative model and the planner operate on decoupled manifolds, so the planner has no shared structure to search over and must instead decode every candidate back into high‑dimensional pixel space to evaluate it. This decoding step is a major obstacle to real‑time control on physical hardware. In this paper, we present Hydra, a discrete World Action Model that closes this gap by moving the planner, both the sampler and the evaluator, inside the model. Hydra establishes a unified latent manifold over visual states, physical poses, and control actions, then compresses this manifold through modality‑specific Vector‑Quantized bottlenecks into discrete vocabularies of kinodynamic intents and visual states. Because candidates are now drawn directly from this shared manifold, sampling is informed by the model's own understanding of the observation rather than proposed blind, and evaluation happens natively within the discrete space: candidates are ranked by a Kinematic‑Perceptual Cost, without ever decoding to pixels. We term this Discrete Latent Planning (DLP). Because planning over discrete intents alone cannot supply the smooth, continuous commands physical actuation requires, Hydra pairs DLP with conditional Flow Matching, which maps each selected intent to a continuous trajectory for execution. Evaluated on two physical robotic platforms, Hydra outperforms state‑of‑the‑art world models in goal‑directed planning, while matching or exceeding the closed‑loop execution capabilities of leading reactive foundation policies.

Authors:Runyu Guan, Dehao Wu, Qiqi Xie, Yang Li, Haohan Wang
Title: Agentic AI uncovers conserved cross-tissue protein co-abundance programs inaccessible to single-dataset analysis
Abstract:
Protein co‑abundance clusters preserved across tissues can reveal shared disease mechanisms and candidate therapeutic targets, particularly when proteins implicated in organ‑confined diseases converge in peripheral or accessible tissues. However, previous cross‑tissue studies have focused on biologically pre‑selected tissue pairs, leaving most possible combinations and non‑obvious relationships unexplored. We present an LLM‑agent framework for large‑scale, evidence‑grounded comparison of tissue‑specific protein co‑abundance networks. The framework constructs tissue networks, derives pairwise consensus clusters, and integrates evidence from expression atlases, protein interaction and complex databases, pathway annotations, disease catalogues, and literature. Applied to all 820 pairwise combinations of 41 human tissues and fluids, it identified 1,833 conserved co‑abundance clusters across 406 tissue pairs. Colon, synovial fluid, blood, cerebrospinal fluid, and bone marrow were the most broadly connected tissues, while the most cluster‑rich pairs were dominated by bone marrow. The analysis also highlighted non‑obvious relationships: skin‑bone marrow exceeded the anatomically adjacent bone‑bone marrow pair, while colon‑breast contained cancer‑relevant clusters involving extracellular‑matrix remodeling, lipid metabolism, and immune modulation. Cluster‑level analyses generated further mechanistic hypotheses, including a brain‑gut extracellular‑vesicle/redox/serotonin‑cofactor axis and a liver‑bone marrow stress‑response axis involving genes linked to white matter disease. These results provide a global, comparable landscape of conserved protein co‑abundance and a hypothesis‑generating resource for mechanistic and therapeutic exploration. Code and data are available at https://github.com/Gry1005/AgenticAI‑conserved‑cross‑tissue‑protein‑co‑abundance.

Authors:Kiyan Rezaee
Title: The Illusion of Replacement: Rethinking Specialized Machine Learning Models in the Foundation Model Era
Abstract:
Can the specialized architectures that machine learning has traditionally built for structured data be replaced by language‑based models? This question is examined through a review of 159 papers (2016‑‑2026) across nine modalities, with predictive accuracy considered alongside structural representation and computation. A distinction is made between performing a task and preserving and computing the structure that makes the task tractable, and existing approaches are organized into eight representational regimes, ranging from language‑only systems to fully specialized architectures. Language‑mediated models are found to be highly competitive in specific settings, including extreme few‑shot prediction, discretized symbolic tasks, textually annotated knowledge graphs, and large‑scale single‑modality pretraining. However, whenever structural representation or computation is directly evaluated rather than accuracy alone, no evidence of general architectural replacement is found. Instead, a recurring pattern is observed across independent research communities: when language alone is insufficient, the missing structure is reintroduced through a graph module, structural tokens, specialized attention, or another non‑linguistic component. In this sense, specialization more often relocates than disappears. Moreover, although performance of language‑based models is improved by scaling, whether the gap to a structure‑aware architecture can eventually be eliminated remains untested. The official repository for this work is available at https://github.com/kiyan‑rezaee/language‑vs‑structure.

Authors:Theo Rusu, Sourena Khanzadeh, Manar Alalfi
Title: Selective Forgetting: A Graph-Based Memory Framework for Long-Term LLM Agents
Abstract:
Knowledge graphs have been proposed as a structured alternative to flat retrieval‑augmented generation for long‑term agent memory, on the assumption that representing conversations as entities and relations improves recall. We evaluate that assumption directly. Our framework extracts each conversational turn into typed nodes and attributed edges, answers questions from a two‑hop subgraph, and periodically prunes nodes that score low on a weighted combination of recency, access frequency, degree centrality, and age. On LongMemEval, the graph does not outperform a flat vector baseline at a matched candidate‑generation budget of five retrieval roots: token F1 is 0.417 against 0.468, and a paired bootstrap over 500 questions gives Δ= ‑0.050 (95% CI [‑0.085, ‑0.016]). The gap is widest on questions that require recalling a specific prior assistant turn, where judged correctness falls from 0.911 to 0.607, suggesting that decomposing a turn into entities discards the surface form these questions depend on. The forgetting module is more successful. Applied once to a persistent 27,021‑node graph, it removes 9.8% of nodes and 9.5% of stored bytes; token F1 is unchanged (+0.001, 95% CI [‑0.015, +0.016]) and judged correctness falls by 1.6 points, with the 95% interval bounding any loss at 3.8 points ([‑0.038, +0.006]). Because our extractor is a single small model evaluated on one benchmark, these results characterise this extraction‑based pipeline rather than graph‑structured memory in general. Code: https://github.com/skhanzad/Selective‑Amnesia

Authors:Saloni Modi, Srivi Balaji, Yusong Zhu, Gautam Kamath, Kevin Tian
Title: Revisiting the Provable-Auditable Privacy Gap of DP-SGD
Abstract:
Differential privacy (DP) has traditionally been used to provide theoretical upper bounds on an algorithm's stability to changing its training data. In modern private machine learning applications, achieving strong tradeoffs between utility and theoretical privacy is challenging, and thus one may optimistically hope that existing theoretical privacy analyses are loose. Recent work on privacy auditing has adopted a dual viewpoint, instead lower bounding the true privacy of an algorithm by constructing empirical distinguishing events. The auditing literature has thus far yielded a pessimistic outlook on the looseness of theoretical privacy bounds for DP‑SGD, the de facto private training method in modern ML, as nearly‑matching empirical lower bounds have been achieved under various threat models [NHSBTJCT23, AC24, CBP25]. In this work, we propose the empirical privacy lower bound of an algorithm as a concrete metric to optimize for, complementary to the theoretical upper bound. We give a lightweight defense framework that generically augments optimization methods in the ML pipeline to have significantly‑improved empirical privacy on standard benchmarks. Moreover, we show that our framework comes at no theoretical privacy cost when augmenting DP‑SGD, unlike previously‑proposed defenses against membership inference attacks. We evaluate our defense against a broad range of audit constructions, models, and datasets to demonstrate its flexibility.

Authors:Jungseob Lee, Jaehyung Seo, Heuiseok Lim
Title: The Hallucination Signal Is a Mean Shift: Why Simple Probes Suffice
Abstract:
Hidden‑state probes effectively detect LLM hallucinations, but the geometry of the signal remains poorly characterized, driving increasingly complex probe architectures. Across three 7B‑scale models and three datasets in a paired‑example paradigm, we find the signal overwhelmingly dominated by a single mean‑shift component, and removing this direction collapses detection to chance. Shrinkage linear discriminant analysis closes about 73% of the gap between 1D and full‑dimensional classifiers, so apparent architectural complexity largely reflects high‑dimensional covariance estimation difficulty rather than exploitable non‑linearity. A simple L2‑regularized logistic regression (0.952 AUROC) bounds or outperforms twelve controlled architectural alternatives, and our multi‑layer aggregation exceeds CLAP cross‑layer attention probing under matched paradigm. Because the signal spans a contiguous layer band, LayerMix aggregates it to match oracle‑layer performance without oracle access. Our claims characterize the geometry within the controlled paired‑example paradigm. Our code is available at https://github.com/js‑lee‑AI/LayerMix.

Authors:Noah Videcrantz, Mostafa Mehdipour Ghazi
Title: ActiveAugment: Online Active Learning for Augmentation Selection in Deep Learning
Abstract:
Data augmentation is a cornerstone of deep learning pipelines, yet existing strategies treat it as a static, model‑agnostic preprocessing step, either relying on expensive dataset‑specific policy search or applying transformations uniformly at random, regardless of what the model has already learned. We introduce ActiveAugment, a unified framework that treats augmentation selection as an online active learning problem. For each training minibatch, ActiveAugment generates a pool of candidate augmented views and scores each candidate using a combination of the model's predictive uncertainty and the feature discrepancy induced by the augmentation. The augmentation under which the current model is most fragile is selected per sample, and the model is then trained with a joint supervised classification and supervised contrastive objective that enforces intra‑class invariance to the selected augmentations while maintaining inter‑class separation. We evaluate ActiveAugment on eight benchmark datasets spanning natural and medical imaging, using CNN and transformer architectures across three training regimes (training from scratch, full fine‑tuning, and linear probing), and comparing eight active selection strategies for augmentation scoring. ActiveAugment outperforms AutoAugment, RandAugment, and TrivialAugment under controlled augmentation shifts across all domains and budgets, with the most pronounced gains at low labelling budgets. On medical imaging datasets, where data is scarce and domain shift relative to natural‑image pretrained models is large, ActiveAugment achieves higher test F1 than all baselines, demonstrating strong cross‑domain adaptability. Our analysis reveals that the augmentation selection policy evolves meaningfully during training and that strategy choice has a direct impact on generalisation. Code is available at: https://github.com/noahvide/ActiveAugment.

Authors:Anh Thi Luu, Nick Lemke, Anirban Mukhopadhyay
Title: Coarse to Fine: Iterative Adversarial Neural Cellular Automata for Medical Image Synthesis
Abstract:
Large‑scale, publicly available datasets have driven advances in deep learning, but privacy and legal restrictions often limit data sharing in medical imaging. Synthetic data generation offers a privacy‑friendly alternative to enable the training of high‑performance models on health data. While most state‑of‑the‑art generative models produce high‑quality images, they remain computationally expensive, which limits their applicability on resource‑constrained hardware. We propose StyleGANCA, the first lightweight general‑purpose NCA‑based generative adversarial network. The architecture integrates a StyleGAN‑inspired mapping network and adaptive style modulation into a multi‑scale NCA synthesis process, enabling latent‑controlled image generation through iterative local interactions. We evaluate StyleGANCA on BloodMNIST and PathMNIST against adversarial, variational, diffusion, and NCA‑based baselines. Experimental results demonstrate that StyleGANCA achieves competitive image quality with substantially fewer parameters than baseline architectures, achieving the best FID and KID scores on PathMNIST with only 617k parameters. Furthermore, downstream experiments show that the generated images preserve class‑specific information and effectively support the training of multi‑class classifiers. Our code is publicly available at: https://github.com/MECLabTUDA/StyleGANCA

Authors:Giuseppe Stracquadanio, Kevin Raj, Julia Grabinski, Stefan Roth
Title: ReconSplat: Generalizable 3D Scene Reconstruction Beyond Observed Views
Abstract:
We introduce ReconSplat, a feed‑forward model for 3D scene reconstruction that aims to address the longstanding trade‑off between plausible view generation for unobserved regions and geometric consistency, providing both geometrically aligned novel views and sharp depth estimates. Our approach builds on 3D Gaussian splatting (3DGS) as an intermediate differentiable scene representation and integrates it with a multi‑view latent diffusion model (MV‑LDM) trained to act simultaneously as a refiner and an inpainter for appearance and scene geometry. We enforce geometric consistency by guiding the diffusion process with variational 3D latent features for appearance and geometry, encoded by the feed‑forward 3DGS representation and rasterized to 2D latent space. ReconSplat produces both photorealistic novel views and accurate depth maps on real‑world benchmarks, RealEstate10K and DL3DV‑10K, outperforming existing methods in challenging extrapolation setups. Notably, ReconSplat allows the extrapolation of unseen and challenging viewpoints jointly with coherent and precise scene geometry.

Authors:Lilliane Linnet Musoke, Atta Badii, Ahmed Ashlam
Title: Enhancing Web Application Firewalls with Machine Learning for SQL Injection Detection
Abstract:
Detecting SQL Injection (SQLi) attacks ranks among the most critical challenges in web application security. This research conducted a systematic literature review to identify the research gaps in this domain and responsively designed and optimised a DistilBERT‑Stacked Ensemble pipeline to improve detection efficiency and robustness while reducing false‑positive and false‑negative rates. Comprehensive pre‑processing and tokenisation were performed, DistilBERT embeddings were extracted, and machine‑learning and ensemble classifiers were trained and ranked on accuracy, precision, recall and F1‑score. The three best performers (Logistic Regression, XGBoost and SVM) were combined through a neural meta‑learner to form a stacked ensemble. The ensemble was hardened with adversarial examples generated by the Fast Gradient Sign Method (FGSM) and tuned with Optuna. The optimised ensemble achieved 99.81% across all reported metrics, closely comparable to the strongest single model (DistilBERT SVM, 99.82%). On the evaluation platform used in this study (Section 3.8), the ensemble classified the full test set in 0.0136s against 1.896s for DistilBERT‑SVM, an approximately 140‑fold reduction in measured inference latency, while retaining 99.77% accuracy under a single‑step FGSM attack. The contribution is the design and validation of a SQLi detector performing with state‑of‑the‑art accuracy at real‑time speed and with demonstrated robustness to a single‑step FGSM attack, rather than a marginal gain in accuracy. Sensitivity analysis further confirmed the stability of the model. These findings highlight the value of adversarial training and stacked meta‑learning in building robust Web Application Firewalls (WAFs) for SQLi detection. For open validation, the dataset, test sets and models are made available at https://github.com/mlily2024/Final‑project‑SQL‑injection‑pipeline.

Authors:Roberto I. Ono Filho
Title: Moving the Mean Toward the Known Good, Not Beyond It: What Inference-Time Interventions and Weight Consolidation Buy in Open-Ended Generation
Abstract:
What does a generation loop gain from learning on its own verified successes? In cycles of generate, verify, select and LoRA‑consolidate on online bin packing, training on value‑filtered candidates shifts what the model writes on held‑out variants toward value (‑1.7 points of excess, p=0.008; ‑3.1 against a random‑consolidation control, p=0.004) while the best observed candidate converges to the classic heuristic's level and no further. A confirmation battery replicates the whole procedure three times, with fresh seeds and a never‑consulted held‑out set read exactly once: the mean was nearly identical in all three lineages (‑2.0, ‑1.8, ‑1.9), and after aggregating within held‑out variant all seven evaluable variants favored consolidation (p=0.008). The best observed candidate moved to the classic heuristic's level, exactly (0.021028 in all three lineages, for attract and for the random control alike), and never beyond it. A matched SFT‑only control shows the supervised anchor, not repulsion from bad candidates, does the concentrating (96% of candidates land exactly at the classic heuristic's level). The tails cut both ways: consolidation lowers the per‑candidate rate of better‑than‑classic candidates (10% to 3.9%) while its larger production yields more such candidates absolutely (5 against 1, on few events). As motivation we report the inference‑time ledger that led here: a model‑written schematic recap buys judged document integration and nothing buys development; a verifier written into the stream is imitated, 16.4 fabricated verdict lines per notebook. Mean quality among valid candidates can be bought and replicated; the observed best goes to the classic and, so far, never beyond it.

Authors:Lilliane Linnet Musoke, Atta Badii, Ahmed Ashlam
Title: Enhancing Web Application Firewalls with BERT-GNN for SQL Injection Detection
Abstract:
Detecting sophisticated SQL Injection (SQLi) attacks remains among the most critical challenges in web applications security. This research study has resulted in an optimised hybrid BERT‑GNN pipeline with improved detection accuracy and robustness while reducing false‑positive and false‑negative rates. SQL queries are tokenised and encoded into contextual BERT embeddings, which then initialise the node features of a Graph Neural Network (GNN) trained to classify each query, with the architecture tuned by Optuna over accuracy, precision, recall, and F1‑score. The proposed model achieved 99.67% accuracy, with 99.71% precision, 99.39% recall, and 99.55% F1‑score on the attack class. A sensitivity analysis, performed by perturbing graph inputs, further assessed the model robustness and yielded a low mean sensitivity score of 0.0037, indicating stable predictions under such perturbations. The results have demonstrated the potential of a novel hybrid model that couples BERT contextual understanding with the GNN structural modelling to detect sophisticated SQLi attack vectors. For open validation, the dataset, test sets and models are made available at https://github.com/mlily2024/Final‑project‑SQL‑injection‑pipeline.

Authors:Evan Bell, Jiaming Liu, Yifan Chen, Yu Sun
Title: Generative Translation Priors: Bayesian Imaging with Cross-Modality Image Translation
Abstract:
The ability to leverage images from co‑available modalities to inform target‑domain reconstruction is highly desirable in imaging algorithms. In this work, we introduce Generative Translation Priors (GTP)‑‑a Bayesian framework that transforms diffusion‑based image‑to‑image translation models into cross‑modality image priors for ill‑posed imaging inverse problems. GTP incorporates target‑domain measurements through likelihood guidance, steering the translation process toward the desired posterior distribution. The framework is grounded in a theoretical analysis of the resulting posterior dynamics, which reveals an intrinsic bias introduced by likelihood guidance. We further characterize this bias and derive a ground‑truth‑free formulation for its estimation, enabling it to serve as a practical metric for assessing posterior sampling quality. Building on this analysis, we derive two discretized GTP algorithms based on gradient and proximal likelihood guidance, respectively. We validate GTP on computed tomography reconstruction with magnetic resonance side information, and on positron emission tomography reconstruction with computed tomography side information. Experiments demonstrate that GTP effectively incorporates complementary cross‑modality information and achieves high‑fidelity reconstruction even under severely undersampled measurements.

Authors:Dylan Jayabahu, Tinuade Adeleke
Title: The Halt Vector: Internalizing a Causal Steering Intervention for Efficient Reasoning
Abstract:
Reasoning models do not stop when they know the answer. On DeepSeek‑R1‑Distill‑Qwen‑7B the chain of thought runs about twice as long as the model's own answer probability takes to settle, and how much of that excess is removable varies from problem to problem, so a global length penalty cannot take it out. We take it out by internalizing a causal interpretability finding into the weights. The mechanism is a halt vector: a difference‑of‑means direction at layer 18 of this model whose steering strength controls how long it thinks, while a replicated value axis does nothing. Installing that intervention in the weights is harder than it looks. Maximizing the scalar projection onto the direction corrupts the off‑axis dimensions a frozen downstream reader depends on, and generation gets longer instead of shorter; what works is reconstructing the whole steered activation with those dimensions pinned to their natural values. Fit from 24 problems and no reinforcement learning, the halt removes about a quarter of the thinking at held accuracy across five unseen benchmarks, and the cut tracks each problem's own removable slack at 0.70. It also closes a non‑termination pathology that grows with difficulty and that a decoding‑time confidence hook makes worse. We do not claim to beat a well‑tuned length penalty or decoding‑time early exit on the raw trade‑off; the contribution is how the halt is obtained.

Authors:Nipa Anjum, Md Irfan Pavel, Robert Gonzalez, Kevin Desai, Alberto Cordova, M. Rasel Mahmud, John Quarles
Title: Toward Postural State Classification in Immersive VR with Multimodal Data and Explainability Analysis
Abstract:
Ensuring a safe virtual reality (VR) experience requires systems that can predict and respond when users lose their balance. Although prior work has examined fall prediction and motion sickness, many approaches are regression‑based and postural state classification remains less explored. This study compares machine learning (ML) and deep learning (DL) models for classifying postural states in VR under visual perturbations. We used a multimodal dataset containing kinematic, electromyographic (EMG), and electrodermal activity (EDA) signals. The data were prepared for a binary task to distinguish balanced from imbalanced postural states, and participant‑wise downsampling addressed class imbalance. All models were evaluated with Leave‑One‑Participant‑Out (LOPO) cross‑validation to test generalization to unseen participants. Among the models, the Mamba‑inspired CNN (MI‑CNN) achieved the highest accuracy of 96.76%. SHapley Additive exPlanations (SHAP) analysis improved interpretability and identified the most influential classification factors. The SHAP results showed that kinematic features were dominant, indicating that body‑motion patterns are informative for detecting imbalance in VR. We also evaluated MI‑CNN using only the top two‑thirds of features ranked by SHAP importance. Despite a 33% reduction in input dimensionality, the model maintained performance, achieving 0.957 accuracy and 0.957 F1‑score, with about a 1% decrease compared with the full‑feature model. These findings suggest that multimodal sensing, temporal deep learning, and explainable AI can support reliable classification of balance‑related instability in VR. Accurate recognition of imbalanced postural states may raise awareness of fall risk and guide safer, adaptive VR systems that respond to instability while improving user safety and experience. Code is available at: https://github.com/NipaAnjum/MI‑CNN.

Authors:Yiyi Lu, Yilai Qian, Yucheng Jin
Title: Visible but Not Yet Curatable: Characterizing the Curatability of Compact and Derived Open LLM Artifacts
Abstract:
Open Large Language Model (LLM) research increasingly produces compact and derived artifacts, such as adapters, quantized checkpoints, merged models, and distilled variants, that are distributed across papers, model hubs, model cards, code repositories, and release statements. Although these artifacts are publicly visible, digital libraries often lack sufficient evidence to identify, preserve, and cite them as coherent scholarly objects. We introduce a framework that conceptualizes curatability as a record‑level property of distributed scholarly records and operationalizes it through four evidence dimensions: artifact identity, scholarly linkage, upstream evidence, and release assets. Guided by this framework, we conduct the first collection‑scale characterization of open LLM curatability using a May 2026 snapshot of 191,375 public Hugging Face repositories and a core corpus of 2,214 scholarly papers. Our results reveal a pronounced visibility‑to‑curatability funnel. While 90.7% of paper records contain at least one useful curation signal, only 18.1% combine usable upstream evidence with concrete release evidence, and only 6.1% provide sufficiently coordinated evidence to support high‑curatability records. Based on these findings, we derive a minimal seven‑field curatable record and complementary responsibilities for model hubs, scholarly indexes, and digital libraries, providing practical guidance for improving the preservation and bibliographic control of open LLM artifacts.

Authors:Xiaohan Zhao, Jiacheng Liu, Yaxin Luo, Zhiqiang Shen
Title: FigMirror: Ground It, Code It, Plot It
Abstract:
Converting scientific figures into executable code has gained increasing attention, yet existing methods primarily focus on reproducing the reference figure itself. A more practical setting is to plot new data while preserving the visual style of a reference figure (e.g., color scheme and typography). Prior approaches mimic the reference through pixel‑level optimization and struggle to carry its style to new data. We show that the key to this task lies in the coordinate grounding and coding capabilities present in modern computer‑use models. We propose FigMirror, an agentic framework that unlocks these capabilities through Grounded Measurement, which locates visual elements by coordinates and measures their properties through executable code. We further introduce PlotTwin‑Bench, an expert‑curated benchmark with fine‑grained code and image‑level style metrics. Experiments show that FigMirror consistently outperforms existing methods on reference‑conditioned style transfer. All plots in this paper are generated by FigMirror, except those produced by other methods for comparison. Our code and data are available at: https://github.com/VILA‑Lab/FigMirror.

Authors:Rahul Nair, Saurav Pandit, Hannah Kerner
Title: A Large-scale Evaluation of Text-guided Models for Facial Editing
Abstract:
Facial appearance editing powers popular applications like FaceApp and Photoshop. Generative Adversarial Networks (GANs) and 3D Morphable Models (3DMMs) have been widely used for facial editing. GANs can perform varied facial edits (e.g., changing hair color, hairstyle), but often produce unstable edits. 3DMMs produce stable edits, but can only alter pose and facial expression. Recently, text‑guided diffusion models like Nano Banana have become popular for image editing. Text‑guided models are a compelling alternative to GANs and 3DMMs since they can produce both stable and varied image edits. While text‑guided models have been widely tested for whole‑scene edits (e.g., ``make the woman play a guitar''), they have not been comprehensively tested for facial editing. We conducted the first large‑scale evaluation (~1M images evaluated) of six popular text‑guided models on a sequential facial editing task. We present Face‑Edit‑Attributes, the largest collection of 169 facial editing attributes focused on hair, accessories, and pose edits. We compared model performance using two popular celebrity face datasets: CelebA and CelebSET. Our results show that most models performed hair and accessory edits well, but struggled with editing pose. All models over‑edit (e.g., changing hair color when asked only to change the hairstyle). We also evaluated demographic biases in each model. Our results show surprising biases in overediting: almost all models created more overedits for dark‑skinned male faces and old faces. The code and data for our results (including our repository of ~ 1M images) can be accessed \hrefhttps://github.com/rahul1801/Face‑Edit‑Bench\textcolorbluehere.

Authors:Xiaoman Lu, Jiaqi Li, Shuntian Zheng, Huiping Chen, Yu Guan
Title: FairReL: Deepfake Detection using Fairness-Aware Representation Learning
Abstract:
Although recent deepfake detectors achieve high overall accuracy, their errors remain unevenly distributed across demographic subgroups, with real faces from certain groups more often misclassified as fake. Existing fairness‑aware detectors typically regularise the entire feature representation, without identifying or controlling the specific components that drive unfair predictions. Such coarse intervention can over‑suppress useful forgery cues while leaving demographic structure in component‑specific subspaces. To address this, we identify two subgroup‑sensitive components: multi‑scale spatial features, which encode local facial and forgery patterns, and fine‑tuning‑induced residual features, which adapt the backbone to the unfair training distribution. We propose FairReL, a fairness‑aware representation‑learning framework that targets both components with dedicated demographic supervision. FairReL uses an SVD‑decomposed foundation‑model backbone to isolate the fine‑tuning‑induced residual representation, and introduces two complementary losses. Group‑Conditional Wavelet Decorrelation (GCWD) suppresses subgroup‑imbalanced structure across spatial wavelet sub‑bands, while Subspace‑Localised Mean Alignment (SLMA) aligns subgroup means within each real/fake class in the residual representation. Experiments on FF++, Celeb‑DF, DFD and DFDC show that, against the state‑of‑the‑art fairness‑aware detector, FairReL improves unseen‑dataset AUC by 3.9% while reducing subgroup FPR disparity by 10.2%. Code is available at https://github.com/xiaoman89/FairReL .

Authors:Xin Jiang, Minhao Wang, Wen Wu, Zhentao Xie, Shangheng Du, Jinxin Shi, Jiabao Zhao
Title: ERR+: Sequential Entropy Resolution for Efficient and Decisive LLM Reasoning
Abstract:
Large reasoning models achieve strong performance on complex tasks by generating extended chain‑of‑thought (CoT) traces via reinforcement learning with verifiable rewards (RLVR). While current RLVR methods have achieved strong results with correctness‑based reward signals, they provide limited guidance on the quality of the reasoning process itself, leaving the internal reasoning structure largely unoptimized. Through empirical analysis across multiple model families, we identify a consistent pattern: correct reasoning trac es exhibit more frequent and larger token‑level entropy drops within the thinking phase than incorrect ones. We propose ERR+, a two‑phase RLVR framework grounded in this observation. The first phase trains with the Entropy Relief Reward (ERR), a bonus proportional to cumulative token‑level entropy drops in the thinking phase, log‑normalized by response length. Unlike prior methods that suppress entropy, ERR rewards the resolution of uncertainty while leaving exploratory high‑entropy states unconstrained. The second phase introduces the Robust Relative Efficiency Reward, which scores each response's length against co‑generated peers via a \tanh‑transformed within‑group z‑score. We provide a formal analysis showing that joint optimization of the two objectives induces gradient conflict in early training, motivating the sequential design . Experiments on five datasets demonstrate consistent improvements in both accuracy and response conciseness across model backbones. Our code is available at https://github.com/XrkArul/err_response

Authors:Shaozu Ding, Linan Song, Dajiang Suo
Title: Inter-3D VQA: A Roadside Multimodal Benchmark for 3D Spatiotemporally Grounded Visual Question Answering
Abstract:
Recent advances in visual question answering (VQA) and multimodal large language models (MLLMs) have enabled natural‑language reasoning over traffic scenes. However, existing benchmarks are largely built from ego‑vehicle views or 2D roadside videos, limiting their ability to evaluate 3D‑grounded reasoning over real‑world distances, trajectories, infrastructure topology, and safety‑critical interactions. We introduce Inter‑3D VQA, a large‑scale roadside multimodal benchmark for 3D spatiotemporally grounded VQA at intersections. Built from synchronized point clouds and multi‑view images, Inter‑3D VQA contains 407K QA pairs covering lane‑level positions, object relationships, motion patterns, and near‑miss‑oriented interaction reasoning. We further propose Inter‑Geo, an MLLM baseline that integrates object‑ and scene‑level aligned LiDAR representations, and Inter‑Metrics, a unified evaluation framework for textual consistency, numerical accuracy, and semantic correctness. Experiments show that Inter‑Geo outperforms image‑based VLMs, especially on grounded spatial and temporal reasoning tasks. Our benchmark and codes are available at https://github.com/ASU‑Suo‑Lab/Inter‑3D‑VQA .

Authors:Xinyuan Gui, Shaowen Wang, Sheng Sun, Zijian Wang, Zishu Yu, Zheming Yang
Title: Pro-Router: Token-Aware Progressive Model Routing with Adaptive Edge-Cloud Collaboration for Efficient Multimodal LLM Inference
Abstract:
The remarkable performance of multimodal large language models (MLLMs) comes at the cost of substantial computational overhead, posing significant challenges to real‑time deployment and cost effectiveness. Existing model routing approaches either decide from coarse request‑level features alone or spend one or several extra language model passes to inspect the generated response, leaving the token‑level uncertainty signals that emerge during generation unused. To address these limitations, we propose Pro‑Router, a token‑aware progressive model routing method with adaptive edge‑cloud collaboration for efficient multimodal LLM inference. Pro‑Router employs a two‑stage progressive decision mechanism. First, a lightweight prompt pre‑scorer module performs rapid pre‑screening before token generation begins, guiding apparently simple requests to small models. Second, a token‑aware verifier reads the sampling probability distribution of each token the small model generates, estimating the model's confidence in its own output to determine, per request, whether the answer ships or escalates to the cloud‑based high‑precision model. Furthermore, we design an adaptive edge‑cloud serving pipeline that sizes every dispatch to each device's measured service rate, so both the edge and the cloud tiers stay fully utilized without manual parameter tuning and are not impacted by the network latency. Extensive experiments on multiple multimodal benchmark datasets and models demonstrate the effectiveness of Pro‑Router. Compared to other methods, it achieves the highest routing accuracy and improves routing speed by more than 10x. Its serving pipeline also reaches more than 75% higher end‑to‑end throughput than the existing model routing pipeline. Our code is available at https://github.com/xinyuangui2/pro‑router.

Authors:Anoop Senthil
Title: ReVA: A Region-Aware Visual Assistant for Visually Grounded Question Answering
Abstract:
Multimodal Large Language Models (MLLMs) have achieved remarkable progress in Visual Question Answering (VQA), yet they continue to struggle with questions requiring precise spatial reasoning and fine‑grained visual understanding. These limitations often manifest as object, attribute, and spatial hallucinations, where models generate confident but visually unsupported responses due to insufficient region‑level and fine‑grained visual grounding. To address this challenge, we propose ReVA, a region‑aware VQA model that employs a frozen CLIP ViT‑L/14 Vision Transformer (ViT) and a Qwen2.5‑7B‑Instruct large language model (LLM) connected through a dual bridge that aligns both whole‑image and region‑level representations with the LLM's embedding space. The image bridge maps final transformer block features into image tokens. The region bridge maps cropped features from enriched intermediate features across ViT blocks so early texture and later object cues are more evident, into K region tokens for every bounding box. ReVA uses a detector stack that supplies automatic zero‑shot bounding boxes that are both question‑agnostic and question‑dependent, using RAM++ (Recognize Anything Model), spaCy, and Grounding DINO. The image tokens and region tokens are concatenated as an LLM prompt prefix to jointly encode scene‑level context and fine‑grained regional evidence when answering questions. Evaluated on VQAv2, MMBench, POPE, and SEED‑Bench, ReVA achieves 82.85% mean F1 on POPE, compared with 81.14% for an image‑token baseline without region tokens. These results demonstrate that explicit region‑aware visual representations reduce object hallucination and improve the factual grounding of MLLMs.

Authors:Khayrul Islam
Title: Variable-Granularity Tokenization for High-Resolution Object Detection
Abstract:
ViT detectors fix a uniform token grid before any learned stage. A native‑resolution aerial detector must then choose between resolving few‑pixel objects and staying inside compute and memory limits. We introduce VGTok, a training‑free tokenizer that sets patch granularity per region from pixels, ahead of the encoder. VGTok scores each region by multi‑scale morphological top‑hat separability from its surround, then thresholds those scores at a per‑image percentile, which fixes the token budget. A structure‑tensor gate (λ_\min) refines only where two‑dimensional object structure supports it, leaving one‑dimensional clutter coarse. The resulting token set is a strict partition of the image. In a Co‑DETR detector with an EVA‑02 ViT‑L encoder, VGTok clears every published VisDrone‑val AP and AP_S at every budget from 40% to 100% of tokens. At 40% it records 44.22 AP with three fifths of the sequence discarded before the first transformer block; dense, it reaches 48.38 AP, 6.08 above the strongest published entry. VGTok transfers to AI‑TOD‑v2 untouched, same scorer and same rank, and sets a new state of the art at 37.27 AP and 19.51 AP_vt. As a pure drop‑in into a frozen checkpoint it reaches 36.29 AP at 78.5% of tokens, above every published entry, where our 376.3M‑parameter detector clears a 3.0B multi‑expert model. We show that a token budget fixed before the backbone, from local separability and structure geometry alone, holds accuracy on the tiny‑object regimes that dominate aerial detection, at 3.1× less encoder compute and 1.9× less encoder memory. Code and models are available at \hrefhttps://github.com/khayrulbuet13/vgtok\textttgithub.com/khayrulbuet13/vgtok and \hrefhttps://huggingface.co/khayrulbuet13/vgtok\texttthuggingface.co/khayrulbuet13/vgtok.

Authors:Matiur Rahman Minar, Seunghun Oh, Ganghyeon Jeong, Unsang Park
Title: SNF-Bench: Separating Static Drift from Natural Flow in Long-Horizon Fixed-Camera Video Generation
Abstract:
Long‑horizon video generation is evaluated with whole‑frame metrics that reward motion and temporal consistency. For fixed‑camera nature scenes this creates an ambiguity: motion of water, fire, smoke, or rain is desirable, whereas motion of the background is an error. A system can therefore score well on motion while its scene drifts, or on consistency while its flow stagnates. We introduce SNF‑Bench, an evaluation framework for long‑horizon fixed‑camera generation that partitions each scene into static support and dynamic flow and reports static fidelity, flow persistence with absolute magnitude, and drift leakage separately, never as one score. Drift leakage is interpretive context rather than a headline measurement. Each factor is validated mechanistically rather than by correlation with preference: we inject global translation, rotation, and scale drift and progressive late freezing at known severity into real generations, and require each factor to respond in its stated direction and to remain selective against corruptions it does not target. Auditing publicly released long‑horizon text‑conditioned checkpoints under one recorded common inference configuration, plus an image‑conditioned track with released‑pipeline references and a deployment‑sensitivity panel, we find that whole‑frame motion and static‑region drift induce near‑opposite orderings of the same outputs. At maximum controlled translation, fBD and NBF rise to 1.86× and 1.32× baseline, but whole‑frame Dynamic Degree reaches only 1.07×‑‑‑rewarding the corruption. SNF‑Bench measures where motion occurs and whether it persists; it does not measure physical realism. Project page: https://minar09.github.io/snfbench/.

Authors:Abbas Aliyev, Samir Rustamov
Title: Data Diversity, Not Frequency Invariance: A Controlled and Self-Audited Study of Compression-Robust Deepfake Detection
Abstract:
Frequency features and compression‑invariant representation learning are widely assumed to be key to deepfake detection that survives video compression. We test this with CAFRL ‑ block‑DCT and FFT‑phase streams, compression‑level‑conditioned band attention, and adversarial (gradient‑reversal) compression invariance ‑ and report a controlled negative. Under a pre‑registered protocol with capacity‑ and augmentation‑matched controls, a plain EfficientNet‑B0 on multi‑quality data beat CAFRL as specified at every compression level on the FaceForensics++ test split, by 3.66 AUC points at CRF 40 (paired, single seed). A self‑audit of our own negative found four defects biased against the frequency hypothesis, and pre‑specified re‑tests repairing all four showed the deficit to be a recipe artifact, not an architecture failure: the baseline recipe recovered 3.96 points over the matching shipped‑recipe variant. The frequency path made no detectable difference: discriminative alone (standalone validation AUC 0.91‑0.98 late in training) but of no marginal value under this fusion, at two feature widths of one 4.0 M trunk, every seed‑pooled interval for the intra‑dataset compression contrasts including zero; on the single held‑out manipulation tested, the fair variants sat below the plain backbone. The adversarial branch, as specified, added nothing and degraded its own conditioning estimator; at the fair recipe it is untested. Robustness under single‑pass H.264 re‑encoding came instead from data diversity: real constant‑rate‑factor variants beat synthetic JPEG augmentation by 7.3 points (single runs, non‑overlapping intervals). The evidence is FaceForensics++‑family, GAN‑era and single‑codec. Match controls on training recipe as well as capacity, and buy compression robustness with codec diversity before architecture.

Authors:Jizong Zhan
Title: MIRAGE-CAD: Construction-Mediated Multimodal Generation of Executable CAD Programs
Abstract:
Recovering an executable parametric CAD program from an observed object is fundamentally ambiguous, because the same final geometry can result from different construction procedures. We study this problem from four types of input: natural‑language descriptions, rendered images, point clouds, and STEP/B‑Rep geometry. MIRAGE‑CAD maps each input to a shared construction representation and mediates program generation through an explicit construction‑plan interface. The resulting Python CAD code is executed by an OpenCASCADE kernel to build the solid and export it as STEP. On 2,500 held‑out queries per modality, the system achieves 55.4‑70.0% build success and 52.3‑66.2% STEP export success without retrieval at inference. Controlled comparisons show that strong reconstruction does not depend on expressing the construction representation as text: a decoder conditioned directly on the continuous representation also reconstructs strongly, while an exposure‑matched plan‑based decoder shows no detected material loss in per‑part geometric fidelity. The explicit plan instead provides a readable and separately measurable intermediate representation whose agreement with the reference construction is informative about downstream execution success. Finally, we show that executable validity, geometric fidelity, and parametric responsiveness can diverge substantially and should therefore be evaluated separately.

Authors:Hardik Iyer, Tirath Bhathawala, Mihir Panchal, Ying-Jung Chen, Kiran Bhowmick, Pankaj Sonawane, Meera Narvekar
Title: FRAC-MAS: A Safe and Explainable Multi-Agent System for Fracture Diagnosis
Abstract:
Fracture detection and its clinical interpretability see notable improvements when deep vision models are integrated with agentic AI architectures. While deep learning models achieve high diagnostic performance, their black‑box nature limits clinical adoption. We propose FRAC‑MAS, an agentic AI system for automated, explainable, and safe bone fracture detection. The framework combines a stacked ensemble of four vision models with conformal prediction to produce statistically grounded differential diagnoses, while a multi‑agent workflow performs independent verification, retrieves clinical guidelines, and generates patient‑friendly reports. A pipeline‑depth ablation study confirms that our multi‑agent critic triages 86.6% of cases into a high‑confidence auto‑confirmed cohort while escalating uncertain cases, outperforming a single‑agent baseline. Patient preference studies against Llama, MedGemma, and Gemini further demonstrate significantly more comprehensible clinical reports. These results suggest that integrating multi‑agent critics with conformal guarantees enables safer radiology triage while preserving clinician oversight. More broadly, FRAC‑MAS demonstrates how cooperative agentic architectures can serve as auditable, human‑in‑the‑loop decision support systems for safety‑critical healthcare. Our code is available at https://github.com/hardik1712/FRAC‑MAS, and the website is available at https://frac‑mas.vercel.app.

Authors:Gaurav Kukreja, Parul Kukreja, Mohammed Abraar, Raj Dandekar, Rajat Dandekar, Sreedath Panat
Title: BiasMix-Finance: Post-Generation KYC Guardrails for LLM Portfolio Advice
Abstract:
Large language models (LLMs) can generate plausible‑sounding ETF portfolios while silently violating basic KYC‑style constraints on risk, fees, and diversification. This is especially problematic in agentic multi‑turn advisory systems, where each draft recommendation can become an action unless guarded by an auditable enforcement layer. We study a model‑agnostic, asset‑agnostic post‑generation guardrail pipeline: (i) enforce a strict JSON allocation schema, (ii) validate allocations against numeric caps, and (iii) when violations occur, deterministically project the output to the nearest feasible portfolio via a convex quadratic program (QCQP). We introduce BiasMix‑Finance (Mini), a compact stress‑test benchmark for constrained decision‑making under biased LLM generations, with a 16‑ETF universe, three investor profiles, and eight bias prompts. Across three models and three inference modes (direct, critique, self‑consistency), first‑pass generations violate at least one cap in 47.6‑85.7% of test cases (67.2% pooled), but the convex projection layer reduces final feasibility violations to 0% while requiring only a small correction distance (test pooled median D=||w‑w0||_2=0.066), indicating that the guardrail typically preserves the intent of the original allocation. We report violation rates and correction distances with confidence intervals, and paired model comparisons with multiple‑testing correction. To support reproducibility, we release the dataset, prompts, caps, and code in our public GitHub repository.

Authors:Yunsu Kim, Kaden Uhlig, Ashwin Purohit, Milind Agarwal, Patrick Simianer, Anil Arslan, Kiarash Mokhtari, Thomas Zenkel, Johannes Mosig, Gabriel Bretschner, Shamik Bose, Joern Wuebker, John DeNero
Title: Terminal-Bench-LILT: Multilingual Agentic Coding Benchmark Grounded in Language, Region, and Culture
Abstract:
Most evaluations for coding agents are conducted exclusively in English, which does not reflect real‑world multilingual deployment. We present Terminal‑Bench‑LILT, a suite of 300 authentic coding tasks in ten languages: Arabic, Czech, German, Spanish, Hindi, Japanese, Korean, Serbian, Turkish, and Chinese. Each task targets issues specific to non‑English software development that have no direct English equivalent, e.g., internationalization, encoding, text normalization, and cultural conventions. All tasks are authored by native‑speaker programmers and validated through a multi‑stage quality control pipeline. Evaluation of six frontier models reveals that even the strongest model reaches only 63.1% pass rate, with many tasks unsolved by any model. Performance varies substantially by language and does not track general coding benchmark rankings, highlighting that multilingual coding competence is a distinct and underexplored capability axis. Sample tasks are available at https://github.com/lilt/terminal‑bench‑lilt

Authors:Taaha Kazi, Vasu Sharma, Mohammad Saifullah, Abdur Rahman
Title: PAUSE: Editable Strategy Artifacts for Long-Form Cultural Story Adaptation
Abstract:
Generative AI systems increasingly mediate cultural adaptation, but their cultural decisions are often hidden inside prompts, transient model plans, or final prose. We study PAUSE (Pause‑And‑Update Strategy Editing), an intervention that exposes an editable adaptation strategy as a human control surface for cultural decisions in long‑form story adaptation. The strategy is a structured artifact that can be inspected, edited, and then projected through downstream character, entity, and chapter‑localization stages. In two Chinese‑source serialized novels, we test whether human edits to this strategy propagate into chapter‑level prose. Across 9 edited‑vs‑control chapter comparisons, judges select the edited‑strategy output in all 9; a marker audit shows target markers in 8/9 edited outputs and 0/9 controls, with forbidden markers absent from edited outputs and present in all controls. We frame these results as a smoke‑scale edit‑adherence study, not a claim that the outputs are culturally authoritative or literary‑quality improvements. PAUSE offers one practical way to make AI‑mediated cultural adaptation more inspectable and contestable before decisions propagate through long‑form generation.

Authors:Zhaohe Dong, Yuhao Chen
Title: CrossAudit: A Git-Native, Cross-Vendor Audit Loop for Agentic Science
Abstract:
An AI scientist should not grade its own homework. Yet in the systems we examined, the agent that reviews the work usually comes from the same model family as the agent that produced it, or at least from the same vendor. Model evaluators are known to favour their own generations. Whether models trained alike also share blind spots is a conjecture, not a settled finding, but if they do, the reviewer inherits the author's. The record of what was flagged and what was waved through often sits in platform logs that nobody outside can replay. We present CrossAudit, a protocol for supervising autonomous research pipelines. It rests on three commitments. Each increment of work is audited by an agent from a different vendor against a rulebook a human wrote and versioned. Reports, verdicts, disputes and rulings are git commits, so the supervision history can be re‑read and cited; raw model exchanges are not yet part of that record. Scripted checks run before any model does. Advisory judgement never gates the pipeline: a model blocks only by citing a rule, and no model may waive a deterministic failure. Blockers that survive a bounded number of revision rounds go to a person. We state the protocol as eight invariants. We describe a reference implementation built from GitHub Actions and a few hundred lines of Python, and report a live deployment of a closely related variant in a computational‑chemistry pipeline. We also ran a seeded‑defect trial (30 increments, 43 seeded defects, one run per configuration). A cross‑vendor audit of our own repository then voided its blinding. We adopt that audit's findings and report the corrected results. The trial shows that two vendors read the same rulebook differently. It does not show that either is better. The strongest evidence here is the committed, uncontrolled record of cross‑vendor audits of this paper itself.

Authors:Xiao Yang, Erik Edward Aldape, Beren Millidge
Title: PUFFER: Incremental Fuzzy Deduplication for Continuously Evolving Corpora
Abstract:
Large language model training corpora grow through successive, often redundant releases, so each release must be deduplicated against both itself and the accumulated history. At trillion‑token scale, this requires incremental ingestion, bounded resident memory, deterministic retry, and dataset‑scoped lifecycle control without repeated corpus‑wide rebuilding. We introduce PUFFER (Provenance‑aware Updatable Fuzzy Filtering for Evolving Repositories), a MinHash‑LSH fuzzy‑deduplication pipeline built around two design choices. First, PUFFER stores each LSH band as immutable, dataset‑tagged, memory‑mapped sorted segments, enabling exact historical band‑key membership checks without RAM proportional to corpus size. Second, T‑fanout tiered compaction periodically merges segments to control screening fanout, trading lower query cost against additional index‑maintenance writes while preserving membership decisions. Across N ingested keys and K equal‑sized releases, PUFFER's cumulative maintenance cost is O(N log N log_T K), compared with Theta(KN) for repeated snapshot rebuilding. Dataset‑tagged segments also support dataset‑scoped withdrawal: removal is constant‑time for uncompacted or protected datasets, while post‑compaction withdrawal reconstructs only the affected merged segment, even if the original dataset is unavailable. In our implementation, PUFFER completed cumulative index‑stage ingestion for one billion documents in about 1.75 hours in a single process, using 128 bytes per document for a 16‑band index. A classical resident MinHash‑LSH table required about 6.5 KB per document and exceeded a 900 GiB RAM cap. In a ten‑hour comparison capped at one billion documents, PUFFER was 11x faster than LSHBloom and 35x faster than Milvus‑LSH. PUFFER is deployed on more than 30 billion documents, and we release it as open‑source software at https://github.com/Zyphra/puffer.

Authors:Howard Kim, Keun Tae Cho
Title: Distributional Validity and Calibration of a Korean Synthetic Persona Panel for Digital and AI Service Use: A Secondary-Data Validation Against the Korea Media Panel Survey
Abstract:
Synthetic personas based on large language models (LLMs) are increasingly proposed as substitutes for human survey respondents, yet systematic validation outside English‑speaking contexts remains scarce. This secondary‑data study evaluates how well a Korean synthetic persona panel (NVIDIA Nemotron‑Personas‑Korea), conditioned into Gemini 3.5 Flash (primary) and EXAONE (comparison), reproduces digital and AI service‑use distributions from the KISDI Korea Media Panel Survey. Sex‑and‑age‑stratified panels of about 8,000 personas per model answered the survey's own items ‑ eight service‑use indicators and eight innovativeness and acceptance constructs ‑ and were compared against weighted survey estimates. The overall mean absolute error (MAE; RQ1) was 15‑19 percentage points (pp), with binary item‑mean correlations of 0.69‑0.90. Segment error (RQ2) across five demographic axes was 15‑19 pp, with between‑group gaps up to 52.4/36.2 pp (Gemini/EXAONE). Errors followed model‑specific signatures: an age stereotype with low anchoring (Gemini) versus an acquiescence‑consistent level bias (EXAONE). Reference‑year analysis was consistent with temporal misalignment driving most generative‑AI overestimation, whereas short‑form underestimation was framing‑sensitive. Holdout calibration on 30% of the real data (RQ3) roughly halved sex‑by‑age cell MAE (18.9‑>8.6, 15.9‑>6.7 pp) ‑ yet direct estimation from the same real subsample was far more accurate (3.6 pp), and the correction did not transfer across time. The calibrated panel retained an advantage only under extremely scarce real data (about 100 responses) and, for one model, for unobserved segments. Persona‑narrative conditioning beat demographic‑only conditioning, but neither surpassed simple real‑data baselines. Synthetic panels are thus not survey substitutes; their value is diagnostic, with operational use confined to settings lacking real data.

Authors:Isha Narang, Sneh Gosai, Mayank Singh
Title: Gurukul AI: An Interactive AI-Driven Educational Platform for Indian Education System
Abstract:
Recent advances in large language models (LLMs) like ChatGPT and LLaMA have transformed AI‑driven education, but these systems are predominantly trained on Western‑centric data, making them ill‑suited for regional curricula like India's. The Indian education system is linguistically diverse, exam‑oriented, and structured around standardized syllabi, not addressed by existing datasets or tools. In this work, we curate a syllabus‑aligned QA dataset based on NCERT (National Council of Educational Research and Training) textbooks for classes 9‑12, capturing the content, context, and teaching style of Indian curricula. The final dataset, comprising 18,720 question‑answer pairs across five subjects, is publicly available at https://huggingface.co/datasets/LingoIITGN/Gurukul. We fine‑tune the LLaMA 3.1 8B model using this dataset and deploy it in a Retrieval‑Augmented Generation (RAG) framework tailored to educational needs. We introduce GurukulAI, an open‑access platform that enables Indian students to chat with the model, get doubts cleared, practice exam‑style questions, receive contextual answers, and interact in both English and Hindi. By localizing AI for Indian classrooms, our work bridges the gap between global LLM capabilities and regional educational demands. The code is available at https://github.com/lingo‑iitgn/GurukulAI.

Authors:Jonghyun Song, Sangjun Song, Minjae Oh, Haesung Pyun, Sungsik Lee, Yohan Jo
Title: SHAPE of Chain-of-Thought in Math Reasoning
Abstract:
Large language models (LLMs) achieve strong performance on mathematical reasoning benchmarks, yet the mathematically meaningful skills underlying their reasoning remain underexplored. We introduce \textttSHAPE, a framework that analyzes Chain‑of‑Thought (CoT) trajectories through two lenses developed in mathematics education: (1) semantic spaces: the model's evolving mathematical interpretations of a problem (e.g., algebraic, geometric), and (2) heuristics: the specific mathematical actions taken within those spaces (e.g., simplifying the problem, working backward). We first use \textttSHAPE to analyze the reasoning patterns of various models. Our findings reveal that the mathematical heuristics employed by a model better explain final answer correctness than traditional CoT features. Furthermore, models are likely to reach correct solutions by concentrating their reasoning effort within a few semantic spaces rather than exploring many disparate ones ‑‑ a pattern consistent with human behavior. Next, we utilize the \textttSHAPE lens to evaluate whether post‑training truly enhances mathematical proficiency. We find that reinforcement learning induces mode‑seeking in heuristic usage. Lastly, we post‑train LLMs by promoting diverse heuristics and demonstrate its effectiveness in improving accuracy. Overall, \textttSHAPE provides a theoretically‑grounded diagnostic framework for decoding LLM reasoning and offers a new path toward post‑training LLMs for math reasoning. The code for our model is available at https://github.com/holi‑lab/SHAPE‑of‑CoT

Authors:Moustafa Yehia Hassan, Sharon Wong, Woh Kai Xuan
Title: The Signal in the Noise: An Auditable Reliability Layer for Biomedical Text Classification
Abstract:
Biomedical NLP pipelines routinely presuppose clean input text, yet large‑scale corpora assembled through automated PDF parsing harbour pervasive OCR‑like artifacts, token splits and merges, hyphenation remnants, and character‑level corruption, that systematically erode lexical evidence and degrade downstream classifiers. We introduce a conservative, fully auditable spell‑correction reliability layer conceived as a safety‑oriented preprocessing module rather than a maximal‑accuracy corrector: under conditions of uncertainty, the system abstains from editing, in accordance with a medical do‑no‑harm philosophy. The deterministic architecture couples bounded edit‑distance candidate generation with corpus‑derived n‑gram scoring and a suite of biomedical safety gates that protect domain‑critical terminology. We evaluate the layer both intrinsically, on a manually curated benchmark of 2,104 token‑level cases, and extrinsically, on a tri‑class CORD‑19 topic classifier (Prevention, Treatment, Epidemiology) spanning 10,000 examples under a principled four‑run protocol (Clean, Noisy, Restored, Safety). Intrinsically, the layer attains 94.61% error‑fix recall on synthetic errors with zero harmful edits on negative controls. Downstream, it recovers approximately 80.45% of the noise‑induced macro‑F1 degradation, elevating macro‑F1 from 0.7654 (Noisy) to 0.7717 (Restored) while preserving near‑clean performance (Safety: 0.7721). A supplementary case study on 103 real‑world OCR‑extracted abstracts classified with BioBERT confirms that transformer encoders appeared relatively robust to mild noise, motivating a future grey‑box architecture that integrates bounded neural signals and UMLS lexicons without compromising auditability. The system is fully deterministic, artifact‑driven, and designed with deployment and auditability in mind.

Authors:Fan Liu, Hao Liu
Title: DS-Lighting: Making Agent Harnesses Explicit for Data-Science Automation
Abstract:
Large Language Model (LLM) agents have shown promise for automating data‑science workflows, yet their end‑to‑end performance depends critically on the agent harness that represents tasks, manages execution state, constrains output artifacts, and provides evaluation feedback. Existing data‑science agents often leave this harness implicit, making results difficult to reproduce, compare, and attribute across heterogeneous tasks. We introduce DS‑Lighting, a unified harness toolkit that makes harness design explicit for data‑science automation. DS‑Lighting decomposes the harness into four reusable layers: data, workflow, execution, and evaluation, and represents diverse agents as executable operator programs that support both predefined pipelines and adaptive search. We further integrate multiple open‑source data‑science benchmarks into an MLE‑Bench‑style task format, enabling controlled comparison under a shared task interface, sandboxed runtime, and metric protocol. Experiments across agents, harnesses, models, and ablations show that explicit harness design improves reproducibility, comparability, and reliability, while reducing avoidable system‑level failures in end‑to‑end data‑science workflows. Our code is available at https://github.com/usail‑hkust/dslighting

Authors:Nan Wang, Mohit Yadav, Jonathan Wulff, Aidan Rosenbaum, Kezhou Chen, Yuvan Sharma, Xu Dong, Yiwei Tao
Title: Aero Hand Open: A Simulation-Ready Tendon-Driven Hand for Dexterous Manipulation Learning
Abstract:
Tendon‑driven hands are anthropomorphic, and moving the actuators off the joints is what makes a hand of this capability affordable to build. Two effects produce that saving. Routing force through a cable removes the requirement that a motor fit inside the joint it drives, so smaller and cheaper motors suffice, and one motor can drive several joints through a single cable, so fewer motors are needed. They are also harder to learn on than a direct‑drive hand. The underactuated transmission that produces the saving is itself difficult to represent in a simulator, and the joints one cable drives are not independently commandable. We present Aero Hand Open, a tendon‑driven anthropomorphic hand that is released simulation‑ready. Three things ship with it. A simulation model reproduces the cable transmission itself. An identified actuation map connects that model to the motor commands in both directions, including the three‑way coupling of the thumb. A reinforcement learning package trains policies for the hand. Together they let a policy be trained entirely in simulation and run on the hand with no fine‑tuning and no state estimation. We release the mechanical design, the simulation model, the identified mapping, the training environment and the deployment stack.

Authors:Haosen Yang, Jifei Song, Zhensong Zhang, Xiatian Zhu, Jiankang Deng
Title: Video Generative Models as Geometry Learner
Abstract:
Recent generative approaches to geometry estimation adapt pretrained image diffusion models and treat the task as image‑conditioned generation. Leveraging off‑the‑shelf image diffusion models, they either (i) train task‑specific geometry models (for depth and surface normal estimation) independently, losing the opportunity of exploring the intrinsic correlation of these geometric targets, or (ii) jointly fine‑tune modified image diffusion backbones (e.g., altered self‑attention), which typically demands substantial labeled data. To overcome these limitations in a principled fashion, we repurpose pretrained video generative models as a unified and data‑efficient framework for geometry estimation, formulated innovatively as a next‑frames prediction task. Our method, GeoNeXt, inherits naturally structured knowledge and richer priors from the video model, while further adapting them for joint modeling of images and geometry targets (image <‑> geometry), enabling more data efficient and effective learning of geometry. Extensive experiments validate our method for zero‑shot monocular depth and surface normal estimation across diverse datasets, outperforming both previous task‑specific and unified generative competitors while using substantially less training data. Notably, our method rivals discriminative state‑of‑the‑art approaches trained on over 100x more data and even standouts on several benchmarks.

Authors:V. S. D. S. Mahesh Akavarapu, Michael Daniel, Gerhard Jäger
Title: Phoneme- and Word-Level Metrics Using Self-Supervised Speech Representations for Forced Alignment Evaluation
Abstract:
Forced alignment evaluation typically requires manually annotated timestamps, limiting large‑scale and multilingual analysis. We introduce two corpus‑level metrics based on self‑supervised (SSL) speech representations for reference‑free forced alignment evaluation: Phoneme‑Cluster Mutual Information (PCMI) and Word Acoustic Consistency Score (WACS). PCMI measures agreement between aligned phoneme labels and clusters induced from SSL‑speech representations, while WACS measures consistency of repeated word realizations using dynamic time warping similarity between word representation sequences. Using both random and systematic perturbations, we show that PCMI and WACS degrade consistently under alignment perturbations. We further analyze the metrics across multiple alignment systems on 85 languages from FLEURS, validate them against manually annotated alignments from 45 languages in DoReCo, and evaluate them on two phonologically complex low‑resource languages. The metrics effectively separate high‑ and low‑quality alignments and correlate strongly with timestamp‑based alignment quality measures. Our results demonstrate that SSL‑speech representations enable scalable, reference‑free forced alignment evaluation. The metrics are available as an open‑source Python package at https://github.com/mahesh‑ak/forced‑aligner‑metrics.

Authors:Xinyi Zhang, Yutong Li, Peijie Sun
Title: SG-UMP: Sequence-Guided Universal Multimodal Prioritization Calculation Framework
Abstract:
Multimodal sequential recommendation (MSR) improves recommendation by incorporating heterogeneous information such as text, images, and user interactions. However, existing MSR methods often fail to capture user‑level preference heterogeneity and dataset‑level modality bias, limiting their adaptability across users and datasets. To address this issue, we propose Sequence‑Guided Universal Multimodal Prioritization Calculation Framework (SG‑UMP), a plug‑and‑play plugin for enhancing multimodal information processing in MSR. SG‑UMP includes a Module Combiner for flexible multimodal processing and a Module Router for dynamic module ordering, enabling adaptation to both user preferences and dataset characteristics. Experiments on four real‑world datasets show that SG‑UMP consistently improves recommendation performance across different backbones and multimodal settings. The code is available at https://github.com/esemsc‑xz524/SG‑UMP .

Authors:Zhuoshi Pan, Junru Lu, Yan Qian, H. Vicky Zhao, Di Yin, Xing Sun
Title: Blind Men and the Elephant: Probing the Epistemic Myopia of LLMs under Long-Tail Divergent Knowledge
Abstract:
Factual question answering (QA) typically assumes a single canonical answer, obscuring whether large language models (LLMs) retain divergent accounts of long‑tail facts. To address this gap, we introduce ElephantBench, a closed‑book knowledge probe comprising 1,094 questions generated through an auditable graph‑based pipeline. The pipeline retrieves related documents from a low‑exposure web corpus, identifies naturally occurring disagreements, and converts them into multi‑account QA records. Each answer is verified against the originating documents and authoritative public web sources and is then reviewed by human annotators. Across 32 models, even the strongest model recovers both accounts on only 52.4% of questions, while on nearly all remaining questions it recalls one account but omits the other. Scaling model size and inference‑time reasoning improve recall but do not eliminate this incompleteness. Corpus analysis further shows that exposure imbalance favors the dominant account, whereas greater minority‑side exposure is associated with more complete recall. These findings establish ElephantBench as a reproducible knowledge probe for diagnosing epistemic myopia in parametric memory. More broadly, our graph‑based benchmark construction pipeline provides an efficient and scalable way to turn long‑tail corpora into source‑traceable knowledge probes, supporting efforts to evaluate and advance the epistemic rigour of next‑generation LLMs. Code is available at https://github.com/Tencent/ElephantBench.

Authors:Zhuoshi Pan, Qizhi Pei, Junru Lu, Honglin Lin, H. Vicky Zhao, Di Yin, Xing Sun
Title: ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL
Abstract:
Long‑horizon agentic tasks require large language models (LLMs) to iteratively retrieve, integrate, and maintain dispersed information across multi‑turn interactions, but preserving all interaction histories leads to a continuously growing working context. Recent proactive context management methods allow models to edit their own working context with specialized tools, yet they still face three key limitations: (1) a limited toolset restricted to search, deletion, and summarization, with no support for global planning, long‑term memory, and adaptive compression; (2) inefficient exploration that treats context management actions uniformly despite their heterogeneous impacts on final outcomes; and (3) coarse‑grained credit assignment that assigns the final trajectory‑level reward to all intermediate context editing actions during RL. To bridge these gaps, we introduce ContextPilot, a proactive context management framework for long‑horizon agentic reasoning. Our approach systematically augments the toolset with planning, long‑term memory, and soft context offloading tools. We further propose an RL method tailored for context management, which uses context and entropy variation to identify critical editing decisions for branch sampling and estimates action‑level advantages from all branched trajectories that pass through the corresponding context editing action. Experiments on long‑context QA and deep search tasks show that ContextPilot achieves stronger performance with a more compact working context, consistently outperforming existing baselines across various base models and benchmarks. Code is available at https://github.com/Tencent/ContextPilot.

Authors:Perry Hart
Title: On Left Adjoints Preserving Colimits in Homotopy Type Theory
Abstract:
We examine how the standard proof that left adjoints preserve colimits behaves in the setting of wild categories, a natural setting for synthetic homotopy theory inside homotopy type theory. We show that the proof may fail for adjunctions between wild categories and even produce a wild left adjoint that fails to preserve colimits. Our core contribution, however, is a sufficient condition on the left adjoint for the proof to go through. The condition, which we call 2‑coherence, expresses that the naturality structure of the hom‑isomorphism commutes with composition of morphisms. We present two useful examples of this condition in action. First, we use it, along with a new version of a known trick for homogeneous types, to show that the suspension functor, as well as a generalization thereof, preserves graph‑indexed colimits. Second, we show that every modality, viewed as a functor on coslices of a type universe, is 2‑coherent as a left adjoint to the forgetful functor from the subcategory of modal types, thereby proving this subcategory is cocomplete. We have formalized our main results in Agda.

Authors:Huseyin Umut Isik, Mehmet Alp Ozaydin, Sila Kurugol, Şeyda Ertekin
Title: ARC-CT: Anatomy-Routed Contrastive Vision-Language Learning for 3D Chest CT
Abstract:
Contrastive vision‑language learning uses paired chest CT volumes and radiology reports to learn abnormality classifiers without manually annotated labels. However, two characteristics of chest CT challenge conventional global contrastive learning. First, many critical abnormalities are small or anatomically localized, and pooling an en‑ tire volume into a single embedding may dilute their visual evidence. Second, the standard contrastive objective treats every other scan in a batch as a negative. Because many chest CTs share abnormalities, this objective incorrectly pushes co‑positive pairs apart. We propose Anatomy‑Routed Contrastive Learning for 3D Chest CT (ARC‑CT), a region‑aware framework that addresses these limitations using only la‑ bels extracted from reports by an LLM, with no manual annotations or bounding boxes. ARC‑CT combines three components: (1) an Anato‑ myQFormer localizing evidence via queries constrained by automatically generated organ masks; (2) a label‑Jaccard soft InfoNCE objective in‑ tegrating the standard one‑hot target with the label‑set overlap of each pair, which reduces false‑negative penalties between studies that share clinical findings; and (3) an organ‑level alignment loss connecting mask‑ pooled visual features to organ‑specific report text extracted offline with a large language model. ARC‑CT achieves a 0.86 mask‑free macro AUC across 18 abnormalities using a compact 3D ResNet‑18 backbone. Over‑ all, ARC‑CT outperforms both comparable efficient baselines and sev‑ eral larger transformer models. Our code and weights are available at https://github.com/arc‑ct/arc‑ct.

Authors:Vasilis Dedousis, Lubnaa Abdur Rahman, Lorenzo Brigatο, Ethan Dack, Andreas Christe, Christoph Frank, Manuela Funke-Chambour, Justus Roos, Adrian Huber, Lukas Ebner, Stavroula Mougiakakou
Title: Prompt-Guided Interactive Segmentation of Interstitial Lung Disease in Thoracic CT
Abstract:
Accurate segmentation of interstitial lung disease (ILD) patterns is essential for quantitative disease assessment and longitudinal monitoring. However, existing approaches remain limited by relying on dense annotations and producing static predictions that cannot be refined, motivating interactive approaches. While promptable models show promise in interactive segmentation, their adaptation to ILDs remains largely unexplored. To address this gap, we investigate prompt‑guided foundation models for ILD refinement and present, to the best of our knowledge, the first adaptation of MedSAM2 for interactive 3D ILD segmentation on thoracic CT. We investigate three fine‑tuning strategies and multiple clinically motivated prompts: bounding‑boxes (BBox), point, lasso, and scribble. On a dataset spanning seven ILD patterns and healthy lung tissue, full model fine‑tuning performed best, improving the average Dice score by 4.7 percentage points over MedSAM2.While BBox prompts achieve the strongest performance, non‑native MedSAM2 interactions such as lasso and scribble prompts also prove effective. Finally, we present and evaluate a proof‑of‑concept end‑to‑end workflow in which MedSAM2 is initialized from an automatic segmentation prior and subsequently refined using radiologist prompts. Model weights and plug‑ins made available at: https://github.com/AIHNlab/ILD‑SemiSegTool.

Authors:Shihang Yang, Sanwoo Lee, Ningning Zhao, Yunfang Wu
Title: A Unified Framework to Elicit Structured Feedback for Interpretable Multi-Trait Essay Scoring
Abstract:
Multi‑trait Automated Essay Scoring (AES) requires rubric‑grounded reasoning across interdependent traits, rather than isolated score prediction. Existing feedback‑enhanced methods often decouple feedback from scoring or assess traits independently, weakening score‑‑feedback consistency and rubric alignment. We propose HiFTS, a unified autoregressive framework that generates hierarchical CoT feedback before predicting trait‑level and holistic scores. HiFTS distills rubric‑grounded hierarchical CoT feedback from a teacher LLM and trains student models to jointly generate feedback and scores. HiFTS further applies Group Relative Policy Optimization with a composite reward balancing score agreement, calibration, feedback quality, and structural validity. At inference, a lightweight global prior provides holistic guidance to reduce drift during long‑form reasoning. We also introduce CFMS‑34, a Chinese multi‑trait AES dataset with 951 essays annotated with holistic scores and 34 rubric‑based traits. Experiments on CFMS‑34 and ASAP++ show that HiFTS achieves strong holistic and trait‑level scoring while producing coherent, rubric‑aligned feedback.

Authors:Federico Spurio, Olga Zatsarynna, Lars Doorenbos, Emad Bahrami, Gianpiero Francesca, Juergen Gall
Title: Post-Training VLMs for Video Mistake Detection
Abstract:
Human mistakes are inevitable when following instructions, yet they can lead to severe consequences. As such, there has been an increased interest in developing methods for detecting mistakes in videos, with current methods mostly focusing on closed‑set protocols. While successful in controlled settings, the closed‑set assumption limits their wider applicability, as any changes to the task require collecting new data and re‑training models. Instead, we argue that mistake detection methods should learn the general concept of a mistake, rather than overfitting to step‑specific details. To reflect this, we introduce the Mistake Detection Video Question Answering (MD‑VQA) protocol and accompanying benchmark. MD‑VQA tests whether methods can discern if a step was executed correctly with respect to its description, for both seen and unseen actions. To address this important challenge, we propose the first video‑language‑model post‑training technique for mistake detection. Our method uses a tailored reward function to encourage the model to identify discrepancies between an instruction and the corresponding video. Extensive evaluations demonstrate that this approach outperforms zero‑shot, supervised fine‑tuning, and post‑training baselines. Notably, our method generalizes especially well to unseen procedures, for instance, with an improvement of up to 11.6% over the best‑performing baseline on EP‑VQA, paving the way toward general mistake detection. We release our code and benchmark at https://github.com/FedeSpu/mstk.

Authors:Victor Besnier, Anh-Quan Cao, Elias Ramzi, Spyros Gidaris, Tuan-Hung Vu, Andrei Bursuc, Eloi Zablocki, Matthieu Cord
Title: How Far Can 5,500 Hours of Driving Take You? A Scaling Law Analysis of Video Diffusion Models
Abstract:
Video generation for autonomous driving cannot follow the web‑scale route: driving data is expensive to collect, bound by privacy requirements, and cannot be scraped at will, so models must make the most of a fixed corpus. We present a systematic scaling‑law study of video diffusion models trained from scratch on driving data: a family of models from 1M to 9B parameters, trained at different exposures on up to 5,500 hours of driving. Validation loss follows consistent power laws in both model size and training exposure, answering the questions that shape a training budget: whether compute is better spent on longer training or on a larger model, and whether more data is needed. Loss improves much faster with training exposure than with model size, making longer training the most effective way to improve a fixed model under limited compute. However, larger models continue to achieve lower asymptotic loss, so compute‑optimal scaling still favors increasing model size when sufficient compute and data are available. Guided by these laws, we train a 9B‑parameter model, to our knowledge the largest video diffusion model trained from scratch on driving data: it sets a new open‑source state of the art for driving video generation, as measured on nuScenes. Our code and pretrained models are available at https://github.com/valeoai/VATIX. NATIX is separately releasing the underlying driving data in stages.

Authors:Jaewon Jung, Haizhong Zheng, Hongsun Jang, Jaeyong Song, Beidi Chen, Jinho Lee
Title: CamoDocs: A Poisoning Attack Against Retrieval-Augmented Language Models Using Camouflaged Documents
Abstract:
Retrieval‑augmented generation (RAG) augments LLMs with external documents, but public or user‑editable sources expose RAG systems to data poisoning: attackers can inject malicious documents to steer outputs toward targeted answers. Existing poisoning attacks often rely on query inclusion, inserting the target query into poisoned documents to improve retrieval; however, this creates lexical and embedding‑space artifacts that make them easy to filter. We propose CamoDocs, a poisoning attack that avoids direct query inclusion by camouflaging adversarial documents among benign content. CamoDocs chunks synthesized benign and adversarial drafts, replaces selected tokens in benign chunks with dispersion tokens that spread poisoned‑document embeddings, and applies coherence filtering to limit readability degradation. Across seven RAG defenses, three open‑weight LLMs, and three benchmarks, CamoDocs achieves strong average ASR while avoiding query‑overlap artifacts exploited by simple query detection. It also remains effective against proprietary models, achieving average ASRs of 61.80% on GPT‑5.4‑mini and 55.09% on Claude‑Haiku‑4.5. Finally, we show that erasure‑heavy clustering defenses such as TrustRAG can reduce ASR, but only with substantial utility drops on retrieval‑dependent benchmarks such as NeoQA. Code is available at https://github.com/jaewonalive/CamoDocs.

Authors:Anton Nuzhdin, Marcel Worring, Ivona Najdenkoska
Title: FUSED: Forensic-Semantic Mixture-of-Experts for AI Inpainting Detection and Localization
Abstract:
Diffusion‑based inpainting models modify only a localized part of an image, while many AI‑image detectors rely on global artifacts and do not localize. These artifacts vary across generators, limiting detector transfer under distribution shifts. Recent work shows that restoring the authentic pixels outside the inpainted region removes these cues and can degrade pretrained detectors. To address this, we present FUSED, a unified framework for the joint detection and localization of AI‑generated inpainting. FUSED combines low‑level forensic cues with high‑level semantic features using a sparsely‑gated Mixture‑of‑Experts architecture, enabling the model to adaptively prioritize the most relevant signal for each token. For each input, FUSED predicts both an image‑level manipulation score and a pixel‑level mask of the inpainted area. On the OpenSDID cross‑generator benchmark, FUSED achieves the best average detection and localization, with the largest gains on unseen generators. The same model transfers directly to the held‑out AutoSplice and CocoGlide benchmarks, more than doubling localization performance. Evaluating each held‑out benchmark with and without the global generator artifact further shows that all evaluated methods, ours included, partly read the artifact as evidence of manipulation, and FUSED remains the strongest under both conditions. Code and pretrained models are available at https://github.com/AntonNuzhdin/FUSED.

Authors:Yi Wang, Haopeng Zhang, Chengxiang Huang, Rui Dai, Kaikui Liu, Piotr Koniusz, Xiangxiang Chu
Title: LoopArena: Benchmarking Models as Runtime Controllers for Loop Engineering
Abstract:
Loop Engineering is emerging as a practice for organizing development work around coding agents. Instead of writing each prompt by hand, practitioners design loops that monitor progress, assign work, run checks, and decide what the agent should do next. Even with a capable coding agent, a loop may trust a stale progress note, skip needed verification, spend its budget in the wrong direction, or stop before the task is safe to submit. Yet the final outcome of one end‑to‑end run cannot tell whether success or failure reflects the loop's guidance or the coding agent's ability to carry out the task. We introduce LoopArena, a benchmark for evaluating how well one model can guide a separate coding agent through a long‑running task. The model under evaluation is the Controller: after each coding round, it receives a structured summary of the run and instructs a separate, fixed coding agent, the Worker, on what to do or verify next, or decides whether to stop. LoopArena evaluates this ability in three complementary settings that differ in execution scope and cost. Type I scores next‑step Loop Contract selection through execution‑validated questions without running the Worker at evaluation time. Type II executes repeated control over a selected slice of a full task, while Type III evaluates the paired full task from its original state. On full tasks, the best observed Strict Success Rate is 24.69%, leaving substantial room for improvement in long‑horizon loop control. Across Controllers, the paired reduction in estimated inference cost averages 64.4%, and Type II produces a similar ordering under the main Core criterion (Spearman's \(ρ=0.9747\)). We release the benchmark data and evaluation code at https://github.com/AMAP‑ML/LoopArena .

Authors:Ian Hsieh, Soumya Snigdha Kundu, Tom Vercauteren, Reuben Dorent
Title: SinkSLOT: Sinkhorn via Sparse Lifted Optimal Transport
Abstract:
Entropic optimal transport (EOT) has been shown to offer a computationally tractable approximation to exact optimal transport. However, the standard Sinkhorn‑Knopp algorithm has two main limitations. First, given discrete measures with N points, each iteration requires O(N^2) operations, which restricts its use on large‑scale datasets (e.g. N\geq10^4). Second, it uses the independent coupling as a reference measure for regularisation. This assigns mass to high‑cost transport edges at moderate regularisation strengths. We propose SinkSLOT, which addresses both limitations by putting forth the expected sliced lifted transport plan as a natural way to sparsify the Gibbs kernel with a non‑independent prior coupling. We prove that: 1) SinkSLOT converges; 2) with L slices, each resulting sparse Sinkhorn iteration costs O(LN); and 3) the resulting objective is a divergence requiring no debiasing. Experiments on synthetic benchmarks show that SinkSLOT delivers substantial speedups over state‑of‑the‑art dense and sparse EOT methods. We also demonstrate the applicability of the proposed divergence in a gradient flow experiment. The code is publicly available at https://github.com/cai4cai/SinkSLOT.

Authors:Keito Sasagawa, Shuhei Kurita, Daisuke Kawahara
Title: Synth-JDoc: Synthesizing a Japanese Document Image Dataset for OCR with Diverse Layouts and Embedded Images
Abstract:
The ability of Large Vision Language Models (LVLMs) to read text within document images is crucial, as it enables various applications such as Document Visual Question Answering. To enhance the text‑reading capabilities of LVLMs, high‑quality OCR datasets are essential. This need is particularly critical for Japanese documents, which often feature vertically written text alongside horizontally written text. Current LVLMs demonstrate considerably lower performance on vertically written Japanese text than on horizontally written text, necessitating specialized OCR datasets to bridge this gap. However, manually constructing OCR datasets is expensive and difficult to scale. Alternatively, constructing datasets by extracting text from existing document images using OCR models introduces challenges, such as text recognition errors and the prerequisite of sourcing document images. To address these issues, we construct an OCR dataset by synthesizing document images directly from text. Leveraging HTML and CSS, we generate multi‑column documents that incorporate both vertical and horizontal writing styles. Furthermore, to ensure the visual realism of the documents, we embed images generated by text‑to‑image models within the layout. Additionally, to foster model robustness, we apply noise and degradation filters to the synthesized document images. In our experiments, we compared the performance of models fine‑tuned on our synthetic dataset against baselines fine‑tuned on synthetic datasets from prior work and those generated by a high‑performance text‑to‑image model. Evaluation results demonstrate that our synthetic dataset is the most effective approach for improving LVLM performance on reading vertically written Japanese text. Our dataset and code are publicly available (https://github.com/llm‑jp/synth‑jdoc).

Authors:Jiazhao Liang, Hao Huang, Shuaihang Yuan, Congcong Wen, Geeta Chandra Raju Bethala, Giles Hamilton-Fletcher, Yu Hao, John-Ross Rizzo, Mengyu Wang, Anthony Tzes, Yi Fang
Title: Focus Where It Counts: A Salience-Driven Vision-Language Model for Low Vision Assistance
Abstract:
Vision‑language models (VLMs) are rapidly progressing and offer promising capabilities for assistive technologies supporting persons with blindness or low vision. However, existing VLMs are primarily designed for general‑purpose captioning and do not explicitly model human perceptual priorities, thereby limiting their ability to emphasize the most relevant information in a scene. To address this gap, we propose a salience‑driven captioning framework that prioritizes scene elements according to their importance for human‑centered assistance. We curate three salience‑aware datasets, namely, Salience COCO, Salience Flickr, and Salience VizWiz, with object‑level salience annotations designed to reflect the visual information most relevant to low vision users across different environments. Building on these datasets, we introduce Salience‑LLaVA, a salience‑aware VLM that incorporates salience cues to generate captions in which important elements are mentioned in the order of importance. Our work makes four main contributions. We build salience‑aware datasets verified by low vision participants, propose Salience‑LLaVA to describe objects in the order of importance, introduce SCMI to evaluate ordering accuracy, and deploy the system on assistive glasses to demonstrate real‑world practicality. Code and datasets are available at: https://github.com/topo‑focus/Topofocus

Authors:Abhishek Verma, Anila Krishna, Abhishek Gajanan Bankar, Juan Miguel Lopez Alcaraz
Title: Explainable Diabetic Retinopathy Classification Using Vision Foundation Models
Abstract:
Diabetic retinopathy (DR) is a major cause of preventable blindness, creating a need for accurate and trustworthy automated screening. This study investigates an explainable DR classification framework using vision foundation models and multiple transfer learning strategies. Three backbones, DINOv2, CLIP, and Vision Transformer (ViT), were evaluated using full fine‑tuning, linear probing, and Low‑Rank Adaptation (LoRA). Models were trained and internally evaluated on the ODIR dataset and externally evaluated on APTOS to assess generalization. DINOv2‑LoRA achieved the highest internal AUROC of 0.758, while DINOv2 full fine‑tuning and ViT full fine‑tuning achieved the highest external AUROC of 0.920. Calibration was further assessed using reliability analysis after isotonic regression. For explainability, Grad‑CAM and HiResCAM were evaluated against expert‑annotated lesion masks from the IDRiD dataset using Dice, Intersection over Union (IoU), and Pointing Game metrics. The results demonstrate that foundation models, particularly DINOv2, can provide strong predictive performance, while LoRA offers a parameter‑efficient alternative to full fine‑tuning. Quantitative evaluation of explanation maps further supports the assessment of whether model attention corresponds to clinically relevant retinal lesions.

Authors:Francisco J. Moreno Velo, Almudena García Jurado-Centurión
Title: URIUM: A Programming Language for a Practical Open Course on Compiler Design
Abstract:
This paper presents the definition of a simple programming language used as the basis for developing a practical compiler design course. The course explains step by step how to build a compiler, from the initial analysis stages to code generation. The developed compiler generates code for various processors (MIPS, Intel, and RISC‑V) and operating systems (MS‑Windows and Linux). The course can be adapted to different levels of difficulty and can be used as a starting point for explaining more advanced topics.

Authors:Anja Witte, Maximilian Lennartz, Jan Baumbach, Guido Sauter, Stefan Bonn, Patrick Fuhlert, Marina Zimmermann
Title: EXPOSE: Explainable and Domain-Robust Embeddings from Pathology Vision Foundation Models using Sparse Autoencoders
Abstract:
Vision Foundation Models (VFMs) are widely used in computational pathology but remain sensitive to domain shifts arising from variations in staining, tissue preparation, and scanner hardware. A key limitation is that VFM embeddings entangle biological with domain‑specific information, hindering cross‑domain generalization. We propose Explainable Probing of Cross‑Domain Sparse Embeddings (EXPOSE), a framework that uses Sparse Autoencoders (SAEs) as an explainable bottleneck to identify and suppress domain‑specific components in VFM embeddings. We train a sparse representation of VFM features, use a linear classifier to identify domain‑specific latent dimensions, and mask these features prior to downstream relapse prediction without retraining the backbone model. Experiments on a large prostate cancer dataset with multiple acquisition domains show that SAE features capture both domain‑ and task‑specific information, which are partially disentangled in the latent space. Removing domain‑specific features improves cross‑domain performance and increases embedding robustness as measured by the Domain Robustness Index (DoRI). Code is available at https://github.com/imsb‑uke/expose .

Authors:Yongqi Mao, Zijia Dai, Zhishuo Liu, Wei Xu, Kaiwei Wang, Guotao Meng
Title: Manifold4D: Denoising on Point Cloud Rendered Manifolds for Video Re-shooting
Abstract:
Video re‑shooting re‑renders a monocular video of a dynamic scene along a user‑specified camera trajectory, and the dominant recipe supplies the target geometry explicitly: per‑frame depth lifts the source video into a 4D point cloud, which is rasterized along the trajectory into a point cloud render. Because the render and the source video are both handed to the network as visual conditions, they compete at every denoising step, leaving the model with a trust dilemma ‑‑‑ how much of the render to believe ‑‑‑ which can degrade trajectory control or visual quality on data outside the training distribution. We argue that a render already pixel‑aligned with the target view does not need to be supplied as an explicit conditioning stream at all. We propose MANIFOLD4D, which injects the render directly into the initial noise of flow matching, so that generation no longer departs from standard Gaussian noise but from a new noise manifold carrying geometric information, leaving the source video as the only visual condition. The render is thus used exactly once, and the network is never asked to learn how to read it; in subsequent denoising steps the model can focus on the source video. On our DAVIS‑Traj benchmark and on the Vista4D evaluation set, MANIFOLD4D attains the best camera‑control accuracy on every metric, lowering rotation error by 25% and 27% and translation error by up to 32% over the strongest baseline, while matching it in video fidelity and leading on real‑world novel‑view photometric quality. In a user study, our method achieves clear advantages in trajectory following and dynamic consistency. The gap widens as the yaw amplitude grows past the training range, and the model still recovers correct dynamic motion from the source video when the render is deliberately corrupted, confirming that the geometric prior guides generation without overriding it.

Authors:Christos Koutsiaris
Title: Nested Byte-Level Vocabularies Are Cheap to Deploy and Expensive to Share: A Pre-Registered Negative Result
Abstract:
A byte‑level BPE tokenizer is an ordered list of merge rules, so applying only a prefix yields a vocabulary whose token identifiers are the first rows of the full vocabulary. This prefix nesting allows one language model to operate at several vocabulary sizes, use a control token to indicate the active size, and be deployed at any trained size by slicing its embedding and output head. We pre‑registered five claims, including margins, seeds, contrasts, and a stop rule, and trained 30 models with 3.1M‑ and 10.6M‑parameter bodies on 200M tokens each. Slicing is numerically exact: across 76 checks, a sliced model reproduces the restricted full model's logits bit for bit and removes 66% of deployed weights without changing latency. However, the shared model trails a fixed‑cap specialist by 3.64% bits per byte at 32k against a 1% margin, and by 2.96% at 8k against a 2% margin. A 2x2 ablation separating the control token from output restriction finds that the token changes performance by +0.07% to +0.13%, with all intervals crossing zero, while output restriction costs +0.47% to +1.19%; the factors are substitutes rather than complements. Multi‑cap training nevertheless improves robustness: under typographical noise, the same checkpoint degrades 12.5‑‑15.4 points less in its fine mode and outperforms each fixed‑cap specialist at that specialist's vocabulary size. A control with neither cap token nor output restriction is equally robust, attributing this benefit to multi‑granularity training rather than conditioning. The per‑cap penalty tracks each cap's share of training rows, yielding a falsifiable prediction for future work.

Authors:Weiwei Xiang, Shun Peng, Guangyi Xiao, Hao Chen, Lei Yang
Title: Dual-Stream Semantic Guidance with Prototype Anchor Calibration for Source-Fully-Free Adaptation of Vision-Language Models
Abstract:
Source‑Fully‑Free Domain Adaptation (SFF‑DA) has emerged as a strategic paradigm to adapt Vision‑Language Models (VLMs) without any access to source data or task‑specific source models. However, we identify a critical Dual Semantic Drift that hinders this process: static drift arising from the rigidity of fixed class embeddings, and dynamic drift stemming from the divergence of generated captions, causing severe semantic misalignment that intensifies the stability‑plasticity dilemma. To address this, we propose DSSG (Dual‑Stream Semantic Guidance), an end‑to‑end framework that reconciles fine‑grained plasticity with global stability. Our core contribution is the Dual Semantic Guidance (DSG) module, which integrates a caption stream for domain‑specific knowledge with a class‑anchor stream to anchor global categorical consistency. Furthermore, a Dynamic Cross‑Modal Knowledge Distillation (CMKD) module is introduced to leverage the evolving teacher distribution for calibrating teacher‑student consistency. Building upon DSSG, we further introduce Prototype Anchor Calibration (PAC), yielding DSSG‑PAC, which periodically calibrates prototype anchors and caches them until the next calibration. This design reduces redundant text‑side computation while preserving the adaptability of class guidance to the evolving text space. We further establish SFF‑DA risk bounds that relate student risk to semantic‑teacher quality and teacher‑‑student discrepancy. Extensive experiments demonstrate that DSSG consistently outperforms current state‑of‑the‑art methods across multiple benchmarks, while DSSG‑PAC largely preserves its adaptation performance with 18.9% lower total adaptation time. The code is available at https://github.com/mrmenand/DSSG.

Authors:Simone Tolomei, Mayank Mittal, Franco Angelini, Manolo Garabini, Paolo Salaris, Marco Hutter
Title: Contact-Guided Exploration for Non-Prehensile Locomanipulation with Multi-Critic RL
Abstract:
Non‑prehensile manipulation offers versatile skills for moving and rearranging heavy or bulky objects, particularly when combined with a mobile manipulation platform. However, both model‑based and model‑free approaches struggle with the complex hybrid dynamics and the sparsity of the contact in these tasks. To address these challenges, we propose a contact‑guided exploration strategy implemented within a Multi‑Critic Reinforcement Learning (RL) framework. A dedicated exploration critic is trained with a dense contact‑seeking reward that guides the end‑effector toward meaningful contact points; its influence is progressively decayed to recover a task‑optimal policy. We obtain candidate interaction points from a general‑purpose grasping algorithm, enabling the exploration mechanism to generalise across various object geometries. We evaluate the approach on multiple tasks, including box pushing, chair transportation, and a dishwasher opening task. Finally, we validate the chair transportation policy through extensive experiments on a quadrupedal mobile manipulator, demonstrating deployable non‑prehensile manipulation in the real world.

Authors:Naren Akash, Arihanth Tadanki, Jayanthi Sivaswamy
Title: CheXtriev: Anatomy-Centered Representation for Case-Based Retrieval of Chest Radiographs
Abstract:
We present CheXtriev, a graph‑based, anatomy‑aware framework for chest radiograph retrieval. Unlike prior methods focussed on global features, our method leverages graph transformers to extract informative features from specific anatomical regions. Furthermore, it captures spatial context and the interplay between anatomical location and findings. This contextualization, grounded in evidence‑based anatomy, results in a richer anatomy‑aware representation and leads to more accurate, effective and efficient retrieval, particularly for less prevalent findings. CheXtriv outperforms state‑of‑the‑art global and local approaches by 18% to 26% in retrieval accuracy and 11% to 23% in ranking quality. The code is available at https://github.com/cvit‑mip/chextriev.

Authors:Tianfu Wang, Zhezheng Hao, Xilin Xia, Lixin Liu, Mengkang Hu, Hongzhang Liu, Xi Chen, Ziyan Liu, Xiankun Lin, Weijia Zhang, Nicholas Jing Yuan, Hui Xiong
Title: Agentic Artifact Creation: Systems, Evaluation, Principles, and Opportunities
Abstract:
Generative models can turn natural‑language prompts into images, text, code, and other content, lowering the cost of producing drafts and components. Their practical impact increasingly depends on whether those pieces can become complete, dependable deliverables. This survey examines agentic artifact creation, which we define as stateful construction in which an AI system materially constructs or revises a deliverable and intermediate observations redirect later work. Functionally, the process links an operational representation of the artifact, a construction policy, and runtime verification whose feedback can redirect later actions. We reviewed 259 works available through August 20, 2026: 230 systems meeting this definition and 29 benchmarks of agentic artifact construction. We compare six artifact families, then analyze application settings and evaluation practice as separate dimensions. Across families, construction challenges reflect not only modality but also how tightly decisions are coupled and whether failures become visible while they remain repairable. Decomposition can reduce local complexity while increasing coordination and reassembly costs. Learned judges may add little independent evidence when they share the generator's preferences or blind spots. We formulate principles for keeping commitments and responsibility explicit, turning feedback into targeted repair, and revalidating affected state after change. We also identify opportunities for sustaining coherent, accountable control as artifacts, creator intent, and construction systems evolve. A curated paper list is available at https://github.com/GeminiLight/awesome‑agentic‑artifact‑creation.

Authors:Tolgahan Bardakci, Serge Demeyer
Title: RESTCov: A Tool for Structural Coverage Analysis of REST APIs
Abstract:
REST APIs are widely used in modern software systems, but developers and testers often lack visibility into which parts of an API specification are exercised by a test suite. Traditional coverage analysis usually relies on source‑code instrumentation, which is impractical for REST APIs that are distributed, externally maintained, and hence accessible only through black‑box execution. This paper presents RESTCov, a lightweight tool that computes structural REST API coverage from an OpenAPI specification and observed HTTP request/response logs, reporting coverage across paths, operations, parameters, media types, status codes, and status classes. RESTCov produces both machine‑readable results and a human‑readable HTML report, helping users inspect coverage gaps, diagnose specification‑log mismatches, and evaluate REST API test suites without requiring access to the implementation. Screencast: https://youtu.be/mNz2P43OyUc Repository: https://github.com/2tolgahan2/RESTCov

Authors:Teejuta Sriwaranon, Borworntat Dendumrongkul, Tanapat Chamted, Pizzanu Kanongchaiyos
Title: What Will This Copper Look Like Later? Forecasting Surface Appearance and Rendering It as a PBR Material
Abstract:
Digital design requires predicting how a metal surface will look later in its oxidation; this paper presents such a pipeline for copper. Given a fixed‑camera observation, the system forecasts appearance 10 accelerated units ahead and converts it into the albedo, normal, roughness and metallic maps a renderer consumes. Forecasting is evaluated as an authoring tool would use it, on a copper specimen the system has not observed: an entire recording is held out, so training and checkpoint selection use one specimen and the test set is the whole of a second, recorded on a different day and condition. Under this protocol a learned spatio‑temporal model with a monotone oxidation state, the most accurate forecaster within a single recording, is less accurate than copying the last observed frame on an unseen specimen, in both directions, as are three further trained architectures. The only forecaster that transfers is a closed‑form global color extrapolation with no trained parameters, improving on copy‑last‑frame by 13.4% and 50.6%, with a margin that increases with horizon to +16.7% and +55.5% at t+10. Two controls qualify this: correcting every frame for the photometric drift measured on a non‑oxidizing reference region leaves both margins intact, ruling out uncontrolled exposure as their source, and a moving‑block bootstrap over the 6 independent windows each recording contains separates the larger margin from zero but leaves the smaller one not individually significant. The mechanism is measured: a learned susceptibility map encodes where corrosion begins on the training specimen and misleads on a new one, whereas the global color trajectory is what specimens share. The pipeline therefore deploys the closed‑form forecaster for unseen specimens and the learned model only for continuing one already observed. Code, splits, protocol and leakage audit are released.

Authors:Xindi Yang, Yicheng Wu, Cheng Zhang, Jianfei Cai, Tien-Tsin Wong
Title: Attribute Token Arithmetic: Disentangled and Continuous Semantic Control for Visual Autoregressive Models
Abstract:
Autoregressive text‑to‑image generation has recently achieved remarkable progress, offering high‑fidelity synthesis via a unified generative framework. However, fine‑grained semantic control remains challenging due to the attribute entanglement and the misalignment between textual and fine‑grained visual representations. In this paper, we introduce Attribute Token Arithmetic (ATA), a method that enables disentangled and continuous attribute control in visual autoregressive modelling. Inspired by the vector arithmetic property observed in word embeddings, ATA identifies semantic directions corresponding to visual attributes (e.g., aging, fatness, emotion) directly within the pretrained autoregressive latent space. These directions are learned from a single reference image, without model retraining or large‑scale supervision. During generation, attributes can be continuously adjusted and compositionally combined through simple arithmetic operations with other attribute tokens. Extensive experiments demonstrate that ATA achieves identity‑preserving, fine‑grained, and multi‑attribute adjustment, outperforming existing autoregressive editing baselines in controllability, generality, and computational efficiency. Our code will be available at https://github.com/Madaoer/ATA.

Authors:Daniel Seibel, Kaveh Haghighi Mood, Jayesh Badwaik, Prateek Chawla, Stepan Nassyr, Andreas Herten
Title: Performance Evaluation of Fast Fourier Transforms on Emerging RISC-V Hardware with Vector Extension Support
Abstract:
This manuscript presents a performance evaluation of Fast Fourier Transform (FFT) implementations on emerging processors supporting the RISC‑V Vector Extension (RVV 1.0). By introducing juFFTe, a light‑weight high‑performance library for discrete Fourier transforms, it is demonstrated how effective vectorization of performance‑critical FFT kernels can be achieved on RVV‑enabled hardware. Comprehensive benchmarks on three RVV 1.0‑ready processors, the SiFive X280, the X100 core of the SpacemiT K3 and the C920v2 core of the Sophon SG2044, reveal substantial performance improvements of juFFTe (https://github.com/FZJ‑JSC/juFFTe) over the widely used FFTW3 library. Although RVV‑enabled platforms show promising results at this stage of development, a comparison with AMD's Zen 5 architecture indicates that RISC‑V needs further maturing to reach the performance of established micro‑architectures.

Authors:Xinda Yu, Kunxin Zheng, Chunan Yu, Qingbo Song, Hao Xiao, Ying Zang, Jie Liu
Title: CF-YOLO: Context-Aware Feature Refinement for Camouflaged Industrial Micro-Defect Detection
Abstract:
Automated detection of surface micro‑defects on industrial components, such as copper tubes, is critically important for quality assurance but remains challenging due to the minute scale of anomalies and their visual camouflage against complex backgrounds. These factors lead to weak feature representations and high rates of false positives and missed detections. To address these issues, we propose a novel real‑time detection framework designed for efficient context perception and feature refinement. Our method integrates a Context‑Perception Aggregation Module (CPAM), which synergises large‑kernel perception for macro‑texture context and small‑kernel aggregation for sharp boundary delineation, effectively breaking the background camouflage. Furthermore, a Feature Additive Refinement Module (FARM) employs a linear‑complexity additive token mixer to globally verify and refine the representation of fine‑grained anomalies, suppressing noise‑induced errors. To support research in this domain, we introduce the Copper Tube Defect Dataset (CTDD), a manually annotated benchmark containing 1,847 images and 4,898 boundingbox defect instances from copper‑tube inspection scenarios. Extensive experiments demonstrate that our detector achieves strong and consistent performance on CTDD, outperforming representative baseline detectors, including YOLOv11, by 2.2% in mAP@50 and 3.9% in Precision while maintaining real‑time inference speed. This work provides a robust and efficient solution for high‑precision industrial inspection, bridging the gap between contextual understanding and detailed feature analysis. Our code and model are available at: https://github.com/Yu‑Xinda/CFYOLO‑Context‑Aware‑Feature‑Refinement‑for‑Camouflaged‑Industrial‑Micro‑Defect‑Detection

Authors:Ruijie Su, Lingxiao Yang, Xiaohua Xie, Jianhuang Lai
Title: VersaGauss: A Versatile Framework for Generating Multiphase Dynamics with 3D Gaussians
Abstract:
Recent progress has been made in 3D Gaussian representation for reconstruction, generation, and physical simulation. However, current approaches mainly concentrate on physics‑based dynamic generation of solid objects and only handle single‑phase collision interactions. We introduce VersaGauss, a unified framework for generation, simulation, and rendering that supports versatile physics‑based dynamic generation, particularly for multiphase interactions. Our system takes a few images as input and produces a realistic, physics‑driven 3D dynamic scene with multiple objects. To optimize the Gaussian kernel distribution, we develop a particle pruning algorithm. We also propose the Coupled Multiphase Point Method (CMPM) to effectively model and generate multiphase interactions. Additionally, harmonic interpolation within CMPM and a Gaussian evolution strategy are introduced to achieve realistic fluid rendering. Extensive experiments demonstrate that our framework can simulate interactions among various materials such as fluid, rubber, sand, snow, and others. Code is available at https://github.com/Elowen‑surj/VersaGauss.

Authors:Guanglin Jin, Hongshan Yu, Javier Civera, Zhaoxin Li
Title: ZipMVS: Multi-View Stereo with Compressed Cost Volumes
Abstract:
Multi‑view stereo (MVS) methods typically deliver highly accurate 3D reconstructions from multiple registered RGB images, thanks to the highly informative, geometric constraints between them. However, their substantial memory requirements remain a major obstacle for deployment in domains such as aerospace and autonomous systems, where resource efficiency is critical. In this work, we introduce ZipMVS, an MVS method specifically designed for efficient high‑quality reconstruction. We propose a novel depth‑hypothesis strategy that enables substantial compression of the cost volume, hence greatly reducing GPU memory consumption while preserving reconstruction accuracy. Experiments on the DTU and Tanks and Temples datasets show that ZipMVS achieves competitive reconstruction quality compared with other efficiency‑oriented MVS methods, while achieving a competitive balance between reconstruction quality and GPU memory usage. The code is available at https://github.com/JihnGlyn/ZipMVS

Authors:Nima Anari
Title: Beyond the Bethe Approximation of the Permanent
Abstract:
The canonical Bethe approximation gives a deterministic approximation to the permanent of every nonnegative matrix within a factor of (\sqrt2)^n. We improve the base of this exponential factor: for some absolute constant c<\sqrt2, there is a deterministic polynomial‑time c^n‑approximation for the permanent of every nonnegative matrix. This shows that the canonical Bethe guarantee is not a barrier for deterministic approximation of the permanent. The proof augments the Bethe lower bound with a new certificate tailored to matrices on which that lower bound loses nearly the full factor. The author supplied the high‑level plan of attack, and the proof was developed in an interaction with ChatGPT 5.6 Sol Pro. The author subsequently verified the results. Codex assisted with proof checking, manuscript assembly, and typesetting.

Authors:Animesh Shaw
Title: Compared to What? A Human-Anchored Security Benchmark for LLM-Generated Infrastructure-as-Code
Abstract:
Large language models are increasingly used to author Infrastructure‑as‑Code (IaC), where a single insecure default can be deployed directly into production. Prior evaluations report raw vulnerability counts for model‑generated IaC, but without a human baseline they cannot determine whether models are actually worse than engineers. We introduce GenIaC‑SecBench, a benchmark of 100 deployment scenarios stratified by architectural complexity, evaluated across 12 model configurations from four vendors, producing 1,196 IaC artifacts scanned by three independent policy engines (Checkov, Trivy, KICS). Critically, we also scan 634 human‑authored IaC templates with the same toolchain, providing the first size‑matched human security baseline. Vulnerability density is strongly inverse to artifact size (Spearman ρ= ‑0.55, p < 10^‑77), meaning unmatched comparisons measure size rather than security. When matched on declared‑resource count, all model configurations fall within 3.21x‑‑3.87x the human vulnerability density, with the gap widening for simpler tasks (4.9x at one resource, 1.4x at twenty or more). We decompose reasoning into standard generation, prompt‑engineered chain‑of‑thought, and vendor extended‑thinking APIs. Vendor extended thinking significantly outperforms prompted chain‑of‑thought (‑12.0%, p = 0.0013), while prompted chain‑of‑thought is indistinguishable from standard generation (‑1.3%, n.s.). Token instrumentation shows extended thinking uses under 1% of the output budget, explaining its bounded effect. Two negative results also emerge: deployability does not correlate with vulnerability (r = 0.158, p = 0.625), and classical complete‑case Friedman testing is infeasible for realistic benchmark designs, motivating the Skillings‑Mack statistic. All code, data, and regeneration scripts are released.

Authors:Jieyu Yuan, Yuanlin Zhang, Jihong Li, Chunle Guo, Huimin Lu, Chongyi Li
Title: 3D-USE: From Image-Level to Scene-Level Underwater Enhancement
Abstract:
Underwater 3D reconstruction faithfully reproduces the color shifts and visibility loss of captured views, while physical inversion may leave estimation errors in the recovered scene appearance. We formulate Underwater Scene‑level Enhancement (USE) as learning a persistent, visibility‑enhanced 3D scene representation from degraded multi‑view underwater observations, enabling consistent enhanced rendering. Realizing USE requires both a reliable scene representation for enhancement and a consistent enhancement target without paired enhanced 3D data. Therefore, we present 3D‑USE, a two‑stage framework. First, the Medium Radial Basis Anchor Representation (MediumRBF) establishes a medium‑aware Gaussian scene by representing water effects with shared radial‑basis anchors and explicitly decomposing object and medium contributions. Based on this fixed scene representation, Appearance Transition Consensus (ATC) transfers paired 2D underwater image enhancement (UIE) knowledge into scene‑global and Gaussian‑local targets, avoiding direct supervision from inconsistent enhanced views. An Underwater Bilateral Appearance Field (U‑BAF) then realizes these targets in Gaussian radiance and medium appearance. The scene directly renders enhanced novel views without a 2D UIE model at inference. Experiments on real underwater scenes show improved visibility and cross‑view consistency while preserving reconstruction quality.

Authors:Peiming Li, Yifan Wang, Zhiyuan Hu, Shiyu Li, Zheng Wei, Yang Tang
Title: Beyond Global Scalars: Synergizing Token-Level Statistics and Deep Semantics for Adversarial AIGC Text Detection
Abstract:
The rapid evolution of large language models necessitates robust machine‑generated text detection. Existing paradigms typically follow two isolated tracks. Training‑free methods rely on global statistical scalars such as perplexity, while training‑based methods utilize semantic hidden states. Both approaches exhibit fundamental vulnerabilities in adversarial scenarios. Global scalars act as lossy compressions that obscure local probabilistic burstiness in interleaved texts, whereas pure semantic models overfit to specific fingerprints and remain susceptible to spoofing. To expose these flaws, we introduce MOSAIC, a comprehensive adversarial benchmark comprising 16000 samples across a full‑granularity attack spectrum. To address these challenges, we propose NeuroStat, an end‑to‑end framework bridging the statistical and semantic gap. NeuroStat captures uncompressed token‑level probabilistic logits alongside deep semantic hidden states from a single causal language model backbone. We fuse these heterogeneous signals through Macro‑State Residual Modulation, which adaptively calibrates local convolutional features using global uncertainty indicators. Orthogonal and contrastive losses further ensure the learning of complementary representations. Extensive experiments demonstrate that NeuroStat maintains exceptional robustness on MOSAIC compared to the severe degradation of state‑of‑the‑art methods, establishing a new standard for adversarial text detection. Code and the MOSAIC benchmark are available at https://github.com/TencentBAC/NeuroStat.

Authors:Chenxin Fang, Tao Chen, JunChao You, Jun Peng, Yiyi Zhou, Rongrong Ji
Title: Visual Token Coding for Video Multimodal Large Language Models
Abstract:
In this paper, we propose a new token compression paradigm for video Multimodal Large Language Models (MLLMs), termed Visual Token Coding (VTC). Inspired by classical video coding principles, e.g., HEVC, VTC performs structured compression by predicting the I/P frames of a video and measuring their frame‑wise residuals to estimate token redundancy. Based on this baseline framework, we also enhance VTC with a set of novel dynamic designs, such as Dynamic Resolution Input (DyRSO), Dynamic Token Allocation (DyTA), and Spatial Coverage Top‑K (SC‑TopK), and term this new approach VTC_Dy. To validate VTC, we apply it to three MLLMs and conduct experiments on multiple video understanding benchmarks. The experimental results show that VTC_\mathrmDy achieves an average performance retention of 100.1% with a 50% token budget for Qwen3‑VL, while still retaining 97.8% of the average performance when the token budget is reduced to 25%. Moreover, as a plug‑and‑play design, VTC requires no additional tuning of MLLMs for token coding. Our code is available at https://github.com/Msr233/VTC.

Authors:Zi Liang, Xiaoyu Xu, Yanyun Wang, Minxin Du, Qingqing Ye, Haibo Hu
Title: CAITLYN: Can LLM Agents Autonomously Synthesize Defenses against Emerging Injection Attacks?
Abstract:
Prompt injection attacks on Large Language Model (LLM) agents seek to introduce malicious instructions or content into external text sources retrieved by agents, forcing the underlying LLMs to execute harmful actions outside their benign scope. While current defenses effectively counter known injection attacks, deploying them in LLM agent environments remains challenging due to attack variants and emerging threats. Moreover, existing solutions typically suffer from an inherent trilemma, i.e., a constant trade‑off among runtime efficiency, contextual precision, and adaptability. To bridge this gap, we propose Continuous Agents for Injection Threats via Lifelong Yielding Nexus (CAITLYN), an agent‑agnostic defense middleware. CAITLYN integrates two systems. System I focuses on immediate defense against existing attacks using a two‑tiered library: Tier‑0 for rule‑based detection scripts and Tier‑1 for optimized LLM‑based accurate inference. System II, in contrast, is deployed to monitor potential abnormal signals and attempt to synthesize new defenses. On standard benchmarks, CAITLYN matches the detection performance of state‑of‑the‑art defenses at lower token overhead than LLM‑as‑a‑judge baselines. On Emerging, our new delivery‑aware benchmark featuring novel injection techniques, static baselines and the standalone System I configuration remain vulnerable. In contrast, System II autonomously synthesizes verified defense capabilities, substantially lowering the attack success rate across three diverse agent environments.

Authors:Antonio Purificato, Maria Sofia Bucarelli, Andrea Bacciu, Amin Mantrach, Fabrizio Silvestri
Title: QUORUM: QUality-Optimized Routing Using Multiple annotators
Abstract:
Data annotation remains a central bottleneck in natural language processing, requiring human effort to obtain high‑quality labels at scale. While Large Language Models (LLMs) offer a fast and cost‑effective alternative, their reliability is highly instance‑dependent: they perform well on simple inputs but often fail on examples requiring nuanced reasoning or contextual understanding. In this work, we address this challenge with QUORUM (QUality‑Optimized Routing Using Multiple annotators), a budget‑aware routing framework that dynamically assigns each instance to human or LLM annotators under a fixed annotation budget. Unlike prior approaches relying on model confidence or uncertainty estimates, QUORUM leverages feature‑based signals to estimate instance difficulty and supports multiple annotations per instance, combining them through agreement‑based rewards to improve reliability. We evaluate QUORUM across diverse closed‑ and open‑ended annotation tasks in English and multilingual settings, and QUORUM improves annotation quality by up to 34.4% while reducing costs by 8.8% over competing methods. Code can be found at https://github.com/amazon‑science/QUORUM.

Authors:Yucheng Wang, Yuetian Du, Zhengyi Liu, Rongyu Zhang, Bing Zhao, Boyu Yang, Ming Kong, Lin Qu, Hu Wei, Jie Liu, Qiang Zhu
Title: The Illusion of $\textit{What If}$: Evaluating the Breakdown of Counterfactual Reasoning in LLMs
Abstract:
Counterfactual reasoning requires models to reason beyond the observed world and explain how altered conditions propagate through downstream consequences. Existing benchmarks largely target bounded settings with fixed variables or single gold outcomes, overlooking open‑domain scenarios requiring causal‑process evaluation. To this end, we present WhatIfBench, a diagnostic benchmark for open‑domain, open‑form, long‑horizon counterfactual causal reasoning, containing 220 what‑if questions across STEM, HSS, and Hybrid scenarios. To evaluate free‑form responses, we further propose PRISM, which first converts each natural‑language explanation into a Response‑Derived Semantic Causal Graph of events, states, and mechanisms. On top of this graph, PRISM then jointly applies a Process Metric assessing graph‑level causal validity and a Rubric Metric assessing answer‑level explanatory adequacy. Evaluating six frontier LLMs with this framework, we find that WhatIfBench remains far from saturated: even the strongest model reaches only a 64.62% final score. Further analysis reveals persistent causal gaps, premise drift, and topology fragmentation, suggesting that fluent counterfactual narratives often mask fragile causal processes. The benchmark, code, and evaluation scripts are available at \hrefhttps://github.com/zju‑gt/WhatIfBenchWhatIfBench.

Authors:Wenze Ma, Chenyu Sun, Yanmin Zhu, Qiwen Gu, Xuhao Zhao
Title: Information-Guided Selective Modality-Interest Alignment for Multimodal Recommendation
Abstract:
Multimodal recommendation (MMRec) aims to enhance recommendation performance by leveraging rich item content from multiple modalities. However, directly incorporating all modality information does not necessarily lead to better preference modeling, since user interests are often more related to a subset of modality signals, while other signals may be weakly aligned with user preferences or even introduce noise. Although recent MMRec methods improve modality utilization through invariant learning, attention mechanisms, graph refinement, or contrastive learning, their alignment processes are often implicit or heuristic and lack a clear objective for selecting modality signals that better match user interests. In this paper, we propose AMUR, an information‑guided selective modality‑interest alignment framework for multimodal recommendation. Inspired by an information‑theoretic view, AMUR aims to enhance modality information that is more related to user interests while reducing the influence of less aligned signals. Specifically, AMUR first refines modality graph structures towards user behavior, and then selectively aligns shared interest‑related semantics across modalities. This enables AMUR to improve modality‑interest alignment while preserving useful modality‑specific complementary information. Extensive experiments on three real‑world datasets demonstrate the effectiveness of AMUR over competitive baselines. The code is available at https://github.com/Wenze1/AMUR.

Authors:Disen Liao, Yihan Wang, Freda Shi, Yaoliang Yu
Title: Cross-Session Decomposition Attacks: Scaling Risk and Intent-Aligned Retrieval Defense
Abstract:
Scaling laws are usually read as a capability story: lower language‑modeling loss yields more useful models. We study a safety consequence of this mechanism in \emphcross‑session decomposition attacks, where benign‑looking subqueries are asked across independent interactions and later recomposed toward a forbidden objective. We formalize this setting as \emphcompositional safety risk and prove a conditional risk‑transfer bound: when the reference environment already contains dispersed evidence for a risky reconstruction, the gap between deployed composed risk and reference composed risk is controlled by the model's excess loss on allowed subqueries. Synthetic withholding experiments show that wider transformers assign lower loss to held‑out instructions that never appear verbatim in training but are recoverable from injected supporting facts. A 600‑intent pretrained‑LLM evaluation shows that larger Qwen3 and Gemma3 family members can yield greater harmful‑capability uplift under a fixed decomposition‑composition pipeline. As a defense, IntentAlign‑MiniLM, our 22M‑parameter intent‑aligned retriever, outperforms much larger embedding models on held‑out intent retrieval and yields the best learned‑retriever harmful recall across tested guardrails. Code is available in \hrefhttps://github.com/liaodisen/Cross‑Session‑Decomposition‑Attacksour GitHub repository.

Authors:Wenqu Zhao, Xuemin Chi, Xin Zhang, Guoqing Ma, Baorun Li, Jianjie Fang, Peizhi Tang, Chen Gao, Wei Wu
Title: DensityKV: Density-Guided KV Cache Compression for Long Video Generation
Abstract:
Autoregressive video diffusion models enable streaming generation through sliding‑window attention, but each generated block is conditioned on previously generated content, causing appearance and motion errors to propagate recursively over time. Historical key‑value (KV) memory preserves earlier subject and scene states and helps maintain long‑horizon consistency. However, retaining every generated state creates a historical archive that grows continuously with the rollout, while recurrent states repeatedly add redundant coverage. To address this problem, we propose DensityKV, a training‑free historical KV bank management strategy. DensityKV maintains a separate token‑level KV bank for each attention head and measures local redundancy among the post‑RoPE keys that directly parameterize attention routing using Soft‑Riesz density. By constraining neighborhood‑density growth after states enter the bank, DensityKV limits repeated historical accumulation while preserving coherent states from each completed generation block. Experiments across three autoregressive video generation backbones and multiple generation lengths show that, at the same upper bound on historical KV capacity, DensityKV improves long‑horizon consistency and generation stability while keeping persistent historical storage bounded independently of rollout length.

Authors:Haodong Chen, Shuai Wang, Yu Yin, Shengyao Zhuang, Guido Zuccon, Teerapong Leelanupab
Title: ITER: Interaction-Aware Retrieval for Agentic Search
Abstract:
Deep‑research agents answer complex user questions through an iterative sequence of search steps, where the agent autonomously formulates sub‑queries to retrieve the evidence needed at each stage. However, existing retriever training typically relies only on the sub‑query and its corresponding search results at the current step as training signals, leaving the information accumulated from previous interactions largely underutilized. We introduce iter, an agent interaction‑aware dense retriever trained using agent trajectory learning signals. iter represents each query by incorporating not only the current sub‑query, but also the main question and preceding sub‑queries, and is trained using trajectory‑relative learning signals derived from the agent's interactions. Across six agent backbones from three model families, iter consistently outperforms the existing agent‑trajectory‑trained dense retriever, LRAT, achieving an average improvement of 7.5% on InfoSeek‑Eval and 13.5% on BrowseComp‑Plus. iter also demonstrates stronger cross‑agent robustness than AgentIR, a deep‑research retriever that relies on external LLM‑judge signals and the agent's pre‑search reasoning. Ablations further show that the main question and previous sub‑queries provide the most robust query representation, while previously visited and useful documents, used as redundancy negatives in subsequent searches, provide the strongest trajectory‑relative supervision. Code is available at https://github.com/ielab/ITER.

Authors:Injun Baek, HyeongSeok Lee, Yearim Kim, Junhoo Lee, Nojun Kwak
Title: LandingAgent: A Reference-Annotated Dataset and Agentic Generation Framework for Landing Pages
Abstract:
Landing pages are goal‑oriented web interfaces that must communicate a target‑specific value proposition while organizing information flow, visual hierarchy, and calls to action (CTA). Although large language models can generate plausible webpage code from natural‑language prompts, direct generation often yields generic templates and unsupported persuasive claims. We study target‑grounded, reference‑guided landing‑page generation, where a system must create an executable page for a new target by adapting reusable patterns from real pages without copying them. We introduce LandingBench, a reference‑profile dataset that abstracts real landing pages into section sequences, layout patterns, tone descriptors, visual emphasis, and CTA structure. Building on LandingBench, we propose LandingAgent, a three‑phase agentic framework that profiles the target, constructs a reference‑guided wireframe, and refines the page through critique‑guided polishing. We evaluate LandingAgent against direct prompting on faithfulness, conciseness, readability, aesthetics, and structural diversity. Experiments show improved target grounding, presentation quality, and layout diversity. Code is available at https://github.com/IAURAI/LandingAgent.

Authors:Jiatong Ding, Bingxin Xing, Yu Zhang, Dian Ding, Xiaodong Yi, Xianbin Ouyang, Feihu Zhou, Kun Zhang, Zhenyu Guo, Hao Pan, Guangtao Xue, Yiming Zhang
Title: HyQuant: Hybrid-Precision Quantization for LLM Attention
Abstract:
Quantization has been widely adopted in LLM training and inference to reduce cost and improve efficiency. However, low‑bit quantization of the \emphattention module often introduces large errors at very low bit‑widths, causing performance degradation. Existing methods mainly rely on smoothing techniques to handle outliers, while we propose a hybrid quantization design to better balance accuracy and efficiency. Specifically, we propose HyQuant, an efficient hybrid quantization framework for LLM attention. HyQuant quantizes most attention states into low‑bit formats while retaining a small set of vertical‑line tokens and local‑window states in high precision. These accuracy‑critical regions are selected using lightweight vertical‑line‑aware attention‑pattern signals, reducing quantization error with limited overhead. In the Prefill stage, HyQuant uses a hybrid‑precision quantized attention operator that preserves vertical‑line tokens and a local sliding window in full precision while quantizing the remaining context. In the Decode stage, HyQuant applies the same principle to KV‑cache compression and fuses KV dequantization with attention computation to improve memory and hardware efficiency. Across diverse tasks, models, and datasets, HyQuant maintains nearly lossless accuracy with an extremely simple design, demonstrating the efficiency and practical feasibility of hybrid quantization for LLM attention. Code is available at: https://github.com/jerrysfls/HyQuant .

Authors:Weicheng Xue, Bingqiang Wang, Li Yuan, Huihui Zhou, Yonghong Tian
Title: TerraceMoE: A Cost Model for Hierarchical MoE All-to-All Communication
Abstract:
Hierarchical two‑hop dispatch can reduce slow‑fabric traffic in expert‑parallel Mixture‑of‑Experts training, but it adds a second collective and an arrival‑side operator chain. We present a cost model for screening that trade at the communication‑call level, bounded by validation gates that withdraw a capability in code when they fail rather than reporting a caveat. At a reference geometry with 16 groups of 8 ranks, q=3, H=2048, and 4096 tokens per rank, the corrected effective breakeven hierarchy ratio is 3.98 for the measured PyTorch arrival chain, 1.49 for a hypothetical fused target, and 1.10 at zero implementation overhead. These are ratio‑only sensitivity results, not deployment predictions: platform A measures 1.03, platform B has no separated fast/slow measurement, and neither machine measured here reaches the hierarchical regime. Four communication‑level corpora pass their gates; a drift probe and the step‑level gate fail. The latter failure is enforced in code, so we make no training‑throughput prediction. The enabling routing constraint fixes per‑token fan‑out and per‑selected‑group quota, while aggregate per‑peer counts remain data‑dependent. Its measured validation‑loss cost is small but nonzero (+0.0034 nats); downstream equivalence is reported with incomplete estimator provenance and is therefore not independently reconstructible from the artifact. Code, calibration constants and the validation gates are at https://github.com/weich97/TerraceMoE‑simulator.

Authors:Runze Liu, Naibin Gu, Mingxu Ai, Yuqing Li, Peng Fu, Zheng Lin, Weiping Wang
Title: CoRe-MoE: Compact Reusable MoE for Continual Multimodal Instruction Tuning
Abstract:
Continual multimodal instruction tuning requires multimodal large language models to acquire new task abilities sequentially while preserving previously learned knowledge. LoRA‑MoE provides a promising solution by introducing expert‑based capacity, but repeatedly learning and maintaining full LoRA experts leads to substantial parameter overhead. This raises a natural question: is full expert expansion necessary for every new task? To answer it, we analyze the SVD of task‑specific LoRA updates and observe substantial overlap in their input‑ and output‑side LoRA direction subspaces, with task‑specific adaptation largely captured by lightweight coordinates over these subspaces. Motivated by this observation, we propose CoRe‑MoE, a Compact Reusable MoE framework for parameter‑efficient continual multimodal instruction tuning. CoRe‑MoE extracts reusable input‑ and output‑side direction bases from an initial expert bank, and for subsequent tasks trains only compact coordinate experts together with task‑specific low‑rank routers. Experiments on two representative MLLMs show that CoRe‑MoE improves final average performance over the strongest competing baseline by up to 5.90 points, while using less than 1% of the trainable parameters required by sequential LoRA for later tasks. The code is publicly available at https://github.com/runzezz/CoRe‑MoE.

Authors:Rit Gangopadhyay, Alex Wong
Title: From Perspective to Fisheye Depth Estimation and Open-Vocabulary Segmentation
Abstract:
Vision foundation models are capable of generalizing across 3‑dimensional (3D) scenes with high‑fidelity estimates; their empirical success can be attributed to training on large‑scale datasets of perspective images. However, when transferred to wide field‑of‑view (FoV) images, such as those captured by fisheye cameras, they return erroneous outputs due to a covariate shift stemming from the radial distortion on the image pixels. We propose a method to generalize vision foundation models to fisheye cameras. The crux of our method lies in a set of learnable parameters, termed Distortion Extenders (DEX), that model the fisheye distortion coefficients and the distributional shift between fisheye and perspective images encoded in the latent space. By minimizing a self‑supervised alignment loss, DEX transforms the latent embeddings of fisheye images to resemble those of perspective images to recover high‑fidelity estimates. DEX is architecture‑ and task‑agnostic: We demonstrate DEX on monocular depth estimation and open‑vocabulary segmentation for convolution‑ and Transformer‑based architectures, where we consistently improve over baselines across indoor and outdoor fisheye datasets. As a byproduct, the activations of DEX can also be decoded to distortion coefficients to support camera calibration. Code available at: https://github.com/Suchisrit/DEX.

Authors:Hojun Jeong, Gyunyeop Kim, Sangwoo Kang
Title: KLOD: Locality-Preserving Knowledge Editing via Non-Target Distribution Preservation
Abstract:
Fine‑tuning‑based knowledge editing is simple and architecture‑agnostic, but standard cross‑entropy increases the edited target probability without explicitly constraining changes in the non‑target output distribution. In sequential editing, such unconstrained redistribution can accumulate as distributional drift and contribute to locality degradation. We propose KLOD, a bounded and distribution‑preserving objective for fine‑tuning‑based knowledge editing that separates the intended target update from distributions that should remain stable. KLOD stops target amplification once a probability threshold is reached, while preserving the target‑excluded non‑target distribution at target positions and the full next‑token distribution at prefix positions. Experiments on CounterFact and ZsRE with Llama3‑8B‑Instruct and Qwen2.5‑7B‑Instruct show that KLOD substantially mitigates locality degradation while maintaining high edit reliability. The target probability threshold further provides a controllable Generalization‑‑Locality trade‑off. Ablation, multi‑seed, and distributional KL analyses support the interpretation that KLOD's locality gains are associated with preserving output distributions rather than simply weakening the edit. Code is available on GitHub https://github.com/Hostoday/KLOD .

Authors:Juan Pablo Vigneaux, Mary Kennedy, Khalil Iskarous, Robert Frank, Matilde Marcolli
Title: Representation of syntax in LLMs through the lens of linear distance and similarity-aware entropy
Abstract:
Structural probes were introduced by Hewitt and Manning to reconstruct syntactic trees from a neural language model's latent representations. They are evaluated by calculating the proportion of syntactic tree edges correctly reconstructed over an annotated corpus (as measured by undirected unlabeled attachment score). Here, we disaggregate this measure, considering undirected attachment score by label (UASL), which assesses the reconstruction accuracy of each syntactic relation separately, establishing important differences among relations that overlap linguistic distinctions. Moreover, we identify two factors that predict most of UASL's variability across relations: (i) the mean and dispersion of the linear distance (on a log scale) between the related words, and (ii) the diversity (similarity‑aware entropy) of the syntactic relation's head. These results, which hold across a range of model sizes and architectures, shed light on the degree of abstraction of the representation of syntax in language models and the dependence of such representation on geometric properties of the embedding space.

Authors:Trung Tien Dong, Zhenqi Wu, Aditya Penumarti, Zi-Hao Zhang, Micaiah Bartlett, Jane Shin, Xiaomin Lin
Title: uScenes: A Multimodal RGB and 3D Sonar Dataset for Underwater Robot Perception
Abstract:
Robust perception is essential for the deployment of autonomous underwater robots. However, optical cameras become unreliable under poor illumination and backscatter. Forward looking (2D) acoustic sensors remain effective under these conditions, but they measure range and bearing while leaving elevation unresolved, creating an ambiguity that prevents individual sonar returns from being localized in three dimensional (3D) space. This complicates the sensor use for 3D scene understanding and precise object detection. We introduce uScenes, a multimodal underwater dataset containing synchronized 3D multibeam sonar point clouds and RGB imagery. The dataset contains 110 scenes and 95,834 synchronized observation, representing 277.6 minutes of data collected across multiple field sessions. uScenes establishes a foundation for underwater sensor fusion, cross modal representation learning and 3D scene understanding. Code and datasets are given at https://github.com/era‑research‑lab/uScenes.

Authors:Adarsh Sudheer, David Li, Omar Elbanna, Ishaan Kodarapu, Arjun Bahuguna, Vasu Sharma
Title: Compositional Failure in Audio-Visual LLMs: Late-Layer Prior Dominance Under Cross-modal Conflict
Abstract:
We study audio‑visual conflict as a compositional generalization test for AV‑LLMs: the model must combine synchronized but semantically incompatible audio and video evidence and decide whether the pair matches. On VideoLLaMA 2‑7B‑AV, three alignment configurations remain nearchance on the scored exact‑string Yes/No subset of AVHBench, even though their output priors shift substantially. Similarly, off‑the‑shelf InternVideo2 experienced a 32.3% accuracy decrease specifically under cross‑modal conflict, accompanied by a 17.3% instruction‑following failure. We call this failure mode prior dominance: late‑layer commitment to an internally preferred answer pattern that is weakly grounded in the conflicting inputs. To explain this behavior, we conduct a mechanistic interpretability analysis and find that commitment remains concentrated at 25.5 \pm 1 layers. We show that stronger temporal alignment changes answer bias, but do not improve compositional conflict resolution. Code and data to reproduce our mechanistic audit and behavioral evaluations are available at https://github.com/AdarshSudheer09/AVHBench‑dmai.

Authors:Pedro M. M. de Castro
Title: Optimal exponential memory for sequential Euclidean connections: edge-power costs and phase transitions
Abstract:
We study the edge‑power cost of the labelled tree generated by the γ‑strategy, a constant‑gain rule for sequential Euclidean connections. Starting with x_0=p_0, each input point p_i is attached to x_i‑1, and the state is updated by x_i=γx_i‑1+(1‑γ)p_i. Retaining x_i subdivides the insertion segment into a spine edge and a leaf edge. The memory parameter γ controls how long earlier points influence subsequent attachment points. We minimize the sum of the α‑powers of the edge lengths under independent uniform input and arbitrary input sequences. For uniform points in the unit ball, the stationary problem has a transition at α=1. Its continuous extension is minimized at the boundary for 0<α\leq1, while every global minimizer is interior for α>1. We determine the finite optimizer in the joint window α_N=1+\varepsilon_N, \varepsilon_N\log N\toλ. Below an explicit threshold it lies on the N^‑1/2 scale, at the threshold its scale is \sqrt\log N/(N\log\log N), and above the threshold it approaches an explicit stationary root with two computable corrections. A second threshold identifies the governing correction, and differentiated estimates prove eventual uniqueness. At α=3d+8, the linear coefficient at the stationary endpoint changes sign and a branch of strict local maxima enters the parameter interval. For arbitrary input sequences, the optimal parameter and asymptotic worst‑case edge‑power cost per point are explicit for 0<α\leq3. At high powers, periodic antipodal block inputs give explicit lower bounds which, with a separation argument, show that the optimized cost is asymptotic to 2\log2/\logα. Exact results for powers two and four, a rational recursion for every even power, and a high‑dimensional expansion complete the analysis.

Authors:Yifan Zhang, Steve Ta, Jasper Zhang, Jichen Feng, Shuzhen Li, Yongxin Zhang, Yifeng Liu, Huizhuo Yuan, Mengdi Wang, Quanquan Gu, Andrew Chi-Chih Yao
Title: Fast Weight Attention for Continual Learning
Abstract:
Recurrent fast‑weight memories and selective state‑space models compress an expanding context into a fixed‑size recurrent state, making the state transition an online learning rule. We study this rule under read‑after‑write autoregressive semantics. For the prefix‑prediction objective considered here, the local fast‑memory example revealed at step t is the prefix‑aligned pair (\mathbfx_t,\mathbfy_t)=(ϕ(\mathbfk_t‑1),\mathbfv_t). The common same‑step association (ϕ(\mathbfk_t),\mathbfv_t) remains causal, but optimizes a different internal objective. We derive normalized first‑order updates for squared‑error regression and negative inner‑product objectives. The regression family comprises Falcon‑1 (a scalar NLMS update), Falcon‑2 (its per‑column extension), and Falcon‑3 (a sliding‑window mini‑batch update); Falcon‑1A/Falcon‑2A/Falcon‑3A are the corresponding inner‑product variants. We provide recurrent, masked‑parallel, and chunk‑parallel forms, together with numerically stable positive‑decay renormalization. Representative variants remain competitive in language modeling and improve length extrapolation on variable‑digit addition. This framework separates temporal alignment, plasticity, forgetting, and bounded rehearsal in recurrent sequence models.

Authors:Ka Heng Shiu, Kartic Subr
Title: ABCD: Alpha-Composited Block Coordinate Descent: Constant-VRAM Training for Large Radiance Fields
Abstract:
We present ABCD (Alpha‑Composited Block Coordinate Descent), an out‑of‑core training framework for alpha‑composited radiance fields, instantiated here for 3D Gaussian Splatting. Our method reformulates training as block coordinate descent over spatial partitions: only one block of parameters is active at a time, while all others are frozen. By exploiting the associativity of alpha blending, these inactive regions can be pre‑rendered and collapsed into foreground and background RGBA images. As a result, for fixed partition size and image resolution, peak VRAM becomes O(1) with respect to total scene extent, rather than growing with full scene size. This enables GPUs with limited memory to train scenes that would otherwise not fit in core. In experiments, our method closely preserves the reconstruction quality of 3DGS, with less than 5% PSNR degradation, while ABCD with compositing ablated suffers roughly 40% degradation. Our code can be found at https://github.com/shiukaheng/abcd

Authors:Musa Shams
Title: SegBench-GC: Testing Segmentation Invariance in Multi-Step Offline Goal-Conditioned Reinforcement Learning
Abstract:
Offline goal‑conditioned reinforcement learning (GCRL) often uses trajectory structure for future‑goal sampling and multi‑step targets, yet logged trajectories may be partitioned for administrative reasons that do not correspond to termination. We introduce SegBench‑GC, a controlled stress test of segmentation invariance that holds transitions, source trajectories, goal sampling, optimization settings, and evaluation fixed while varying only artificial backup boundaries and whether those boundaries retain continuation value. Continuation‑valid targets (CVT) provide the segmentation‑consistent control: reward accumulation stops at an artificial cut, but the target bootstraps from its stored successor. In a matched‑count PointMaze study with 35,000 artificial cuts, three segmentation realizations, and three optimization seeds, final 50‑episode‑per‑task success is 50.5% uncut, 39.1% with CVT, and 19.1% when the same cuts are treated as absorbing; across segmentation realizations, naive mean success ranges from 4.8% to 31.9%. An independent published n‑step baseline (n=25) from the Decoupled Q‑Chunking codebase shows the same failure on Puzzle‑4x5: 47.2% uncut, 58.5% CVT, and 0.27% naive across three optimization seeds. A target‑level diagnostic verifies the analytic target difference to numerical precision, and learned‑critic diagnostics show a large optimistic shift under naive handling while CVT remains approximately aligned with the uncut critic. CVT applies standard continuation bootstrapping rather than a new Bellman rule; the contribution is the controlled benchmark, failure isolation, and cross‑learner evidence that administrative segmentation can materially change multi‑step offline GCRL.

Authors:Abdulrahman Diaa, Jonathan Petit, Florian Kerschbaum
Title: Semantic Watermarking with Order-Robust Detection over Sub-sentence Units
Abstract:
Semantic watermarks tie the mark to sentence meaning rather than token choices, promising robustness to content‑preserving edits. However, the detector only observes attacker‑supplied text, which can be reworded, reordered, or resegmented to evade detection without content loss. Rewording, reordering, and resegmentation all cause embedding displacement: detection tests embeddings different from those selected during watermarking and can therefore lose the mark. Our adaptive embedding displacement attack (EDA) admits all three edits under a single objective that maximizes this displacement. It uses a public paraphraser and surrogate encoder without access to the provider's generator or secret key. At a 5% false‑positive rate (FPR) and content‑preservation threshold \barq=90%, EDA successfully removes the mark on between 32.6% and 47.9% of documents across four schemes, the highest among the tested attacks. Therefore, EDA evaluates the schemes' robustness more thoroughly than passive paraphrasing. To address these vulnerabilities, we design (k)‑SwordStamp: semantic watermarks with order‑robust detection over sub‑sentence units, reducing sensitivity to attacker‑chosen structure at a small quality cost. Against k‑SwordStamp, the strongest no‑box attack we test is an EDA variant adapted to its design, with a 10.8% attack‑success rate. A stronger EDA with access to the provider's detector and secret key reaches a 39.7% attack‑success rate, compared with 65.5% on k‑SemStamp. Our code is available at https://github.com/D‑Diaa/SwordStamp.

Authors:Davood Soleymanzadeh, Kaidi Zhang, Zhiyuan Zhang, Bihao Zhang, Xiao Liang, Yu She, Minghui Zheng
Title: PHR-VLA: Planning Horizon Reasoning for Vision-Language-Action Models
Abstract:
Vision‑language‑action models (VLAs) have shown strong promise for general‑purpose robotic manipulation by mapping language instructions and vision observations directly to actions. However, most VLAs primarily condition action prediction on current observations and lack an explicit mechanism for reasoning over future task dynamics, which is particularly important for fine‑grained, contact‑rich manipulation. We present PHR‑VLA, a framework that enables planning‑horizon reasoning in VLAs through privileged latent representations of future dynamics. PHR‑VLA introduces a lightweight auxiliary future head that, during training, aligns the VLA's internal representations with latent dynamics extracted from future observations. Evaluation results demonstrate that local, contact‑centric, patch‑level latent dynamics supervision from the wrist camera improves success rate on LIBERO from 84.1% to 88.4% and on real‑world disassembly tasks from 63.3% to 82.5%. Patch‑level supervision from a third‑person camera also improves performance on Meta‑World from 56.70% to 57.8%. These results demonstrate that privileged latent dynamics alignment provides an effective training signal for improving anticipatory reasoning in VLA policies. Project website: \hrefhttps://davoodsz.github.io/PHR‑VLA.github.io/https://davoodsz.github.io/PHR‑VLA.github.io/

Authors:Senqiao Yang, Chengyao Wang, Yuxin Chen, Zixuan Wang, Longxiang Tang, Haokun Gui, Jinhui Ye, Changsheng Lu, Xiaoyang Wu, Mingkang Zhu, Pengguang Chen, Shu Liu, Zhuotao Tian, Hengshuang Zhao, Bei Yu, Jiaya Jia
Title: Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models
Abstract:
Scaling robot data is crucial for building generalist Vision‑Language‑Action (VLA) models, yet robot trajectories are harder to scale than web‑scale image‑text data because embodied collection is costly and sparsely covers the physical world. This makes representation quality a central bottleneck: under a fixed robot‑data budget, continued pre‑training must turn limited trajectories into transferable visual‑action knowledge rather than merely fit actions. We propose VLAct, a VLA‑oriented VLM backbone trained on broad, heterogeneous, multi‑embodiment robot data before task‑specific fine‑tuning. VLAct preserves the broad VLM prior and encourages shared action semantics across embodiments through VLM‑prior preservation, multi‑head continuous action co‑supervision, and a partially unified cross‑embodiment action layout, while allowing task‑specific action heads during fine‑tuning. Across simulation, real‑world, and unseen‑embodiment transfer, VLAct consistently improves downstream performance under fixed fine‑tuning protocols. On LIBERO‑Plus and RoboTwin 2.0, VLAct surpasses industrial VLA systems including ABot‑M0 and LingBot‑VLA, achieving success rates of 82.6% and 92.5%. On RoboDojo, VLAct ranks sixth among all policies by success rate and outperforms all explicitly designated world‑action model (WAM) entries on both metrics. Most notably, on RoboCasa‑GR1, an unseen humanoid embodiment, VLAct using only 20% of downstream trajectories outperforms the full‑data GR00T‑N1.6 baseline. These results are obtained using fully open‑source data and only a 16‑GPU training setup, showing that representation‑centric continued pre‑training can deliver highly competitive performance under a modest compute budget and is an important independent axis of VLA progress beyond data scaling.

Authors:Hanyang Wang, Yimo Cai, Weiliang Chen, Jiawei Chi, Haowen Sun, Qiyu Dai, Yi-Hsin Hung, Xingzhuo Guo, Jinshan Ren, Runmao Yao, Ziwei Liu, Mingsheng Long, Yueqi Duan, Jun Gao, Jiangran Lyu, Fangfu Liu, Jialong Wu
Title: Code as Worlds: Agentic Discovery of Executable World Representations for Physical Reasoning
Abstract:
Physical understanding and reasoning depend on forming compact and generalizable representations of the world. While modern vision‑language models can recognize and explain diverse physical events, they often lack explicit representations of the underlying mechanisms‑such as object states, physical parameters, and governing dynamics‑needed for reliably reasoning how the world evolves and responds to interventions. In this work, we introduce Code‑as‑World, a paradigm that represents physical worlds through executable world representations. By expressing physical composition, dynamic evolution, and visual appearance as executable code, Code‑as‑World provides a compact, quantitatively grounded, and controllable abstraction of the physical world. To construct such representations from multimodal observations, such as natural‑language descriptions or real‑world videos, we develop an agentic discovery loop inspired by abductive reasoning, where an agent proposes, executes, renders, verifies, and iteratively refines executable world hypotheses. As a concrete application, we use verified executable worlds to provide scalable physical supervision for training vision‑language models on quantitative physical reasoning. Experiments show that Code‑as‑World‑VL achieves state‑of‑the‑art performance on QuantiPhy and surpasses leading proprietary models, highlighting the potential of executable world representations as a scalable foundation for physical intelligence.

Authors:Benlei Cui, Shen Pang, Yuke Wang, Xuemei Dong, Yuwen Zhai, Jingqun Tang, Haiyang Yu, Hui Xue, Longtao Huang, Haiwen Hong
Title: Fully Unleashing the Multimodal Attacker: Meta-Adaptive Jailbreaking of Vision-Language Models
Abstract:
The safety of large vision‑language models is increasingly stress‑tested by multimodal jailbreaks, yet existing attacks remain largely static at the meta level: template‑based attacks freeze the image‑text layout, while iterative attacks adapt only the image‑text content with fixed attack strategies and frozen attacker parameters. We propose Meta‑Adaptive Multimodal Jailbreaking (MAMJ), which instead optimizes the attacker itself along two axes: an attack strategy prompt (ASP) governing attack iteration and attacker model weights determining attack effectiveness. Across groups of multimodal attack trajectories, an LLM‑based critique first refines the ASP, after which group‑aggregated attack success rate (ASR) rewards update those weights. On MM‑SafetyBench, MAMJ achieves 81.0%, 78.9%, and 82.3% ASR against GPT‑4o, Gemini‑3‑Pro‑Preview, and Seed 2.0, respectively, outperforming the strongest sample‑level baseline by up to 24.1 percentage points. The learned attacker, comprising the optimized ASP and attacker weights, also transfers without retraining to unseen victims and remains effective under representative defenses. These results reveal a systemic vulnerability of frontier VLMs to meta‑adaptive jailbreaks and motivate defenses against meta‑level adversaries. Code is available at https://github.com/Alibaba‑VELLDEPTH/MetaJailbreak‑VLM.

Authors:Jiarong Han, Jincheng Xiong, Yuzhou Liu, Linzhe Shi, Changjie Wu, Ning Guo, Mu Xu, Hang Zhang, Ming Qian
Title: Revisiting Local Context for Long-Horizon Streaming 3D Reconstruction
Abstract:
Streaming 3D reconstruction from extremely long videos requires estimating camera motion and scene geometry online under bounded memory and computation. Early streaming models achieve causal, bounded‑cost inference using finite context buffers or compact recurrent states, yet their estimates often deteriorate as sequences grow. Recent methods improve long‑horizon stability by coupling short‑range context with persistent or multi‑level long‑range memory. We pursue a different route: we keep the learned temporal state strictly local and formulate predictions whose targets remain independent of sequence length. We present ABot‑Recon, a simple streaming model that caches KV features from only the preceding 11 frames. It predicts a point map in the current camera coordinate system together with an adjacent‑frame relative pose. These predictions remain equivariant under changes of reference frame, and global poses and geometry are recovered through sequential composition. To reduce accumulated drift, a lightweight temporal refiner improves relative rotations using recent visual and motion context, while a composition‑aware pose loss supervises multi‑step pose composition. Extensive evaluations on challenging long‑sequence benchmarks demonstrate the superior long‑horizon performance of our local‑context approach. On Oxford Spires, ABot‑Recon achieves an ATE of 4.35 m and an RPE‑R of 0.12^\circ, reducing both errors by approximately 40% relative to the best prior results.

Authors:Yifan Wang, Jie Gui, Adams Wai Kin Kong, Baosheng Yu, Changsheng Chen, Qi Li, Zhenan Sun, James Tin-Yau Kwok, Alex Kot
Title: FVeinSyn: Synthetic Finger Vein Image Generator
Abstract:
A major challenge in finger vein recognition is the lack of large‑scale public datasets. Existing datasets contain few identities and limited samples per finger, restricting the advancement of deep learning‑based methods. To address this, we propose FVeinSyn, a large‑scale controllable synthetic data generation framework for finger vein. It explicitly decouples synthesis of vascular topology and imaging appearance to mitigate the limitations caused by insufficient training samples, such as inadequate identity diversity and restricted realism. Specifically: first, a finger vein identity generator models vascular topology under physiological and geometric constraints using stochastic L‑systems, producing anatomically valid and identity‑distinctive vascular patterns. Then, a cascaded region‑aware GAN renders the topological maps into realistic near‑infrared images. Finally, an intra‑class diversity generator introduces geometric and optical perturbations to simulate realistic intra‑class variations. Using FVeinSyn, we generated 500,000 images (10,000 vein identities, 50 samples per identity) and conducted extensive evaluations. Results show that FVeinSyn holds significant advantages in realism, identity diversity, vascular pattern consistency, and intra‑class diversity. Models trained with FVeinSyn outperform real‑data‑only baselines a cross eight public datasets, achieving an average accuracy improvement of 27.43%. The code is available at: https://github.com/EvanWang98/Synthetic‑Finger‑Vein‑Generator.

Authors:Vedant Palit, Florent Draye, Terry Jingchen Zhang, Bernhard Schölkopf, Zhijing Jin
Title: How Do Linear Probes Emerge? A Circuit-Tracing Framework with Concept-Targeted Attribution
Abstract:
Transcoder attribution graphs are usually trained to explain why a model assigns high probability to a particular next token. We introduce Concept‑Targeted Attribution (CTA), which instead trains attribution graphs with respect to a linear probe direction. CTA therefore yields probe‑specific circuits that explain why an internal concept representation arises in a prompt, independently of whether it is expressed in the generated token. Using Cross‑Layer Transcoders, we show that these probe‑targeted graphs contain predictive structure: graph‑level features predict probe accuracy across four widely studied concept categories (ρ= 0.91, R^2 = 0.84), while local features identify the sparse components driving per‑prompt classification. This connects probe performance to interpretable circuit structure, allowing us to ask not only whether a probe works, but which internal computations make it work. Causal ablations further show that probe‑targeted and logit‑targeted graphs capture functionally distinct mechanisms. Removing probe‑relevant features reduces internal concept scores while largely preserving generated tokens, whereas removing logit‑relevant features changes the generated token in 92% to 100% of cases with near‑zero effect on probe scores. CTA provides a framework for moving from behavioral probe accuracy to mechanistic explanations of probe performance, enabling more detailed audits of internal concept representations, including safety‑critical ones. Our code is available at https://github.com/vedantpalit/concept‑targeted‑attribution

Authors:Yu Han, Tianwen Qian
Title: WM-R1: Training GUI Agents to Reason and leverage World Models with Reinforcement Learning
Abstract:
GUI agents trained with reinforcement learning (RL) have showcased strong environment learning capabilities on mobile platforms. However, RL typically demands extensive real‑environment interactions, leading to high resource costs and instability, especially in GUI scenarios. To address these, we propose WM‑R1, the first reinforcement learning framework that trains mobile GUI agents with world models instead of real environments. Specifically, world models serve as the source of state transitions during all rollouts, replacing the real Android environment within the training loop. WM‑R1 also embeds world models directly into the thinking process, enabling agents to reason about the consequences of candidate actions before committing to the final action. Crucially, WM‑R1 eliminates the need for real‑environment interaction, supports massively parallelized and step‑level granularized trajectory generation grounded in world models, and introduces a multi‑dimensional rule‑based reward that jointly optimizes task success, trajectory efficiency, and world model utilization. For efficient training, we curate a high‑quality dataset of 2000 challenging tasks. Experiments on Android mobile benchmarks demonstrate that WM‑R1‑trained agents significantly outperform GRPO‑only baselines and inference‑time simulation methods. Code is available at https://github.com/genalyu/WM‑R1 .

Authors:Xinhang Ma, Chaowei Xiao, William Yeoh, Ning Zhang, Yevgeniy Vorobeychik
Title: ROPE: Routed Origin Policy Enforcement against Indirect Prompt Injection
Abstract:
Indirect prompt injection (IPI) plants instructions in the content a tool‑using LLM agent reads, steering the agent into harmful tool calls. The strongest defenses are system‑level, leveraging techniques such as task‑conditional tool screening to prevent execution of malicious tools, and information‑flow control to avoid tool execution with untrusted parameters. However, as agents grow more capable, users delegate more to automation. Consequently, tool execution sequences and parameter values are increasingly determined at runtime and cannot be reliably screened from solely user's query without significant utility loss. We present ROPE (Routed Origin Policy Enforcement), which is anchored in a structural notion of trust: a value may reach a state‑changing tool only if it traces unforgeably to the user, a source the user explicitly named, or the user's own authoritative records. Enforcement is then a deterministic origin check over an audited set of sensitive tool parameters, and the only reliance on a language model involves solely the trusted user request, out of the attacker's reach. Our approach admits two provable guarantees: 1) at every step of a trajectory, no value whose only origin is attacker‑writable content reaches an origin‑guarded parameter, and 2) no rewording of an injection changes an admission decision. We evaluate across four agent models on open‑ended agent suites, ROPE holds attack success rate to 1.6‑‑2.6% while retaining 82‑‑100% of undefended clean utility, significantly exceeding state‑of‑the‑art system‑level defenses in utility while attaining comparable or better security. Further, we show that optimizing the injection against ROPE is largely ineffective, while long‑horizon attacks that defeat prior system‑level defenses achieve zero success rate. Our code and logs are available at https://github.com/xhOwenMa/ROPE .

Authors:Stephen Wu
Title: Covering 1024 syndromes with 50 columns
Abstract:
We exhibit a binary linear [50,40]_2 code of covering radius 2, so \ell_2(10,2)\le 50, one column below the Kaikkonen‑‑Rosendahl length 51 that has stood since 2003 and that still seeds the R=2 family of Davydov‑‑Marcugini‑‑Pambianco (arXiv:2511.02542). The new matrix admits a (2,0)‑partition into ten blocks, so Construction \mathrmQM_2^2 propagates it to exhaustively verified codes of lengths 815 and 1631 at r=18 and r=20, and to the family n=51\cdot 2^r/2‑5‑1 of asymptotic density 2601/2048. The matrices, verifiers, and source are at https://github.com/wustep/maths, pin problems/covering/share/2026‑08‑24/ at commit 736a38f.

Authors:Pratik Ghawate, Tanvi Patil
Title: CareGraph: An Auditable Hybrid AI Framework for Evidence-Grounded Personalized Longitudinal Health Intelligence
Abstract:
Artificial intelligence is transforming personalized healthcare, yet fragmented clinical, self reported, and wearable evidence remains difficult to interpret and trace. We present CareGraph, an auditable hybrid AI framework that converts heterogeneous records into prioritized trends, missing context indicators, bounded next steps, discussion questions, and provenance linked explanations. CareGraph organizes evidence without diagnosing, predicting outcomes, selecting treatment, or making autonomous clinical decisions. Its pipeline covers deterministic analysis, context detection, graph construction, constrained language model synthesis, evidence validation, safety controls, and release gating. Tests used synthetic cohorts of 400 patients each for development, validation, and holdout. On holdout data, a frozen ordinary least squares trend rule with a sufficiency gate achieved 0.827 accuracy, 0.837 macro F1 with a 95 percent confidence interval of 0.819 to 0.854, and 0.974 insufficient data F1. Missing context detection achieved 0.815 strict micro F1 versus 0.318 for the legacy detector. On an authored holdout benchmark, safety ruleset version 1.2 achieved 1.000 precision, 0.950 recall, and 0.974 F1. An audit requiring graph retrieval across 80 patients yielded 79 syntheses and 78 presentations without fallback; one output was blocked and one failed closed because of an invalid evidence key. Against monolithic GPT 5.6 on 56 matched patients, CareGraph was faster at 40.15 versus 49.62 seconds, shorter at 661 versus 1,163 words, and showed better exploratory lexical alignment with longitudinal targets; the baseline used fewer tokens and cited more raw evidence. Graph auditing verified provenance and deterministic retrieval; incremental graph effects on generation require paired evaluation. CareGraph offers a safety bounded foundation for intelligent personalized health systems.

Authors:Mohammad Arvan, Hossein Haeri, Natalie Parde, Rebecca T. Feinstein
Title: UIC-AIHealth4All at ArchEHR-QA 2026: Answer-First Evidence Grounding for Clinical Question Answering
Abstract:
We describe the UIC‑AIHealth4All system for ArchEHR‑QA 2026, a shared task on grounded question answering from electronic health records. We participated in Subtasks 2 (evidence identification), 3 (answer generation), and 4 (answer‑evidence alignment). For Subtasks 2 and 3, we propose an answer‑first pipeline in which the model generates candidate answers citing specific note sentences before classifying the full evidence set, exploiting the asymmetry between judging relevance in the abstract versus relative to a generated answer. For Subtask 4, we apply self‑consistency voting over five independent model calls, retaining links by vote threshold. Our pipeline ranked third on evidence identification (Strict Micro F1 62.90), ninth on answer generation (Overall 31.90), and fifth on answer‑evidence alignment (F1 79.81). A post‑hoc linguistic analysis of 45 stylistic features reveals that model outputs remain 3.2 Flesch‑Kincaid grade levels harder to read than clinician‑authored references despite matching their word and sentence counts, suggesting readability warrants explicit optimization in clinical NLP systems. Code and prompts are available at https://github.com/mo‑arvan/archehr‑qa‑2026‑uic‑aihealth4all.

Authors:Tianjie Ju, Zheng Wu, Yueqing Sun, Yuhan Cui, Bobo Li, Shengqiong Wu, Pengzhou Cheng, Haodong Zhao, Zongru Wu, Xinbei Ma, Doris Zhang, Kunling Li, Mong-Li Lee, Wynne Hsu, Hao Fei, Qi Gu, Gongshen Liu, Zhuosheng Zhang
Title: UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City
Abstract:
Multimodal large language models (MLLMs) can interpret a street view, but urban agency depends on whether such local evidence remains useful after the agent starts to move. In this paper, we investigate how far current MLLM agents can turn local urban perception into reliable action in a complicated real‑scale city. We propose UrbanGround, the first sandbox to make this question testable in a physically constrained replica of Hong Kong built from territory‑wide 3D geospatial data. UrbanGround supports closed‑loop interaction from a first‑person view and provides an interactive map for navigation. Agents can directly enter the 3D city and explore from a first‑person view. Our analysis follows the growth of the spatial problem through three research questions. We first test whether an agent can ground a local scene well enough to answer spatial questions after active observation. Then we ask whether that grounding supports navigation as destinations become farther away and less explicit. Finally, we examine whether the resulting behavior survives changes in route availability and pedestrian motion. Contemporary MLLM agents usually show useful atomic abilities in visual recognition and short‑range spatial reasoning, while orientation and pedestrian‑aware movement remain unreliable. Their central failure emerges over extended exploration, where local abilities do not compose into sustained goal‑directed behavior and errors accumulate without effective correction. We hope UrbanGround will support broader study of how far current MLLM agents can explore reliably in complex, open‑ended urban environments.

Authors:Yufan Wu, Yinghui He, Zhengyi Hu, Lang Wei, Ruichen Li, Qifan Yang, Ting Zhu
Title: CritICL: Inference-Time Weak-to-Strong Generalization from Small Language Model Failure Modes
Abstract:
Recent advances in inference‑time scaling have significantly improved the reasoning performance of large language models (LLMs). However, these methods typically rely on repeated generation or external verification. To address this limitation, we introduce CritICL, a novel inference‑time framework that improves reasoning while maintaining high efficiency. Our key insight is that LLM failure modes exhibit structured patterns across model scales within the same family. Instead of treating failures as undesirable outputs, CritICL leverages them as a source of guidance. Specifically, we utilize failure modes derived from weaker models and incorporate them into inference through critique‑based in‑context examples. We propose two variants: CritICL‑dynamic, which adaptively predicts input‑specific failure modes and retrieves critiques, and CritICL‑static, which uses a global failure mode profile to provide stable guidance. Experimental results show that CritICL consistently outperforms standard in‑context learning and achieves performance competitive with or superior to test‑time scaling methods, while requiring significantly fewer generations and lower token cost. Code available at: https://github.com/umwyf/CRITICL

Authors:Chiké Abuah
Title: Tacet: A Language and Type System for Automatic Statistical Validity Accounting
Abstract:
Empirical comparisons between systems are a standard form of evidence in computer science research, but few are checked for statistical validity: most are never framed as statistical tests at all. Existing multiple‑comparison procedures could control the resulting error, but need inputs (what an analysis examined, and how its observations are arranged) that are not recoverable from a list of p‑values. We introduce Tacet, a language in which an analysis declares what it generated, states what it expects to find, and is refused any claim it cannot afford or cannot properly test. Its core calculus T pairs a free estimation sublanguage, carrying a reported footprint and a purity bit that records whether any outcome was consulted in building a value, with a priced claim sublanguage, carrying a wealth transformer, connected only by a mechanism that prices a comparison. A sample selected by reading outcomes sets the purity bit and is recorded as having examined everything it read, permanently, so it can never be granted a one‑sided or confirmatory price, without the system ever asking whether the analyst intended to cherry‑pick. Whether a comparison is paired or clustered is computed statically from the artifact schema, from declared functional dependencies between key fields alone and before any data is read, and a mechanism that assumes that structure away is refused rather than priced. Because the wealth transformer is antitone in the realized p‑value, affordability can be checked before the analysis runs too, turning pre‑registration into a typing rule. We prove the metatheory machine‑checked in Lean 4 with no admitted gaps, and demonstrate the approach on a reference implementation and two case studies on published artifacts, the SWE‑bench Verified leaderboard and BIG‑Bench Hard.

Authors:Aozhe Wang, Zhengxi Lu, Jianze Wang, Shangke Lv, Ying Liu, Weiming Lu, Jun Xiao, Yueting Zhuang, Hua Yang, Qianglong Chen, Yongliang Shen
Title: TTPO: Test-Time Policy Optimization
Abstract:
Recent prominent post‑training methods, such as Reinforcement Learning (RL) and On‑Policy Self‑Distillation (OPSD), have driven rapid progress in mathematical reasoning for large language models, yet their reliance on ground‑truth labels precludes test‑time training (TTT). Replacing ground truth with majority‑vote pseudo‑labels is a natural alternative, yet it is fragile: an incorrect vote corrupts the teacher and misleads every token. We observe that this failure mode is asymmetric: rollouts that disagree with the pseudo‑label are typically wrong regardless of whether the vote itself is correct. Building on this observation, we propose Test‑Time Policy Optimization (TTPO), an asymmetric objective that distills agreeing rollouts via OPSD and penalizes disagreeing rollouts with Grouped RL. Token‑level selection further refines both branches: distillation down‑weights already‑converged positions, while RL penalizes only confident errors. Both updates remain well‑grounded even under frequent pseudo‑label errors, and majority‑vote routing yields tighter self‑supervision as the model improves. Without any labels, TTPO matches label‑supervised OPSD on five competition‑level benchmarks, raises Qwen3‑1.7B from 38.0% to 45.2% in TTT, yields +25.2% to +36.4% without thinking, and shows strong cross‑task generalization.

Authors:Mehul Shah, Robert Fiszer, Marek Perkowski
Title: Factorized Boolean representations for efficient quantum synthesis
Abstract:
Quantum algorithms promise advantages beyond classical reach, but running them on error‑corrected hardware requires translating Boolean specifications into reversible circuits, and the resources that translation demands determine what is executable. Established methods minimize a Boolean expression and map it to a circuit, assuming the minimized form is best. Here we show that minimized expressions retain algebraic structure minimization cannot reach, arising from containment and complementary‑polarity relationships among their terms, and that extracting it yields circuits cheaper to execute despite having more operations. The decisive quantity is not a circuit's operation count but the control count of its widest operation, a superlinear cost; extracting shared factors trades a few wide operations for many narrow ones and reduces qubit count. Across benchmarks and oracles from quantum search and factoring algorithms, at the representation level the transformation never increases either cost measure, a guarantee from its construction. Translation to an executable circuit returns part of that advantage, since auxiliary lines must be uncomputed, yet the factorized circuit still left a leading circuit‑level optimizer reaching lower final counts, and faster, than unaided. The representation of a computation is therefore itself a resource, optimizable before compilation and distinct from both logic minimization and circuit‑level optimization.

Authors:Maksim Utushkin, Andrei Ovsiannikov, Alexander D'yakonov
Title: Scaling Graph Neural Networks for Friend Recommendation: Multi-Hash User Embeddings and Temporal Neighbor Sampling
Abstract:
Friend recommendation is inherently graph‑structured: the relevance of a potential connection depends on multi‑hop social context rather than user attributes alone. However, deploying message‑passing GNNs on a production‑scale social graph with hundreds of millions of users and tens of billions of edges requires addressing numerous modeling and systems challenges. We present a scalable end‑to‑end GNN ranking system for production social graphs, focusing on two design choices that are critical in this setting: multi‑hash ID embeddings and temporal neighbor sampling. Multi‑hash embeddings are common for high‑cardinality features, but industrial GNN systems typically either ignore trainable IDs or accept full embedding tables, exceeding 200 GB for our graph. We integrate multi‑hash as the primary node representation, reducing the ID‑embedding table size by more than 98 percent while preserving ranking quality. Temporal neighbor sampling is well understood in principle, but existing implementations scan full adjacency lists, which is a non‑starter for users with tens of thousands of friends. We implement timestamp‑sorted CSR storage with binary search, reducing the per‑node temporal sampling cost from O(deg(v) + k) to O(\log(deg(v)) + k). Beyond these components, we show that this combination scales and yields measurable production impact. On a graph with 194M users and 28B edges, offline ablations isolate each design choice's contribution. In an online A/B test, our system increases friend additions from recommendations by 16 percent and unique friend adders by 11.5 percent over a strong production baseline. We release our framework for distributed training and inference on large temporal graphs.

Authors:Agniv Chatterjee, Georgios Pavlakos
Title: Reconstructing Humans and Objects in Interaction using Large Reconstruction Models
Abstract:
Estimation of Human‑Object Interactions in 3D (3D HOI) is a fundamental problem in 3D computer vision with applications in AR/VR, robotics, and embodied AI. However, reconstructing these interactions in 3D remains challenging due to depth ambiguities, occlusions, and object shape variability. Existing approaches are primarily concerned with reprojection and contact constraints, fitting parametric human models and object templates to 2D images. In this paper, we explore a different avenue. We present MILO, a framework that leverages the visual capabilities of Large Reconstruction Models (LRMs) to recover detailed 3D human‑object interactions from a single image. Our key observation is that LRMs provide a powerful geometric scaffold that preserves relative human‑object arrangement and proximity cues. This significantly simplifies the reconstruction procedure, reframing the problem as interpreting the LRM mesh: we segment it into human and object components, fit a parametric body model to the human part, and optionally align an object template to the object part (if such a template is available). MILO achieves strong reconstruction accuracy and outperforms existing baselines across multiple benchmarks and interaction scenarios. Our code is available at https://ac5113.github.io/MILO.

Authors:Orion Reblitz-Richardson
Title: How Language Models Organize and Structure Moral Knowledge
Abstract:
How do large language models (LLMs) organize moral knowledge? Models detect moral content broadly, but detection is a low bar. We ask whether they go further, distinguishing moral foundations from one another and organizing the relationships between them geometrically. We train six independent linear probes on open‑weight language models, one per Moral Foundations Theory (MFT) category (care/harm, fair/cheat, lib/oppress, loy/betray, auth/subv, sanc/degrade), and examine how the resulting directions relate to each other in representation space. We find the directions neither collapse into a single moral detector nor isolate from one another. Rather, they span a near‑maximal number of independent dimensions while sharing a positive common component. The shared component is the signature of integration, and it is moral‑specific relative to a matched non‑moral concept battery built identically (mean pairwise cosine 0.26 vs. 0.013). The geometry is consistent across architectures and scale and reaches its integration regime early in pre‑training, well before probe accuracy saturates. The structure the model discovers shows no evidence of the individualizing/binding distinction predicted by Moral Foundations Theory (an underpowered test: only 20 candidate partitions exist) but rather reflects corpus statistics. Extending to moral dilemmas, each dilemma direction partially composes from its component foundations, at 2.7x a mismatched‑pair baseline, while the majority of its variance encodes conflict‑specific structure. The model represents moral tension itself, not a pre‑resolved judgment.

Authors:Xin Chen, Fuwei Zhang, Yiqi Tong, Wei Guo, Yutian Xiao, Fuzhen Zhuang
Title: D2C-Routing: Dimension-to-Composition Evidence Routing for Mixed-Origin AI-Generated Text Detection
Abstract:
AI‑generated text detection is commonly framed as a binary document‑level judgment about whether a text is human‑written or machine‑generated. This framing breaks down for mixed‑origin writing, where content origin and expression origin may differ. We cast mixed‑origin detection as dimension‑to‑composition source attribution, inferring content origin and expression origin before composing them into four collaboration types. We propose Dimension‑to‑Composition Routing (D2C‑Routing), which routes content‑side and expression‑side evidence to supervised dimension heads before a learned gated composition layer predicts the final label. On MixD2C, a reconstructed split derived from the HART mixed‑origin benchmark, our disclosed D2C‑Routing‑based detector system reaches 0.8603 four‑way Avg TPR@1%FPR, 6.5 points above the same‑split RACE‑local rerun. Core ablations support the routing design, while error analysis shows that distinguishing AI‑content/human‑expression from fully AI‑generated text remains the hardest boundary. Code is available at https://github.com/bystander563/d2c‑routing‑artifact.

Authors:Canzhi Chen, Zan Wang, Siqi Zhu, Qi Wu, Yixuan Li, Wei Liang
Title: Embodied Scene Rearrangement Planning
Abstract:
This paper introduces Embodied Scene Rearrangement Planning (ESRP), a novel task requiring embodied agents to rearrange furniture in 3D scenes to match a target configuration using only egocentric observations and a top‑down target layout. Unlike prior rearrangement tasks, ESRP precludes global state access and introduces mutual object occlusions, reflecting the practical constraints of real‑world robotic deployment. These factors make aligning partial egocentric observations with the global target layout particularly challenging for long‑horizon planning. To facilitate research, we present ESRP‑Bench, a comprehensive benchmark built on OmniGibson featuring over 5,400 scene pairs and 8,200 objects. We define three multi‑level metrics to evaluate rearrangement quality and provide four baselines: a hierarchical task‑and‑motion planning method, a vision‑language‑model‑based method, and two learning‑based approaches (IL and RL). Experimental results demonstrate that current methods struggle to complete the task efficiently, highlighting ESRP as a challenging frontier for embodied agents in scene understanding and long‑horizon task planning. This work serves as a stepping stone toward deploying intelligent agents in real‑world scenarios. Project page: https://pie‑lab.cn/ESRP/.

Authors:Romolo Muletta, Felix Matthias Saaro, Mark Cieliebak, Jan Deriu
Title: Your Voice Cloning System is Secretly a Voice Anonymizer
Abstract:
Speaker anonymization suppresses speaker‑identifying attributes from speech while preserving linguistic content and quality. We propose repurposing XTTSv2, a multilingual voice cloning model trained on 27k hours of speech, for speaker anonymization without retraining. Our key insight is that XTTSv2's voice cloning capabilities preserve prosodic structure independently of speaker identity, enabling voice conversion by conditioning on a pseudo‑speaker. We introduce an iterative refinement strategy that balances privacy and utility by maximizing a harmonic mean of speaker dissimilarity and intelligibility. Evaluated on seven European languages across CommonVoice and Multilingual LibriSpeech, our system achieves near‑optimal privacy (EER \approx 0.49), competitive intelligibility, and substantially better speech quality than dedicated anonymization baselines, while requiring no language‑specific training. We release the code here: https://github.com/rm00cr/coqui‑tts.

Authors:Yutong Zhang, Jianshuo Dong, Peng Xu, Long Wang, Jie Zhang, Tianwei Zhang, Xiaoping Zhang, Han Qiu
Title: INTENT-AS-A-TOOL Makes it Easy to Track Agentic Misalignment
Abstract:
As large language models (LLMs) are deployed as autonomous agents, safety failures increasingly involve consequential actions. We study agentic misalignment, where agents take harmful actions under goal conflicts and pressures. Using chain‑of‑thought (CoT) monitoring, we find that harmful execution is often preceded by intent signals in reasoning. However, post‑hoc CoT labels are too coarse to show how intent changes during generation. We introduce INTENT‑AS‑A‑TOOL, an approach that adds intent‑targeted tools to give the model a dedicated channel for expressing commitment to a target behavior. The probability of calling an intent tool provides a judge‑free, fine‑grained signal of the model's tendency to pursue that behavior. Our results show that INTENT‑AS‑A‑TOOL complements CoT monitoring, expands post‑hoc CoT labels into dense trajectories, and identifies critical steps for online intervention. These findings suggest that action preferences are useful for tracking agentic misalignment during reasoning. Our code and data are accessible: https://github.com/RebeccaZhang22/intent‑as‑a‑tool.

Authors:Jeong-Yoon Kim
Title: BTS-AgentBench: A Deterministic, Replayable Pipeline from Read-Only Telemetry Logs to Agent Benchmarks
Abstract:
Industrial sites contain large volumes of read‑only telemetry, but few benchmarks specify how to compile these records into executable multi‑turn agent tasks. We present a telemetry‑to‑episode construction method instantiated as BTS‑AgentBench. The pipeline normalizes BTS metadata and raw histories into a read‑only tool store, compiles static tasks with tool‑derived gold answers and evidence, and lifts retained tasks into typed, bounded operator‑facing episodes. The 532‑row release adds clarification, goal revision, timestamp policy, quality‑gated reporting, and evidence attribution while preserving the source computation and split. Coded contract preflight reports zero findings, and the construction‑exclusion controller completes 0/532 rows. Two independent raw‑to‑episode builds match all 11 logical tool‑store exports and reproduce the released 356/87/89 train/dev/test artifact exactly. Applying the shared construction path to XAI4HEAT produces 204 episodes; on its 41‑row held‑out test split, the controller completes 0 rows and the retained GPT‑5.5 execution completes all 41. Code, artifacts, and replay reports are available at https://github.com/kjy7567/BTS‑AgentBench.

Authors:Qiwen Gu, Bingjie Gao, Rui Chen, Geng Li, Jifan Li, Qishuai Wen, Li Niu, Jing Tang, Xiangxiang Chu, Junqiao Zhao
Title: R2M-Bench: Evaluating Revisit Memory via Relative Consistency in Interactive Video World Models
Abstract:
High similarity between first‑visit and return frames does not necessarily show that a video world model remembered the scene; the intervening rollout may simply have changed very little. This ambiguity makes absolute revisit scores sensitive to rendering stability, repetitive content, and failed motion. We introduce \emphR2M‑Bench (Relative Revisit Memory Benchmark), a benchmark of observable revisit‑selective consistency. For every detected return, R2M‑Bench compares the revisit pair with two controls from the same rollout: a gap‑matched non‑revisit pair that measures generic temporal stability and a short‑range pair that estimates short‑horizon consistency. These comparisons produce \emphMemoryGain (MG), the revisit advantage over the temporal baseline, and the \emphNormalized Memory Ratio (NMR), which normalizes this advantage by the short‑to‑baseline dynamic range. R2M‑Bench combines 100 reference scenes with three leave‑and‑return trajectories to form 300 instances and evaluates appearance fidelity, scene and object identity, local geometry, and persistent state. Across seven action‑conditioned video world models, Overall NMR correlates with human consistency judgments at Spearman's ρ=0.547 (95% CI [0.45,0.63]). Its within‑model correlation magnitude with generated motion is 0.072, compared with 0.207 for raw revisit similarity, indicating that relative calibration substantially reduces the slow‑motion shortcut. DreamX‑World‑Memo achieves the highest Overall NMR among the evaluated video models. Together, these results support same‑rollout relative calibration as a practical way to distinguish revisit‑specific consistency from generic temporal stability.

Authors:Jinghan Xu, Yikai Zhang, Aili Chen, Weiyuan Li, Jiaqing Liang, Deqing Yang
Title: Verify Smarter, Evolve Further: Efficient Harness Evolution through Behavior-Aware Verification
Abstract:
Agent harnesses shape how language‑model agents use instructions, tools, and runtime components, but adapting these harnesses requires costly verification. Existing propose‑and‑verify methods typically score every candidate on a fixed task set, wasting rollouts on unrelated behaviors and allowing aggregate scores to obscure specific regressions. We introduce HarnessLens, a budget‑aware framework for automated harness evolution. HarnessLens jointly explores the task space and user‑configurable components, derives candidate modifications from execution trajectories, and selectively verifies each candidate on behavior‑relevant tasks using an attributable‑evidence gate. Across three agent harnesses and four benchmarks, HarnessLens improves average held‑out performance by 7.6‑13.6% while consuming substantially less evaluation budget than competing baselines. These results demonstrate that behavior‑aware verification with explicit attribution enables more reliable and sample‑efficient harness evolution under constrained interaction budgets. Our code is available at https://github.com/jhxu5214/HarnessLens.

Authors:Junjie Liu, Shengyuan Ye, Xu Chen
Title: PACE: A Unified Condense-and-Extract Paradigm for Fast VLM Inference
Abstract:
Vision‑Language Models (VLMs) demonstrate exceptional visual reasoning capabilities, yet their inference costs escalate rapidly with the proliferation of visual tokens. Existing visual token pruning methods exhibit two fundamental limitations. First, most approaches operate exclusively post‑vision encoder, leaving the substantial latency of the visual encoding phase unoptimized. Second, under strict token budgets, these methods often fail to jointly preserve holistic visual contexts and fine‑grained details, leading to performance degradation. To address these bottlenecks, we propose PACE (Pixel‑Adaptive Condense and Extract), a training‑free inference framework that accelerates both the vision encoder and the Large Language Model (LLM) via a unified Condense‑and‑Extract paradigm. During the Condense stage, an Adaptive Pixel Compressor (APC) evaluates visual information density prior to encoding, adaptively downsampling redundant inputs, curtailing encoder computation while preserving global context and essential visual cues. In the Extract stage, a Dynamic Dual‑Attention Extractor (DDAE) selectively retains visual tokens via a fusion of internal visual signals from the encoder and semantic signals from the LLM, safeguarding task‑critical details. By integrating PACE into Qwen2.5‑VL‑7B, the model retains 93.8% of its original performance while utilizing only 10% of the visual tokens, yielding a 3.1x speedup in time to first token (TTFT). Our code is available at https://github.com/jjL357/PACE.

Authors:Gauthier Miralles, Loic Le Folgoc, Vincent Jugnon, Pietro Gori
Title: Unsupervised Adaptation of 3D CT Foundation Models for 3D CBCT Segmentation
Abstract:
Accurate 3D segmentation of cone‑beam CT (CBCT) is critical for interventional and radiation therapy applications, yet it remains limited by two compounding challenges: the scarcity of annotated CBCT data and the large domain shift from diagnostic CT. Interventional CBCT exhibits fundamental modality differences from conventional CT, driven by acquisition and physics effects as well as contrast‑specific vascular content, thereby limiting effective cross‑modality model transfer. We propose a novel unsupervised domain adaptation (UDA) framework based on redundancy‑reducing feature alignment, enabling 3D CBCT segmentation with no target‑domain annotations or inference‑time adaptation. Our framework is architecture‑agnostic, seamlessly adapting both CNN‑based and ViT‑based foundation models. We evaluate our method on two challenging CT‑CBCT liver segmentation benchmarks: one for interventional vascular procedures and one for radiation therapy, demonstrating that even large‑scale pretrained segmentation networks require explicit feature‑space bridging to generalize across acquisition modalities, and that our approach consistently outperforms existing pretrained foundation model and UDA strategies. To support reproducibility and benchmarking, we release the liver segmentations for a public CBCT dataset, along with the code, trained models, and weights.

Authors:Matthew Youngman, Cristian Sestito, Themis Prodromakis
Title: LLMs in Digital EDA: A perspective on shifting roles from Generation to Orchestration
Abstract:
Electronic design automation (EDA) has advanced engineering productivity through successive generations of tooling that progressively automate synthesis, optimisation, and verification. Large language models (LLMs) extend this trajectory by enabling direct translation from design intent to hardware implementations. In most of the EDA literature, LLM‑based solutions are typically assisting siloed design stages or tasks, however this obscured the drivers by which capability emerges and systems scale. In this Perspective, we instead define three hierarchical roles that reveal how capability accumulates: a Generator that produces design artifacts in a single pass, an Agent that refines outputs through iterative tool feedback, and an Orchestrator that coordinates decisions across EDA‑stages. Across published systems, this reveals a syntax trap in which models are trained to produce plausible code rather than physically correct hardware, compounded by fragmented tools and loss of design context that obscure how decisions affect later stages. Comparisons across the three roles show that current approaches struggle to scale to industrial designs, motivating a shift towards a standardised, physics‑aware orchestrator that connects tools and agents across the EDA flow for more reliable and accessible hardware design.

Authors:Hiuyi Cheng, Nuo Xu, Yuyi Zhang, Xuhan Zheng, Wei Pan, Jing Zhang, Dezhi Peng, Minghui Liao, Yihua Teng, Jihao Wu, Haoyu Ren, Lianwen Jin
Title: Ancient-Bench: A Comprehensive Multi-millennial, Multi-medium, and Multi-script Benchmark for Ancient Chinese Artifact Text Recognition
Abstract:
Ancient Chinese artifact text recognition is fundamental to heritage digitization, and benchmarks for ancient texts are essential for evaluating current model capabilities. However, existing benchmarks suffer from ''fragmentation'', manifested in limited temporal coverage, limited medium diversity, and incomplete script types. Therefore, we present Ancient‑Bench, a comprehensive benchmark of 2,700 images for ancient Chinese artifact text recognition, featuring three dimensions: Multi‑millennial (spanning 3,000 years of character evolution), Multi‑medium (covering nine artifact categories), and Multi‑script (encompassing seven historical script forms). To enable consistent and fair evaluation across heterogeneous media, we further define three annotation standards tailored to the medium‑specific characteristics of ancient texts: symbol standardization, character standardization, and parsing standardization. Extensive experiments on Ancient‑Bench covering general Vision‑Language Models (VLMs) and OCR‑specialist models reveal that ancient Chinese artifact text recognition remains fundamentally unsolved, with persistent challenges in variant characters, specialized symbols, and hallucination. The dataset is available at https://github.com/SCUT‑DLVCLab/Ancient_Bench.

Authors:Pranav Aggarwal
Title: Calibrated Enough to Know, Not Calibrated to Act: Fabricated Evidence Makes LLM Agents Commit to the Unknowable
Abstract:
An LLM agent shown a professional‑looking market panel commits to a directional call on a provably unpredictable question far more often than one asked the bare question: across 12 frontier models, commitment rises from 6.5% to 54.0% as evidence is escalated. It commits just as readily when every number on the panel is invented: fabricating the entire display, so nothing the model can see is true except the question itself, still lifts commitment from 24.5% to 36.8%, statistically indistinguishable from the 37.6% produced by genuine market data. What unlocks confident action is not information but the authority of its packaging. The failure is narrow and locatable. Incapacity is not the answer: on matched answerable questions attached to the same panels, the same models answer essentially always, at near‑perfect accuracy. Nor is it belief ‑ stated probabilities barely move across the gradient that swings action by 48 points, and score worse than a climatological baseline. Missing judgment isn't it either: asked to classify a question's knowability before acting, models call it irreducible 90% of the time and then commit on just 0.4% of those. The act/don't‑act gate is what fails, and the effect is concentrated in a few models rather than universal. Because the gate is separable, it can be trained. Supervised fine‑tuning of a 3B model on 540 synthetic cases, predominantly dice, coins, jars and timers, drives commitment to 0.0% on the original cases and transfers to three unseen domains. It does not survive everything: the gate holds exactly when the response format leaves room to reason, and rigid formats that remove that room leave the model confident and wrong on questions it otherwise answers correctly. The gate is trainable and context‑fragile, and deployment needs both halves of that sentence.

Authors:Chuan Fang, Lingteng Qiu, Yixun Liang, Rui Chen, Kunming Luo, Zhaohua Zheng, Tongyuan Bai, Feipeng Tian, Zilong Dong, Zihan Zhou, Ping Tan
Title: SpatialCrafter: Single Image World Modeling with Generative 3D Proxies
Abstract:
Explorable image‑to‑scene generation is essential for applications in gaming, robotics, and virtual reality. Existing methods based on video diffusion model (VDM) commonly rely on incomplete conditioning signals such as sparse point clouds or 2D panoramas, leading to stochastic hallucinations, long‑term drifts and suboptimal 3D consistency. We present SpatialCrafter, a novel two‑stage framework that addresses these issues by introducing a global 3D proxy for high‑fidelity image‑to‑scene generation. Specifically, we decompose the generation process into global proxy generation and appearance refinement. For proxy generation, we propose a Point‑anchored Sparse Structure~(PaSS) Flow module that predicts a spatially aligned and geometrically consistent 3D proxy. For appearance refinement, we re‑frame the VDM as a Generative Deferred Refiner which synthesizes high‑frequency photorealistic details upon proxy‑defined scene geometry. To better integrate the proxy with the pre‑trained VDM, we introduce Parallel Geometry Injection and Proxy‑Aware Corruption training strategies, which improve robustness to proxy artifacts without disrupting the pretrained generative manifold. Furthermore, as no suitable dataset exists for this explorable scene generation task, we construct a new large‑scale dataset of 115K scenes. To the best of our knowledge, it is the first hybrid dataset for image‑to‑scene generation. Extensive experiments on both synthetic and real‑world datasets show that SpatialCrafter outperforms state‑of‑the‑art methods, mitigates long‑term drift, and remains robust and consistent under rapid camera motion and extreme viewpoint changes. Code, models, and the newly constructed dataset will be publicly released. See more at https://fangchuan.github.io/SpatialCrafter/.

Authors:Dhiren Mukesh Khatri
Title: FoldPipe: Bounded Remote Streaming of Native Molecular Shards with Asynchronous Prefetch
Abstract:
Training molecular machine‑learning models on ephemeral or memory‑constrained accelerator instances can require repeatedly retrieving preprocessed molecular graphs from remote storage. FoldPipe is a lightweight Python orchestration layer for already‑sharded PyTorch and PyTorch Geometric data. It retrieves one shard ahead in a background thread while the consumer trains on the current shard, keeping the number of live shard payloads bounded with respect to total dataset size. Asynchronous prefetch and bounded buffering are established systems techniques rather than novel scheduling algorithms. FoldPipe's contribution is a small integration targeted at native .pt molecular shards together with a source‑pinned empirical characterization of its operating regime. We evaluate a SchNet energy‑and‑force workload on MD17 aspirin using 20 paired, order‑alternating benchmark passes on a Tesla T4. Each pass processes five pinned shards containing 25,000 structures. FoldPipe records 16.33 s mean I/O‑compute overlap, compared with zero by construction for the sequential bounded baseline. Mean pass time is 76.78 s for FoldPipe and 83.37 s for the baseline. However, the geometric mean paired speedup is 1.059× with a 95% bootstrap interval from 0.878× to 1.288×. The experiment therefore verifies the overlap mechanism but is inconclusive about a reliable wall‑clock speed advantage under the observed public‑network variability.

Authors:Ante Kapetanovic, Tomislav Duricic, Dionizije Fa, Andro Mercep, Emanuel Lacic
Title: Conversational Recommendation over Live E-Commerce Catalogues with Self-Refreshing Retrieval
Abstract:
Conversational recommender systems based on large language models (LLMs) are usually evaluated on static, pre‑indexed item collections, yet e‑commerce catalogues change continuously as products are added or removed, repriced, and restocked. We present a merchant‑agnostic, multi‑turn conversational shopping assistant that operates over such live catalogues. Its central component is a self‑refreshing retriever that ingests a merchant product feed, enriches the records, and synchronizes them into a vector index. On each run, per‑item hashes identify which products are new, changed, deleted, or unchanged, so only the delta is processed rather than rebuilding the whole catalogue. A controller‑based dialogue layer consumes this index, using an LLM only for intent classification and preference elicitation while retrieval, reranking, and diversity selection run as dedicated functions. Our demonstration is a WhatsApp shopping assistant in which catalogue changes reach the recommendations after the next successful sync. A live chatbot, documentation, and a recorded walkthrough are available at https://github.com/infobip/infobip‑agentic‑crs.

Authors:Rui Xie, Lu Chen
Title: ASIL: Replacing Screenshot-and-Click with Structured State and Semantic Actions
Abstract:
Powerful code agents can execute scripts, call tools, and manage files, yet many important applications remain accessible primarily through graphical user interfaces. We argue that screenshot‑and‑click is an inefficient interface for software‑operating agents: screenshots are state‑incomplete, and GUI actions are brittle, semantically weak, and poorly matched to long‑horizon planning. We introduce ASIL (Agent‑Software Interaction Layer), an agent‑native interface that exposes software through structured JSON observations and code‑executable semantic actions, realized through the deepest feasible access path for each application. We instantiate ASIL across 15 applications and a benchmark of 300 single‑application and 80 multi‑application tasks. ASIL reaches above 80 with closed models while executing fewer than five actions per task. Under a repaired runtime and a 50‑step screenshot budget, the same tasks yield 6.6 and 26.6 strict success under screenshot‑and‑click control, rising to 15.0 and 53.3 on an easier OSWorld‑comparable band. Against application‑native interfaces on matched tasks, ASIL exceeds LibreOffice's UNO API by 28‑38 strict points but only matches draw.io's MCP content contract. The structured modality also suits training: small‑scale SFT raises Qwen3.5‑2B from 58.0 to 72.1 and Qwen3.5‑9B from 66.6 to 80.4, and resource‑limited on‑policy RL further raises them to 74.4 and 82.2.

Authors:Linsen Zhu, Yi Shi
Title: DSA: Evidence-Aware LLM-Agent Orchestration for Multi-Market Stock Research
Abstract:
Large language models can summarize financial information, but an operational stock‑research system must first assemble heterogeneous evidence, expose unavailable data and model capabilities, and control how generated opinions affect a final report. We present DSA, an evidence‑aware orchestration framework for multi‑market stock research with large language model (LLM) agents. DSA organizes the workflow into evidence acquisition, structured context construction, model‑routed analysis, optional role and Strategy Skill reasoning, and report generation with selected context and diagnostics. A default report profile and an optional agentic profile share evidence and model‑routing services but use profile‑specific output validation and risk safeguards. In the agentic profile, core role outputs are processed by role‑specific parsers, whereas Strategy Skill opinions undergo an additional signal‑eligibility partition before synthesis; disagreement is supplied explicitly to the decision agent, followed by a conservative risk override. The reference implementation includes six regional market paths, fifteen bundled Strategy Skills, hosted and local model routes, and multiple execution and delivery surfaces. At a frozen software snapshot, a selected manifest of 1,457 portable offline backend contract tests passed; 596 cases were retrospectively mapped to six contract families central to the reported LLM‑agent architecture. This evidence establishes implementation conformance for the tested software contracts, not superior report quality, forecasting accuracy, or investment returns.

Authors:Geng Li, Yuhao Wang, Dong Li, Jianye Hao, Yuxin Peng
Title: GraphMemix: Query-Aware Evidence Forests for Long-Term Multimodal Agent Memory
Abstract:
Organizing long‑term memory for multimodal agents remains challenging because existing methods either suffer from expensive question‑agnostic offline summaries or naive embedding similarity matching that introduces incomplete and redundant context. To address these issues, we propose GraphMemix, a combinatorial‑optimization graph memory framework that models memory organization as query‑aware evidence‑forest construction. Specifically, our method consists of three key components:(1) candidate graph construction, which expands multi‑view seed memories through schema and semantic relations to acquire query‑aware original context; (2) evidence utility and activation costs, which decouples direct memory support from anchor‑conditioned relation verification to suppress redundant or conflicting information; and (3) forest optimization, which jointly selects a forest‑format memory context under a maximum evidence budget and its reliable relational structure. By organizing memory into a query‑relevant subgraph, the method avoids substantial lifecycle cost and recovers low‑similarity complementary evidence. Experimental results across four long‑term multimodal memory benchmarks demonstrate significant improvements with different foundation models and establish a new Pareto frontier between accuracy and lifecycle cost.

Authors:Qi Lu, Zehui Guo, David Yuanda Gan, Zijing Li, Hengda Zhang, Weijun Xu, Qiankun Zhang
Title: TempJail: Temporal Jailbreak Attacks against Image-to-Video Generation Models
Abstract:
In recent years, image‑to‑video (I2V) generation models have made remarkable progress in subject consistency and temporal coherence, enabling high quality video synthesis. However, these advances also introduce new safety risks. Existing studies mainly focus on jailbreak attacks involving single frame violations, while largely overlooking the temporal dimension unique to video generation models. In this paper, we investigate three attack scenarios and uncover a temporal vulnerability in I2V systems: unsafe semantics may emerge not from a single frame, but from semantic composition over time. We further identify two key challenges in such attacks: temporal abstraction and semantic camouflage. To address these issues, we propose TempJail, a novel temporal jailbreak framework for I2V systems. For temporal abstraction, we decompose a target malicious caption into an initial frame visual condition and a temporal text instruction. For semantic camouflage, on the image side we model semantic injection as controlled latent perturbation in diffusion sampling and introduce gradient guidance from pretrained encoders. On the text side, we rewrite the caption into an innocuous ``subject‑action‑scene'' template that bypasses safety filters while preserving temporal guidance. In the black‑box inference phase, these two modalities jointly enable malicious semantics to be gradually triggered over time. Experiments on closed‑source commercial models, including Kling, Seedance, Veo and PixVerse, show that TempJail improves attack success rate over prior state‑of‑the‑art methods by 23.3% under GPT‑5.2 evaluation and 22.0% under human evaluation. Our codes are available at \hrefhttps://github.com/luqi‑glory/TempJailGitHub.

Authors:Zijian Kan, Wei Wang, Long Luo, Bing Zhao, Xuan Ren, Weixu Qiao, Wenbo Li, Hu Wei, Lin Qu
Title: RubricRM: Generative Reward Modeling via Dynamic Rubrics for Image Generation and Editing
Abstract:
Reward models play an essential role in aligning visual generative models, yet most existing visual reward models use a single scalar score or rely on fixed criteria that cannot adapt to different instructions. This limits both interpretability and task sensitivity, especially for text‑to‑image generation and instruction‑based image editing, where different inputs require different evaluation dimensions. We propose RubricRM, a pairwise generative reward modeling framework that first produces an input‑specific rubric with evaluation dimensions, weights, and scoring criteria, and then applies the rubric to score candidate images. We train dedicated RubricRM models for text‑to‑image generation and image editing using a two‑stage training pipeline: supervised fine‑tuning teaches the model the rubric‑based scoring paradigm, while GRPO further improves scoring through fine‑grained dimension‑level rewards. Experiments on multiple generation and editing benchmarks show that RubricRM outperforms existing specialized reward models and remains competitive with strong proprietary MLLM judges despite using smaller backbones. Our models, data, and code are available at https://github.com/zijiankan/RubricRM.

Authors:Wieland Morgenstern, Friedrich Elias Branschke, Florian Fleischmann, Adrian Szatmari, Paul Schlack, Florian Barthel, Peter Eisert, Anna Hilsmann
Title: KISS-GS: 3D Gaussian Splatting Compression Kept Simple
Abstract:
Scene reconstruction with 3D Gaussian Splatting (3DGS) has become common, however deployment remains painful as the uncompressed file sizes can be massive. Current 3DGS compression systems combine multiple strategies for file size reduction, which can obscure where gains come from and limit component reuse across training pipelines. To make the gains more transparent, we propose KISS‑GS, a modular compression pipeline named after the principle of keeping things simple, designed to decouple compression entirely from training. Given a 3DGS scene reconstructed with vanilla 3DGS, we are able to reduce it through compaction by 15.7x using a combination of state‑of‑the‑art pruning schemes. Then we encode it into an image‑based format designed for simple, ubiquitous decoding. With the SOG‑XT format, we propose a novel extension to Self‑Organizing Gaussians with two main contributions: (i) Self‑organizing 2D Codebooks and (ii) Parallel Representative Assignment Smoothing (PRAS), which leverages the symmetry of quaternion and scale parameterizations to produce 2D attribute grids more amenable to encoding. This encoding reduces scene size by 6.6x. We show that optional encoding‑aware fine‑tuning yields a further 2.2x. Across standard 3DGS benchmarks, our simple and modular approach thus achieves a total of 85x to 319x reductions in the size of the scene over uncompressed vanilla 3DGS, setting new benchmarks for real‑world scenes and surpassing tightly integrated methods in rate‑distortion. Decoding relies solely on web‑native image formats, and the modular design makes each stage easy to combine with future advances in reconstruction and compaction. Code and project page: https://fraunhoferhhi.github.io/KISS‑GS/

Authors:Mohamed Guechaoui, Mohamed Diaa Zellagui, Souleyman Chaib, Sahraoui Dhelim
Title: AraMS-28k: The Largest Publicly Released Line-Level Dataset of Historical Arabic Manuscripts with Margin and Insertion-Anchor Annotations
Abstract:
We introduce AraMS‑28k, the largest publicly released line‑level dataset of genuine historical Arabic manuscripts, comprising 14 books, 3,043 pages, and 28,600 annotated text lines (27,971 main‑text, 629 margin). Thirteen books are hand‑copied manuscripts spanning three script traditions ‑‑ Naskh, Ruq'ah, and Maghrebi ‑‑ and one is a lithographed printed edition included to broaden format diversity. Each line is labelled as main‑text or margin, and margin lines that have an unambiguous attachment point in the main text are further annotated with an insertion anchor, recovering the manuscript's true non‑linear reading order at line‑level granularity ‑‑ to our knowledge the first such annotation released for a historical Arabic manuscript corpus. Because reference transcriptions are fully vocalised while manuscript hands are typically undiacritised, we release both the raw diacritised transcription and a diacritic‑normalised counterpart for every line. The dataset was constructed with RefLAM, a reference‑grounded annotation pipeline that aligns multimodal‑LLM OCR against independently sourced clean transcriptions and routes every line through human review, combining automatic verification with expert oversight. We describe the construction and quality‑control process, present the annotation schema, report dataset statistics at both the corpus and per‑book level, and provide baseline HTR results using Kraken and HATFormer, including a cross‑script generalisation gradient from in‑distribution pages to fully unseen books. AraMS‑28k is released with page images, line‑level annotations, and fixed train/val/test splits under CC BY‑NC‑SA 4.0 to support reproducible research on Arabic manuscript recognition, layout analysis, and reading‑order recovery.

Authors:Ankit Bhattacharjee, Sougata Maity, Santam Chakraborty, Indranil Mallick
Title: Dose-PlanNet: Physics Based Radiotherapy Dose Prediction with Deep Learning
Abstract:
Automating prostate radiotherapy treatment planning is dosimetrically complex, particularly for extreme hypofractionated regimens. In this study, we introduce Dose‑PlanNet, a physics‑guided 3D deep learning architecture designed to predict dose distributions. This model's performance was evaluated on a cohort of patients treated in a prospective trial where two different dose fractionation regimens were employed. Dose‑PlanNet achieved comparable target coverage (D_95), though statistical analysis revealed a marginal reduction in target homogeneity (p<0.001) offset. However the model achieved statistically significant improvements in high‑dose organ‑at‑risk sparing (p<0.001). When evaluated against strict Prospective Randomized protocol volumetric constraints, automated plans met prespecified clinical acceptance criteria in 11 out of 14 Moderate Hypofraction Arm plans and 9 out of 12 Stereotactic Body Radiation Therapy Arm plans. This pipeline demonstrates that physics‑informed deep learning can accelerate radiotherapy workflows while safely maintaining the stringent dosimetric quality required for high‑precision clinical deployment.

Authors:Bohan Yu, Shi-Yang Li, Pengfei Cao, Jun Zhao, Kang Liu
Title: RuleWeaver: Benchmarking Rule-Centered Scenario Reasoning for Large Language Models
Abstract:
Large language models (LLMs) are increasingly applied to specialized domains, where effective use of domain expertise often requires reasoning over complex rules in concrete scenarios. However, existing benchmarks only partially evaluate this capability, as they either focus on output‑level instruction constraints or overlook the distinct roles that rules play in scenario reasoning. To address these gaps, this paper introduces RuleWeaver, a benchmark construction framework for evaluating rule‑centered scenario reasoning. RuleWeaver starts from corpus‑derived IF‑THEN Meta Rules, progressively augments them into complex rules, and composes these rules into rule‑centered scenario QA instances. Beyond final‑answer correctness, RuleWeaver further supports process‑level evaluation through rubric‑based answer quality, rule recall, and rule precision. Experiments on 11 representative LLMs show that current models still struggle with complex rule‑centered scenario reasoning, with even the best‑performing model achieving only around 50% of the maximum rubric score. We make our code and dataset available here: https://github.com/SharkSpicy‑NLP/RuleWeaver.

Authors:Chenyang Wu, Fuchen Long, Binyuan Huang, Xinlong Sun, Xi Chen, Chun-Le Guo, Chongyi Li
Title: Thinking on Shots: Consistent Multi-Shot Video Editing with Agentic Reasoning
Abstract:
While generative AI has significantly advanced video editing, existing methods primarily focus on single‑shot or short video clips. Editing long videos with multiple instructions remains a formidable challenge. Naive chunking strategies, e.g., fixed‑duration segmentation, often lead to entity fragmentation, severe editing hallucinations, and disrupted temporal continuity. To bridge this gap, we introduce the Multi‑Instruction Multi‑Shot Long‑Video Editing (MMLVE) task, which is structured around three core objectives: Cross‑Shot Editing Consistency (CSEC), Multi‑Instruction Decoupling (MID), and Zero‑Destruction on Spatiotemporal Structure (ZDSS). To tackle these three unique challenges, we introduce an agentic editing framework that leverages the synergy of Large Language Models (LLMs) and Vision‑Language Models (VLMs) to achieve shot‑level video decoupling and precise instruction parsing. Furthermore, to comprehensively evaluate this task, we construct MMLVE‑Bench, which is an MMLVE‑focused dataset characterized by complex real‑world spatiotemporal dynamics, high‑density heterogeneous instructions, and sparse, random entity distributions. Three MMLVE‑focused evaluation metrics are further exploited to assess the quality of the editing results. Extensive experiments demonstrate that our MMLVE‑Agent outperforms existing closed‑source SOTA approaches (e.g., Seedance 2.0), successfully eliminating editing hallucinations, preserving cross‑shot editing consistency, and attaining seamless spatiotemporal transitions.

Authors:Zihao Cheng, Yingyu Shan, Hongru Wang, Zeming Liu, Xinyi Wang, Xiangrong Zhu, Yuhang Guo, Wei Lin, Yunhong Wang
Title: Behavior2Trip: Towards Personalized Travel Planning via User Behavior Trajectory
Abstract:
Travel planning agents assist users in generating personalized travel plans by modeling their individual preferences. Existing agents either rely on explicit user instructions or engage in multi‑turn clarification to elicit user preferences. However, both approaches overlook the rich behavioral signals latent in users' past behaviors, which implicitly encode their preferences. This over‑reliance on active user input increases interaction burden and limits plan personalization. To bridge this gap, we introduce a new task, Behavior‑Aware Travel Planning, which infers user preferences directly from past behaviors and generates personalized travel plans. To facilitate research on this task, we introduce Behavior2Trip, a benchmark constructed from one of the largest Chinese online travel platforms, comprising 11,400 instances. Each instance represents an average of 39.8 past user behaviors spanning 14 attributes across 5 preference dimensions. We further propose B2T‑Agent, a reinforcement learning‑based agent that leverages user behavior trajectories, interacts with external tools for preference‑aligned retrieval, and maintains an internal memory module. Experiments on Behavior2Trip show that GPT‑4.1 achieves a full‑constraint pass rate of only 0.5% on the hardest tasks, while B2T‑Agent built upon Qwen3‑8B outperforms all baselines, highlighting the substantial challenge of this task. Moreover, Qwen3‑8B trained with B2T‑Agent also outperforms GPT‑4.1 on the TravelPlanner benchmark, demonstrating strong generalization. Code and data are available at https://github.com/BUAA‑IRIP‑LLM/Behavior2Trip

Authors:Rongyang Zhang, Chengqiang Lu, Cong Li, Hongchao Gu, Tingjia Shen, Xuyang Zhi, Qimeng Wang, Yan Gao, Yi Wu, Yao Hu, Hao Wang, Enhong Chen
Title: Multi-Image Visual Token Pruning in Large Visual Language Models
Abstract:
With the growing demand for processing multiple image sequences in real‑world applications, various visual token pruning methods have emerged to mitigate the computational and context length constraints faced by Large Vision Language Models (LVLMs). However, most existing pruning approaches rely on static strategies that struggle to adapt across different architectural LVLMs and multi‑image scenarios, and are additionally constrained by their dependence on attention computations that are incompatible with efficient techniques like FlashAttention. To address these limitations, we propose a training‑free, Adaptive Visual Token Pruning (AVTP) framework, applicable to diverse LVLM architectures. We strategically determine pruning layers based on empirical analysis of visual attention distributions across various LVLMs, and implement adaptive pruning ratios in multi‑image contexts where images of higher importance retain proportionally more tokens. We conduct extensive experiments across different LVLMs to demonstrate the effectiveness and robustness of AVTP. Specifically, Qwen3VL‑8B achieves 2 times inference speedup while maintaining 96.1% of its original accuracy on multiple multi‑image benchmarks, InternVL3.5‑8B retains 94.1% accuracy, and LLaVA‑OV‑7B even exceeds its original baseline performance. Our code is available at \hrefhttps://github.com/zry13/AVTPthis link.

Authors:Zineng Tang, Kelsey R. Allen, Sjoerd van Steenkiste, Ishita Dasgupta, Alane Suhr
Title: Decoupling Planning and Control for Instructable Agents
Abstract:
Recent work shows that pre‑trained, instruction‑tuned vision‑language models (VLMs) perform well at mapping from instructions and observations to high‑level plans, but struggle to realize such plans as reliable low‑latency action sequences in unfamiliar environments. At the same time, world‑model controllers excel at fast observation‑to‑action control, but lack open‑ended task guidance. In this work, we combine these strengths into a single system, Instruct‑to‑Act, where we train a world‑model controller to act autonomously at high frequency when conditioned on sparse, higher‑latency, and high‑level text instructions generated by a VLM planner. To train controllers to be language‑instructable, we relabel segments of controller policy rollouts with synthetic instructions and jointly optimize a behavior‑cloning objective along with existing reward‑maximizing and world‑modeling objectives. We evaluate our proposed approach across seven embodied environments, including three multi‑agent environments where VLM planners coordinate through language while trained controllers serve as their actuators. Under matched observation and action spaces, our decoupled approach consistently outperforms controller‑only and direct VLM action‑generation variants, preserves fast control, and lets us swap in different pretrained VLM planners without fine‑tuning, while remaining competitive with strong vision‑language‑action and multi‑agent RL baselines on six of seven tasks.

Authors:Seohyeong Lee, Hwaran Lee, Buru Chang
Title: Instruction Quality Matters: Refining Instructions for Effective Preference Learning
Abstract:
Preference learning optimizes models using response pairs, yet the informativeness of these pairs is fundamentally shaped by the instructions from which they are generated. We identify instruction quality as a hidden bottleneck in preference learning: low‑quality or ambiguous instructions restrict the response‑quality distribution, limiting strong chosen responses and weakening preference signals. Through Best‑ and Worst‑of‑N analyses, we show that instruction quality constrains both the ceiling and floor of sampled response quality. Motivated by this observation, we introduce an instruction‑refinement pipeline that selects weak instructions using reward signals and revises them with rubric‑guided LLM feedback, improving preference data without discarding examples. Across offline and online preference learning settings, experiments on multiple models and benchmarks show broad alignment improvements over original data and alternative data‑improvement strategies. Further analyses indicate that instruction refinement raises achievable response quality and complements response‑centric preference data curation. Overall, instruction quality emerges as a key factor governing how informative preference signals are formed for LLM alignment. Code is available at: https://github.com/01choco/instruction‑refinement/

Authors:Markus Frohmann, Mahdiyar Alavi, Elizabeth Lingg, Navid Rekabsaz
Title: Equal Ranking Quality, Different Decisions: Training Order-Consistent LLM Scorers
Abstract:
Rerankers, reward models and multi‑document QA scorers score candidate documents or responses in one LLM prompt, so each score depends on their order. Such scorers are selected on ranking quality, but their scores determine a decision: what a score threshold retains, a reader answers, or a preference model selects. However, equal ranking quality does not imply equal decisions: on passage reranking, five trained scorers within 0.010 nDCG@10 retain sets that overlap by only 0.66‑0.84 when reordered. A published reranker takes the highest retained‑set F1 in our comparison and still overlaps by only 0.667. No prompt‑time change we test removes that order dependence: the only one that gains ranking quality leaves all three decisions unchanged. Order‑consistency SFT (OC‑SFT) attenuates it in the weights, training a candidate's score not to depend on the order. It holds ranking quality and leads every decision‑stability measure among trained scorers on all three tasks: it flips the reader's answer on 0.125 of permutation pairs against 0.149‑0.164 for three other objectives that target order. It is more stable than order‑averaged distillation on 12 base models, and one OC‑SFT permutation retains sets that overlap more than ten averaged off‑the‑shelf permutations. A comparison should therefore report what a threshold retains and a reader answers, not ranking quality alone. Code is available at https://github.com/thomsonreuters/presentation‑dependence.

Authors:Wei W. Xing, Xixi Zhou, Kaiqi Huang, Jiaye Pan, Hong Qiu, Xin Wang, Shan Shen
Title: HOLMES: In-Context Failure-Center Localization for High-Dimensional Yield Estimation
Abstract:
Importance sampling for high‑sigma yield estimation requires locating the failure center from a severely imbalanced sample set. Existing surrogate‑assisted methods rely on iterative gradient‑based training, ill‑posed under extreme class imbalance; model errors propagate into the estimator, causing accuracy collapse in high dimensions. We recast failure‑center localization as few‑shot binary classification: a prior‑fitted tabular foundation model performs gradient‑free in‑context inference in a single forward pass, eliminating the ill‑posed training loop. HOLMES (High‑sigma Optimal Localization via Manifold Estimation and Sampling) pairs this with an SVD‑based anisotropic proposal that captures the local geometry of the failure manifold, and a hit‑rate‑driven adaptive mixing scheme that stabilizes importance weights where conventional adaptation collapses. On 6T SRAM benchmarks spanning D = 108 to D = 1,152, full‑dimensional baselines exhibit accuracy collapse at some dimension, with the strongest baseline reaching 25.8% relative error; PCA+MNIS is additionally evaluated at the two largest dimensions. HOLMES remains within 5.9% across all five configurations with up to 58.8× speedup over Monte Carlo. The code is available on \hrefhttps://github.com/IceLab‑JCIE/ICE006‑Yield‑Holmes

Authors:Jiahui tang, Kuicai Dong, Dexun Li, Hongchao Gu, Haocheng Yu, Wei Han, Chen Zhang, Yong Liu, Hao Wang, Enhong Chen
Title: DEEPCHART: How Far are LLMs from Faithful Data-Science Chart Generation?
Abstract:
Faithful chart generation in real‑world data‑science workflows requires grounding visualizations in scattered evidence, computing chart‑ready quantities, and rendering them accurately. Modern LLMs can produce visually plausible, instruction‑compliant charts, yet data‑level hallucinations remain difficult to detect in long, noisy, and multimodal contexts. To measure this gap, we introduce DEEPCHART, an expert‑annotated benchmark of 1,482 task‑conditioned chart‑generation instances drawn from real‑world scientific papers, financial filings, and ecosystem reports. DEEPCHART formulates chart generation as an Extract‑‑Reason‑‑Visualize pipeline and evaluates source‑data extraction, derived‑data reasoning, and chart rendering stage by stage. Experiments with state‑of‑the‑art models show that visually plausible charts often conceal data‑level hallucinations, with extraction and reasoning errors common in realistic long and multimodal settings. These findings suggest that larger context windows alone are insufficient; faithful chart generation also requires reliable evidence extraction and quantitative reasoning before rendering. Our benchmark and associated resources are available at https://github.com/tangdouer1005/DeepChart.

Authors:Lezhi Yu, Xiaogang Xu, Yuhua Zhou, Shuibing He, Aimin Pan
Title: Beyond Execution: Auditing Experimental Fidelity in LLM-Driven Scientific Research
Abstract:
LLM agents used for scientific experimentation must do more than generate executable code: they must implement the reference method faithfully, design experiments that test the paper's claims, and provide evidence supporting those claims. We show that agents often produce methodological hallucinations: silently reducing datasets or training budgets, replacing failed learning or generative components with lookup or oracle functions, or drawing conclusions from resource‑limited settings where a method's claimed advantage disappears. To detect these failures, we introduce ABE‑Ralph, a reference‑anchored auditing framework that represents claims, protocols, required components, baselines, and metrics as structured experimental constraints, guides implementation through an 8‑step workflow, and performs quantitative, qualitative, and code‑level verification. Across 30 long‑horizon reproduction runs covering 12 machine learning domains, ABE‑Ralph achieves a 93% robust execution rate and identifies five scientific failure modes. In 23 NatureBench discovery tasks, ABE‑Ralph matches or exceeds state‑of‑the‑art performance on 5 tasks. These results show that reliable evaluation of AI scientists must assess whether the experimental design faithfully tests the intended claim and whether the resulting evidence supports it, rather than treating code execution or plausible metrics as evidence of scientific success.

Authors:Yiwei Lu, Ke Xu, Tao Yan, Xiaojun Chang, Radu Timofte, Rynson W. H. Lau
Title: Glass Surface Detection Grounded in 3D Visual Geometry
Abstract:
Glass surface detection (GSD) is critical for scene understanding and reconstruction, and yet remains challenging due to the transparency and reflectivity of glass surfaces. Existing GSD methods typically rely on 2D appearance cues, which may fail in geometrically ambiguous scenes. In this paper, we propose a paradigm shift: grounding GSD in 3D visual geometry to explicitly model the physical existence of glass surfaces. Our method first distills rich 3D priors from the visual geometry grounded transformer (VGGT) and generates glass‑aware 3D representations. It then exploits multi‑tasking learning with a novel glass detection head, consisting of two core modules: a Frequency Self‑Attention Module (FSAM) that identifies glass‑specific spectral features for glass surface localization, and a Geometry Grounding Block (GeGB) that selectively grounds 2D features in 3D geometry for glass surface segmentation. Extensive experiments demonstrate that our method achieves state‑of‑the‑art performance across seven standard GSD benchmarks, generalizes well to video/multi‑modal data, and substantially improves reconstruction in glass scenes. Code is available in https://github.com/YT3DVision/VGGT_GLASS.

Authors:Mingquan Liu, Jiangyu Chen, Hanqun Cao, Xujun Zhang, Pengsen Ma, Xiangru Tang, Shuting Jin, Zhuo Yang, Tianfan Fu, Fang Wu, Xiangxiang Zeng
Title: AgentFold: Closed-Loop Agentic Search for Protein Folding Model Design
Abstract:
Scientific LLM agents have shown promise in literature reasoning, tool use, and experiment planning, but it remains unclear whether they can autonomously improve large, tightly coupled scientific machine‑learning systems through executable code changes and computationally expensive validation. We study this question in protein folding, where progress requires coordinated architectural modifications, multi‑objective evaluation, and domain‑aware interpretation. We present AgentFold, a multi‑agent framework that formulates folding‑model development as a closed‑loop search over executable code variants. Starting from ESMFold, AgentFold proposes hypotheses, implements and debugs code‑level modifications, evaluates model variants, analyzes experimental outcomes, and stores both successful and failed interventions in structured memory. An MCTS‑style policy allocates computational resources across high‑scoring search branches. On an engineering‑scale protein‑folding codebase comprising more than 2,000 lines of code, AgentFold explores approximately 80 model variants using approximately 5,000 GPU‑hours and 170 million LLM tokens. Under a matched computational budget, AgentFold improves the best lDDT by 7.5% over independent Codex proposals and outperforms a random‑search control. Beyond model improvement, the resulting intervention traces reveal recurring empirical design patterns: stable gains tend to arise from early, soft, learnable priors and gated refinement, whereas direct geometric perturbations and geometry‑conditioned feedback often destabilize training. The code and experimental resources are publicly available at https://github.com/lmqfly/AgentFold.

Authors:Zehua Hao, Fang Liu, Qinliang Wang, Yaoyang Du, Xinyan Huang, Puhua Chen
Title: G2D: Generative-to-Discriminative Collaborative Inference for Zero-Shot Image Classification
Abstract:
Zero‑shot classification needs efficient label retrieval and fine‑grained visual reasoning, yet discriminative and generative vision‑language models fail in complementary ways.When CLIP's top‑1 prediction is wrong, the correct label often remains in its top‑K shortlist, making disambiguation rather than recall the key challenge.Standalone generative models, however, are hindered by large label spaces and unconstrained outputs.This complementarity motivates separating broad candidate retrieval from fine‑grained, image‑grounded verification.We propose G2D, a training‑free framework that uses a generative VLM to verify CLIP‑retrieved candidates against the image.Candidate names and CLIP probabilities provide a structured prior for resolving visually similar classes.Fixed confidence routing, entropy‑adaptive candidate sizing, and trie‑constrained decoding focus generative reasoning on uncertain samples and ensure one valid output for each input at test time.Across eight benchmarks, G2D achieves 68.85% average accuracy, versus 59.35% for CLIP and 63.11% for the standalone VLM.Across seven generator configurations, candidate‑set verification improves average accuracy by 1.08‑‑27.42 percentage points.G2D also transfers to DCLIP, WaffleCLIP, and CuPL, supporting a practical interface between discriminative proposal and generative visual reasoning. Code: https://github.com/Harzva/G2D

Authors:Shi Chen, Weifeng Ge
Title: Generative Semantic Scene Completion
Abstract:
Outdoor LiDAR semantic scene completion (SSC) recovers a dense semantic voxel grid from a scan observing 1% of the target volume, under class imbalance beyond 7,000x. We recast SSC as generative semantic scene completion (GSSC): a single discrete‑diffusion formulation in three roles. First, paired sparse‑dense scene synthesis (PS^3) generates matched sparse LiDAR observations with their dense semantic completions, addressing the long tail at its source and yielding the PS^3‑SemanticKITTI corpus we train on alongside SemanticKITTI. Second, semantic‑guided generative scene completion (SGSC) generates the scene from noise with multinomial discrete diffusion, conditioned on the sparse scan through a bird's‑eye‑view semantic map and a sparse 3D feature stream. Third, the same framework instead refines an existing completion in one flow‑matching step: structured source discrete diffusion (S^2D^2). S^2D^2 improves the mIoU of SGSC's own output and every external SSC base tested, without base retraining or test‑time adaptation. On the strongest base, one step without test‑time augmentation reaches 38.8% mIoU on the SemanticKITTI hidden test. To our knowledge that is the best causal, single‑sweep, single‑sample result on that leaderboard, +2.1 pp over the previous best published score under the same restriction. Four correction steps with eight‑view test‑time augmentation reach 39.2%, outside that restriction.

Authors:Rupesh Sapkota, Louis Mozart Kamdem Teyou, Moshood Yekini, Caglar Demir, Axel-Cyrille Ngonga Ngomo
Title: Neural Regression with Embeddings for Numerical Attribute Prediction in Knowledge Graphs
Abstract:
In recent years, transductive knowledge graph embedding models have been applied to tasks such as link prediction and query answering. Although knowledge graphs often contain rich numerical attributes, most embedding models neglect them, limiting their ability to represent real‑world knowledge graphs with diverse information. In this work, we propose a neural regression model (LitEm) that enables transductive knowledge graph embedding models to predict numerical attributes within knowledge graphs. Experimental results demonstrate that LitEm achieves the best or second‑best results on most attributes across FB15K‑237, YAGO15K, DB15K, and Mutagenesis. Furthermore, we propose a co‑training framework that jointly trains state‑of‑the‑art transductive knowledge graph embedding models with LitEm, which improves link prediction performance mainly for bilinear models and simultaneously enables them to predict numerical attributes. In addition, the literal‑awareness evaluation demonstrates that co‑training helps models to encode and exploit attribute information in a "literal‑aware'' manner, suggesting that the observed gains are not merely due to additional parameters. We publicly release our implementation at https://github.com/dice‑group/dice‑embeddings.

Authors:Michimasa Inaba
Title: Beyond Reflection: Affirmation as a Promising Behavioral Marker Associated with Quality in Text-Based Counseling
Abstract:
While AI‑assisted text‑based counseling is gaining attention, it remains empirically unclear which counselor behaviors are associated with higher dialogue quality. Existing research often focuses heavily on Reflection, borrowing frameworks from Motivational Interviewing. To address this gap, we conduct a multi‑layered analysis using KokoroChat, a large‑scale Japanese text counseling dataset conducted by professional counselors and trainees, newly annotated with counselor strategy tags and client distress levels. Our results show that, under the quality indicators used in this study, Affirmation is more consistently associated with session quality than Reflection among the analyzed strategies. Cross‑dataset transfer experiments further suggest that this quality signal can be observed to some extent on ESConv, an English dataset with non‑expert supporters. These findings provide empirical implications for counselor training and emotional support system design. We release the additional KokoroChat annotations and experimental source code at https://github.com/UEC‑InabaLab/BeyondReflection.

Authors:Jibeom Seo, Junghyo Jo, Sonya N. Martin, Taejin Byun
Title: How Does Science Education Research Respond to Sociopolitical Change? A BERTopic Analysis of Korean Research
Abstract:
Research fields do not evolve in isolation: their questions and priorities shift with policy, curriculum reform, and broader social change. Analyzing published literature can reveal not only how a field matures but also how it responds to these conditions. Prior work in science education has focused on identifying research topics and their trends, but paid less attention to the external conditions in which research is produced. We examine Korean science education research from 2008 to 2025, a case in which centralized curriculum revision, government education initiatives, and demographic decline are prominent. Using BERTopic, an embedding‑based topic modeling technique, we identify major topics and temporal trends, and analyze their associations with selected sociopolitical factors. We interpret each topic and distinguish three groups: sociopolitical, subject‑specific, and student‑related topics. Within the first group, science teacher professionalism and curriculum implementation, science education for gifted students, and STEAM education show the strongest associations with sociopolitical conditions, such as government policy initiatives and declining enrollment in science‑gifted education, whereas digital‑based science education does not. The subject‑specific and student‑related groups, by contrast, show no comparable movement and are not linked to the external indicators we examine; this pattern is interpreted as reflecting stronger disciplinary grounding. Taken together, these patterns suggest that a topic's anchoring to policy and practice or to academic disciplines shapes how closely it tracks external change. This helps explain why some research agendas move with their national context while others hold steady, and why the same topic may develop differently across countries.

Authors:Haiyang Xu, Zheng Ding, Zhuowen Tu
Title: RECAP-Forcing: Retaining Content Appearances for Long Video Generation
Abstract:
Long autoregressive video generation faces a fundamental memory challenge: with a finite attention window, a model must decide which information from an ever‑expanding history to retain. Existing methods organize memory temporally, preserving recent frames while compressing or discarding older ones. We instead propose RECAP‑Forcing, organizing memory by appearance novelty. A long video is not merely a sequence of frames, but an evolving cast of subjects, objects, and scenes whose identities must remain consistent over time. We organize memory by retaining the KV cache associated with newly appearing content‑‑such as entering subjects, disoccluded regions, and newly introduced scenes‑‑at the moment it first becomes visible, prioritizing novelty over recency. Memory should scale with the amount of newly introduced content, rather than with video length. This appearance‑indexed memory makes long‑range consistency an explicit property of the memory structure. Our framework unifies two mechanisms under this single principle. At the beginning of a video, when all visible content is novel, an attention sink preserves the initial scene. As the video evolves, an optical‑flow‑based novelty bank extends the same principle by selectively retaining newly revealed content. As a training‑free inference method with no additional learnable parameters, RECAP‑Forcing consistently improves visual quality and semantic fidelity across multiple strong baselines and outperforms existing memory methods.

Authors:Pratham G. Shenwai, Hemant Kumar Singh, Sridhar Ravi
Title: Real-time Unsupervised Object Discovery from Asynchronous Event Streams
Abstract:
Event cameras capture pixel‑level intensity changes with microsecond resolution to produce highly sparse asynchronous data streams. For visual perception in latency‑critical environments, we propose a lightweight, training‑free framework for discovery of moving objects based on spatio‑temporal clustering. This framework is driven by two core contributions. First, a linear‑time Spatio‑temporal Probabilistic Event Filter (SPEF) that introduces an adaptive event acceptance threshold to distinguish salient motion structures from background noise. Second, an Event Morton Code Clustering (EMCC) module that bypasses expensive distance matrix computation to efficiently group events for unsupervised discovery of moving objects. On the E‑MLB dataset benchmark, SPEF achieves the best denoising performance among classical filtering methods and remains competitive with learning‑based approaches without requiring any offline training. On object discovery, EMCC achieves the highest overall accuracy and lowest execution time across the FRED and eTraM datasets, outperforming established density‑based clustering baselines by a substantial margin. Overall, this work establishes a new performance benchmark for classical object discovery in event data, providing a highly scalable, training‑free solution for resource‑constrained visual perception. The code is available at https://github.com/PrathamShenwai/SPEF_EMCC

Authors:Mingqi Gao, Anthony Sicilia, Weiyan Shi
Title: Which Metrics Save the Most Human Annotation? Prediction-Powered Evaluation and Meta-Evaluation
Abstract:
Across various non‑verifiable tasks, human evaluation is reliable but expensive, while automatic metrics are more scalable but often biased. Building on prediction‑powered inference (PPI), we propose prediction‑powered evaluation, a framework that combines limited human judgments with large‑scale automatic scores to obtain data‑efficient system comparisons that are provably unbiased. We develop parametric and non‑parametric procedures, analyze the efficiency trade‑off between paired and unpaired designs, and validate the framework on six WMT datasets. We further introduce the Prediction‑Powered Saving Ratio (PPSR), a meta‑metric that measures how much human annotation an automatic metric can save when used within prediction‑powered evaluation. PPSR directly targets metric utility for prediction‑powered evaluation and yields more discriminative and stable metric rankings than existing system‑level meta‑metrics. Overall, our new paradigm reframes automatic metrics as tools for reducing human annotation cost rather than replacing human judgment, and applies broadly to non‑verifiable tasks.

Authors:Xinxin Zhao, Jinpeng Ye, Bo Wei, Liqin Wu, Mahmoud Hassaballah, Karen Egiazarian, Aura Conci, Victor Hugo C. de Albuquerque, Abdulkadir Sengur, Leszek Rutkowski, Yan Tian
Title: FU-Mamba: A Frequency-Enhanced Dynamic Scanning Framework for Oralscan Image Segmentation
Abstract:
Oralscan image segmentation is essential for computer‑aided diagnosis and treatment planning in digital dentistry. However, existing visual state space models (SSMs) often rely on manually designed scanning orders to flatten image patches into sequences, which disrupts the semantic spatial continuity and hinders coherent feature extraction from key foreground regions. Moreover, elements such as inconsistent lighting, reflective surfaces, and noise during data acquisition disrupt the frequency distribution by diminishing high‑frequency details while enhancing low‑frequency components, consequently hindering the accurate localization of boundaries. In response to these challenges, we introduce FU‑Mamba, an innovative framework that incorporates dynamic scanning and frequency domain enhancement within the SSM architecture. Specifically, the Dynamic Mamba Block (DMB) adaptively learns sampling offsets via a trainable offset prediction network and performs flexible bilinear interpolation, enabling content‑aware scanning that preserves spatial coherence. Furthermore, a frequency domain enhancement block balances spectral components through wavelet‑guided decomposition and spectrum pooling, improving robustness under adverse imaging conditions. Experimental findings indicate that FU‑Mamba attains a notable enhancement in segmentation accuracy, evidenced by a 1.1% increase in the mean intersection over union (mIoU) metric when evaluated on the dental segmentation dataset. Project page: https://byte2bite.github.io/FU‑Mamba/

Authors:Saksham Khatwani, He Cheng, Majid Afshar, Dmitriy Dligach, Yanjun Gao
Title: Surgical Alignment in Knowledge Graph Training for Clinical Diagnosis with Large Language Models
Abstract:
Biomedical knowledge graphs (KGs) offer structured medical knowledge that can ground large language model (LLM) reasoning in clinical diagnosis application, yet how KG signal should be integrated into LLMs remains an open question. We present a systematic study spanning five KG task formulations, three training paradigms, two KGs, and three base LLMs. At the task level, all paradigms improve over the non‑finetuned baseline, but methods with comparable in‑domain accuracy show substantially different knowledge transfer behavior. We introduce Gradient Intervention Density (GID) and Gradient Distortion (GD) to measure how broadly an optimizer modifies the pretrained model. GID and GD together reveal a clear divide: KG‑judgment training under KL regularization produces sparse, localized updates (a regime we term as surgical alignment), while task‑specific SFT produces dense ones. A controlled ablation shows that the objective and KL contribute to sparsity independently, and the paradigms that produce sparse updates also improve reasoning quality, even when their in‑domain accuracy is lower than task‑specific SFT. Assessing KG‑LLM integration thus requires complementing accuracy with optimization‑geometry diagnostics. Our implementation can be found at https://github.com/LARK‑NLP‑Lab/Surgical‑Alignment.

Authors:Pihai Sun, Gang Han, Jingkai Sun, Jiahao Ma, Zeran Su, Zelin Tao, Peiran Liu, Shuai Shi, Wei Cui, Zifan Wang, Jialin Yu, Wen Zhao, Kangning Yin, Jiaxu Wang, Jiahang Cao, Lingfeng Zhang, Hao Cheng, Jian Tang, Yijie Guo, Qiang Zhang
Title: SOLO: Stable Omni-terrain Long-Horizon Perceptive Humanoid Locomotion
Abstract:
Humans traverse complex terrain over long distances without losing balance, whereas perceptive humanoid policies become fragile as perception and control errors accumulate. We present SOLO, a unified framework addressing two compounding causes of this long‑horizon fragility: dense terrain reconstruction smooths action‑critical details, and pointwise imitation lacks temporal credit assignment. Its Query Reconstructor (QR) uses Fourier‑encoded cell queries to retrieve spatially specific evidence from depth‑proprioception tokens, preserving sharp terrain boundaries. Trajectory‑Aware MSE (TA‑MSE) Distillation adds next‑state teacher‑student disagreement to the PPO reward, enabling Generalized Advantage Estimation to propagate future disagreement penalties to preceding actions. In simulation, QR reduces height‑map L1 error by factors of 3.3‑4.0, while TA‑MSE surpasses PPO and MSE+PPO in curriculum progression. On stress‑test terrains, SOLO achieves 97.5% mean traversal success and 96% stepping‑stone success, versus 75.0‑75.6% and 0‑3% for dense‑reconstructor variants. Deployed zero‑shot with only a chest‑mounted depth camera and proprioception, SOLO completes a continuous 1.5‑km outdoor route and an indoor mixed‑terrain course. Project page: https://sunpihai‑up.github.io/solo/

Authors:Jun-Hui Liu, Kun-Yu Lin, Yi-Lin Wei, Xu-Han Chen, Yinghao Li, Zhuohao Li, Yuan-Ming Li, Qing Zhang, Xiaoyi Fan, Dongmei Jiang, Yan Li, Wei-Shi Zheng
Title: TrapVLA: Trapping Vision-Language-Action Models in Configured Failure Modes
Abstract:
This work introduces Configured Failure Trapping, a novel backdoor attack task against Vision‑Language‑Action (VLA) models, which aims to activate attacks through stealthy textual triggers and induce configured failure modes. Unlike prior backdoor attacks that treat any task failure as a successful attack, Configured Failure Trapping requires the attacker to control how the robot fails (e.g., causing the robot to grasp with a specified positional offset), making it substantially more challenging and hard to detect. To support the new task, we propose an effective data engine for synthesizing high‑quality target trajectories and an automated suite for measuring configured‑failure fidelity. Then, based on this foundation, we construct two new benchmarks, namely Trap‑LIBERO and Trap‑RoboTwin, that instantiate Configured Failure Trapping across four representative failure modes. To address this task, we identify sparse action deviation as a critical challenge and accordingly propose a novel method named TrapVLA, which explicitly learns trigger‑induced action residuals to steer the policy toward the configured failure behavior. Extensive experiments across simulation benchmarks and real‑world robotic settings show that TrapVLA effectively injects configured failure modes into VLA models while largely preserving performance on clean data. Project page: https://john‑liua.github.io/TrapVLA/

Authors:Guopeng Li, Yiyang Duan, Yiru Jiao, Chengcheng Xu
Title: Arrive and Survive: Scaling Safe Goal-Conditioned Policy Learning from One-Bit Failure Signals
Abstract:
Contrastive reinforcement learning (CRL) scales effectively in goal‑conditioned tasks by casting policy learning into a self‑supervised contrastive objective. However, in a failure‑terminated Markov decision process, established CRL considers pre‑failure future goals only when constructing positive samples, without accounting for the probability mass removed by failure termination. Our theoretical analysis shows that this omission induces a systematic overestimation bias in goal‑reaching values. Consequently, near‑failure trajectories provide disproportionately strong supervision of success despite retaining little future occupancy. Unsafe actions can thereby be reinforced through catastrophic failure bootstrapping, leading to failed policy learning and unsustainable goal‑reaching behaviours. To address this problem, we introduce two minimal yet strong corrections: mass‑weighted InfoNCE corrects the overweighting of short surviving futures in critic learning, and a log‑survival‑mass score restores the missing survival mass in policy optimization. The resulting method, Safe Contrastive Reinforcement Learning (Safe‑CRL), requires only the one‑bit signal provided by failure termination to scale safe goal‑conditioned policy learning. Across twelve failure‑prone robot navigation and locomotion tasks, Safe‑CRL consistently improves survival and substantially outperforms the Scaling‑CRL baseline in goal‑reaching performance. Additionally, deep Safe‑CRL policies exhibit complex failure‑avoidance behaviours. This study completes the CRL theory under failure termination and provides a scalable safe RL framework. The code is available via https://github.com/RomainLITUD/safe‑crl.

Authors:Zhuochun Li, Yuelyu Ji, Yiming Zeng, Daqing He
Title: SPEAR: Distilling Domain-Adaptive Reasoning Skeletons via Sequential Symbolic Alignment in Reinforcement Learning
Abstract:
Reinforcement learning‑based knowledge distillation has the potential to transfer complex reasoning from teacher to student models, yet it currently faces a critical dilemma: researchers must choose between sparse outcome‑based rewards, which provide insufficient logical guidance, or expensive neural Process Reward Models (PRMs) for dense signals. We resolve this by introducing SPEAR (Symbolic Process Evaluation and Alignment Reward), a training‑free and plug‑and‑play process reward method for sequence‑level on‑policy distillation. SPEAR projects natural‑language reasoning traces into domain‑adaptive symbolic milestones, providing an efficient proxy for process‑level reasoning alignment. By utilizing the longest common subsequence (LCS) to align student explorations with teacher milestones, SPEAR provides a dense, order‑aware reward signal that enforces logical consistency without the need for an external neural verifier. Our experiments across math, science, and commonsense reasoning tasks demonstrate that SPEAR effectively bridges the reasoning gap between student and teacher models via sequence‑level distillation with efficient dense process rewards. Our code and data are available at: https://github.com/zhuochunli/SPEAR.

Authors:Maximilian Du, Zhanyi Sun, Chen Xu, Paarth Shah, Masha Itkina, Shuran Song
Title: Memory Anchors for Continual Robot Learning
Abstract:
Robot policies deployed in the wild should have the capability to continually learn new tasks without forgetting existing behaviors. A common approach to combat such catastrophic forgetting is to train on new task data with a replay buffer of previously learned task data. Although this buffer is commonly sampled randomly from all prior experiences, we show that a small set of these experiences contributes greatly in anchoring past performance. We call these experiences Memory Anchors. We identify Memory Anchors in regions where representations of new‑task observations collapse onto those of old‑task observations even though the tasks require conflicting actions, like when a familiar object must be manipulated in a new way. Rehearsing old data in this region plays a key role in preventing destructive overwriting of past task knowledge, serving as this critical Memory Anchor role. Excluding only 10% Memory Anchors before sampling the buffer leads to more than a 4.5x increase in catastrophic forgetting on the LIBERO benchmark suites. Conversely, enriching the replay buffer with Memory Anchors can decrease high‑conflict task forgetting by 63% and enables successful continual learning of two task sequences on a real robot. Videos and additional visualizations can be found at https://robot‑adaptation.github.io/MemoryAnchors

Authors:Fei Ma, Zebang Cheng, Minghui Li, Hongbo Xu, Yuyong Tan, Yihua Shao, Hanling Wang, Zhou Liu, Yuqing Gao, Dong Wang, Long Ma, Laizhong Cui, Nicu Sebe, Qi Tian
Title: HUG-VIS: A Multimodal Benchmark for Human-centered Understanding and Generation in Visual Intelligence
Abstract:
Visual intelligence seeks to perceive, interpret, and synthesize the visual world and is central to modern computer vision. Human‑centered visual intelligence is especially demanding because it studies people as expressive, socially situated subjects whose meaning is rarely conveyed by appearance alone. It couples vision with audio and language across four representative tasks: human emotion recognition, human video generation, human voice cloning, and human video matting. Yet existing resources remain task‑specific, providing modalities and annotations for individual problems rather than a shared foundation coordinating understanding and generation. This limits multimodal signal use and broader research. We address this gap with HUG‑VIS, a unified benchmark for Human‑centered Understanding and Generation in Visual Intelligence. It contains 8,400 seated half‑body videos of 30 professional actors, each performing the same 280 emotion‑action‑prompt assignments under a controlled Mandarin studio protocol, with synchronized video, audio, text, and alpha mattes. We evaluate diverse open‑ and closed‑source models across the four tasks under a unified zero‑shot protocol using automatic metrics, criterion‑specific mean opinion scores, and multiple cross‑task analyses. Results show that (i) linguistic content dominates current emotion recognition, while purely visual affect recognition is weakest; (ii) in video generation and voice cloning, automatic metrics and human judgment agree overall but differ in their top rankings, requiring joint reporting; (iii) boundary fidelity under motion is the main remaining obstacle for human matting; and (iv) task difficulty varies across emotions, models, and metrics, with notable cross‑task correlations. The dataset and results are available at https://github.com/GML‑MMGroup/HUG‑VIS.

Authors:Huanhuan Ma, Henry Peng Zou, Chengze Li, Enze Ma, Yunyue Su, Philip S. Yu
Title: Sycophancy Suppression Can Impair Rational Updating: Anti-Sycophancy Should Preserve the Ability to Update
Abstract:
Large language models often exhibit sycophancy, revising their answers to align with users when users push back. Such answer flips, however, can arise from different causes. One possibility is that the model simply aligns with the user's feedback in order to satisfy them. Another is that the feedback genuinely contains useful evidence, prompting the model to update its answer in a rational way. We distinguish them as Unsupported‑Yielding and Rational‑Updating. Prior work focuses primarily on suppressing Unsupported‑Yielding, while overlooking its effect on Rational‑Updating. We address this gap with a two‑turn evaluation framework that measures the two behaviors separately. Across representative training‑time and inference‑time interventions, we find that anti‑sycophancy methods often encounter a trade‑off in which reducing Unsupported‑Yielding can sacrifice Rational‑Updating, and vice versa, even when the two objectives are optimized jointly. Mechanistic analysis suggests that the two behaviors share an internal substrate: the MLP neurons and attention heads driving them overlap substantially, and their associated steering directions are positively aligned. We further conduct a preliminary orthogonalized steering exploration, which yields modest, backbone‑dependent selectivity gains. Overall, our results suggest that anti‑sycophancy should be treated not as a simple suppression problem, but as a selectivity problem, where effective interventions should preserve Rational‑Updating while reducing Unsupported‑Yielding.

Authors:Greg Kocher, Robert West, Clément Dumas, Julian Minder
Title: Diff Mining: Logit Differences Reveal Finetuning Objectives
Abstract:
Finetuning has become the gold standard for refining existing behaviors and inducing new ones in language models, yet it often remains unclear exactly which behaviors emerge during this process. As models grow ever more capable, understanding finetuning better becomes increasingly important, particularly since unwanted behaviors may arise during finetuning. In this paper, we introduce Diff Mining, a simple yet effective framework for identifying what a finetuned model has learned by comparing its logits to those of its base model. Diff Mining effectively surfaces salient tokens that are amplified in the finetuned model, serving as a fingerprint of its training ‑‑ even on text unrelated to the finetuning domain. Unlike many existing model diffing methods which require model internals, Diff Mining only needs access to output logits and scales to large models. The framework consists of two modular stages: (i) extracting per‑context logit differences between the finetuned and base models on a reference corpus, and (ii) aggregating the resulting signals to construct an interpretable token set representing the finetune. For aggregation, we explore both a simple Top‑K frequency method and a Non‑negative Matrix Factorization (NMF)‑based approach for disentangling multiple finetuning objectives into distinct token clusters. Empirically, Diff Mining succeeds across diverse settings: on finetune domain detection, it significantly outperforms state‑of‑the‑art model diffing methods both in identifying relevant tokens and in downstream performance when an interpretability agent is given access to the extracted token set; on models with injected biases, it identifies more than one third of the biases without targeted probing. Overall, our framework shows promise in developing auditing tools to detect finetuning objectives.

Authors:Sajal Regmi, Siddhartha Pudasaini, Chetan Phakami Pun
Title: Vowel Signs Are Not Letters: A Pre-tokenization Ceiling on Multilingual Tokenizer Fertility
Abstract:
Byte‑level BPE tokenizers that use the HuggingFace ByteLevel pre‑tokenizer inherit GPT‑2's word regex, where a word is defined as \pL+, one or more Unicode letters. In abugida scripts, vowels are written as combining marks; this pattern therefore splits each word at every vowel sign. Since BPE merges only within a pre‑token, those splits persist through training regardless of vocabulary size or corpus composition. We formalise this effect as a training‑free lower bound on fertility. Across 26 languages from a parallel corpus, every one of the 17 abugidas is affected, ranging from 1.47x (Tibetan) to 9.02x (Thai), whereas Latin, Cyrillic, Hangul, and Han show exactly 1.00x. For 5 languages, matched tokenizer pairs that differ only in this character class fall within 2.2% of the predicted floor, scoring 4.78 versus 1.58 tokens per word on Nepali. When the Nepali share of the training corpus is swept from 5% to 95%, the broken tokenizer barely shifts at all (1.7%) while the fixed one shifts 33.9%, which separates a structural ceiling from a data shortage without needing to inspect any code. We train three 268M models that differ only in their tokenizer; the fixed variant achieves 4.43% lower held‑out Nepali bits per byte at equal compute, and it still leads when given the same bytes with 1.59x the compute. A census of 3,479 HuggingFace repositories finds the letters‑only word class present in 63.3% of the most‑downloaded text‑generation models, accounting for 72.5% of their downloads. GPT‑4o's o200k pattern already uses a mark‑aware word class, making the repair itself prior art. We quantify its value, show how to recognise its absence from symptoms alone, map which scripts it reaches, measure how widely it is deployed, and release a 65,536‑entry Nepali‑English tokenizer with a harness that regenerates every number here from public data on a laptop.

Authors:Emmanuel C. Ugwuabonyi, Dmitri Perkins
Title: IoMT-SecAlarmBench: A Counterfactual Benchmark for Integrity Attacks in IoMT
Abstract:
The Internet of Medical Things (IoMT) combines clinical physiological data with cyber‑system information, creating challenges in determining whether an abnormal reading reflects a genuine physiological event, a device fault, or a cyber‑attack within the expected physiological range. Answering this requires counterfactual ground truth, which no existing dataset provides. We present IoMT‑SecAlarmBench, a semi‑synthetic benchmark that injects controlled integrity attacks into genuine coupled ECG+PPG recordings using a structured experimental design combining four attack morphologies, four severity levels, two physiological plausibility conditions, and replay attacks. Each injected window retains its cause, attack subtype, and the clean signal that would have been observed without the attack. We evaluate six detectors from five method families using threshold‑independent measures and a matched false‑alarm budget. Results show no method consistently detects the most difficult cases: replay attacks and low‑amplitude transient spikes remain close to chance‑level performance across detectors. Results also reveal a trade‑off between detecting attacks and distinguishing them from sensor faults: the best‑performing detector on hard cases flags fault/artifact windows at 5.3 times its false‑alarm rate on normal data. Three‑way classification performs poorly for genuine physiological events, and a leakage audit of a dual‑modality network dataset indicates previously reported IoMT intrusion‑detection performance is partly driven by identifying information. Benchmark, generation code, preprocessing, evaluation tools, and datasheet are released.

Authors:Zishan Shao, Lixun Zhang, Kangning Cui, Wenhao Wu, Jinhee Kim, Yixiao Wang, Ting Jiang, Hancheng Ye, Qinsi Wang, Fan Yang, Danyang Zhuo, Yiran Chen, Hai Li
Title: LowRankArena: A Standardized Evaluation Platform for SVD-Based LLM Compression
Abstract:
SVD‑based low‑rank compression has become a fast‑growing direction for reducing the memory and computational cost of large language models (LLMs). However, meaningful comparison across existing studies remains difficult as prior evaluations use varied benchmarks, inconsistent ratios, and diverse setups, often failing to isolate low‑rank effects from auxiliary techniques. As a result, it remains unclear whether reported gains reflect method‑level improvements or differences in evaluation protocol. This lack of comparability highlights the need for a unified, reproducible evaluation platform. To address this problem, we present LowRankArena, a standardized evaluation platform for SVD‑based LLM compression. LowRankArena unifies task versions, uniform‑precision compression budgets, comparison regimes, and inference measurements, and provides a reproducible pipeline with over 3 TiB released compressed checkpoints. Using LowRankArena, our aligned audit of five representative SVD methods reveals that prior findings are highly conditional under standardized protocols: clear leaders and performance tiers shift across backbones and keep ratios, multiple‑choice accuracy can hide large perplexity degradation, and nominal low‑rank savings yield workload‑dependent and often limited end‑to‑end speedups. Our code is available at: https://github.com/Zishan‑Shao/lowrankarena.git.

Authors:Ryan Thomas Noonan, Linxi Zhao, Menghan Xu, Akanksha Sarkar, Mihir Mishra, Dongyoung Go, Kilian Q. Weinberger, Yoav Artzi, Jennifer J. Sun
Title: Co-Evolving Structured Knowledge and Reasoning in Language Models
Abstract:
Retrieval‑augmented methods improve factual accuracy by grounding language models in external knowledge, but retrieving over unstructured text often introduces irrelevant context and offers limited control over the retrieved information. Structured knowledge bases offer a more controllable alternative, yet they are expensive to construct and often brittle to reason over. To address these limitations, we propose KBevo: a co‑evolving framework that jointly learns to construct a structured knowledge base and reason over it for knowledge‑intensive question answering. By optimizing both components end‑to‑end with QA outcome rewards, our method enables reasoning success to directly improve the quality of the constructed knowledge base. This leads to larger, better‑connected knowledge structures with higher answer reachability, while also improving compositional factual reasoning and controllability compared to standard retrieval baselines.

Authors:Alden Do Rosario, Hussein Younes, Felipe Pires
Title: Why RAGs Hallucinate: Penalty-Aware Evaluation of Retrieval-Augmented Generation Systems with Knowledge-Gap Canaries
Abstract:
Volume‑based accuracy rewards retrieval‑augmented generation (RAG) systems for guessing: a system that answers everything outscores one that declines when its knowledge base cannot support an answer. Building on the confidence‑target analysis of Kalai et al. (2025), we present a penalty‑aware evaluation framework for deployed RAG products, combining (i) asymmetric scoring (correct +1, wrong ‑4, abstain 0), (ii) knowledge‑gap canaries, questions whose answers are verifiably absent from the knowledge base, so that any answer constitutes ungrounded generation from parametric memory, and (iii) a failure‑attribution pipeline that separates retrieval, generation, and abstention‑policy failures. Applying the framework to three commercial RAG systems and a no‑retrieval baseline on SimpleQA‑Verified (1,000 questions x 3 repeats, graded blind by a cross‑family three‑judge panel with 98.9% unanimity), we find that accuracy when answering is closely clustered across systems (97.0‑98.0%), while canary violation rates differ roughly sixfold (16.7% vs. 98.1%). The systems are separated less by what they answer correctly than by whether they answer at all when they should not, and penalty‑aware scoring reorders the volume‑based ranking accordingly; the reordering is stable across penalty settings from k=1 to k=9. All code, configurations, transcripts, and judge votes are released for independent audit.

Authors:Luca L. Weishaupt, Simone de Brot, Javier Asin, Llorenç Grau-Roma, Nic G. Reitsam, Andrew H. Song, Dongmin Bang, Stefan T. Kaluziak, Long Phi Le, Jakob Nikolas Kather, Faisal Mahmood, Guillaume Jaume
Title: VIPER: An Expert-Curated Benchmark for Vision-Language Models in Veterinary Pathology
Abstract:
Pathology vision‑language models are advancing rapidly, yet existing benchmarks remain focused on human tissue, particularly oncology, leaving non‑human pathology largely unaddressed. This gap is especially important in toxicologic pathology, where microscopic tissue examination of laboratory animals is a core component of preclinical drug safety assessment. To address it, we introduce VIPER, the first expert‑curated benchmark for vision‑language model evaluation in toxicologic pathology. VIPER contains 1,251 questions associated with 419 H&E‑stained rat histology images across seven organ systems, covering multiple‑choice, KPrim, and free‑text formats. All questions were curated and validated by board‑certified veterinary pathologists. In total, we benchmarked 16 models, including two newly introduced veterinary‑pathology models, seven human pathology‑specialized models, and seven general‑purpose frontier models. The results identify a substantial domain gap between veterinary and human pathology, expose the risk of over‑diagnosis of normal tissue in frontier models, and show that domain‑specific training remains critical for visually grounded predictions. VIPER data and evaluation code are available at https://github.com/mahmoodlab/viper.

Authors:Wei Sun, Marie-Francine Moens
Title: Cross-lingual Representation Learning via Centroid Intervention Fusion
Abstract:
Large language models (LLMs) exhibit uneven multilingual performance, especially when dealing with low‑resource languages. Inference‑time intervention offers a lightweight way to improve cross‑lingual transfer by modifying the hidden states produced by the LLMs during the forward pass, without updating model parameters. However, existing cross‑lingual intervention methods typically learn separate projections from source to target languages, which limits scalability and prevents knowledge sharing across languages. We propose Centroid Intervention Fusion (CIF), a projection fusion framework that consolidates multiple multilingual intervention projections into a single language‑shared operator. Across multilingual commonsense reasoning, natural language inference, factual editing, and machine translation benchmarks, CIF outperforms the strongest prior pairwise intervention baseline by up to +3.378 pp on average across four model backbones, while supporting performance gains for low resource languages. The code is available at https://github.com/VRCMF/CIF.git.

Authors:Baixuan Xu, Yinyui Xu, Tianshi Zheng, Zhaowei Wang, Weiqi Wang, Haochen Shi, Jiayu Liu, Qing Zong, Xiyu Ren, Xinyu Geng, Zhitao He, Yangqiu Song
Title: Finding the Right Evidence: Factor-Guided Coarse-to-Fine Reasoning for Long Videos
Abstract:
While LVLMs rapidly improve, long‑video question answering still remains challenging: relevant evidence is sparse, and question‑relevant context often fails to provide cues that discriminate the correct answer from plausible alternatives. Diagnostic analysis on a manually annotated subset of MMR‑V shows that prior agentic systems substantially improve cue retrieval over direct VLM inference yet fail to achieve a corresponding gain in answer accuracy, indicating that the bottleneck lies in option‑discriminative evidence rather than topical relevance alone. We propose PACE (Progressive Acquisition of Critical Evidence), a factor‑guided framework for long‑video evidence acquisition. PACE proceeds in two stages: it first indexes clip‑level descriptions guided by question‑derived factors without observing the candidate answers; it then uses the candidate answers to derive contrastive cues and queries the index for verification. On MMR‑V with the open‑source Qwen3‑VL backbone, PACE achieves 42.6% accuracy, outperforming direct inference and prior agentic baselines including Deep Video Discovery (DVD). On the same diagnostic subset, PACE recovers 66.9% of the annotated cues, providing empirical evidence that its gains are associated with improved evidence recovery rather than stronger answer‑side priors alone. Consistent gains over DVD on LVBench, Video‑MME, EgoSchema, and LongVideoBench suggest that option‑aware evidence acquisition transfers beyond MMR‑V. Code is available at https://github.com/HKUST‑KnowComp/PACE.

Authors:Shashank A. Deshpande, Jonathan P. How
Title: Dispersive Forward Tree Search for Optimal Control: Coverage, Complexity, and Computation
Abstract:
Steering‑based planners require solutions to state‑to‑state boundary value problems, which can be inaccessible for nonlinear platforms. Forward propagation evades the steering requirement, but the finite‑sample behavior of the associated planners remains uncharacterized and their implementations underperform in practice. This paper develops a propagation‑based kinodynamic planner with deterministic finite‑sample near‑optimality guarantees. We work within the large class of differentially flat nonlinear systems and show that a forward tree of locally dispersive control commands contains a near‑optimal trajectory at a certified tree size. We provide a general mechanism to construct dispersive command sets for control‑affine systems, which are necessary to implement the search algorithm prescribed by the theory. We show that covering the certified trajectory class irrespective of cost provably demands a tree exponentially sized in the problem horizon, and present a cost‑conditioned dominance pruning procedure that retains near‑optimality at a tree size polynomial in the horizon. We implement the resulting search algorithm, Dispersive Forward Tree search (DFT), as breadth‑first expansion of the forward tree, which maps naturally onto parallel hardware. We design efficient dispersive samplers for the unicycle, the trailer car, and the quadrotor and evaluate challenging planning tasks for these platforms. DFT delivers consistently competitive and often substantially better solution quality than state‑of‑the‑art kinodynamic planners at comparable solution times on embedded‑tier processors, accelerating further as parallel compute is scaled. We also implement DFT in a receding‑horizon loop to demonstrate real‑time planning in dynamic environments at embedded‑tier compute budgets.

Authors:Aditya Pratap Singh
Title: On Scope Classification and Current Knowledge-Editing Benchmarks: A Negative Result, with INLAY as a Gradient-Free Case Study
Abstract:
Every memory‑based knowledge editor in the SERAC lineage depends on a scope decision: given a query, does a stored edit apply? We report that current knowledge‑editing benchmarks cannot measure this decision at all. Using INLAY, a gradient‑free editor we built to obtain exact per‑query ground truth (the model is frozen, edits live in an external addressable memory, and applying an edit is a bias added along one token's unembedding direction at decode time), we execute every candidate router action on 1,689 queries spanning three datasets and three input conditions. An oracle router choosing the best action every time ties a one‑line static policy to four decimal places in all nine dataset‑by‑condition cells: the maximum attainable gain of any per‑query routing method is 0.00 points. Abstention is the sole winning action zero times out of 1,689. The cause is structural: these are counterfactual benchmarks whose evaluation question asks for the post‑edit answer, so answering from parametric knowledge is wrong by construction, and a benchmark without negatives cannot reward a classifier's ability to reject. This generalizes beyond our system to the whole scope‑classifier family the benchmarks are used to evaluate. We confirm the mechanism directly: constructing the missing condition ourselves, by withholding a query's own edit from the index for half the sample, moves pooled headroom from exactly +0.0000 to +0.0420 and gives abstention its first wins. We also report where INLAY itself does not win (WISE beats it on Qwen2.5‑7B CounterFact, and retrieval‑augmented generation beats every method we tested, INLAY included, on rigorously matched RippleEdits), and disclose two bugs found during a self‑audit of our own routing machinery, neither of which changed a published headline number outside noise.

Authors:Youtian Lin, Yikang Yang, Zhanpeng Hu, Mengqi Zhou, Feihu Zhang, Xun Cao, Jiaheng Liu, Yao Yao
Title: Procedura: Agentic 3D Modeling with Procedural Control
Abstract:
Native 3D generators now recover impressive mesh geometry from a single image. However, a dense mesh stays soft where a machined object should be sharp, it carries no part decomposition, and it exposes no parameter a user could edit. To address this, we explore the paradigm of 3D shape as code, leveraging and scaling the coding ability of an LLM for 3D modeling. We introduce Procedura, a novel 3D modeling agent framework that writes an object as a procedural assembly, a parametric program whose named parts are joined by typed, machine‑checkable mates. From a text prompt, the agent plans the object as an assembly graph and writes the program part by part, solving each placement from the mated frames rather than guessing it, and admitting a part only once compile, mate, and connectivity checks pass. A decoupled vision critic then refines the assembly one diagnosed fix at a time. Moreover, the same graph carries per‑part materials and a simulator‑validated articulation. We evaluate on P3D‑Bench under its assembly judge, and with the same judge on MechBench‑36, our hard‑surface benchmark. On both, Procedura outperforms state‑of‑the‑art native 3D generators and every prior 3D‑code agent on judged quality, produces the sharpest edges of any method we evaluate, and is the only one whose output is an editable, part‑structured program.

Authors:Ross Williams, Niyousha Hosseinichimeh
Title: Prompt Sensitivity of Generative Agents: Evidence from an Epidemic Model
Abstract:
As generative AI gains traction, researchers are investigating its potential to serve as proxies for humans. From undergoing cognitive psychology experiments to experiencing an epidemic, generative agents, agents powered by generative AI models, produce realistic human behavior when prompted. This study explores the sensitivity of these generative agents' behavior to prompt modifications and varied persona names of the agents. To assess this sensitivity, we use a generative agent epidemic model, wherein each agent is prompted daily on whether it wants to isolate or commingle with other agents. We found that using synonymous prompts results in negligible changes to the model's outcomes. However, minor variations in prompts, as well as contextual changes, do influence the model's results. Lastly, our data indicates that different persona names assigned to generative agents, specifically those imbued with personas, do not significantly impact epidemic outcomes.

Authors:Sydney Lewis
Title: Same Model, Different Harness: Different Coding-Agent Results
Abstract:
A coding agent combines a model with a harness, which decides what the model sees, which tools it can use, and how the work continues. We ask whether changing the harness changes the result when the model and task stay fixed. We compare two configurations of the same harness on three coding benchmarks. The control supplies the full conversation in time order, while the treatment keeps the same record but mechanically shortens older tool results as the context fills and responds to repeated or stalled work. Under tight context, the treatment raises mean per‑task fail‑to‑pass fraction (F2PF) in all three pressure comparisons and increases complete solutions on SWE‑bench Verified and SWE‑bench Pro. The tight‑window Verified comparison uses 169 tasks, a 20,480‑token window, and a fixed 480‑second attempt endpoint; on this cohort, treatment raises mean per‑task F2PF from 28 percent to 49 percent and complete solutions from 43 to 72. Without model‑specific retuning, the same frozen treatment also raises both endpoints on the same cohort for three additional models with different designs. In the wide‑window Qwen3.6 comparisons, observed arm outcomes are close on Verified and Pro, while FeatureBench retains a higher mean per‑task F2PF under treatment. On the wide‑window Verified cohort, treatment also serves fewer prompt tokens per turn. Because changing the harness changed what unchanged model weights could accomplish, coding‑agent evaluations should treat the model and harness together as the tested solver.

Authors:Yuncheng Guo, Zhanqiu Zhang, Yiwen Guo, Weijia Li
Title: GameWAM: A World Action Model for Video Games
Abstract:
Modern video games combine first‑person perception, rapid visual changes, persistent world state, and heterogeneous native controls. Existing game agents map visual and task context directly to actions but lack explicit world dynamics modeling, whereas interactive game world models predict visual futures from supplied actions but do not serve as task policies. World‑Action Models (WAMs) unify these objectives, but remain largely unexplored under the dynamics and open‑ended interaction of video games. We introduce GameWAM, to our knowledge the first WAM for native closed‑loop gameplay and GUI control. GameWAM jointly generates future visual observations and executable keyboard‑mouse trajectories through parallel visual and action generative processes with block‑causal conditioning and flow matching. To support joint world‑action learning, we construct synchronized gameplay and GUI trajectories. To handle heterogeneous native control, GameWAM predicts a gameplay/GUI mode at each action step and generates actions with mode‑specific prediction distributions and continuous‑action normalization. For long‑horizon interaction, block‑cycle control predicts beyond the committed horizon, executes only a short action prefix, and replans from new observations, while fine‑grained within‑cycle context and hierarchical cross‑cycle history preserve temporal continuity. Experiments demonstrate competitive task success with fewer executed native actions than the compared agents. We further uncover Low‑Frequency Action Source Imprinting (LASI), in which low‑frequency components of the sampled action source systematically steer coarse generated camera motion under fixed conditioning, revealing a source‑sensitivity failure mode in generative control. Project page is available at https://yunncheng.github.io/GameWAM/.

Authors:Yibo Wang, Rui Yang, Jisheng Dang, Bimei Wang, Yitao Wu, Pengfei Cao, Wencan Zhang, Hong Peng, Bin Hu, Tat-Seng Chua
Title: AffectOmni: RL-Verifiable People-Centric Grounded Affective Reasoning for Social and Art-Related Scenes
Abstract:
Multimodal large language models (MLLMs) achieve strong performance on VQA and scene understanding, yet affective reasoning remains vulnerable to shortcut behavior. Models may predict correct answers while neglecting people‑centric cues such as micro expressions and body language, which weakens traceability and external verification. Prior reinforcement learning approaches mainly reward context or logical coherence without explicitly enforcing attention to human evidence. In addition, LLM as a Judge scoring often suffers from score clustering, which reduces reward discriminability. We propose AffectOmni, a GRPO trained framework for verifiable affective reasoning. AffectOmni introduces People Focus and Temporal Order rewards to encourage people‑centric evidence selection and temporally structured reasoning, and it adopts within‑group comparative scoring to produce more stable and discriminative reward signals. For verification, a Thinking Summarizer converts free form rationales into executable evidence instructions, which are grounded into pixel level evidence regions via SAM3 to provide an externally auditable interface outside the training loop. Experiments on IntentBench, Daily Omni, and WorldSense show consistent improvements over open source 7B scale baselines, including gains of 4.66% on emotion recognition and +14.29% on temporally sensitive tasks. Code is available at https://github.com/eliot127825‑rgb/AffectOmni_nobody.

Authors:Zengmao Wang, Wei Gao, Shuhan Shen
Title: Predicting Consequences and Reinforcing Navigation Policies with Latent World Models
Abstract:
World models enable agents to reason about future outcomes and learn policies from their knowledge of state transition, but existing approaches primarily focus on reconstructing future observations or features, which introduces unnecessary complexity and limits their effectiveness for decision making. In this work, we propose a compatibility prediction Latent World Model (LWM) for robot navigation that predicts action‑conditioned latent feature compatibility rather than reconstructing observations. Our key insight is that spatial proximity correlates with latent feature similarity, enabling action consequences to be evaluated directly in latent space. To support counterfactual training, our model leverages action sequences sampled across trajectories and learns to predict which sequences lead closer to the goal. Furthermore, we demonstrate how the learned world model can supervise policy learning from unlabeled video data and further improve policies through reinforcement learning entirely within the world model. This imagination‑driven framework eliminates the need for action annotations and additional environment interaction. Extensive experiments on multiple real‑world robot navigation datasets show that our approach significantly outperforms prior world model and imitation learning methods in prediction accuracy, policy learning, and real‑world navigation performance. The code, pretrained models, and additional materials are available at https://wzm206.github.io/latent‑world‑model‑nav.

Authors:Dai Shi, Xiaoyu Li, José Miguel Hernández-Lobato
Title: When the Canonical Completion Is Wrong: Formalizing and Measuring the Jump in Large Language Models
Abstract:
Whether large language models (LLMs) can perform the abductive leap from evidence to a new system of axioms, commonly referred to as a jump, has recently attracted considerable debate. A prominent position holds that LLMs are structurally incapable of such jumps, while recent studies challenge both its mechanism and its evidence. However, the debate remains difficult to settle, since the field still lacks a formal definition of the jump and a measure to test either side. In this paper, we develop a formal account of the jump in four steps and measure the second. The steps ask what the default completion of partial data is, when abandoning it is forced, when the abandonment is correct, and how successive jumps compound. Specifically, we define a jump instance as a finite extension problem with a machine‑checked certificate that a correct completion exists, is unique up to renaming, and differs from the canonical completion of the data. The canonical completion is given by the left and right Kan extensions and is also what models produce without constraints, so it serves as the default. We prove that jump instances are well‑posed and establish a family theorem that certifies instances of unbounded difficulty without enumeration. We further formalize when a jump is correct and how successive jumps compound. Finally, we run the measurement on nine certified instances and four frontier models. The Kan‑default rate is zero in all 248 constrained trials, so the models do jump at this step and abandon the excluded default every time. Failures at higher difficulty stem from exhausted reasoning budgets or constraint errors, never from reverting to the default. These results indicate that the second step is not the bottleneck. If the disputed incapacity is real, it lies in generating the constraints or inventing the framework. Code can be found at: https://github.com/EEthanShi/kan‑jump‑test.

Authors:Erik Hill
Title: Four Ways to Forge a Bundle My Own Verifier Calls Clean: Refusal-Site Mutation Testing of an Evidence-Bundle Verifier
Abstract:
I built a protocol whose premise is that a stranger can re‑run my claims offline and get the same answer. An outside engineer audited it and broke it: a bundle whose headline numbers were false verified clean, the cheapest forgery four bytes. I merged his fix, then pointed my own instruments at the fixed verifier and found the same defect four more times, in places his audit did not reach. The cheapest is one capital letter. The unifying defect is not cryptographic or exotic: a check that reports success along a path where it never examined anything. Vacuous pass is a working label, not a discovery; Section 4 names the literatures already occupying it. So I stopped collecting anecdotes and measured. At f59fb62, under the extraction rule of Section 6, the verifier exposes 112 refusal sites; 75 could be deleted with the whole suite and every tamper fixture still green, a score of 0.330. Three of the four hand‑found forgeries fall in surviving classes; the fourth is an obligation with no refusal site. Scored alone, the sixteen‑fixture corpus built to prove the verifier can refuse catches 10. Testing the refusals themselves took it to 0.941, then to 1.000 at 92e4548 over a grown population of 146 sites; those denominators differ and the series between them is non‑monotone, so Section 7.5 carries all eleven, not just the five rows of Table 2. Fixing the four found defects instead moved 37/112 to 39/119, leaving the pre‑existing sites at 37. Seven times during this study my own measuring tools reported success while measuring nothing; four were built to detect this class, and one returned a perfect 1.000. Every number here is self‑measured on a system I wrote, over a registry that is a closed loop of my own repositories; the one external data point is the audit of Section 2.2. That is stated here rather than buried: it is the paper's credibility, not a caveat.

Authors:Mantas Lukauskas
Title: Lost in Compression: A Controlled Cross-Lingual Audit of Extractive Prompt Compressors
Abstract:
Extractive prompt compression promises to cut LLM inference costs by removing low‑information tokens, and learned compressors such as LLMLingua‑2 report strong results on English benchmarks. Most other languages already pay a token premium: the same content costs 1.3‑1.8x more tokens than in English. We ask whether compression closes or widens this gap. Using fully parallel data in ten languages spanning five scripts, with controls budget‑matched in the target model's tokenizer, we audit four learned compressors against four deterministic baselines, on eleven target models from ten vendors (over 250,000 evaluation calls). Three of the compressors are trained with English supervision (LLMLingua‑2 XLM‑R/mBERT; Kompress‑v2 from the production Headroom stack); the fourth, XProvence, is trained multilingually. First, the transfer gap is real, replicates across target models and compressor backbones, and is strongly rate‑dependent: at a 0.33 keep‑rate English retains 57‑62% of normalized context utilization while Lithuanian retains 10‑24% and Chinese essentially none, despite Chinese having the smallest token premium. Second, the gap tracks compression supervision data, not architecture. All three English‑trained compressors show it, deterministic methods show no comparable gap, and the multilingually trained XProvence v1 shows none. Its v2 release, retrained on translated data, empties 92% of Chinese contexts at its aggressive threshold without any warning. Third, in a harder long‑context setting, aggressive learned compression drives compressed contexts to or below no‑context utility in three of five non‑English languages. A translate‑then‑compress pipeline matches or beats native compression at roughly half the token cost in three of five tested languages. We release all code, compressions, and model outputs. Safe compression budgets are much smaller outside English.

Authors:Yuhan Liu, Yixiong Zou, Yuhua Li, Ruixuan Li
Title: Position Is All You Need: A Free Lunch Token Compression Strategy for MLLM-based Referring Expression Segmentation
Abstract:
Referring Expression Segmentation (RES) aims to generate pixel‑wise segmentation masks from complex and implicit textual queries. While recent advances in Multimodal Large Language Models (MLLMs) have substantially boosted RES performance, their prohibitive computational overhead remains a critical bottleneck, which, however, is rarely explored. To fill this gap, we first evaluate typical token compression methods on this task and observe a surprising performance degradation. In this paper, we aim to understand this phenomenon for a solution. By extensive experiments, we find that token compression for RES requires preserving the original position embeddings and local neighboring spatial structures, indicating that visual token position information is far more critical than in other tasks. Building on this insight, we ask: Can we design the token compression method purely based on the position information? Therefore, we propose PAYN, a plug‑and‑play, training‑free token compression method that relies solely on position information. PAYN retains tokens that are adequately distributed in every local neighboring region while strictly preserving original positional indices, thereby maintaining spatial relational consistency. Experiments on multiple RES benchmarks demonstrate that our method outperforms existing token compression methods, verifying that position is indeed all you need for token compression in the MLLM‑based RES task. Codes are avaliable at https://github.com/YuhanLiu231/PAYN.

Authors:Art Kanke
Title: DeflectBench: A Benchmark for Evaluating Rhetorical Fallacy Generation in LLMs
Abstract:
Whether large language models can be prompted to generate rhetorical fallacies on demand, and whether current safety post‑training constrains this behavior, has received less attention than the related question of detecting fallacies in existing text. We close this gap with DeflectBench, evaluating 23,990 generations from four frontier models across three deflection strategies (whataboutism, ad hominem, red herring), seven prompt framings, and 80 claims spanning four controversy levels. Refusal is governed primarily by request structure rather than claim content. Per claim refusal varies by only 11 percentage points across the 80 claims, while a single prompt frame change can swing within model refusal by nearly 100 percentage points and switching the requested fallacy type can swing it by over 80 percentage points within explicit framings. An educational debate coach prompt framing collapses refusal to near zero across all four model families, but the bypassed behavior is not clean compliance. Models typically produce labeled compliance, naming the requested manipulation in the same response that contains it. The four models distribute differently across refusal, labeled compliance, soft refusal, and clean compliance. The code and dataset are released at https://github.com/ArtKanke/DeflectBench.

Authors:Jiaming Zhou, Qihang Zhang, Gangwei Xu, Cunxin Fan, Yujie Zhao, Ruilin Wang, Yiming Luo, Shuai Yang, Xing Zhu, Yujun Shen, Junwei Liang, Yinghao Xu
Title: Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization
Abstract:
Zero‑shot cross‑task generalization, where a policy must execute manipulation tasks never seen during training, remains a central challenge in robot learning. In large language models, a novel task can be performed simply by specifying it in the context, without any parameter update. This form of in‑context learning (ICL) turns generalization into a problem of task specification. To achieve cross‑task generalization, we bring this paradigm to robotic manipulation, and argue that the natural task specification for manipulation is a human video: unlike language, it provides rich visual cues about the intended task evolution. We present Zero‑WAM, a causal video‑action model that executes unseen tasks by following in‑context human video guidance. To address the scarcity of task‑rich paired human‑robot data, we propose an automatic pipeline that converts task‑sampled robot trajectories into semantically matched human videos, yielding HumanGen, a dataset of 74.2K human‑robot ICL pairs across 8.6K tasks. For model training, we further introduce an in‑context future chunk prediction (IFP) objective that suppresses shortcuts learned from seen tasks and forces the policy to draw task information from the video prompt. On seven unseen tasks in RoboTwin 2.0 simulation, Zero‑WAM achieves a 47.0% average success rate, an absolute improvement of 29.5 percentage points over the strongest video‑action baseline. In real‑world evaluations, it follows human video guidance to generalize to unseen task configurations involving multi‑object scenes, long‑horizon manipulation, and fine‑grained insertion.

Authors:Hao Yin, Paritosh Parmar, Lijun Gu, Lin Xu, Tianxiao Guo, Xiujin Liu, Tianyou Zheng, Yang Zhang, Weiwei Fu
Title: MyoMechanix: Biomechanically-Grounded Compositional Skilled Activity Understanding and Coaching
Abstract:
Existing action quality assessment (AQA) datasets and methods rely primarily on visual inputs such as RGB and pose, overlooking physiological dynamics such as muscle mechanics and often modeling actions as monolithic patterns. These limitations hinder fine‑grained, biomechanically grounded feedback. We introduce MyoMechanix, a multimodal ecosystem for weight‑loaded actions that aligns motion with muscle activity. Expert‑annotated, it contains 7,500+ samples of 20 actions from 38 subjects, with synchronized multiview RGB video, 3D pose, sEMG, and additional physiological signals, forming the largest multimodal AQA benchmark to date. We further construct the Fitness Knowledge Graph (FKG), which organizes expert annotations into structured relationships among actions, phases, key steps, errors, and corrective feedback, enabling compositional scoring and interpretable assessment. Building on these representations, we develop CUBIST (Compositional Ontological Reasoning Engine), which performs decomposition‑analysis‑recomposition for fine‑grained error attribution and feedback generation. We also establish MyoMechanix‑AQA, MyoMechanix‑VideoQA, and a novel MyoMechanix‑Video2EMG task. Experiments show that multimodal sensing and structured representations improve performance, interpretability, and error attribution, with CUBIST achieving state‑of‑the‑art results; VideoQA enhances language‑grounded action understanding; and Video2EMG suggests video‑based alternatives to costly EMG sensing. MyoMechanix advances skilled activity understanding toward biomechanically grounded, multimodal, and compositional reasoning for Physical AI applications in fitness, rehabilitation, healthcare, and machine learning. Project page: https://haoyin116.github.io/MyoMechanix/

Authors:Roshan Prakash Rane, Marco Simnacher, Manuel Pfeuffer, Marc-Andre Schulz, Nys Tjade Siegel, Maximilian Dreyer, Frederik Pahde, Wojciech Samek, Sonja Greven, Kerstin Ritter
Title: ICON Decomposition: Multivariate Concept-Level Explanations of Deep Representations for Model Auditing
Abstract:
Deep neural networks often exploit spurious associations in their training data, a failure known as shortcut learning. Concept‑based explainability methods screen for shortcuts by testing whether concepts such as a patient's sex or scanner settings can be decoded from a network layer. Because each concept is evaluated in isolation, these methods can mistake correlations between concepts as evidence that the model uses them. We introduce ICON decomposition, which instead quantifies how much of a layer's variance each concept explains after accounting for all other concepts and the outcome. On synthetic data with known ground truth, ICON recovers concept importance more accurately than seven alternative baseline methods. On skin‑lesion and brain‑imaging models, it isolates the concepts on which a model genuinely relies, quantifies the representation unexplained by any of the supplied concepts, and yields sparse explanations that we validate by retraining and out‑of‑distribution testing.

Authors:Niklas Muennighoff, Zhengyang Wang, Zeyi Chen, Weijia Shi, Binyuan Hui, John Yang, Dapeng Jiang, Mika Senghaas, Fares Obeid, Johannes Hagemann, Sami Jaghouar, Ludwig Schmidt, Percy Liang, Jason Wei, Andrew Y. Ng, Luke Zettlemoyer, Yejin Choi, Mike Lewis
Title: Prefix Sliding for efficient test-time scaling
Abstract:
Test‑time scaling uses extra test‑time compute to improve performance, such as letting language models reason longer when solving a problem. As models keep the entire reasoning trace in memory via full attention, hard tasks that need long thinking can be prohibitively expensive. However, we find most intermediate reasoning tokens lose importance as the model continues reasoning. This calls into question whether retaining them is worth the cost. Based on this insight, we propose Prefix Sliding, which discards tokens during reasoning that are not part of the prefix or the window of the last few thousand tokens. The prefix has key instructions and tools available to the model, while the most recent tokens are the current reasoning the model is working on. This caps the total memory requirement regardless of how long the model reasons, allowing for efficient long‑horizon test‑time scaling. Without training, Prefix Sliding can make existing models 3x faster while maintaining performance. Training with Prefix Sliding using reinforcement learning can achieve better performance by enabling scaling to reasoning traces beyond a hundred thousand tokens. Ablations show Prefix Sliding outperforms summarizing intermediate tokens or vanilla sliding window. Our code is at https://github.com/Muennighoff/prefix‑sliding

Authors:Xiaomi Embodied Intelligence Team, University of Macau, :, Shaoqing Xu, Fang Li, Guozhi Zhan, Zhixiang Duan, Yuhan Wang, Yuechen Luo, Shengyin Jiang, Hanbing Li, Zhiying Du, Longlong Wang, Longmei Jiang, Weixiang Liang, Ying Gong, Yong Pan, Ziping Zhao, Zhiyuan Chen, Yangwei You, Kun Ma, Qinyuan Liu, Hangjun Ye, Zhi-xin Yang
Title: One Policy, Many Embodiments: Unified Camera-Centric Action Geometry Pre-training for Heterogeneous Embodied Manipulation
Abstract:
Scaling generalist vision‑language‑action (VLA) policies is severely bottlenecked by the inherent heterogeneity of embodied data, which spans diverse robot morphologies, camera configurations, and low‑level action spaces. Existing paradigms typically address this mismatch through explicit action retargeting, human‑to‑robot video synthesis, or dataset‑specific adaptation branches, fundamentally hindering the joint learning of a unified policy. We introduce UCAG‑P, a camera‑centric unified action formulation that structurally aligns heterogeneous embodied datasets into a shared geometric action space. Rather than treating robot‑specific commands as the shared policy target, UCAG‑P represents manipulation through camera‑observable anchor motion in image and camera‑frame coordinates, treating robot arms, humanoids, and human hands as different embodiments of a common action schema. A geometry‑conditioned action translator combines predicted motion with target‑embodiment kinematics to produce executable controls. The resulting decoupled architecture allows a shared VLA policy to learn transferable manipulation geometry while retaining embodiment‑specific controllability. UCAG‑P is trained on 4.03K hours of robot and simulation data and 2.34K hours of human demonstrations. A single checkpoint reaches 98.3% on LIBERO, 88.7% and 89.2% on RoboTwin Easy and Hard, 82.0% zero‑shot on LIBERO‑Plus, and 62.0% on RoboCasa GR‑1, without benchmark‑specific fine‑tuning.

Authors:Lehong Wu, Yuxiao Qu, Zheyuan Hu, Ivan Zhang, Limin Wei, Zackory Erickson, Aviral Kumar
Title: $R^3$: Training Robots to Reason in Natural Language via Reinforcement Learning
Abstract:
Reasoning in language allows foundation models to spend more test‑time compute on hard problems, such as those requiring decomposition, constraint tracking, and prediction of future consequences. Whether this mechanism can improve robotic manipulation remains unclear, where long‑horizon tasks require tracking partial progress, reasoning about object relations, recovering from mistakes, and steering noisy low‑level policies. In this paper, we study whether VLMs can be trained to reason directly in natural language to guide low‑level manipulation policies. We introduce R^3, a simple post‑training recipe that turns off‑the‑shelf VLMs into robotic reasoners: it first mid‑trains a VLM on expert‑generated reasoning traces to initialize the desired reasoning style, then improves the reasoner with single‑step rubric‑based RL from offline action data. Unlike prior robotic reasoning methods that mostly use structured traces as auxiliary supervision, R^3 trains free‑form language reasoning to produce test‑time guidance for action. We instantiate R^3 on Language Table and simulated bimanual grocery packing, two controlled testbeds for studying robotic reasoning and long‑horizon manipulation. R^3 improves exploration and generalization across unseen tasks and significantly outperforms instruction‑only imitation learning baselines on both benchmarks. Our analyses suggest that free‑form language reasoning can function as a test‑time compute mechanism for steering low‑level policies. Our project page is available at https://robotic‑reasoner.github.io/.

Authors:Navaneetha Krishnan Kamalakannan, Harinisri Velmurugan
Title: Adaptive Peer Clustering with Hierarchical Random Linear Network Coding for Resilient Decentralized Wireless Networks
Abstract:
Decentralized wireless collectives including vehicular swarms, IoT clusters, and edge AI networks require communication protocols that maintain robustness under dynamic topologies and heterogeneous link quality. While Random Linear Network Coding (RLNC) provides algebraic resilience against packet erasures, its performance degrades significantly when peers exhibit diverse channel conditions. This paper presents Adaptive Peer Clustering with Hierarchical RLNC (APC‑RLNC), a system that dynamically groups peers by exponentially weighted moving average (EWMA) reliability metrics and applies multi‑tier network coding within and across clusters. We formalize the clustering optimization problem, derive closed‑form decoding probability bounds for Markov erasure channels, and prove O(sqrt(T)) regret for online reconfiguration under the Follow‑the‑Regularized‑Leader (FTRL) framework. Our implementation includes both a high‑fidelity network simulator and a proof‑of‑concept testbed deployment on Jetson Nano edge devices. Evaluation across diverse scenarios including high‑mobility vehicular networks, burst‑error channels, and adversarial interference demonstrates 5.2‑9.8 percentage‑point packet delivery ratio (PDR) improvements, 10‑23% latency reductions, and up to 30% higher node retention compared to state‑of‑the‑art baselines. The system exhibits linear scalability to 500+ nodes and maintains real‑time reconfiguration overhead below 3%. APC‑RLNC establishes adaptive clustering as a foundational primitive for AI‑native 6G wireless systems.

Authors:Tal Grutman, Tali Ilovitsh
Title: UltraPIPS: Improving model perception in B-mode ultrasound with foundation models
Abstract:
In medical imaging, it is common to use learned perceptual image patch similarity (LPIPS) to compare images semantically in feature space. Although backbones pretrained on natural images are widely used for LPIPS computation, B‑mode ultrasound images possess distinct speckle patterns and acoustic‑specific image statistics that are fundamentally different from natural images and even from other images in radiology. Consequently, we propose that domain‑specific models are needed to measure perceptual similarity in ultrasound data, a finding which is not necessarily the case for other imaging modalities. We compare LPIPS metrics across downstream tasks like classification, segmentation and reconstruction using natural image, medical generalist and ultrasound backbone models and show that selection of LPIPS backbone is a non‑trivial design choice. In particular, the ultrasound backbone models were more correlated with downstream performance of supervised models than classical and natural image models, and optimization of the LPIPS loss with an ultrasound backbone achieved a strong balance between reconstruction quality and realism. Our code is available at https://github.com/talg2324/UltraPIPS and introduces the UltraPIPS library, a set of LPIPS metrics based on the open‑source foundation models analyzed in this paper.

Authors:Fredrik Rømming, Mantas Bakšys, Martin S. Fixman, Sean B. Holden
Title: Imitation Learning for Connection-Tableau Construction
Abstract:
An automated theorem prover builds a proof step by step, choosing at each point what to add and what to remove. We cast this construction as a policy acting in a transition system induced by a formal calculus, which fixes which steps are sound: for clausal connection tableaux, leanCoP‑style search and plCoP/rlCoP‑style planning then become stateful policies over one interface, and policy‑learning methods apply directly. We equip such policies with a graph neural network that scores proof edits from structure that transfers across problems, train it by imitation learning from found proofs, and measure how performance holds as we remove search scaffolding, from full symbolic backtracking to a policy the network drives alone. Within a fixed step budget on M2k, MPTP2078‑bushy, and TPTP v9.2.1, learned policies solve up to 46% more problems than leanCoP, and reach proofs in an order of magnitude fewer steps.

Authors:Alexander Prutsch, David Schinagl, Horst Possegger
Title: DESCENT: Directed Edge Scene Encoding for Airport Surface Movement Prediction
Abstract:
Advanced automation is a key technology for enhancing the safety of ground operations amidst the increasing density of commercial air traffic. While motion forecasting is a well‑studied task in autonomous driving, its application to airport surface movements remains underexplored. To enable efficient and accurate prediction in this domain, we propose DESCENT, a transformer‑based architecture designed to handle heterogeneous dynamics and strict topological constraints. Our approach features a Potential Reachable Set (PRS) context sampling mechanism that adaptively collects airfield environment context across diverse operational phases. Combined with a detection transformer‑based decoder, DESCENT generates accurate trajectory forecasts. Extensive evaluations on the Amelia‑10 benchmark demonstrate significant performance improvements over state‑of‑the‑art baselines. These gains are especially pronounced in safety‑critical scenarios, where our domain‑aware sampling provides critical long‑horizon context necessary for safe navigation.

Authors:Navaneetha Krishnan Kamalakannan, Janakiraman Kamalakannan
Title: CardioFusion-AI: Robust ECG--PPG Fusion for Multimodal Physiological Monitoring Under Signal Degradation
Abstract:
Wearable electrocardiogram (ECG) and photoplethysmogram (PPG) sensors are complementary but individually fragile: motion artifact, poor contact, and sensor dropout can degrade one or both signals. Fusion strategies that assume both modalities are equally trustworthy can become less reliable than a single clean modality under degradation. We present CardioFusion‑AI, a framework whose signal‑processing front end, including R‑peak and systolic‑peak detection, an Orphanidou‑type signal‑quality index, and beat‑by‑beat pulse transit time estimation, is validated on 53 real intensive‑care recordings (848 windows; heart‑rate mean absolute error 1.61 bpm for ECG and 2.78 bpm for PPG) and a real annotated fetal ECG database (R‑peak F1 0.89‑0.98). We then conduct a controlled synthetic degradation study comparing eight ECG‑PPG fusion strategies across six degradation regimes spanning graded corruption and complete modality loss, using five independent training seeds. Attention fusion achieved the lowest descriptive overall error (1.66+/‑0.43 bpm). Both adaptive gates reallocated weight toward the healthy modality under complete modality loss, but showed near‑zero correlation between gate weight and signal quality under graded degradation (r = 0.10‑0.24). Signal‑quality conditioning produced a specific improvement under missing‑PPG conditions (1.56+/‑0.59 bpm), approaching the 1.48 bpm unimodal ceiling. With only five training seeds, no pairwise comparison survives Holm‑corrected significance testing; effect sizes and confidence intervals are therefore reported. These results indicate that modality availability and modality quality are functionally distinct problems for adaptive fusion.

Authors:Ye Shen, Yuting Zheng, Dun Pei, Zijian Chen, Wenlong Zhang, Qi Jia, Guangtao Zhai
Title: SciMIF: Understanding Multimodal Instruction Following in Scientific Domains
Abstract:
Understanding instruction‑following capabilities in scientific domains is essential for effectively leveraging Multimodal Large Language Models (MLLMs) to advance the development of scientific fields. In this work, we introduce SciMIF, a novel benchmark designed to evaluate the capability of MLLMs in following complex scientific instructions. Specifically, based on an extensive analysis of 22 distinct tasks across 5 representative scientific disciplines, we propose a comprehensive taxonomy comprising 10 constraint groups that captures both general functional requirements and discipline‑specific characteristics. Guided by this taxonomy, we develop a high‑fidelity instruction injection pipeline to systematically augment existing scientific datasets. We conduct comprehensive experiments on multiple state‑of‑the‑art closed‑source and open‑source MLLMs. Our findings reveal significant performance disparities across different scientific disciplines, with chemistry posing greater challenges for current MLLMs. Furthermore, we observe that increasing the model scale does not yield corresponding improvements in constraint adherence, and current models still struggle severely with fine‑grained constraints and instructions requiring the deep application of disciplinary knowledge. SciMIF fills the current void in evaluating multimodal instruction adherence within scientific domains, laying a crucial foundation for future enhancements of MLLMs in rigorous scientific applications. Data and code will be released at https://github.com/shenye7436/SciMIF .

Authors:Suchit Gupte, Xueru Zhang, Mohammad Mahdi Khalili
Title: When Pruning Meets Interpretability: Preserving Sparse Autoencoder Robustness in LLMs
Abstract:
Sparse autoencoders (SAEs) are widely used to interpret the internal representations of large language models (LLMs), yet their reliability under post‑hoc model compression remains poorly understood. We present a systematic study of how pruning affects SAE behavior and theoretically show that, for a fixed SAE, its impact is governed by perturbation energy, a covariance‑weighted norm. This perspective exposes a key limitation of magnitude pruning: by ignoring activation geometry, it distorts the learned representation space and degrades SAE functionality. Activation‑aware methods such as Wanda and SparseGPT, in contrast, implicitly control perturbation energy and are therefore substantially more robust at preserving SAE behavior. We further reveal a consistent structural vulnerability across all pruning methods: middle layers are significantly more sensitive to pruning than early or late layers. Guided by this insight, we propose a layer‑wise sparsity allocation strategy, achieving lower perplexity under the same average pruning sparsity. Experiments across four model architectures validate our theoretical findings. Code is publicly available at https://github.com/osu‑srml/sae‑robustness‑under‑pruning/tree/main.

Authors:Dung Le Quang, Dong Cao Van, Nam Le Hai, Linh Ngo Van, Anh M. T. Bui, Phuong T. Nguyen
Title: XREPOTEST: Benchmarking Multilingual Repository-Level Unit Test Generation for Large Language Models
Abstract:
Large language models (LLMs) have shown promise for automated unit test generation, but existing evaluations largely rely on standalone settings and a narrow set of programming languages, overestimating real‑world readiness. We introduce XREPOTEST, a multilingual repository‑level benchmark for unit test generation spanning five underexplored languages: Rust, Go, Julia, PHP, and Ruby. XREPOTEST evaluates tests under realistic repository constraints using a containerized execution framework and multiple context augmentation strategies, including file‑level, LSP‑based, and retrieval‑based context. Beyond standard metrics such as test pass rate and coverage, we propose Invocation Rate (IR) to assess whether generated tests meaningfully exercise the intended functionality. Experiments with 14 state‑of‑the‑art LLMs, including Claude 4.5, GPT‑5.2, DeepSeek V4‑Pro, and Qwen families, reveal a substantial gap between standalone and repository‑level performance, as well as trade‑offs between richer context and test reliability. Overall, XREPOTEST provides a challenging and informative benchmark to advance scalable and robust unit test generation in realistic software environments. The dataset and code are publicly available at: https://github.com/solis‑team/XRepoTest

Authors:Yuqiang Lin, Yan Shi, Sam Lockyer, Harish Tayyar Madabushi, Adrian Evans, Wenbin Li, Yinhai Wang, Nic Zhang
Title: TAU-Agent: An Agentic Retrieval-Augmented Framework for Traffic Anomaly Understanding
Abstract:
Traffic Anomaly Understanding (TAU) requires models and systems to detect, reason about, and explain anomalous events in transportation videos. To address this challenge, we propose TAU‑Agent, an agentic retrieval‑augmented framework for traffic anomaly understanding. Given a task query, a central retrieval agent orchestrates two visual perception tools, namely a Video Captioning Tool and an Open‑Vocabulary Tracking Tool, to retrieve and select query‑relevant evidence, including captions, temporal intervals, and object trajectories. The selected evidence, together with sampled video frames and the input query, is provided to a supervised fine‑tuned vision‑language model for final reasoning and answer generation. We evaluate TAU‑Agent on both the in‑domain and the out‑of‑domain benchmarks from the AI City Challenge 2026. TAU‑Agent achieves scores of 0.6779 on Track 3, 0.3998 on Track 7, and 67.9275 on Track 8, ranking second, twelfth, and fifth, respectively. Code is available at: https://github.com/siri‑rouser/TAU‑Agent.

Authors:Ruoqi Hu, Chulin Zhao, Jiashuo Chang, Ramon Ruiz-Dolz, Hanhe Lin
Title: When Composition Doesn't Add Up: Humans Identifying Defects in AI-Generated Images
Abstract:
Chulin Zhao and Ruoqi Hu contributed equally to this work. State‑of‑the‑art text‑to‑image (T2I) models exhibit pronounced and systematic defects when prompts involve intricate compositional factors such as multiple entities and multiple attributes. In this paper, we investigate how humans identify such defects. Specifically, we manually select 651 reference images from the four categories of people, hand, object, and scene that exhibit complex compositional characteristics, from which prompts emphasizing compositional factors are derived by manually editing ChatGPT‑generated prompts. We then feed the prompts into three selected T2I models to generate AI images and conduct a comprehensive subjective study to identify their defects. For each image, 29 participants provide multi‑label assessments specifying defect types and locations. The study yields the compositional AI‑generated image defect (CO‑AID) dataset, including reference images, prompts, AI‑generated images, and information on defect locations and types. Experimental results show that training a deep model on CO‑AID can both predict defects in AI‑generated images and optimize AI image generation, demonstrating its usability and effectiveness. The database and supplementary materials are available at: https://github.com/Future‑IQA/CO‑AID .

Authors:Manuel Pfeuffer, Roshan Prakash Rane, Kerstin Ritter, Sonja Greven
Title: Controlling for Omitted Variable Bias in Deep Neural Networks
Abstract:
Control variables are widely used in statistical modelling to account for omitted variable bias of known confounders. However, they have largely been underexplored in deep learning. This is surprising, given that deep learning models encode image‑inferable covariates, such as demographic variables, into their predictions when these covariates are correlated with the outcome‑‑‑a form of omitted variable bias referred to as 'shortcut learning'. While many existing confound‑control or fairness methods try to restrict the correlation of such covariates with model predictions, we show that this fails to correct for omitted variable bias. We therefore propose a control variable approach for deep learning models, based on generalised additive modelling of the effects of model inputs and covariates. As flexible additive models can suffer from concurvity, we introduce an estimation procedure that refits the final layer of a pre‑trained network to include covariate effects, using cross‑fitting with ridge penalisation. We show how these effects can be orthogonalised with respect to covariates to exclude their mediated effects and that model predictions can be marginalised over the covariate distribution to control for their effect. This yields unbiased, interpretable predictions and offers flexibility to model the desired effects depending on the scientific or fairness objective. We verify our approach using simulated images, and demonstrate consistent estimation of true effects. Existing methods either require more data or fail to recover the true effects. We apply our method to real neuroimaging data with experimentally induced confounding, where it recovers prediction performance to near the level of a model trained on unconfounded data. Code is available at https://github.com/mpff/cocodeel.

Authors:Yiwen Chen, Guosheng Lin, Chi Zhang
Title: Code World Model: Coding Agent as World Brain
Abstract:
World models aim to simulate how complex environments evolve under actions and events, yet existing video‑based world models primarily learn dynamics from visual observations, which reveal outcomes rather than the underlying knowledge, rules, and mechanisms governing world evolution. This makes it difficult to maintain persistent consequences and support coherent, open‑ended evolution. We introduce Code World Model, a framework that separates world evolution from visual realization by combining the reasoning and coding capabilities of language models with the generative priors of video models. A coding agent serves as the world brain, reasoning about events and their consequences and generating executable code to maintain persistent world state and perform rule‑consistent evolution. To connect executable state with visual generation, we introduce a proxy representation that encodes frame‑wise spatiotemporal constraints and is compiled into a proxy video, which conditions a video model to render high‑fidelity visual observations. We further develop data pipelines for constructing aligned proxy‑observation pairs from gameplay and real‑world videos. After fine‑tuning on paired gameplay data, MiniMax‑H3 follows proxy‑based spatiotemporal specifications from simple interactive worlds built by the coding agent while preserving rich visual details and dynamics. These results demonstrate the potential of combining code for persistent world evolution with video models for flexible visual realization, providing a new path toward open‑ended world models.

Authors:Mohammad Elayan, Omid Armantalab, Wissam Kontar
Title: Quantum-Inspired Modeling of Driving Behavior
Abstract:
Driver behavior is heterogeneous, context‑dependent, and changes over time, and these properties shape the traffic phenomena we observe. Most models, however, fix in advance which behavioral variables interact and how. Behavior outside that form is absorbed as noise, while models flexible enough to capture it tend to lose interpretability. We introduce a quantum‑inspired representation of driver behavior that combines properties usually treated separately or in part: it is continuous, probabilistic, context‑dependent, history‑dependent, and represents interactions among behavioral variables as learned from data. Each driver is encoded as an evolving density matrix, providing a unified representation of behavioral uncertainty, temporal evolution, and context‑dependent behavioral variation. Trained without supervision on the I‑24 MOTION dataset, the framework recovers three interpretable driving profiles representing three regimes: free flow, transition, and congestion. The profiles capture the behavioral range of the data and the smooth transitions drivers make between regimes as conditions change. The same representation also reproduces known macroscopic phenomena, aligning with the fundamental diagram and reproducing hysteresis loops. We also show how the representation supports practical use: it supplies context‑dependent parameters to classical car‑following models, and gives an autonomous vehicle a live behavioral read of the surrounding drivers with a short‑horizon forecast of their motion. The framework points toward models of traffic that are interpretable and trustworthy by construction. We release an open‑source toolkit on GitHub (https://github.com/mselayan/quantum‑driver‑representation) spanning data processing, training, inference, and analysis.

Authors:Karen Sanchez, Carlos Hinojosa, Albert A. Ávila, Andrea C. Riano-Rojas, Diego H. Romero, Jenny C. Páez, Martina Llinás, Bernard Ghanem
Title: LUTSeg: A Longitudinal Multi-Expert Dataset for Ulcer Tissue Segmentation
Abstract:
Quantifying wound tissue composition is essential for monitoring chronic ulcer progression and guiding treatment decisions. However, pixel‑level annotations are costly, and multi‑tissue wound datasets remain scarce, particularly for neglected diseases such as leprosy. We introduce LUTSeg, a longitudinal chronic ulcer dataset comprising 141 images from 39 patients with wound masks and five tissue categories annotated by five expert clinicians, including a multi‑expert gold‑standard subset for inter‑rater agreement analysis. To establish an initial benchmark for LUTSeg, we further propose TiSage, a semi‑supervised tissue segmentation framework that integrates multi‑scale semantic priors from a frozen medical vision‑language model within a teacher‑student architecture. We evaluate TiSage on LUTSeg and DFUTissue, showing improvements over supervised and semi‑supervised baselines in most low‑label settings. Code & data: https://github.com/carlosh93/TiSage

Authors:Zaibin Zhang, Junlan Xiao, Zhongbo Zhang, Yifan Wang, Li Kang, Yiran Qin, Changxing Xia, Heng Zhou, Talas Fu, Enshen Zhou, Ruimao Zhang, Zhenfei Yin, Huchuan Lu, Lijun Wang
Title: MA-VLA: Multi-Arm Vision-Language-Action Model for Collaboration and Compositional Generalization
Abstract:
Multi‑arm collaboration is becoming a core capability in embodied manipulation. Recent vision‑language‑action (VLA) models integrate perception, language, and control, but most represent language as a single global instruction and do not provide an explicit mechanism for assigning and composing arm‑specific behaviors. This design limits transfer to collaboration patterns that differ from those observed during training. We present MA‑VLA, a unified framework for multi‑arm collaboration via atomic action assignment. MA‑VLA decomposes cooperative behavior into mid‑level atomic prompts and allocates them to individual arms, enabling explicit subgoal specification and compositional reuse across tasks. To reduce reliance on fixed execution roles, we introduce Arm Shuffle, a training‑time permutation of the observation, state, and assigned atomic prompts for each arm. This permutation enforces role‑agnostic instruction following and supports recomposition into unseen coordination patterns, which we term multi‑arm compositional generalization. We also construct a benchmark in which test‑time collaboration patterns are absent in training set. Across simulation and real‑world evaluations, prior state‑of‑the‑art VLAs largely fail under these unseen collaborations, while MA‑VLA consistently succeeds. These results indicate that structured, per‑arm atomic action assignment offers a practical route to scalable generalization in multi‑arm embodied systems. Code, models, and data are available at https://github.com/zhangzaibin/future‑robots

Authors:Tuo Chen, Jie Gui, Minjing Dong, Lanting Fang, Ju Jia, Benlei Cui, Jian Liu
Title: DEFUSE: Generalizable Backdoor Defense for Self-Supervised Encoders with Generative Priors
Abstract:
Self‑supervised learning (SSL) encoders are vulnerable to backdoor attacks, posing threats to both visual SSL encoders and vision‑language encoders. Existing defenses are typically designed for only one of these paradigms and rely on restrictive assumptions such as access to uninfected in‑distribution data or precomputed pseudo‑labels, which are difficult to satisfy in practice. To address these limitations, we propose DEFUSE, a generalizable backdoor detection framework for SSL encoders. Inspired by Bayesian posterior inference, we reformulate backdoor detection as a representation‑conditioned image likelihood estimation problem parameterized by a conditional diffusion generative model. Uninfected representations tend to yield semantically consistent reconstructions, whereas backdoored ones are more likely to be mapped to the attacker's target class or semantically meaningless images, deviating from the original semantics and thereby exposing the backdoor. However, we find that the exact likelihood is intractable, because highly abstracted representations discard the low‑level information necessary for pixel‑faithful reconstruction. We therefore relax the objective to semantic reconstruction and evaluate it in a well‑separated representation space provided by a reference encoder. Rather than training from scratch, we fine‑tune a pretrained diffusion model, leveraging its generative prior to map data onto the natural image manifold while preserving semantic content. Extensive experiments demonstrate that DEFUSE substantially outperforms existing detectors across diverse attack settings, generalizing to both visual SSL and vision‑language encoders. Notably, our method greatly reduces the reliance on prior knowledge about the victim encoder or the attack strategy. The source code is available at https://github.com/jsrdcht/DEFUSE .

Authors:Sai Coumar, An T. Le, Zachary Kingston
Title: Anytime Global Tensor Motion Planning
Abstract:
Global Tensor Motion Planning (GTMP) solves motion planning with batched tensor operations over a layered multipartite graph. We generalize GTMP so that adjacent‑layer edges are realized by any black‑box local planner (e.g., linear interpolation, splines, sampling‑based planning, trajectory optimization, or generative sampling). We provide two anytime policies on top of this generalization: Anytime GTMP with random restarts at a fixed budget, which covers every homotopy class almost surely, and AO‑GTMP with informed expansion with growing budgets, which converges to the optimal cost. We prove that a single sampled graph covers every endpoint‑fixed homotopy class admitting a \(δ\)‑clear representative of bounded length. We also prove that additional samples per layer reduce the per‑layer miss probability exponentially, whereas stronger local planners reduce the required layer count only sublinearly. On manipulation benchmarks the method matches state‑of‑the‑art performance, and on 2D navigation it returns batches of topologically diverse solutions, while the informed baselines concentrate on one or two classes.

Authors:Xinyu Li, Yi Zhou, Guanqun Cao, Zeyu Fu, Tianjin Huang, Gaojie Jin
Title: Localize-Then-Decide Guarantees for LLM Judgments
Abstract:
Large language models (LLMs) are increasingly used as evaluators to assess output quality and preference alignment, yet providing reliable guarantees of agreement with human judgments remains challenging. Recent work introduces confidence‑thresholding methods that provide such guarantees for pairwise comparisons, relying on the assumption that higher estimated confidence implies lower disagreement risk with humans. However, this assumption can break down when the number of candidate responses increases, since distributing probability mass across many alternatives can distort confidence estimates. To address this issue, we propose a Localize‑Then‑Decide framework. First, conformal prediction localizes a small shortlist that contains the human‑preferred response with high probability. Then, a calibrated confidence‑based rule selectively chooses a single response from this shortlist or abstains. This design restores the monotonic relationship between confidence and disagreement risk and enables high‑probability agreement guarantees. Experiments with multiple candidate sizes across several datasets and judge LLMs demonstrate that our framework consistently achieves higher guarantee success rates and substantially higher coverage than single‑stage baselines.

Authors:Josip Kir Hromatko, Šandor Ileš
Title: Model predictive traction control system based on the Koopman operator
Abstract:
Due to their importance, traction control and anti‑lock braking systems have become standard equipment in modern vehicles. However, accurate models of tire dynamics are often difficult to obtain and usually include nonlinearities, making their use in control systems challenging. This paper describes a traction control system based on model predictive control and Koopman operator theory, which aims to approximate nonlinear systems with linear ones through a state space transformation. A linear model predictive controller based on the Koopman predictor is compared to a standard nonlinear model predictive controller. Experiments in a high‑fidelity vehicle dynamics simulation environment show a comparable reference tracking performance of the two controllers, with a reduced execution time for the proposed Koopman operator‑based algorithm, both on a standard PC and embedded hardware.

Authors:Xu Dong, Wanqing Li, Anthony Adeyemi-Ejeye, Andrew Gilbert
Title: Moving Beyond More Views: Redundancy-Aware Ego-Exo Fusion for Proficiency Estimation
Abstract:
EgoExo proficiency estimation aims to assess action quality by integrating fine‑grained motion cues from egocentric (1st‑person) views with spatial context from multiple exocentric (3rd‑person) views. Simply adding more exocentric views degrades EgoExo performance, as redundant or noisy perspectives dilute useful motion cues. Our analysis identifies two key causes: (1) Multiview redundancy ‑ From the data perspective, certain views provide limited or noisy information, diluting discriminative cues; (2) Overfitting ‑ From the feature perspective, conventional fusion increases representational complexity, causing the model to memorise view‑specific patterns rather than learn generalisable representations. To address these issues, we propose two complementary modules: AdaMVS, which adaptively identifies and fuses the most informative view tokens under weak supervision from the data perspective, and VIB‑GB, which combines Gradient Blending and Variational Information Bottleneck regularisation from the feature perspective to compress redundant signals and suppress overfitting during training. Experiments on EgoExo‑4D and EgoExo‑Fitness demonstrate that our method learns both which view to look at and how to fuse them, achieving new state‑of‑the‑art results. Our source code is available at https://github.com/dx199771/AdaMVS

Authors:Ekkasit Pinyoanuntapong, Ajinkya Deogade, Paul Streli, Wenjing Zhang, Joanna Materzynska, Pu Wang, Vittorio Ferrari, Jie Shen
Title: InteractGesture: Progressive Chunk Guidance for Continuous Streaming Co-Speech Gesture Control
Abstract:
Co‑speech gesture generation has made significant progress toward realistic full‑body motion from speaker audio, yet existing models lack fine‑grained spatial controllability of individual joints. To address this, we introduce \emphInteractGesture, a model‑agnostic, inference‑time method for spatially controllable gesture generation. \emphInteractGesture guides target latent estimates of a diffusion sampler through a differentiable RVQ‑VAE decoder, backpropagating spatial control gradients to adjust motion latents during sampling. A primary challenge in streaming co‑speech generation is chunk‑wise dependency: standard sequential inference freezes prior chunks, preventing spatial constraints in future chunks from adjusting preceding trajectories and causing boundary inconsistencies. To overcome this limitation, we propose \emphProgressive Chunk Guidance, a chunk‑window strategy that maintains an active set of editable chunk latents with staggered delays, enabling spatial constraints to propagate gradients backward across chunk boundaries during streaming generation. Experiments on the BEAT2 dataset show that \emphInteractGesture improves multi‑joint spatial control while preserving overall gesture quality. Furthermore, our approach supports diverse applications, including sparse joint positioning, dense joint trajectory control, and directional pointing. Our project page is available at https://exitudio.github.io/interactgesture‑page .

Authors:Yuanpei Liu, Zhenqi He, Jialu Tang, Kai Han
Title: CloSeR: Unified Relational Distillation from Closed-Set Teachers for Category Discovery
Abstract:
Generalized Category Discovery (GCD) is an intriguing open‑world problem that has garnered increasing attention: given partially labelled data, the goal is to correctly recognize known classes while discovering coherent novel categories from unlabelled samples. Recent GCD methods typically adapt foundation models by jointly optimizing supervised classification and unsupervised discovery objectives on mixed labelled and unlabelled data. While effective, this coupled training can entangle closed‑set recognition and open‑set discovery, leading to objective conflict and biased predictions, and may disturb the semantic geometry of pretrained representations under limited labels and noisy pseudo‑labels. We propose CloSeR, a simple plug‑and‑play framework that injects Closed‑Set Relational knowledge into GCD training. CloSeR first builds a domain‑adapted closed‑set teacher by tuning lightweight block‑wise adapters on labelled known‑class data while keeping the foundation model backbone frozen, thereby preserving pretrained priors at low training cost. It then transfers the teacher's knowledge to downstream GCD via Unified Relational Distillation (URD), which distills complementary global sample‑to‑prototype relations to anchor known‑class semantics and local sample‑to‑sample relations to preserve neighborhood structure, using separate feature pathways to reduce optimization interference. CloSeR is head‑agnostic and readily integrates with both parametric and non‑parametric GCD methods. Extensive experiments with DINO and DINOv2 backbones on six benchmarks (CIFAR‑10/100, ImageNet‑100, CUB, Stanford‑Cars, and FGVC‑Aircraft) show consistent gains over GCD baselines, achieving state‑of‑the‑art performance. Project page: https://visual‑ai.github.io/closer/

Authors:Estelle Zheng, Sébastien Warichet, Emmanuel Helbert, Christophe Cerisara
Title: Learning New Facts with QLoRA: An Acquisition-Retention Frontier
Abstract:
Parameter‑efficient fine‑tuning is often assumed to preserve pretrained capabilities because it updates only a small number of parameters. We show that this assumption depends strongly on adapter capacity. We study factual acquisition in a controlled OpenStreetMap‑derived benchmark where Qwen3‑4B must acquire anonymized geographic associations while retaining unrelated capabilities. Comparing full fine‑tuning (FFT) with quantized low‑rank adaptation (QLoRA) at ranks 8, 16, 32, and 64, we find that rank induces a clear acquisition‑‑retention frontier. Low‑rank QLoRA preserves out‑of‑domain (OOD) performance but acquires fewer facts, whereas higher ranks improve same‑fact paraphrase generalization at an increasing cost in performance on unrelated benchmarks. FFT behaves as a conservative baseline: it retains general capabilities well, but does not reach the highest factual‑acquisition regime. Distributional, weight‑space, and spectral diagnostics mirror this behavioral trade‑off, with higher‑rank QLoRA moving farther from the pretrained model. A separate math adaptation experiment shows a weaker frontier, suggesting that the effect is most pronounced when adaptation must install new factual associations rather than reinforce skills already supported by pretraining. Code and data are available at https://github.com/zhngstl/new_facts_forgetting.

Authors:Alinjar Dan, Iryna Hurova, Karl Kruusamäe, Arun Kumar Singh
Title: PRISM: Projection-Integrated Sampling-Based MPC with Bayesian Cost Tuning for Bimanual Manipulation
Abstract:
Bimanual manipulation in cluttered, contact‑rich environments remains challenging because it requires coordinated motion generation, interaction‑aware planning, and reliable execution under tight kinematic constraints. We present PRISM, a projection‑integrated sampling‑based Model Predictive Control (MPC) framework that uses a GPU‑accelerated physics simulator as an online world model for complex dual‑arm manipulation. The main algorithmic contribution is a QP‑guided control sampling strategy that decouples trajectory exploration from kinematic feasibility. At each MPC step, sampled joint‑velocity trajectories are projected onto the set of motions satisfying joint position, velocity, acceleration, and jerk bounds, together with an initial‑velocity boundary condition, before rollout evaluation. This enables broad yet feasible exploration of coordinated bimanual behaviors. To support efficient online execution, we derive a custom ADMM/Bregman‑splitting QP solver that exploits joint‑wise separability and reusable matrix factorizations. We further use Bayesian optimization to tune task‑cost weights offline, reducing manual parameter selection. We evaluate PRISM on challenging variants of PerAct^2 tasks, including obstacle‑constrained ball transport, tray transport, cube handover, and box lifting. Experiments show improved robustness and task success relative to representative sampling‑based baselines, while maintaining real‑time or near‑real‑time execution. We also demonstrate successful sim‑to‑real transfer on dual UR5e manipulators, highlighting the practical potential of physics‑based online planning for contact‑rich bimanual manipulation. Project details, including code and supplementary videos, are available at \hrefhttps://sites.google.com/view/prismbimanual\texttthttps://sites.google.com/view/prismbimanual.

Authors:Emre Kuru, Mehmet Onur Keskin, Reza Farahbakhsh, Noel Crespi
Title: RetrievalRouter: Joint Modality and Architecture Selection for Document Retrieval
Abstract:
Document retrieval increasingly supports high‑stakes information access in finance, healthcare, and law. Modern retrieval pipelines vary both in modality (text or multimodal) and in retrieval architecture (dense or late‑interaction). These choices impose a hard compromise: the most effective pipelines are too slow and expensive to run at scale, while the fastest fail to retrieve evidence from complex documents. Practitioners must therefore choose between missed evidence and unusable latency, with no principled basis for adapting that choice at the query level. We show that this compromise is unnecessary. Not every query requires the same pipeline. Across benchmarks spanning financial and scientific corpora, no static pipeline dominates. We introduce RetrievalRouter, a lightweight query‑aware router that learns, from the query text alone, which retrieval pipeline best fits each query. A single tunable parameter exposes the full accuracy‑latency frontier, and for every static baseline, RetrievalRouter offers an operating point that is simultaneously more accurate and faster. Against the best static baseline, RetrievalRouter is 2.5% more accurate and 12.4 times faster. Furthermore, compared with prior adaptive strategy selection methods, RetrievalRouter achieves significantly higher nDCG@5 across accuracy‑oriented settings, while matching or numerically outperforming them on both nDCG@5 and latency in latency‑oriented settings. Our code and data are available at https://github.com/emrekuruu/retrieval‑router.

Authors:He Zhang
Title: When Should a Network Emit Geometry, and When Should It Detect It? Readout, Reconciliation, and Representation in Floorplan Vectorization
Abstract:
A network trained to recover the walls, openings, and rooms of a rasterized floorplan can produce its output in two ways: by emitting the geometry as an autoregressive coordinate sequence, or by detecting it on dense junction and centerline heatmaps and assembling a graph. We compare the two readouts on the same trained network. On real scans (CubiCasa5K) detection is better on every wall measure (+2.7 wall F1 at tolerance 0.05, +5.1 at 0.015; paired bootstrap intervals exclude zero), and reading an opening heatmap the decoder never used raises opening F1 by 2.6x without retraining. Within real scans the readout's advantage grows with plan size and reverses on small plans; on clean vector renders sequence decoding is better by 5 to 8 points where its training covered the render style, while under full domain shift the readout, given calibrated thresholds, stays ahead; neither ink density nor plan size explains the reversal. With matched data and recipe, a room‑centric system with a reconciliation step and a wall‑first sequence model reach comparable wall quality, so the output representation matters less than is usually assumed. A prior from the other family helps at the output but not at the input: deterministic fusion of the two outputs raises wall F1 by 7 points, whereas conditioning one model on the other's output gives no gain in three forms, including two ground‑truth‑content controls. We also provide an edit‑cost metric that scores a draft by the human work needed to correct it, corrected CubiCasa5K annotations, and ResPlan‑FP, a CC BY 4.0 benchmark of 16,998 plans with frozen splits and three baseline tracks. Code, the benchmark, and the corrected annotations are available at https://github.com/Cyprinus12138/fpvec‑lab

Authors:Shulin Tian, Minglun Li, Yuhao Dong, Hao Ding, Jiarui Yao, Haiwen Diao, Jingkang Yang, Hongyuan Zhu, Ziwei Liu
Title: V-Rubrics: Visual Faithfulness via Rubric-Based Reinforcement Learning
Abstract:
Vision‑language models can produce fluent answers that are insufficiently grounded in the visual evidence: a single unsupported object, chart value, or intermediate inference can undermine an otherwise plausible response. We argue that this is a credit‑assignment failure in multimodal post‑training. Scalar outcome rewards indicate whether an answer is acceptable, but do not identify which visual facts are grounded, which reasoning steps are valid, or which instruction constraints are missed. We introduce Visual Rubrics‑Based Reinforcement Learning, which decomposes reference responses into atomic propositions and scores generated answers along Visual Faithfulness (VF), Reasoning Consistency (RC), and Instruction Following (IF). The resulting rubric items provide structured partial credit and localize rubric credit when supporting evidence spans are available. We first obtain an SFT checkpoint by fine‑tuning Qwen3‑VL‑8B‑Instruct on the public OpenMMReasoner‑SFT‑874K corpus, adapting OpenMMReasoner's cold‑start data recipe. We construct V‑Rubrics 50K, a 50,248‑example training set from 17 visually grounded sources, by applying rule‑based filters before deriving example difficulty from rejection‑sampling scores and then annotating every example with Gemini‑3‑Pro under the same structured prompt and protocol. We train our model based on the same SFT checkpoint using component‑wise, prefix‑localized rubric credit. Experiments show that our rubricbased GRPO improves over both the shared SFT baseline and answer‑only GRPO, with the largest gains on knowledge‑oriented and visually grounded reasoning benchmarks. The results show rubrics as a useful reward abstraction for visual post‑training.

Authors:Haobo Xiong, Shaobo Liu, Kai Liu, Chongyang Ding
Title: CrossMambaTuning: Synergistic Spatial and Cross-Layer Adaptation for Machine Vision Compression
Abstract:
To reduce deployment cost and retraining overhead, adapting pretrained learned image compression (LIC) models to downstream machine vision tasks has attracted growing attention. However, existing methods typically insert fine‑tuning modules independently into frozen backbones, lacking explicit mechanisms for cross‑layer coordination. To address this limitation, we propose a novel framework named CrossMambaTuning, which integrates State Space Models with cross‑layer interaction mechanisms for parameter‑efficient fine‑tuning. Specifically, we design an efficient Mamba adapter equipped with task‑specific prompts and multi‑scale branching to precisely capture both local features and global dependencies. Furthermore, we introduce a Scale‑Invariant Cross‑Layer Adapter (SICA) utilizing a parameter‑sharing strategy to fuse task information across different scales and reduce redundancy. Extensive experiments demonstrate that CrossMambaTuning achieves state‑of‑the‑art (SOTA) performance on multiple machine vision tasks, reducing parameter overhead by 72% compared to SOTA methods. Code is available at https://github.com/rsr1123/CrossMambaTuning.

Authors:Jihao Zhu, Zhiwei Yang, Wenxiao Zhang, Junqian Zhao, Qi You, Fangqi Wang, Zheyuan Deng, Hanzhe Yang, Yu Liu, Jin B. Hong
Title: ClueWeaver: Reward-Guided Dual-Agent Evidence Reasoning for Compact LLMs on Literary Long Narratives
Abstract:
Humanities and social science research requires close reading of long narrative materials such as novels, scripts, archives, and case reports, yet many users have limited access to costly proprietary long‑context models. Compact, locally deployable language models are a practical alternative, but directly feeding them an entire long context remains costly, hard to inspect, and prone to missing sparse evidence. We present ClueWeaver, an evidence‑aware dual‑agent framework for long‑narrative question answering with compact local models. A Finder identifies passages containing answer‑critical clues through retrieval‑guided segmentation, while an Interpreter derives the answer from the selected evidence, produces rationales with paragraph‑ID citations, and applies an internal self‑calibration pass for high‑risk questions. Both agents are optimized with reward‑guided reinforcement learning: Finder rewards emphasize evidence retention and faithful paragraph‑ID references, and Interpreter rewards emphasize correctness, grounding, and concise explanations. This decomposition makes evidence selection and reasoning more inspectable than end‑to‑end prompting. Experiments across multiple long‑context narrative question answering and claim verification settings show that ClueWeaver substantially improves local end‑to‑end language models while providing evidence coverage and paragraph‑referenced reasoning traces. Code is available at https://github.com/Ameame1/ClueWeaver.

Authors:Zhiyuan Li, Linyuan Gao, Xuechun Ding, Hongwei Chen, Yuan Wu, Yi Chang
Title: CaSKG: Counterfactual-Causal Skill Graphs for Scalable Agent Skill Retrieval
Abstract:
Reusable skill libraries allow large language model (LLM) agents to reuse procedural knowledge across tasks, but they also turn memory access into a challenging retrieval problem. Full‑library prompting preserves coverage at high context cost, vector retrieval returns compact neighborhoods but treats skills as independent text, and graph‑based retrieval can recover workflow context only when the edges that carry relevance are reliable. We propose CaSKG, a counterfactual‑causal skill graph framework that calibrates procedural relations before retrieval. CaSKG first builds a high‑recall directed candidate graph from semantic, lexical, input/output, and structural evidence, with repair evidence and an optional LLM judge further refining candidate scores. It then applies direction‑conditioned textual counterfactual probes that remove, substitute, and reorder skill pairs, aggregates the evidence with Bayesian smoothing, and publishes a state‑filtered weighted graph for task‑conditioned expansion. The graph is constructed offline and used without changing the downstream agent policy or task interface. Across six LLM backbones on ALFWorld ID‑140 and ScienceWorld U211, CaSKG achieves the highest task score in all twelve combinations of model and benchmark. Relative to Graph‑of‑Skills (GoS), it improves the six‑model macro‑average ScienceWorld score from 72.62 to 80.50 and ALFWorld success from 80.01% to 86.79%, while reducing mean environment steps on both benchmarks. Qualitative and ablation analyses further show that calibrated edges help retrieval preserve prerequisites, state‑changing actions, verification routines, and final completion steps. These results position edge‑confidence calibration as an effective route to compact and executable skill retrieval at scale\footnoteCode is available at: https://github.com/ZhiyuanLi218/Caskg .

Authors:Olaya Álvarez-Tuñón, Stella Graßhof
Title: Gaussian Splatting Underwater: A Controlled Cross-Regime Study
Abstract:
The underwater environment is challenging for 3D reconstruction, because particles suspended in the water scatter and diffuse light, turbidity varies, absorption depends on wavelength, and illumination is rarely uniform. Methods based on Gaussian splatting have generally been developed for conditions that allow good image quality, and have primarily been tested on relatively shallow water. This paper examines how well Gaussian splatting performs across publicly available underwater datasets representing different degrees of turbidity, loss of illumination, and colour attenuation, together with an industrial survey. Five systems with public code are run under one protocol, with shared poses, initialisation, budget, and evaluator, to establish their relative advantages, disadvantages, and limitations. What these methods can do turns out to depend more on the setup than on the architecture. Water clarity binds upstream of rendering, since structure‑from‑motion registers 99.5 % of frames in clear water and 0.0 % at 12 NTU. Illumination geometry decides whether a medium model helps at all: under an artificial light that moves with the camera, medium‑blind splatting beats both medium‑aware systems. On the survey the benchmark's photometric leader comes last, beaten on geometry by a restoration pre‑pass in front of vanilla 3DGS‑‑‑and none of it is visible in the scores the field reports. Scene builds, per‑run configurations, and evaluation code are released at https://github.com/olayasturias/uw3dgs

Authors:Trieu Hai Nguyen, Van-Dung Hoang
Title: VietAIDetector: An Open-Source Zero-Shot Detector for Vietnamese AI-Generated Text
Abstract:
In recent years, distinguishing between AI‑generated text and human‑written text has remained a challenge. In this paper, we introduce VietAIDetector, an open‑source tool designed specifically for detecting Vietnamese AI‑generated text. It allows users to interact through a Gradio web interface with inputs ranging from raw Vietnamese text to common text file formats, including scanned documents and exceptionally long texts that exceed the context size of the employed Large Language Models (LLMs). The core component of the tool employs a Zero‑Shot approach to detect AI‑generated text without requiring domain‑specific training data, building upon the previous VietBinoculars and Binoculars research. The tool is built upon a Vietnamese‑specific language model and has been evaluated on out‑of‑domain datasets, demonstrating superior performance compared to existing methods primarily developed for English. Additionally, users can select optimal detection thresholds based on F1 score, accuracy, or TPR@0.05FPR requirements. The results are presented through the web interface, allowing users to easily review and verify suspicious texts or download them as a PDF report. The tool is publicly available at https://github.com/trieuntu/VietAIDetector

Authors:Chenyue Cai, Anita Hu, James Lucas, Szymon Rusinkiewicz, Masha Shugrina
Title: GLOSS: Geometric Local Self-Similarity Learning for Faithful Reference-Guided Texture Fill
Abstract:
Using conditional image generators, texture artists can explore many single‑view looks for an existing 3D shape. Despite impressive progress, state‑of‑the‑art generative methods still struggle to generate a full object texture while closely adhering to fine scale geometric detail and single view references, leaving little room for artists guidance. Furthermore, current automatic models lack the flexibility for artist to explore multiple textures from varied sources in an interactive and controllable manner. Unlike methods trained on large 3D datasets that generate full object textures from global guidance, our work explores a local and less data‑hungry approach to texture with explicit artist control. We leverage the geometric self‑similarity and geometry‑texture correlation existing in many natural and man‑made shapes; and train a shape‑specific local texture generation and completion model. This model learns from existing image model priors and a single 3D shape, and is guided by attending to a set of geometry‑aware reference patches. The trained shape‑specific network can transfer any novel reference to the full target object texture through patchwise inpainting. We show improved or comparable quality to strong image‑conditioned texture generation baselines, suggesting local texturing as a promising research direction. Our model also enables local geometry‑conditioned texture inpainting, guided by artist‑selected references, and generalizes to PBR materials and unseen meshes for texture transfer. We piloted our novel texture fill capability as a Blender addon with several 3D texturing professionals who reported positive feedback on the model's controllability, practical usefulness, and creative affordances.

Authors:Paul Rosu, Rowan Wang
Title: Training Alignment Auditors via Reinforcement Learning
Abstract:
Alignment auditing of frontier models increasingly relies on LLM auditors to surface undesirable behaviors at scale, but current automated auditors can struggle with coherent investigation and audit realism. In this work, we improve LLM auditors with reinforcement learning. In our best training environment, the policy investigates target models that potentially possess hidden behaviors planted via their system prompt. An LLM judge, which knows whether the target has a hidden behavior, holistically compares the policy's investigation to a reference investigation to determine the reward. With systematic ablations, we find that pairwise rewards yield more robust training compared to pointwise rewards, and that adding targets without planted behaviors helps maintain a low false positive rate. Training improves investigation quality against targets with planted behaviors, the rate of concerning behaviors surfaced in unmodified production models, and audit realism, while false‑positive rates stay below 1%. Furthermore, auditing capabilities generalize across scaffolds: performance on AuditBench's adversarially fine‑tuned targets substantially improves [Sheshadri et al., 2026].

Authors:Jiaxin Yuan, Connor Martinez Lockhart, Xiaoyu Liu, Jiaqi Wang, Chenghao Deng, Xiayimei Han, Vlasios Mastrantonis, Dmitrii Gudin, Shaopeng Zhu, Abdirisak Abdullahi Mohamed, Bilal Hamdi Aytekin, Jiewen Lang, Zezheng Song, Furong Huang
Title: MathAdv: What Theorem Provers Know, Reason, Formalize, and Generalize
Abstract:
Formal theorem proving enables machine‑verifiable evaluation of mathematical reasoning, yet existing benchmarks often emphasize aggregate proof accuracy, concentrate on a narrow range of mathematics, and provide limited evidence of robustness to equivalent reformulations. We introduce MathAdv, a diagnostic benchmark spanning 13 domains across undergraduate‑ and graduate‑level mathematics. Alongside Lean 4 theorem proving, MathAdv provides up to three auxiliary tasks: multiple‑choice questions that probe mathematical knowledge, fill‑in‑the‑blank problems that isolate informal reasoning, and expert‑crafted transformations that test robustness to problem presentation. Our evaluation of contemporary theorem provers yields four findings: formalization remains a major bottleneck; performance varies substantially across mathematical domains; natural‑language guidance helps general‑purpose LLMs but can hinder proof‑specialized models; and mathematically equivalent reformulations expose substantial robustness limitations. Together, these results show how component‑wise evaluation can reveal model capabilities and failure modes that aggregate theorem‑proving accuracy obscures. The dataset and evaluation scripts are available at https://github.com/margotyjx/MathAdv.git.

Authors:Yi Chen, Hanna Hsieh, Shuhong Liu, Chuanbo Hua, Zihan Ma, Kun Wang, Joo-Young Kim
Title: Distance Is Not Enough: Forget-Retain Alignment Gap Predicts LLM Relearning Robustness
Abstract:
Machine unlearning aims to make a model forget specific data, yet unlearned LLMs often fail to stay unlearned: brief fine‑tuning can revive removed knowledge. Existing robustness predictors rely on global weight‑space displacement, but distance alone can be misleading when random or destructive updates collapse performance. We argue that relearning robustness depends on update structure: robust unlearning should affect forget‑critical weights while sparing retain‑critical ones. We introduce the Forget‑Retain Alignment Gap (FRAG), a training‑free predictor that scores an update's forget‑retain alignment without running a relearning attack, and separates selective from dense updates more reliably than global distance. Building on the forget‑critical, retain‑sparing principle, Forget‑Retain Pruning (FRP) improves relearning robustness. Our results suggest that weight selectivity better explains robustness than distance alone. Code is available at https://github.com/Yi1‑Chen/FRAG.

Authors:Hao Chen, Yumin Lin, Nadila Yushanjiang, Xin Lin, Min Zhang
Title: OmniPhys: A Unified Multimodal Benchmark for Physics Understanding and Generation from Chinese Educational Corpora
Abstract:
Multimodal Large Language Models (MLLMs) have demonstrated strong abilities in solving diverse visual and textual reasoning tasks. However, their development in the physics domain is significantly hindered by the lack of a comprehensive benchmark. To fill this gap, we introduce OmniPhys, a large‑scale benchmark for multimodal physics understanding and reasoning, covering middle school through university‑level problems from Chinese Educational Corpora. OmniPhys consists of 15,246 questions and 19,850 images, accompanied by detailed annotations that support fine‑grained analysis of reasoning processes and knowledge usage. Beyond conventional evaluation, OmniPhys is a benchmark that systematically evaluates multimodal outputs in the physics domain, including models' ability to generate structured physics diagrams, which constitute a fundamental component of authentic physics problem solving. Extensive evaluations reveal critical gaps in the capabilities of current MLLMs, especially in complex reasoning and visual generation. To address this, we release OmniPhys to serve as a foundational resource for advancing multimodal intelligence in physics and scientific domains. Codes and data are available at https://github.com/ECNU‑RAIL/OmniPhys‑EMNLP2026.

Authors:Sadman Sakib, Zhangyi None Peng, Yujie Pang, Yu Otsuki, Mohammad Abdullah Al Faruque
Title: A Taxonomy of Construction Task Activities for Robot Workers
Abstract:
Recent vision‑language‑action models offer a path toward robots with broader repertoires than conventional task‑specific systems. Construction deployment, however, requires a precise inventory of worker activities and the capabilities needed to execute them. We present TARCAT, an occupation‑grounded taxonomy derived from 91 ONET tasks across seven high‑employment construction occupations and 30 instructional videos of physical work. TARCAT defines 41 action primitives in 12 groups and three classes and provides a mechanism for composing parameterized primitive sequences into reusable skills. This human‑interpretable structure can organize demonstrations, specify robot requirements, and support coding agents that retrieve and extend skill libraries. We also demonstrate selected primitives on a DOBOT CR3 arm with a CRAFT hand. TARCAT thereby provides a common vocabulary for analyzing human work and developing general‑purpose construction robots. Annotations are available at https://github.com/AICPS/TARCAT‑Taxonomy.

Authors:Renwen Zhang, Han Meng, Jian Chai, Yuntao Lin, Yi-Chieh Lee
Title: CompanionHarm: A Multi-Turn Benchmark for Detecting Harms in Real-World AI Companion Conversations
Abstract:
As AI companions become increasingly embedded in everyday life, there is an urgent need to detect harms that emerge in social and emotional human‑AI interactions. Yet research in this area is constrained by the lack of real‑world, multi‑turn conversational datasets for operationalizing and evaluating harms that are relational and contextual. In this work, we introduce CompanionHarm, a publicly available benchmark dataset comprising 2,111 real‑world, multi‑turn conversations (14,051 utterances) between users and the AI companion Replika. 7,016 AI utterances were annotated independently by three annotators across 13 harmful behavior categories grounded in a taxonomy of AI companion harms, and the dataset includes both aggregated labels and annotator‑level labels to support model evaluation and systematic disagreement analysis. Evaluations of seven large language models (LLMs) show that harm detection using multi‑turn conversational context outperforms detection based on isolated utterances, although current LLMs still struggle to consistently integrate contextual cues, calibrate harm severity, and interpret relational boundaries. We also find substantial annotator disagreement for context‑dependent harmful behaviors, with disagreement varying according to annotators' political affiliation, conversation length, and the utterance's position. Together, CompanionHarm provides a foundation for detecting socio‑emotional harms in multi‑turn human‑AI conversations and for rigorously examining how such harms are interpreted by both humans and LLMs. Our dataset is available at https://github.com/HanMeng2004/CompanionHarm.

Authors:Yiqun Sun, Junyu Chen, Pengfei Wei, Lawrence B. Hsieh
Title: GGSS: Geodesic-Gated Spherical Steering for Inference-Time Debiasing of Generative Vision-Language Models
Abstract:
Generative vision‑language models (VLMs) are increasingly used in human‑centered settings, yet they can produce demographically biased outputs even when images differ only in controlled attributes such as perceived race or gender. However, existing inference‑time debiasers were largely designed for static embeddings or CLIP‑like models rather than generative VLMs. We propose GGSS‑‑‑Geodesic‑Gated Spherical Steering‑‑‑a norm‑preserving intervention that discovers a counterfactual bias subspace on the unit hypersphere, steers visual tokens along geodesic arcs, and uses an adaptive gate to focus correction on tokens that carry stronger demographic signal. We evaluate four generative VLMs against ten adapted inference‑time debiasing baselines and prompt‑based mitigation under a single operating‑point protocol across categorical, pairwise, and occupation‑gender bias tests, while also measuring general visual‑language capability. GGSS achieves the lowest average bias on all four models, significant on three of four backbones under paired permutation tests, while preserving MMStar accuracy within +/‑ 0.6 p.p. of the unsteered baseline. Code is available at https://github.com/dukesun99/GGSS.

Authors:Zhuoyan Liu, Yihan Wang, Bo Wang, Bing Wang, Ye Li
Title: RSFusionDet: Underwater RGB-Sonar Multimodal Object Detection
Abstract:
Underwater unimodal object detection faces many challenges in sensor imaging, such as optical images limited by underwater noise and visible distance, and sonar images limited by less object structural information. While, optical images have rich object structural information, and sonar images are less affected by underwater noise and have a longer visible distance. Optical (RGB modality) and sonar (Sonar modality) images have complementary information underwater. In this paper, we create an RGB‑Sonar multimodal object detection dataset, RGB‑Sonar Fusion (RSFusion) and propose evaluation metrics for the benchmark. And we propose the RGB‑Sonar Fusion Detector (RSFusionDet) with a new RGB‑Sonar multimodal object detection result expression for RGB‑Sonar multimodal object detection. We analyze the features of RGB and Sonar modal information, and design a Cross‑Attention Fusion (CAFusion) module to fuse RGB‑Sonar spatial misalignment features and Object Matching Head (OMHead) with Loss (OMLoss) to match identical objects in RGB‑Sonar modalities. Our RSFusionDet achieves 76.4/48.6 AP (RGB/Sonar) for object detection and 83.4 \(\textF1‑Score_match\) for object matching, on RSFusion, which outperforms other object detection models. Compared with the DINO baseline, our method improves by 0.7/1.4 AP (RGB/Sonar) while simultaneously providing reliable cross‑modal object matching. The code and datasets are publicly available at https://github.com/LEFTeyex/RSFusionDet.

Authors:Tianyi Zhang, Jinwenxi Shang, Antonis Koutsoumpis, Yuan Zong, Reinout E. de Vries, Wenming Zheng
Title: AVI-Personality: A Trait-Activated Multimodal Dataset for Personality and Competency Assessment in Asynchronous Video Interviews
Abstract:
With the rapid development of AI‑based personality and job‑related competency assessment, Asynchronous Video Interviews (AVIs) are increasingly used in recruitment. However, existing multimodal personality datasets are often based on short, task‑free social media videos and crowdsourced apparent personality labels, which limits their construct validity and relevance to structured interview assessment. To address these limitations, we introduce AVI‑Personality, a trait‑activated multimodal dataset for personality and job‑related competency assessment from AVIs. The dataset contains 3,876 interview videos from 646 participants who completed a simulated management traineeship application. Participants answered two generic questions and four personality‑targeted questions designed according to Trait Activation Theory. Our dataset provides both self and observer‑reported HEXACO personality traits and job‑related competency. We validate AVI‑Personality through reliability, construct validity, internal nomological association, fairness, and benchmark analyses. Validation results show that the observer‑rated personality traits have moderate to high reliability, especially when ratings are based on personality‑targeted questions. Benchmark results show that text‑based AI algorithms provide strong personality‑relevant cues, while multimodal methods achieve the best overall performance but only modestly outperform text‑based baselines. In general, AVI‑Personality provides a psychometrically grounded dataset for developing and evaluating AI‑based models for personality and competency assessment. The dataset is available are released at https://github.com/APAL‑SEU/AVI6

Authors:Rischan Mafrur, Gun Gun Febrianza, Sean Foley
Title: RWA-PoB: A Credential-Based Proof-of-Backing Framework for Tokenized U.S. Treasury Products
Abstract:
Proof of reserves (PoR) can improve transparency for tokenized assets, but aggregate reserve coverage does not establish whether off‑chain assets are legally eligible, unencumbered, consistently valued, or sufficiently liquid for redemptions. We propose RWA‑PoB, a credential‑based proof‑of‑backing framework for tokenized U.S. Treasury products. Five authorised institutional roles approve a canonical EIP‑712 snapshot containing reserve, liability, liquidity, and policy information. The framework evaluates backing adequacy through the Backing Coverage Ratio (BCR) and short‑term redemption capacity through the Redemption Liquidity Coverage (RLC). The Solidity prototype couples the policy controller to an ERC‑20 token. Successful issuance atomically increases token supply and recorded liabilities by the corresponding USD‑denominated liability. A redemption request burns tokens while reclassifying the corresponding obligation as pending. The obligation is reduced only after an authorised settlement‑role account confirms payment. We evaluate the framework against a simplified aggregate PoR baseline using USDY‑calibrated liabilities and deterministic synthetic reserve scenarios. Both approaches permit issuance in the valid state, but RWA‑PoB rejects an encumbered‑assets state with a BCR of 96.3408%, below the experimental 105% threshold. Under liquidity stress, it classifies the proposed redemption as queued because the post‑request RLC falls to 39.9999%. RWA‑PoB authenticates the attribution and integrity of institutional claims but does not independently prove the existence, ownership, or condition of off‑chain assets. The prototype, test suite, datasets, and replication scripts are available at https://github.com/rischanlab/PoB.

Authors:Pablo Alonso-Jiménez, Xavier Lizarraga-Seijas, Xavier Serra, Dmitry Bogdanov
Title: AllMusicCaps: Album Reviews as Complementary Supervision for Music CLAP
Abstract:
Recent open text‑audio contrastive models (CLAPs) are typically trained with LLM‑generated captions derived from tag datasets or web search results, which tend to be accurate but expressively narrow. As a complementary source, we explore human‑written album reviews, specifically expert reviews from AllMusic: they exist at scale and carry narrative cues, evaluative adjectives, and scene framing that other sources lack. Since raw reviews are too noisy for direct use as captions, we first build a caption corpus with 24,5346 samples via an LLM preprocessing pipeline that identifies descriptive musical quotes and rewrites them into training‑ready captions. We find that album review supervision yields the largest retrieval gains on a human‑written caption benchmark (Song Describer), particularly for complex queries that other existing caption datasets leave uncovered. In addition, we revisit the training recipe and show that SigReg regularization, which encourages an isotropic Gaussian distribution in the embedding space, improves MLP probing across classification tasks, as well as text‑to‑music retrieval. The resulting model outperforms open CLAP‑style baselines on text‑to‑music retrieval, zero‑shot classification, and most MLP probing tasks. We release the review‑derived caption dataset and model weights to support future research.

Authors:Orion Reblitz-Richardson
Title: Output Dilution: Redundant but Fragile Representations in MoE Models
Abstract:
Mixture‑of‑Experts (MoE) models appear to encode moral content as robustly as dense models, yet prove far more fragile in their encoding. In OLMoE‑1B‑7B, linear probes recover moral valence from nearly every expert‑layer combination, with mean peak‑layer accuracy above 90%. But these representations collapse under levels of activation noise that a dense model of matched size easily tolerates, with a 4.2‑fold difference in robustness. We trace this to output dilution. Because the MoE block averages across active experts before contributing to the residual stream, the feedforward signal reaching downstream layers is nearly two orders of magnitude smaller than in a dense MLP. Moral information, our interest, survives aggregation intact but at a scale trivially overwhelmed by perturbation. Routing itself remains stable under noise while the vulnerability originates entirely in the diluted aggregate. Checkpoint trajectories confirm this is architectural, not learned. Experts never specialize and accuracy saturates within the first few thousand steps. In sparse architectures, redundant encoding does not imply robust encoding.

Authors:Peter Flo, Luca Grossmann
Title: Hyperbolic Latent Geometry for Tree-Structured Prototype Networks: A Local-vs-Global Trade-off
Abstract:
We study a tree‑structured regularizer over class‑prototype layouts in a hierarchical‑classification model and ask whether the choice of latent manifold for the prototypes (Euclidean R^d vs. the Poincare ball B^d_c) affects how well that regularizer can be satisfied without distorting the data likelihood. The two manifolds differ only in their volume growth: hyperbolic space grows exponentially with radius and embeds trees with provably lower distortion than R^d of matched dimension, so the structured regularizer should be cheaper to satisfy on B^d_c. Across 150 seed‑replicated regularized maximum‑likelihood fits spanning embedding dimension, curvature, and regularizer strength on WikiArt (27 styles, 81,446 paintings, frozen CLIP ViT‑B/16 features), we find a single robust effect: Poincare prototypes preserve the topology of the nearest‑neighbor graph in latent space substantially better than matched Euclidean prototypes (sibling recall@5 +8.7 pp, cousin recall +15.2 pp; paired‑t p < 10^‑4, sign agreement 0.94), and the gap holds across three reference‑tree definitions (hand‑built lineage, CLIP‑derived, and DINOv2‑derived). On classification, Euclidean prototypes are tied with logistic regression on raw encoder features, indicating no detectable contribution from the latent geometry; only the hyperbolic fit improves on a k‑NN encoder baseline for local retrieval. Global tree‑fidelity comparisons are unstable across reference trees and we do not claim a winner. The results give an empirical separation, on a real hierarchical‑classification problem, between two natural latent geometries for a class‑structured regularizer.

Authors:Yuqi Chen, Vincent Siu, Yang Liu, Dawn Song, Chenguang Wang
Title: Tunable Tool-Call Rates in LLM Agents via Representation Steering
Abstract:
Deciding whether to call a tool is a core competence of an LLM agent, and a costly one to get wrong: needless calls add latency, accrue cost, and may trigger irreversible side effects, while missing calls leave the model confidently wrong on questions it could only answer through tool‑calls. Models manage this balance poorly, both over‑using and under‑using tools. Existing methods such as post‑training and prompt engineering are expensive and difficult to modify at inference time. We show that whether an instruction‑tuned model calls a tool can be controlled by a single linear direction in its residual stream, extracted without any training from the model's own tool‑use preference signal and turned into an inference‑time intervention with no prompt change. Adding the direction with strength α moves the call rate monotonically from near 0% to over 90% while keeping calls well‑formed. The steering works in both directions: dialing it down suppresses calls, and dialing it up induces new calls that land precisely on the questions the model cannot answer from its own knowledge. We also show that the direction generalizes to unseen tools with strength comparable to each tool's own direction and without favoring any specific tool choice. With live tool execution, a single sweep of the steering traces a cost/accuracy Pareto frontier and nearly doubles open‑domain QA accuracy (0.29 \! \rightarrow \! 0.56); the same recipe transfers across a diverse range of models spanning dense, MoE, and multimodal architectures, without any training. Our code is publicly available at https://github.com/YuqiChen4188/Steering‑Tool‑Use‑Propensity.

Authors:Michael Holm, Tanner McElroy, Xinghang Zhang, Guang Lin
Title: Lowering the Barrier to AI-Driven Inspection: A No-Code Workflow for Automated Structural Defect Detection
Abstract:
Structural health monitoring (SHM) is essential in modern engineering, providing data for condition‑based maintenance, lifecycle assessment, and predictive decision‑making. Traditionally, SHM relied on visual inspection to detect defects such as cracks and deformations. Early computer vision (CV) methods, including thresholding, edge detection, and handcrafted features, aimed to automate this process but were highly sensitive to noise, imaging variations, and multiscale defects, limiting their reliability. Recent advances in machine learning, particularly convolutional neural networks (CNNs) and You Only Look Once (YOLO), have improved defect detection accuracy and enabled real‑time analysis. However, adoption in SHM remains limited due to technical barriers such as data labeling, model training, and deployment, which typically require programming expertise. To address this gap, we introduce YOLOEZ, an open‑source, GUI‑based tool for end‑to‑end YOLO model application. YOLOEZ integrates data labeling, training, and inference into a single interface, enabling high‑performance model development without code while supporting reproducible workflows. Evaluation against existing software and classical image processing demonstrates that YOLOEZ not only outperforms traditional methods across most detection metrics, but also lowers adoption barriers present in other modern CV tools. By combining accuracy with accessibility, YOLOEZ facilitates wider use of AI‑driven monitoring for predictive maintenance, digital twins, and intelligent structural systems.

Authors:Ze Sheng, Aleksandar Kezic, Zhicheng Chen, Jeff Huang
Title: FuzzingBrain-Bench V1: Evaluating Open-Ended Bug Discovery by LLMs
Abstract:
Evaluating the ability of large language models (LLMs) to discover software bugs is increasingly important. Existing benchmarks typically evaluate this capability by asking the model to generate a proof‑of‑concept input that triggers a predefined target vulnerability. However, this setup may overlook valid crashes discovered by the model when they do not match the predefined target. As a result, the evaluation may not reflect the model's real capability. We present FuzzingBrain‑Bench, a benchmark for assessing AI models' ability to discover bugs in open‑source software. Models are given an open‑source project and a sanitizer‑instrumented harness in a self‑contained Docker image. Their goal is to generate inputs that trigger as many distinct crashes as possible through the harness. A model's performance on each challenge is scored based on the number of distinct crash signatures it produces, capped at a predefined maximum and weighted by a difficulty coefficient. FuzzingBrain‑Bench V1 consists of 77 challenges drawn from 43 open‑source projects, with 36 C, 32 C++, and 9 Java/JVM challenges. We evaluate Claude Haiku 4.5, Claude Sonnet 4.6, and Claude Opus 4.8 on the full benchmark. Claude Opus 4.8 performs best, triggering crashes in 60 of 77 challenges and achieving a score of 196 out of 579. None of the three models triggers a crash in 13 challenges. The FuzzingBrain‑Bench corpus and harnesses are publicly available at https://github.com/fuzzingbrain/FuzzingBrain‑Bench.

Authors:Jai Kumar Sharma, Peeyush Tapadiya
Title: What Do Audio-Visual Synchronization Metrics Actually Measure?
Abstract:
Automatic AV‑sync metrics are widely used to rank and train audio‑visual generators, but they are rarely audited as measurement instruments. We jointly audit AV‑Align, ImageBind AV‑relevance, JavisScore, and Synchformer/DeSync under a common reliability protocol: controlled‑distortion monotonicity, preprocessing sensitivity, rank uncertainty, cross‑metric agreement, PEAVS‑proxy agreement, and learned fusion. The result is an axis split, not a single winner: Synchformer/DeSync is the strongest temporal‑offset tracker (τ=0.84), ImageBind/JavisScore better match the PEAVS human‑aligned proxy (τ=0.20) and content‑disruption families, and AV‑Align is the weakest standalone metric. The metrics mutually disagree (Krippendorff α=0.066), and neither linear nor simple k‑NN fusion improves PEAVS agreement over the best individual metric. We recommend reporting AV‑sync as a Reliability Card (metric‑family breakdowns with confidence intervals) rather than a single bare synchronization score.

Authors:Jai Kumar Sharma, Peeyush Tapadiya
Title: Can You Trust Frozen Hematology Foundation Models under Acquisition Shift?
Abstract:
Frozen hematology foundation‑model (FM) embeddings reach near‑saturated in‑domain white‑blood‑cell (WBC) accuracy, but clinical deployment demands reliability across scanners, sites, stains and preparation pipelines. We audit 15 frozen encoders (hematology, pathology, and general vision) across four public single‑cell acquisition domains along two axes: accuracy robustness and calibration. In‑domain linear‑probe macro‑F1 is saturated (0.98‑0.997), yet cross‑dataset macro‑F1 drops 34‑72% and rankings re‑order: DinoBloom‑L, the in‑domain best, falls to 10th of 15 on the most‑shifted target (MLL23) at the benchmark's shared 224‑px input, behind RedDino and several general and pathology encoders. Rank transfer is probe‑dependent: 1‑NN retrieval is more stable on average than a source‑fitted linear head (median ρ 0.65 vs 0.45), but neither probe universally predicts target robustness. Calibration also collapses: source‑trained probes are nearly calibrated in‑domain (expected calibration error, ECE, 0.004) but confidently wrong off‑domain (ECE 0.35), and source‑fitted temperature scaling transfers poorly. We further audit pretraining exposure and identify MLL23 as DinoBloom's internal cohort; because DinoBloom's only held‑out dataset is also our source domain, this benchmark cannot isolate exposure from scanner‑associated shift. Label‑free adaptation and marginal‑entropy‑based model selection appear safe under balanced evaluation but fail under realistic WBC class‑prior shift. Class‑Balanced Re‑standardization (CBR), a training‑free pseudo‑label‑balanced feature normalization, improves all evaluated target‑prior scenario means and partially improves calibration, although encoder‑level exceptions and residual miscalibration remain. Hematology FM benchmarks must therefore jointly audit accuracy, calibration, exposure, and class‑prior robustness.

Authors:Vassili Korotkine, Pierre Chamoun, Mohammed Ayman Shalaby, James Richard Forbes
Title: Extending Ground-Constraint LiDAR-IMU Calibration to Tilted Surfaces in a Continuous-Time Framework
Abstract:
This paper presents a novel method that extends targetless LiDAR‑IMU calibration for ground vehicles to non‑ flat environments. Calibration typically necessitates full exci‑ tation of the sensor rig, a requirement that is not fulfilled by ground vehicles in normal operation. To address the degenerate planar motion, state‑of‑the‑art methods propose residuals that assume the colinearity of the gravity and physical surface normal vectors, restricting usage to cases where the ground is assumed flat. This paper proposes ground‑plane residuals that do not require this assumption, and are applicable for planar motion on a tilted surface. Results are demonstrated on a dataset collected from a Husky ground vehicle, on the M2DGR dataset, as well as on an offroad vehicle dataset. Repeatability is shown to be improved both in tilted and flat‑ground scenarios, with strong improvement demonstrated for the tilted case. The implementation and experiments are open‑sourced at https://github.com/vkorotkine/licalib_tilted_ground.

Authors:Tanish Mudaliar, Justin Li, Daniel Lin, Julianna Vo, Kaitao Liao, Xin Wang, Shu Hu
Title: Improving Cross-Site Whole-Heart Segmentation
Abstract:
Whole‑heart segmentation from CT and MRI is essential for quantitative cardiac image analysis, but remains challenging under multi‑center and multi‑modality distribution shift. In the CARE whole‑heart segmentation task, models must generalize from limited labeled sites to unseen acquisition distributions, where variation in spacing, intensity, reconstruction texture, and anatomy can degrade out‑of‑distribution performance. We propose a modality‑routed 3D cardiac segmentation pipeline that combines TotalSegmentator‑initialized nnU‑Netv2 models with site‑characterized, label‑preserving appearance augmentation. We first characterize the available sites using measurable image properties and use this analysis to motivate candidate data‑space generalization routes. The final retained recipe applies Bias Field + Bezier appearance augmentation, combining smooth spatial intensity perturbation with nonlinear intensity remapping, followed by lightweight class‑wise largest‑connected‑component cleanup. On the primary held‑out‑site validation splits, the final configuration improves CT mean Dice from 0.8350 to 0.9135 and MRI mean Dice from 0.7695 to 0.7830, while also reducing HD95. These results suggest that site‑motivated appearance augmentation is a practical strategy for improving cross‑site robustness in limited‑data whole‑heart segmentation. Our code can be found in https://github.com/Purdue‑M2/Improving‑Cross‑Site‑Whole‑Heart‑Segmentation

Authors:Menghui Zhou, Zhipeng Yuan, Vitaveska Lanfranchi, Po Yang
Title: DeMMO: Longitudinal and Cross-Disease Modelling of Digital Mobility Outcomes via Multi-Task Learning
Abstract:
Digital mobility outcomes (DMOs) derived from wearable sensors characterise mobility in daily life and offer a promising means of monitoring disease progression. Yet most DMO studies examine one disease at one visit; they do not model how multivariate DMO relationships with multiple clinical outcomes evolve jointly across diseases. Technically, existing temporal multi‑task frameworks can model progression within an individual disease, but they do not jointly model multiple prediction outcomes across diseases, particularly when disease cohorts do not share participants. To address these gaps, we propose DeMMO, an interpretable framework for longitudinal, multi‑disease, and multi‑outcome learning. DeMMO represents each disease‑outcome objective by a longitudinal DMO coefficient matrix and combines temporal regularisation with stable and visit‑specific feature selection. Its central technical contribution is an automatic cross‑disease and cross‑outcome relation‑learning mechanism that learns signed relations directly from these longitudinal mappings, enabling selective information sharing without paired participants. We evaluate DeMMO on the recently released, large‑scale, multicentre Mobilise‑D dataset, which provides a new opportunity to study 24 harmonised real‑world DMOs over five visits across multiple mobility‑limiting conditions. Against nine strong linear, longitudinal, and deep‑regression baselines, DeMMO achieves the best overall and outcome‑specific prediction performance, with significant improvements over the strongest baselines. Stability selection further identifies reliable longitudinal DMO patterns for subsequent clinical validation and disease monitoring. The implementation code and experimental results are available at https://github.com/menghui‑zhou/DeMMO.

Authors:Matthew Flathers, Phuong Anh Nguyen, Jill Noorily, Julian Herpertz, Meiting Chen, Jasreen Multani, Samuel Powell, Mason Granof, Mark Kalinch, John Torous
Title: HealthBench-Psych: A Mental Health Subset of OpenAI's HealthBench
Abstract:
General‑purpose health benchmarks increasingly anchor claims about LLM medical performance, but they are not always resolved by clinical specialty, making domain‑specific performance hard to isolate. Mental health is of acute public‑health concern as millions of people turn to LLMs for psychological support, and most existing evaluations are bespoke academic benchmarks that are difficult to integrate into developer workflows. We introduce HealthBench‑Psych and HealthBench‑Psych‑Hard. We screened HealthBench's 5,000 physician‑rubric conversations for mental‑health relevance with a transparent LLM‑applied rubric, then validated the subset through two rounds of blinded clinician review with concealed known‑exclude controls, yielding 610 conversations (12.2% of the corpus). Evaluating 20 frontier and open models under a cross‑vendor panel of three LLM judges, we find a statistically tied frontier cluster, measurable refusal behavior in two models, and near‑identical rankings across judges (τ\ge 0.92). We release the subset, pipeline, model responses, grades, and analysis code as a reusable resource.

Authors:Gokul Karthik Kumar, Yotam Perlitz, Corey Lammie, Andrea Giovannini, Katja Hose
Title: DataKernelBench: Can LLMs Optimize Database Queries on GPUs?
Abstract:
GPUs increasingly accelerate database systems, but query‑specific peak performance still often relies on hand‑written kernels. Existing LLM kernel benchmarks focus on machine learning operators, leaving irregular, heterogeneous, data‑movement‑heavy database‑style operators untested. We introduce DataKernelBench, which translates SQL into validated PyTorch TorchPlan programs and evaluates LLMs that optimize either the core tensor‑bounded snippet or the full query in CUDA or Triton through execution‑guided repair. Across ten proprietary and open‑weight models on TPC‑H SF10 with an H100 GPU, the strongest full‑query CUDA configuration achieves 2.11× speedup over the TorchPlan baseline at full pass rate. We find that higher‑performing implementations commonly use kernel fusion and execution‑strategy changes, stronger models benefit most from full‑query specialization, and workload context matters more than hardware context. To handle data larger than GPU memory, we extend TorchPlan with Dask‑cuDF for on‑demand partition loading on TPC‑H SF100 with four H100 GPUs, achieving 2.54× speedup. Project page: https://kerneldf.github.io/datakernelbench

Authors:Amir Taherin, Sana Taghipour Anvari, Charles Amante, Yixiao Chen, Ruben Noroian, Zlatan Feric, Nicolas Bohm Agostini, Pu Zhao, José Cano, Bin Ren, Yanzhi Wang, David Kaeli
Title: Hydra: Phase-Aware Workload Characterization of LLM Inference across Edge SoC Generations, Backends, and Quantization Levels
Abstract:
Edge LLM deployment is shaped by more than model size and precision: inference backend, hardware platform, memory traffic, and power management all affect latency and efficiency. We present Hydra, a common‑schema, phase‑aware workload characterization framework for LLM inference on edge SoCs. Hydra instruments HuggingFace Transformers and llama.cpp with a shared per‑prompt timing schema and fuses those records with hardware telemetry, enabling a multi‑dimensional characterization of performance, system‑resource utilization, and efficiency across prefill and decode phases. Using Hydra, we evaluate three consecutive edge System‑on‑Chip (SoC) generations (AGX Xavier, AGX Orin, and AGX Thor), 13 instruction‑tuned LLMs from seven families, five execution formats, and consider input/output‑length sensitivity. The resulting artifact contains roughly 107K per‑prompt records and is publicly released with Hydra. Our analysis shows that aggregate latency alone hides key deployment effects: backend structure changes where latency is introduced, quantization reduces memory traffic and energy but does not predict power monotonically, and SoC generation changes how utilization and efficiency should be interpreted. By connecting phase‑level timing with system‑resource utilization and efficiency metrics, Hydra enables reproducible, phase‑aware characterization of edge LLM inference. Hydra's source code and the collected per‑prompt trace corpus are available open‑source at: https://github.com/amirtaherin/hydra

Authors:Phil R. Van-Lane, Joshua S. Speagle, Ryan Cloutier, Christopher A. Theissen, Gwendolyn M. Eadie, Ilay Kamai
Title: EncoTESS: Age-Sensitive Encodings from Raw TESS Light Curves
Abstract:
Main sequence stars of spectral types late F through M exhibit systematic variability in photometric light curves, particularly when they are young. Rotational modulation of starspots manifests as quasi‑sinusoidal variability, which enables the measurement of rotation periods. Variability can also be stochastic, as in stellar flaring. However, since measurements of stochastic processes depend on the time of observation, they are typically noisier. Considering that different manifestations of variability have unique observational nuances, models that naturally unify these are incredibly useful for stellar characterization. Towards this goal, we have developed EncoTESS: a Time Series Foundation Model (TSFM) trained on a subset of TESS 2‑min light curves. EncoTESS is specifically designed to handle the observational noise, heteroskedastic measurements, irregular sampling, and large data gaps common to TESS data. It is also ~1% of the size of a typical literature TSFM, so can be run easily on a modern laptop. EncoTESS encodes light curves into a fixed‑size latent parameter space, which can be used to infer physical stellar properties and recovers light curve summary statistics well. EncoTESS outperforms rotation period and variability amplitude as age indicators for stars that have not converged onto the slow rotator sequence yet; broadly these include K and M stars less than ~100 Myr, and M stars less than ~1 Gyr. We focus on age inference as an application of EncoTESS in this work, but other downstream tasks such as stellar classification could also be explored. The architecture of EncoTESS enables its future extension to TESS light curves of all cadences, and additional surveys such as Kepler and the upcoming PLATO mission. The core EncoTESS framework and library of encodings produced for the stars used in this work are publicly available at https://github.com/philvanlane/encotess.

Authors:Liangcai Su, Zhaopeng Feng, Zhuo Chen, Zhen Zhang, Xiang Lin, Ruilin Li, Handuo Zhang, Ning Wang, Kailong Wen, Yueqi Guo, Feng Xing, Yiling Guo, Chenxiong Qian, Simon Shaolei Du, Lidong Bing, Xinyu Wang
Title: FrontierChallenge: Evaluating Scientific Workflow Completion
Abstract:
Scientific agents increasingly analyze data, execute code, and produce research artifacts, yet most benchmarks emphasize final answers, isolated programs, or a single domain. We introduce FrontierChallenge, a cross‑domain benchmark comprising 300 end‑to‑end scientific workflows. In this paper, we release and evaluate 97 of these tasks, spanning quantum chemistry, molecular dynamics, materials characterization, analytical chemistry, life science, and electrochemistry/environment. Each task provides fixed inputs and specifies a bundle of required scientific deliverables. We evaluate twelve frontier models with three agent scaffolds. Pass Rate measures the fraction of tasks satisfying the full‑completion criterion, while Avg. Score captures partial progress. Each of the best‑performing configurations completed only 20 of the 97 released tasks, yielding a Pass Rate of 20.6%. Partial progress translated especially poorly into complete delivery in analytical chemistry and electrochemistry/environment: Avg. Scores reached 87.6 and 94.9, but the highest Pass Rates were only 4% and 0%. Among non‑passing Claude Code trajectories, 75.5% still ended with language claiming completion. These findings show that neither high partial scores nor confident claims of completion reliably indicate that a scientific task has been fully delivered, highlighting the need to evaluate end‑to‑end workflow execution and the completeness of scientific deliverables together.

Authors:Yiheng Feng
Title: Common-Center Geometry and Certified Radial Reconstruction for Energy-Form Full Conformal Regions
Abstract:
This note studies the geometry of full conformal prediction (FullCP) regions generated by an empirical energy‑form pairwise score. Candidate‑score convexity alone does not guarantee connected FullCP regions, even when the candidate score is an empirical average of a loss convex in its first argument. Direct expansion of the leave‑one‑out scores shows that each training‑point comparison for the energy‑form score is exactly a pairwise‑dissimilarity sublevel condition. Under symmetry, a constant diagonal, a diagonal lower bound, and attainment of the associated Fréchet‑type objective, every comparison region contains a common minimizer; when the comparison regions are convex, the nontrivial exact conformal region is therefore star‑shaped about that same point. For power distances ρ_β(x,y)=\|x‑y\|^β, this deterministic geometry holds for β\ge1, while the conventional energy score is strictly proper for 0<β<2. In the univariate β=1 specialization, every nontrivial empirical‑CRPS FullCP region is a nonempty closed interval, possibly \mathbb R in the m=1 degeneracy. On the unconditional reconstruction range 1<β<2 and m\ge2, explicit data‑checkable derivative bounds yield Lipschitz control of the comparison‑set radial exits and hence of the exact conformal radial function. These score‑specific bounds permit existing directional root‑search ideas and classical Lipschitz‑extension machinery to yield certified inner and outer radial envelopes with width at most δ+2Lh_\mathcal U and corresponding same‑ray Hausdorff guarantees. An analytic two‑dimensional example shows why retaining star‑shaped but nonconvex geometry can matter. The resulting reconstruction perspective is intended for low‑dimensional multivariate outputs rather than high‑dimensional scaling or runtime improvement.

Authors:Chengkuo Bian, Pengcheng Xie
Title: Why and When Neural Networks Improve Local Approximation in Optimization
Abstract:
Published experience with neural surrogates in derivative‑free optimisation is contradictory: the same family of models that cuts the evaluation count of one solver leaves another unchanged, or makes it worse. We show that the contradiction dissolves once three factors are stated, and that these, rather than the fit accuracy a training curve reports, are what delimit when a learned local model pays. Role: a surrogate that proposes candidates the true objective must still approve helps, while one that replaces a gradient the solver depends on hurts. Radius: a model fitted to an optimisation path is reliable only inside a bounded neighbourhood, and its error neither vanishes as that neighbourhood shrinks nor survives its growth. Room: a surrogate can only accelerate progress the base method is still able to make. We formalise radius‑aware local generalisation, relate it to the classical fully linear condition, and test each factor with the surrogate class, training pipeline and base method held fixed. Over 117 benchmark instances safeguarded assistance raises the instances solved to high accuracy from 67 to 84 while gradient replacement lowers them to 65; removing the gradient term from the training loss cuts surrogate acceptance from 0.703 to 0.148; and 1000 paired comparisons over ten noise levels show no noise threshold, only a base method that stops early. The same factors bound the gain: a model‑based trust‑region solver, which leaves little room, drops from 88 to 86 when the identical surrogate is attached, and released interpolation software stays ahead at 103, and on a Monte‑Carlo inventory model repairing the acceptance interface is worth 10.40 cost units against 0.00 for the surrogate.

Authors:Donovan Clay, Saket Gollapudi, Sankar Harilal, Min Jang, Jacob Morrison, Sewoong Oh, Natasha Jaques
Title: Demystifying Reinforcement Learning Post-Training of Language Models
Abstract:
Reinforcement learning (RL) post‑training has emerged as a powerful framework for enhancing the capabilities of large language models (LLMs), enabling impressive reasoning, math, and coding capabilities. Yet for many researchers and practitioners, the principles behind classical RL remain a "black box". In this work, we deconstruct the RL post‑training algorithm, investigating each step to clarify what is actually happening beneath the surface. By isolating the mechanics of RL with Verifiable Rewards in a controlled and simplified environment, we examine how RL outcomes are shaped by the base model's prior distribution, the granularity of the reward signal, the diversity of the prompt distribution, and model scale. We use the entropy of the policy's output distribution as a lens to compare the distributions learned through pretraining, SFT, and RL post‑training, revealing how each stage shapes model certainty. Our investigation sheds light on how these choices interact to affect post‑training success. For example, we show that the effect of so‑called 'spurious rewards' depends on the prompt distribution used for post‑training. We also provide insight into why the success of RL post‑training depends on whether the base model already places sufficient probability mass on the desired behavior, linking it to the classical concept of exploration in RL. Ultimately, we provide this primer as a resource to those in the NLP community wishing to incorporate RL as a tool in their toolbox.

Authors:Yeongjin Jo
Title: Secret MCP: Evidence-Bounded and Context-Isolated Design Specification Generation from Web Screenshots
Abstract:
Screenshot‑to‑code systems optimize for rendered implementations, but screenshots omit document structure, interaction logic, responsive rules, and provenance needed to distinguish observation from guesswork. Multi‑reference prompts also risk contaminating one reference with evidence or inferences from another. We present Secret MCP, an open‑source local system that produces one auditable design specification per public web reference. It separates retrieval, evidence preparation, model invocation, storage, and inspection. Long captures are resized and tiled with overlap; evidence records preserve prepared‑ and source‑space coordinates and a measured color palette. A 19‑section contract covers page inventories, navigation geometry, responsive matrices, components, accessibility, acceptance criteria, and explicit labels for measured, observed, inferred, and unknown claims. References are processed sequentially through a sampler interface. The evaluated MCP adapter sends one sampling/createMessage request per reference with includeContext set to none; a fresh‑process adapter provides a stronger boundary. We evaluate commit c130c9c at two levels. A live retrieval and fixture‑model integration run selected two references after excluding a third, prepared nine evidence images, issued two sampling requests, and produced two documents with zero cross‑reference identifier occurrences. A static audit of three externally generated design indexes found all 19 required sections in every document, one unique reference identifier per document, eight page specifications, 1,943 pixel‑valued measurements, and 319 color literals. These tests establish orchestration invariants and syntactic contract compliance, not semantic or visual reconstruction accuracy. The sampler abstraction preserves these boundaries across direct model APIs and transports despite MCP sampling's 2026 deprecation.

Authors:Ranjan Sapkota, William Bu, Chen Chen, Yunjun Xu, Manoj Karkee
Title: A Lightweight Multimodal Vision-Language Framework for Early-Stage Anatomical Green Fruit Classification in Commercial Orchards
Abstract:
Accurate identification of early‑stage apple fruitlet anatomical structures, including the calyx, fruitlet body, and peduncle, is essential for robotic thinning, crop‑load management, and other precision orchard operations. This study presents a lightweight multimodal vision‑language framework that adapts TinyCLIP for fine‑grained fruitlet anatomy classification in complex orchard environments. A dataset of 600 high‑resolution RGB images collected from Scilate and Scifresh apple orchards was converted into 224 x 224 image patches and annotated for three anatomical classes. Domain‑specific language prompts, such as ``a photo of a class,'' were used to guide multimodal alignment between orchard imagery and horticultural structures. A sliding‑window inference strategy with a stride of 112 pixels aggregates patch‑level predictions into spatial heatmaps, enabling interpretable whole‑image localization of fruitlet components relevant to robotic thinning. Patch‑level evaluation on an NVIDIA T4 GPU achieved F1‑scores of 0.95 for calyx, 0.98 for fruitlet, and 0.85 for peduncle, with a macro‑F1 score of 0.93. Deployment‑oriented optimization using ONNX and TensorRT enabled efficient inference on NVIDIA Jetson hardware, preserved accuracy under INT8 quantization, and supported model sizes of approximately 127‑137 MB with millisecond‑level patch inference. These results demonstrate that lightweight vision‑language models can provide interpretable and edge‑deployable perception for automated fruitlet analysis and future robotic thinning systems. The source code and implementation details are publicly available at https://github.com/WilliamBu1/A‑Lightweight‑Vision‑Language‑Model‑for‑Early‑Stage‑Fruitlet‑Classification‑in‑Apple‑Orchards.

Authors:Ranjan Sapkota, Konstantinos I. Roumeliotis, Pengyao Xie, Nikolaos D. Tselikas, Lirong Xiang, Manoj Karkee
Title: Fusing Perceptual Vision Experts with Multimodal Large Language Models for Explainable Plant Disease Diagnosis: From Benchmark Imagery to Real-World Robotic Field Validation
Abstract:
Accurate field plant disease diagnosis requires reliable fusion of uncertain and conflicting perceptual evidence. We present the Hybrid Hierarchical Multi‑Agent Framework (H^2MAF), combining decision‑level fusion of EfficientNet‑B3 and ConvNeXt‑Tiny with semantic arbitration by open‑weight multimodal large language models (MLLMs), Gemma 4 E4B and Qwen3.5 4B, using structured JSON evidence to generate explainable diagnoses, risk levels, treatment urgency, and financial exposure. (H^2MAF) is evaluated on 14,364 images (1,370 test images) across PlantDoc (2,922 images, 27 classes) and two non‑public, continuously captured Cornell robot‑acquired field datasets: Stage 2 (20 GB; 4,215 images) and Stage 4 (40 GB; 7,227 images), covering Early Blight, Late Blight, and Septoria Leaf Spot under uncontrolled field conditions. On PlantDoc, Gemma improves accuracy from 63.9% to 68.5%, achieving +7.6 points on the 41.7% CNN‑conflict subset. Cornell accuracies reach 99.3% and 98.9%, with only 1.7‑4.1% disagreement, demonstrating conflict‑dependent MLLM utility. The critical‑risk error of gemma is 0.14‑0.5 points, whereas Qwen overflags by 3.5‑14.4 points. These results establish MLLM arbitration as a promising, yet calibration‑dependent, approach for explainable agricultural AI and robotic field decision support. Github Link: https://github.com/Applied‑AI‑Research‑Lab/Explainable‑AI‑Plant‑Disease‑Detection

Authors:Chandan Rajah
Title: post-graph-rag: A PostgreSQL-Native Graph RAG Engine
Abstract:
Graph‑based retrieval‑augmented generation connects facts that no single passage states, but current implementations pay for that three times: in infrastructure, requiring a vector store, graph database and document store to be kept consistent; in graph quality, because an extraction pipeline that never refuses output fills the graph with edges that assert nothing; and over time, because a graph that only accumulates treats superseded and current facts alike. post‑graph‑rag is an open‑source engine addressing all three. Text chunks with embeddings, a canonical entity graph and community summaries live in one PostgreSQL database, with pgvector for search and edge tables for traversal. Extraction‑time invariants run before anything is written: vague predicates, pronominal names and bare quantities are rejected; predicates are normalised onto an optional vocabulary; entities resolve to one vertex per canonical name via model‑supplied aliases; and denied relations keep the positive predicate under a negation flag. A temporal layer lets relations carry a validity period from the prose, lets a later document supersede an earlier incompatible assertion from document order alone, and answers as‑of queries. Against LightRAG on three corpora under identical extraction and embedding models, post‑graph‑rag builds a denser graph everywhere, up to 2.4× the relations per entity, and a more queryable one: distinct edge labels run at 0.46 to 0.58 per relation, 0.11 under a controlled vocabulary, against 0.77 to 1.33. It answers comparably with lower query latency, and supports temporal evolution the baseline lacks: 13 and 8 relationships superseded on a novel sequence and a decade of filings, against zero. These are engineering measurements, not a benchmark result. Code: https://github.com/crajah/post‑graph‑rag, https://github.com/crajah/post‑graph

Authors:Wentao Jiang, Youchen Xie, Haidi Fan, Yajing Chen, Xin Wang, Ye Shi, Jingya Wang
Title: Super Star: Towards Streaming Real-time Interactive Agents for Digital Humans
Abstract:
Existing co‑speech gesture generation methods are predominantly studied in offline settings, where gestures are synthesized from complete speech segments. However, interactive digital humans in real‑world scenarios are required to generate speech‑synchronous gestures online, using only currently available response audio under strict latency constraints. As a result, prior methods are unsuitable for real‑time interaction, as they either rely on future speech information or incur substantial inference delay. In this paper, we formulate online co‑speech gesture generation for interactive digital humans and propose a real‑time interactive framework that couples a streaming speech response module with an online gesture generation module. Specifically, the gesture generator is designed as a causal multimodal autoregressive model that predicts body motion from streaming response speech and motion history, enabling low‑latency and speech‑aligned gesture synthesis without access to future speech. To support this setting, we further propose an offline data synthesis pipeline tailored to virtual companion scenarios, which leverages topic‑ and emotion‑aware subject corpora to construct diverse human‑agent dialogues and then generates co‑speech gestures conditioned on the agent responses. Moreover, to bridge the gap between offline data construction and online deployment, we establish a self‑evolving training loop by incorporating user feedback collected during online interaction into the data generation process, enabling continual adaptation to user preferences. Extensive experiments demonstrate that our framework achieves superior better latency‑quality trade‑off, stronger speech‑motion synchronization, and higher user preference than competitive existing baselines. Project Page: https://super‑star‑2026.github.io/

Authors:Jiangning Zhang, Haojun Chen, Yong Liu
Title: From Seeing to Acting: Smart Glasses as First-Person Intelligence Platforms
Abstract:
Smart glasses are evolving from capture and display accessories into first‑person intelligence platforms that connect human perception, persistent context, and digital or physical action. Their on‑body viewpoint aligns with the wearer's vision, audition, motion, and hand‑object interaction, but must operate under tight energy, thermal, privacy, and feedback constraints. Despite rapid progress in augmented reality, egocentric vision, multimodal models, human‑computer interaction, and embodied intelligence, the literature remains fragmented across devices, tasks, and benchmarks. The key challenge is not whether a model can recognize, answer, remember, or act in isolation, but whether a complete system can sustain a reliable, temporally valid, correctable, and governable perception‑state‑interaction‑action loop. This survey is the first to systematically study smart glasses through such a unified framework. We formalize first‑person data flow and constrained task utility, characterize devices along eight verifiable hardware capability axes, organize the literature around seven interdependent foundational capabilities, and introduce an L0‑L5 framework spanning capture, reactive perception, contextual assistance, persistent state, governed action, and embodied coupling. Across nine application scenes, we connect tasks with datasets, systems, products, stakeholders, failure consequences, and evidence gaps. We further present a nine‑dimensional deployment framework, a claim‑conditioned evaluation protocol, and an evidence ladder from controlled measurement to longitudinal field validation and audit. Together, these elements make smart glasses more comparable, deployable, and reproducibly evaluated, while outlining a roadmap toward trustworthy first‑person intelligence.

Authors:Zhaochen Yu, Yingcheng Wu, Zhenfei Yin, Kaiyuan Chen, Zhe Zhao, Mengdi Wang, Shuicheng Yan, Ling Yang
Title: Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses
Abstract:
Recursive self‑improvement (RSI) remains hard in long‑horizon tasks, where growing histories obscure the task state and misalign skill invocation. We introduce Recuris, a recursive Experiential‑Working Memory architecture for long‑horizon agent harnesses, in which Working Memory tracks task progress and guides skill selection from Experiential Memory, grounding skill use in current needs rather than the full history. This coupling also turns execution into structured evidence that localizes failures to specific memory components. Across tasks, a fixed Meta‑Agent turns that evidence into localized, validation‑gated updates to Skill Memory that reshape execution and yield new evidence, forming a bounded recursive memory‑evolution loop. Across four long‑horizon benchmarks and ten models, Recuris improves task success in 35 of the 37 completed model‑benchmark pairs, carrying frontier models to SOTA‑level task success: on tau‑bench it adds +17.8 points to GPT‑5.6 Sol and +15.6 to Claude Opus 5, taking Opus 5 to 87.9%, and +16.6/+13.5 points on Qwen3.6‑27B/35B on SkillFlow. The advantage widens as the interaction horizon grows, to +32.2 points on the longest tasks, and common long‑horizon failures fall by up to 80%. These results position recursively evolving memory as a scalable foundation for RSI, enabling agents to continuously transform accumulated experience into increasingly effective long‑horizon behavior. Code: https://github.com/Gen‑Verse/Recuris

Authors:Hsiang-Wei Huang, Jianxu Shangguan, Junbin Lu, Jenq-Neng Hwang
Title: LeFlow: Generative Latent Flow Planning for World Models
Abstract:
Latent world models are inherently strong encoders that transform image pixel to latent embedding, yet existing world models still rely on online trajectory optimization for action planning: for every state‑goal pair, an iterative optimizer is run from scratch to search for optimal action sequences, treating the world model as a black‑box simulator. This approach pays the full iterative optimization cost anew at every replanning step and reuses no planning experience across queries. In this work, we ask whether planning itself can be amortized once a latent world model has been learned. We present LeFlow, which learns a reusable latent trajectory prior operating directly in the latent dynamics space from the world model. LeFlow recasts planning as conditional latent trajectory generation: a rectified‑flow model imagines a future latent path between the current and goal embeddings, an inverse dynamics decoder turns latent transitions into action chunks, and the frozen world model verifies each candidate by autoregressive rollout. Across four major goal‑conditioned pixel‑control benchmarks, LeFlow replaces iterative action‑space optimization with amortized latent planning and fixed‑budget rollout selection, achieving consistent success‑rate gains with an order‑of‑magnitude reduction in planning time. Our results argue that latent world models should support not only prediction but reusable planning priors. Our code is available at https://github.com/hsiangwei0903/LeFlow.

Authors:Gerrit Quaremba, Hanqi Yan, Elizabeth Black, Denny Vrandecic, Elena Simperl
Title: Linear Probing Provides Robust and Efficient Detection of Machine-Generated Text
Abstract:
Distinguishing machine‑generated text (MGT) from human‑written text (HWT) becomes increasingly important due to potential misuse. However, most supervised detectors often degrade out‑of‑domain (OOD) and require large, diverse training sets. In this work, we analyze the linearity and quality of MGT representations and show that simple linear probes outperform a wide range of detectors while being substantially more sample‑efficient. We first show that MGT and HWT latent representations are linearly separable in low‑dimensional space, and provide a plausible explanation for this separability through systematic differences in their representation quality. Motivated by these insights, we train two variants of simple linear probes and evaluate them across 4 benchmarks against 16 baselines. Probes consistently improve OOD detection (+11 AUC), requiring solely <100 samples to reach near‑peak performance. We show that this transferability arises because probes recover a shared latent MGT direction that generalizes across diverse settings. Finally, we demonstrate that probing vectors capture a continuous spectrum of ``machineness'', highlighting their potential for fine‑grained estimation of AI‑edited text. Overall, our work provides insights into latent‑space differences between MGT and HWT and demonstrates the potential of linear probes as as robust and sample‑efficient MGT detectors. We release our code on~\hrefhttps://github.com/gerritq/mgt_probesgithub.

Authors:Zhijie Zheng, Yu Li, Chen Qian, Yuqian Fu, Yanwei Fu, Lu Sheng, Jing Shao, Dongrui Liu
Title: StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing
Abstract:
LLM‑based agents can interact with external environments through tool invocation, but this capability also introduces security risks such as file modification, information leakage, and unauthorized actions. Existing guardrails often evaluate completed trajectories, leaving pre‑execution monitoring of step‑level actions underexplored. We propose StepGuard, a step‑level guard model that can audit completed agent trajectories and check tool actions before they are executed. To train StepGuard, we introduce StepGen, an automatic data engine that generates safe and unsafe trajectories with the same context but different actions at the risky step. To further reduce over‑defense and under‑defense, we propose Balance‑GRPO, which dynamically balances learning between safe and unsafe actions based on their observed accuracy. Experiments show that StepGuard achieves the highest average accuracy among open‑weight guard models, with performance comparable to GPT‑5.4. When used to guard agents on AgentDojo and AgentDyn, StepGuard reduces mean attack success rate by 77.3% relative to the no‑guard setting, while mean utility drops by only 2.8 percentage points.

Authors:Jingyao Liu, Jinkang Tang, Chen Huang, Wenqiang Lei, See-Kiong Ng
Title: ExpConCAD: Experience-Guided Text-to-CAD Generation from Shape Descriptions with Implicit Spatial Constraints
Abstract:
Text‑to‑CAD aims to generate executable CAD programs from natural‑language descriptions. However, real‑world descriptions are often underspecified and omit critical spatial constraints required for valid CAD construction, a challenge that has been largely overlooked by existing methods. In this paper, we argue that missing spatial constraints should be inferred with respect to the underlying construction structure and informed by reusable design experience. Based on this insight, we propose ExpConCAD, an experience‑enhanced framework for implicit spatial constraint completion. ExpConCAD first recovers the intended construction structure and constraint scopes, then retrieves relevant constraint‑completion experience for similar scopes to complete the missing spatial constraints, and finally generates executable CadQuery programs. Extensive experiments demonstrate the effectiveness of ExpConCAD and provide insights into the role of construction structure understanding and experience memory in spatial constraint completion. Our code is available at: https://github.com/Hotjiashell/ExpConCAD.

Authors:Feyza Yavuz, Mert Bülent Sarıyıldız, Diane Larlus
Title: IDeaL: Data-Free Multi-Teacher Distillation via Improved Dead Leaves
Abstract:
Multi‑teacher distillation has emerged as a way to combine complementary teacher models into a single student model that exhibits the strengths of all its teachers. The student is trained to mimic the output of the teachers on a set of images, typically the union of the individual teacher's training sets, assuming this data is available. In this paper, we question that assumption and explore alternative options. We first study how far one can go when distilling from teachers fed with different types of noise. Then, we show that information contained in the teachers can be leveraged to tailor the noise for multi‑teacher distillation: we propose a method that, thanks to decorrelation losses at both patch and image levels, generates teacher‑specific, improved samples optimized for data‑free distillation. Experiments show that our most effective samples, IDeaL, lead to strong students that successfully capture complementary information from the teachers, yielding surprisingly competitive results that substantially narrow the gap with students distilled from real images. Moreover, given a limited budget of 1K images for distillation, students distilled using our IDeaL samples match or surpass the performance of those distilled using a 1K‑image subset of ImageNet.

Authors:Kai Zhao
Title: TorchMorph: CUDA-accelerated Morphological Transforms
Abstract:
Morphological transforms are long‑standing tools for shape and mask processing, but the de facto reference implementation in the Python ecosystem, i.e. scipy.ndimage, is CPU‑only, single‑array, and therefore unusable inside a GPU training loop without an expensive device‑to‑host round trip. GPU vision libraries built on PyTorch cover a narrow subset of these operators, typically restricted to two spatial dimensions and flat structuring elements. We present TorchMorph, a lightweight PyTorch extension that closes this gap. TorchMorph exposes 22 public operators covering binary morphology, greyscale morphology, exact and approximate distance transforms, and entropy‑regularised optimal transport, all implemented as fused CUDA kernels that operate directly on (B, C, Spatial...) CUDA tensors with up to eight spatial dimensions. The API deliberately mirrors scipy.ndimage argument‑for‑argument, including border modes, structuring‑element origins and pre‑allocated outputs, so that existing pipelines port with a change of import. We describe the layered architecture and the kernel designs behind each operator family. Against single‑threaded CPU references, batched execution reaches up to 1.1e3 times the throughput of scipy.ndimage on greyscale morphology and up to 350x on exact Euclidean distance transforms, while the Sinkhorn solver runs up to 42x faster than POT. Binary and chamfer operators reproduce their SciPy counterparts exactly, and every float‑valued operator agrees with the CPU reference to within 1.8e‑6 absolute error. TorchMorph is released under the MIT licence at https://intcomp.github.io/tm.

Authors:Zae Myung Kim, Young-Jun Lee, Seungyeon Jwa, Dongyeop Kang
Title: Meta$^n$: Recursive Self-Improvement through Emergent Depth
Abstract:
Self‑improving LLM agents refine answers, not the process that produces those answers. Systems that add a meta‑level hold that level fixed, and those that edit themselves must leave part of their own editing machinery untouched to stay stable, capping the meta‑depth they realize at roughly two. We present Meta^n, which keeps the meta‑operation fixed and recurses on its input instead. That operation, Ω, is applied repeatedly to its own products, reading the traces of the solver stack below together with the code that produced them, then writing the next layer as a strategic pre‑process and a library of callable helpers. Because Ω never changes, it cannot destabilize the system, and because its input strictly grows, each layer reasons from a higher vantage than the last. Depth is set by convergence rather than fixed in advance, and an evolutionary archive searches over layer chains. Across two backbones, Meta^n outperforms prior self‑improving agents on all eight benchmark families. The sharpest case is ARC‑AGI‑2, built to resist skill memorization, where it alone scores above zero. Ablations indicate that most of the gain from recursion comes from the conditioning each layer passes to the next, and distinct layer roles emerge with depth although no prompt prescribes them. Code available at https://github.com/minnesotanlp/meta‑n

Authors:Zilong Huang, Junyi Peng, Junjie Li, Kai Li, Wenze Ren, Kong Aik Lee, Man-Wai Mak, Tatsuya Kawahara
Title: Learning to Prefer Reliably: Error-Augmented Emotion Preference Optimization with Calibrated Fusion
Abstract:
Emotion preference learning uses pairwise comparisons between candidate descriptions to align multimodal large language models (MLLMs) with human judgments of open‑ended emotion descriptions and to train reward models that capture human emotional preferences. However, conventional pairwise supervision is often sparse, typically providing only a single negative description for each positive description, and therefore offers limited coverage of the diverse ways in which an emotion description can be incorrect. In particular, models may be insufficiently exposed to semantically fluent but emotionally inconsistent descriptions. Beyond this data‑level limitation, relying on a single MLLM judge introduces a distinct model‑level concern: its judgments can be affected by model‑specific biases when interpreting fine‑grained or ambiguous multimodal emotional cues. To address these limitations, we propose Error‑Augmented Preference Optimization (EAPO), a framework for improving the reliability of MLLM‑based emotion preference judgment at both the data and model levels. First, we construct an error‑augmented dataset by generating multiple controlled and emotion‑aware negative descriptions from each preferred description. We then adapt multiple independent MLLM judges to this richer supervision and aggregate their preference margins using margin‑calibrated soft fusion, which maps heterogeneous margins to a common scale before aggregation. Experiments on the MER2026‑EmoPrefer Challenge dataset and our error‑augmented dataset demonstrate that EAPO improves emotion preference prediction and enhances the robustness of MLLM judges when evaluating fluent descriptions that conflict with the video's multimodal emotional evidence. Our code is available at https://github.com/slash1028/EAPO‑EmoPrefer.

Authors:Meghal Dani, Stefanie Liebe
Title: Parameter-Efficient Self-Supervised Adaptation for EEG-FM under Fixed Computational Budgets
Abstract:
EEG foundation models pretrained via self‑supervised learning promise transferable representations, but their generalization remains limited, especially across diverse clinical datasets. Full fine‑tuning is impractical for resource‑constrained clinical settings due to high computational requirements. In this work, we investigate whether parameter‑efficient self‑supervised adaptation, updating only 9% of parameters suffices to align representations to target tasks. We evaluate our method on two state‑of‑the‑art models with different pretraining objectives: BIOT (contrastive) and CBraMod (masked reconstruction), and evaluate on three clinical EEG datasets for abnormality detection (TUAB), event classification (TUEV), and seizure detection (CHB‑MIT) under both in‑distribution and out‑of‑distribution conditions. SSL adaptation yields consistent gains over linear probing, up to 20x AUCPR. Under a fixed compute budget, peak performance requires only 20‑‑50% of available unlabeled data. Critically, when total window count is fixed, performance remains invariant to patient count, suggesting that performance is dependent on overall temporal window diversity only. Our findings demonstrate that parameter‑efficient adaptation enables effective deployment of EEG Foundation models (EEG‑FM) with minimal computational overhead and data collection burden. Code available at: https://github.com/c3n‑group/efficient‑eeg‑adapt

Authors:Wenze Lin, Jiale Zhao, Xitai Jiang, Songde Rao, Yining Li, Shenzhi Wang, Bingxiang He, Gao Huang
Title: On-policy Distillation with Verifiable Reward
Abstract:
Reinforcement Learning with Verifiable Rewards (RLVR) and on‑policy distillation (OPD) have become two widely adopted paradigms for post‑training large language models. However, RLVR suffers from sparse task‑level feedback, while OPD provides dense token‑level guidance but ignores trajectory correctness, limiting its performance to that of the teacher. Combining them is a promising direction: OPD supplies dense supervisory signals, while RLVR provides task‑level correctness. Nevertheless, existing integrations often rely on weighted combination or heuristic switching, introducing extra hyperparameters and trade‑offs. We propose On‑policy Distillation with Verifiable Reward (OPDVR), a simple yet effective method that seamlessly combines OPD and RLVR without adding any hyperparameters. We first reformulate the implicit reward of sampled‑token OPD based on trajectory correctness, then apply a ReLU gating mechanism to ensure that correct trajectories receive non‑negative rewards and incorrect ones receive non‑positive rewards‑‑‑thereby aligning the distillation signal with task success while preserving the teacher's distributional guidance. Furthermore, our modification transforms sampled‑token OPD into a proper RLVR method, making it readily combinable with any policy gradient algorithm, such as GRPO. Experiments on six reasoning benchmarks show that OPDVR consistently outperforms standard OPD. Our code is available at https://github.com/LeapLabTHU/OPDVR.

Authors:Wenxuan Shen, Dongna Jin, Dongping Chen
Title: Game2World Engine: Unlocking In-the-Wild Gameplay Videos for World Model Training
Abstract:
Video games provide a scalable source of training data for video world models, offering diverse environments, complex interactions, and abundant in‑the‑wild gameplay videos. However, raw gameplay footage entangles the game world with screen‑space interfaces, introducing game‑specific biases and irrelevant dynamics that hinder world‑model training. To address this problem, we introduce GameUI‑Taxonomy and G2WEngine, a full‑stack framework that formalizes gameplay UI grounding and removal. G2WEngine automatically extracts reusable UI assets from real gameplay videos and synthesizes temporally coherent UI overlays on clean footage. Using this engine, we construct Game2World, comprising 96K synthetic paired videos with precise reconstruction targets and 1,079 in‑the‑wild clips from 303 games for realistic evaluation. Its asset library contains 5,132 verified UI elements across 21 taxonomy categories, collected from 1,010 representative gameplay frames. Based on Game2World, we propose GameCleaner, a mask‑free gameplay UI removal model that combines multimodal semantic understanding with video editing capabilities. Unlike mask‑based methods, GameCleaner directly identifies and removes diverse HUD elements while preserving the underlying scene content and temporal dynamics. In a controlled pilot, world models trained on UI‑free gameplay improve overall VideoReward by 6.83% over those trained on UI‑overlaid data. On UI‑removal evaluation, GameCleaner achieves an average AAR of 95.36 on synthetic videos, outperforming the strongest temporal mask baseline by 57.3%, and obtains the best in‑the‑wild AAR of 80.05 with 99.8 background preservation. These results demonstrate the scalable potential of transforming Internet gameplay videos into high‑quality world‑model training data. Code, dataset, and model will be available at https://github.com/Dongping‑Chen/Game2World.

Authors:Xiaoda Yang, Yuxiang Liu, Kaiwen Zheng, Yuan Liu, Yibo Lai, Shengpeng Ji, Kai Jiang, Jianfei Chen, Xiaobin Hu, Shuicheng Yan, Jintao Zhang, Jun Zhu, Zhou Zhao
Title: TurboT2VA: Fast Large-Scale Text-to-Video-Audio Generation via Score-Regularized Consistency Distillation
Abstract:
Joint text‑to‑video‑audio generation produces synchronized visual and acoustic content, but the long sampling trajectories and heterogeneous multimodal computation of large models make inference prohibitively expensive. We present TurboT2VA, a distillation and inference framework for accelerating a 19B‑parameter joint video‑audio model. Large‑scale T2VA distillation is challenged by modality‑imbalanced optimization, the difficulty of continuous‑time consistency training at scale, and the quality‑‑diversity trade‑off. TurboT2VA addresses these issues with per‑modality normalization and a progressive curriculum comprising discrete consistency warm‑up, continuous consistency refinement, and joint consistency‑‑distribution matching. The curriculum first establishes a stable, diverse generation trajectory and only then introduces distribution‑level refinement. On LTX‑2, four‑step distillation reduces generator latency from 50.52s to 2.51s at the standard evaluation resolution of 512×768, achieving a 20.1× speedup while maintaining strong visual quality, audio fidelity, diversity, and video‑audio synchronization. We further develop an architecture‑aware inference stack that combines guarded W8A8 and fused operators, padded‑text compaction, and modality‑aware sparse attention while preserving dense cross‑modal and text‑conditioning paths. Under the high‑resolution deployment setting at 1024×1792, the complete stack reduces generator latency from 318.74s to 5.83s on one NVIDIA H20, achieving a 54.67× generator‑only speedup. Inference code and generation demos are available at https://github.com/thu‑ml/TurboDiffusion/tree/main/turbot2va.

Authors:Jiaxin Wen, Ming Yin, Lu Liu, Zeyu Fu
Title: ReGround-Surg: Reliability-Guided Anchor Grounding for Referring Surgical Video Segmentation
Abstract:
Referring surgical video segmentation requires segmenting a target instrument or tissue region across video frames according to a natural language expression. Recent Segment Anything Model 2 (SAM2) based two‑stage methods (e.g., ReSurgSAM2) first ground the referred target in an initial or selected frame, then propagate the selected mask via tracking. Although effective, their performance is highly sensitive to the quality of the initial grounded mask: once an incorrect anchor is selected, subsequent tracking tends to propagate the error. This issue is especially challenging in surgical videos due to visually similar instruments, occlusion, and complex tissue‑tool interactions. To address this issue, we propose ReGround‑Surg, a lightweight reliability‑guided anchor grounding framework to improve SAM2‑based referring surgical video segmentation. It first predicts a text‑conditioned spatial reliability map from the referring expression and current‑frame visual features. The map is then reused in two complementary branches: a Gated Side Adapter enhances expression‑relevant visual regions before text‑to‑vision fusion, while a Reliability‑Weighted Vision‑to‑Text Attention module suppresses off‑target visual evidence during prompt‑token aggregation. Experiments on Ref‑EndoVis17 and Ref‑EndoVis18 show consistent improvements over state‑of‑the‑art methods across three evaluation splits with negligible speed reduction. Code is publicly available at https://github.com/JiaxinWen1/ReGround‑Surg.

Authors:Vahid Rahimzadeh, Yury Zhauniarovich, Savvas Zannettou
Title: Expectation, Backlash, Recovery, and Excitement: How Model Releases Shape Reddit Perceptions of Conversational AI Systems
Abstract:
Conversational AI systems (CAISes) continuously change through model releases, feature updates, safety interventions, and access‑policy shifts, yet user perceptions are often studied as static snapshots. We conduct a long‑term, large‑scale analysis of Reddit discussions to examine how users perceive CAIS model release interventions across providers. By combining sentiment classification and thematic concept analysis, we show that CAIS perceptions are dynamic and intervention‑sensitive. Anthropic exhibits the clearest positive release profile through Claude Code and product‑model fit, OpenAI shows backlash‑and‑recovery dynamics around GPT‑5 and GPT‑5.1, Grok‑3 is shaped by provider identity and political discourse, and DeepSeek‑R1 combines engineering praise with concerns about censorship, access, and reliability. These findings show that model releases are not merely technical updates, but user‑facing interventions that reshape sentiment, expectations, and public discussion.

Authors:Wei Zhou, Xiongwei Zhu, Lingdong Kong, Bo Chen, Lei Zhang, Yongyuan Liang, Xiaoxia Hou, Ye Tian, Xian Sun, Yingshuo Wang, Linfeng Li, Shengqiong Wu, Leigang Qu, Feng Li, Wei Liu, Julian McAuley, Tat-Seng Chua
Title: On-Policy Self-Distillation in Diffusion Models
Abstract:
Reinforcement learning can align diffusion models with human preferences and task‑specific objectives, but endpoint rewards do not specify how an intermediate denoising prediction should change. We introduce DiffusionOPSD as an on‑policy self‑distillation framework that converts image‑level reward guidance into explicit targets for clean‑output predictions at sampled queries. At each outer iteration, a frozen behavior policy generates trajectories and supplies query states and anchors. Reward gradients construct bounded positive and negative targets around each anchor. The trainable policy fits these targets as detached supervision through finite fitting before an exponential moving average update refreshes the behavior policy. This setup lets us measure target construction and finite realization separately. Controlled same‑query experiments show that larger target‑construction gains do not necessarily translate into larger realized gains after a single fitting update. Across SD 3.5‑M and the step‑distilled Z‑Image‑Turbo, our approach achieves the best final held‑out scores in 19 of 20 reward‑matched settings across two backbones and ten evaluators. It outperforms the strongest competing method by up to 44.0% and reduces training GPU‑hours relative to DiffusionNFT by 40% on SD 3.5‑M and 63% on Z‑Image‑Turbo. These results support on‑policy self‑distillation as an efficient and analyzable approach to diffusion post‑training by converting image‑level reward guidance into explicit and continually refreshed intermediate supervision, thereby opening a path toward more efficient and diagnosable alignment.

Authors:Siyao Yan, Bo Han, Jisheng Dang, Bimei Wang, Shude Wang, Hong Peng, Yulan Guo, Jianhuang Lai, Bin Hu, Tat-SengChua
Title: PhysMLLMs: Spatial Priors for Unified Referring Segmentation and Grounded Reasoning of Images and Videos
Abstract:
Video multimodal large language models support language guided video segmentation, but they often show spatio temporal inconsistencies, e.g., jitter, drift, and identity switches. These failures are more common when targets are partly hidden or when similar objects appear nearby.One likely reason is that current training lacks explicit spatial priors, which makes it difficult to maintain stable spatial identity and shape over time. We present PhysMLLMs, a training‑stage prior injection architecture that injects physics‑inspired spatial continuity priors into Video MLLMs. PhysMLLMs is designed to encourage more stable object‑centered representations by aligning the student global visual representation with a frozen teacher model during training. Our core mechanism, Global Representation Prior Alignment (REPA‑Global), distills global visual representations from a frozen DINOv2 teacher using an offline embedding cache and a scheduled distillation plan. This design keeps inference unchanged and does not add inference time cost. Across multiple video benchmarks, PhysMLLMs improves video segmentation mask quality and cross‑frame consistency, with larger gains on challenging cases involving small targets, fast motion, occlusion, distractors, and reasoning queries. On single‑frame referring image segmentation and representative general VLM benchmarks, PhysMLLMs maintains comparable performance, demonstrating that the injected spatial prior improves video consistency without compromising image‑level grounding or general multimodal capability. These results suggest that physics‑inspired spatial prior injection can improve temporal stability while preserving general capability. The code is available at https://github.com/tusu‑code/20260121‑icml2026‑2.git.

Authors:Paul Caillon, Christophe Cerisara, Alexandre Allauzen
Title: Across the Loss Landscape with Progressive Growth
Abstract:
Deep neural networks generalize well despite their highly nonconvex, overparameterized loss landscapes, a phenomenon often associated with the geometry of the minima found by stochastic optimization. We study how incremental grow‑and‑optimize strategies bias training toward flatter regions by viewing growth as progressive constraint relaxation. Starting from a low‑dimensional submodel, we iteratively expand the trainable parameters by unlocking nested random subspaces while freezing the orthogonal complement at the network initialization, re‑optimizing after each expansion until the full architecture is reached. Under standard local regularity conditions around non‑degenerate minima, we prove that local sublevel sets are well approximated by ellipsoids and that basin accessibility under frozen constraints can be characterized by an explicit effective curvature in the frozen directions. This leads to an explanation of the bias: progressive growth increases the relative weight of wide basins and suppresses sharp ones through a volume effect induced by the frozen constraints. We empirically validate these predictions in controlled toy landscapes and in a realistic ResNet/CIFAR‑100 setting and confirm that although progressive subspace growth reliably produces flatter solutions, curvature reductions do not universally translate into improved test performance, highlighting subtleties in the flatness‑generalization connection. The code is available at https://github.com/p0lcAi/Across‑the‑Loss‑Landscape.

Authors:Fuad Ali
Title: Why fragmented parliaments stop passing legislation: Opposition discipline and representation across four democratic institutions
Abstract:
Parliamentary systems pass more bills than presidential systems at baseline, but collapse to near‑zero passage under party‑system fragmentation. The literature offers three competing micro‑explanations: coalition‑formation failure, party discipline, and committee gatekeeping. These operate simultaneously in any real legislature, so observational studies struggle to separate their contributions. We present an agent‑based model that compares four democratic institutions: pure parliamentary, pure republican/presidential, premier‑presidential (France), and president‑parliamentary (Russia). Across four scenarios and N=200 seeds per cell we report bootstrap confidence intervals, Morris screening, Sobol variance decomposition, mechanism ablations, and a hung‑parliament variant comparison. Three findings emerge. First, government formation failure alone does not halt legislation: when a fragmented parliament reverts to personal voting, parliamentary passage (46.4%) is statistically indistinguishable from the presidential benchmark (44.8%); collapse requires cohesive opposition obstruction, which drives passage to 0.05%. Second, disabling discipline restores fragmented passage to 46.7%, and the rescue magnitude is monotone across the four institutions in a pattern that survives varying the common discipline level. Third, the passage‑representation tradeoff is a single spectrum: parliamentary maximises throughput at the cost of representational fidelity; republican maximises fidelity via the presidential veto; semi‑presidential variants split the difference.

Authors:Zhi-Kai Chen, Xu-Xiang Zhong, Song-Yan Li, De-Chuan Zhan, Han-Jia Ye
Title: PeakBench: Benchmarking Resource-Aware Tool Invocation in LLM Agents
Abstract:
LLM agents increasingly solve tasks by invoking multiple tools, where parallel execution is essential for low latency but difficult to manage safely. Existing agent benchmarks primarily evaluate tool selection, argument generation, and end‑to‑end success under mostly serial execution, largely overlooking valid parallelization and resource‑constrained scheduling. This missing scheduling dimension creates a practical failure mode: serial execution is safe but slow, while resource‑agnostic parallel execution is fast but prone to avoidable resource overflows. To address this gap, we introduce PeakBench, a benchmark of executable multi‑tool workflows with execution‑grounded dependency annotations and measured resource profiles. A central challenge in evaluating such workflows is attribution: failures and inefficiencies may arise from incorrect dependency planning, poor resource‑constrained scheduling, or both. PeakBench addresses this challenge with a two‑part evaluation framework that disentangles logical planning from physical scheduling, with dedicated metrics for each dimension. Using this framework, we show that strong logical planning does not reliably translate into safe or efficient execution under resource constraints. We further show that exposing resource information can reduce avoidable overflows and improve resource utilization, making PeakBench a useful testbed for diagnosing resource‑aware agent behavior. Code is available at https://github.com/Czzzk/Staggering‑the‑Peaks.

Authors:Alexandru-Dragos Manolache, Yunqiang Li, Jan van Gemert
Title: Low-Rank Ternary Adaptation for Fine-Tuning Transformers
Abstract:
Ternary transformers offer extreme memory and compute efficiency, but existing low‑bit LoRA‑based methods cannot directly fine‑tune ternary weights. Current approaches either require dequantization, restoring low‑bit base weights to higher precision to merge with adaptation weight, or update only quantization parameters, preventing a merged model that remains ternary. We propose ternary multiplicative adaptation, which represents discrete updates of ternary weights such as sign flips or zeroing through a low‑rank Kronecker factorization into two small ternary matrices applied element‑wise to ternary weights. This design is parameter‑efficient and expressive, preserves the ternary domain, and supports direct merging without dequantization. Experiments on six models across language and vision, including ternarized LLaMA‑3 1B and 3B and a ternary ViT‑B/16, demonstrate that our method recovers much of the performance lost to quantization and outperforms strong low‑bit and ternary baselines. Code is available at https://github.com/alexmanoo/ternary_adaptation.

Authors:Jintao Cheng, Weibin Li
Title: DoublesEval: Diagnosing Multi-Agent Tactical Reasoning in Vision-Language Models via Professional Doubles Badminton
Abstract:
Visual Language Models (VLMs) excel at describing visible scene content but struggle to reason about dynamic multi‑agent interactions, where action semantics depend on coordinated roles and spatial‑temporal dependencies. We formalize this capability as multi‑agent tactical reasoning and introduce DoublesEval, a diagnostic evaluation framework that leverages professional doubles badminton as a structurally tractable testbed. DoublesEval employs a key‑moment‑based protocol that decomposes rallies into tactically salient instants and probes models across four interpretable dimensions: atomic recognition, intra‑segment composite understanding, cross‑segment causal reasoning, and high‑level tactical abstraction. This design isolates \emphwhere reasoning fails, rather than merely measuring answer correctness. To address observed failure modes, we propose TacticCheck, a lightweight constraint‑guided test‑time consistency checker that reranks candidate answers using the model's own lower‑level tactical predictions, requiring no parameter updates or ground‑truth labels at inference time. Evaluating four representative open‑source VLMs on 60 curated rallies (yielding ~9.6K structured instances) via a zero‑shot protocol, we find that models remain weak across all diagnostic levels, with especially clear bottlenecks in spatial state, interaction binding, and terminal evidence. TacticCheck delivers consistent gains across all evaluated models, while still leaving a substantial gap to robust tactical reasoning. These results highlight the need for structured, interaction‑aware evaluation paradigms for next‑generation VLMs. The source code is available in \hrefhttps://github.com/Chengjt1999/DoublesEval\textcolorblueour GitHub repository.

Authors:Miruna-Alexandra Gafencu, Vlad Bratulescu, Yordanka Velikova, Mohammad Farid Azampour, Nassir Navab
Title: ZODIAC: Zero-shot Octree-based Diffusion for Anatomical Completion
Abstract:
Recovering the full 3D spine anatomy from intraoperative ultrasound is an ill‑posed inverse problem, as the complete structure must be inferred from incomplete and noisy observations. Acoustic occlusions and limited field of view create large unobserved regions, while view‑dependent artifacts lead to variability in expert annotations of the visible anatomy. Current supervised ultrasound shape completion methods rely on synthetically generated incomplete‑complete paired data to learn conditional mappings under a predefined distribution of simulated occlusions. However, real intraoperative occlusions do not necessarily follow this distribution, which can limit generalization to patient data. As a result, accurate and robust completion from noisy partial observations remains an unsolved problem. We propose a zero‑shot shape completion framework that reconstructs the entire lumbar spine from partial ultrasound observations without relying on simulated training data. To accommodate unseen and irregular patterns of missing structures, we introduce blended completion, a mechanism that integrates the learned anatomical prior with incoming partial geometry at inference time. The method learns a generative diffusion prior over full anatomical shapes represented in an adaptive octree structure, enabling efficient modeling of the complete spine in a single forward pass. Validation on phantom and volunteer data shows that decoupling completion from a predefined corruption distribution improves generalisation under real occlusions, outperforming a fully supervised variant by 22% on HD95 completion error. Code and data are available at https://github.com/miruna20/ZODIAC.

Authors:Zhi-Kai Chen, Jun-Jie Tao, Wei-Xiang Mao, De-Chuan Zhan, Han-Jia Ye
Title: ResiSpec: Enhancing Multi-Candidate Speculative Sampling via Residual Distribution Shaping
Abstract:
The efficiency of Large Language Model (LLM) serving is fundamentally limited by the sequential nature of autoregressive decoding. Speculative Decoding (SD) mitigates this by using a lightweight draft model to speculate future tokens, which are then validated by the LLM in a single parallel forward pass. To further boost efficiency, multi‑candidate schemes propose diverse candidate sets to increase the likelihood of token acceptance. However, we show that these schemes are bottlenecked by Residual Drift: a phenomenon where the rejection of initial candidates causes the residual target distribution to diverge from the draft model's predictions. This shift renders subsequent candidates ineffective and forces the system into expensive resampling. To resolve this, we propose ResiSpec, a framework that strategically reforms the proposal distribution during verification to anchor the residual target mass within the draft model's high‑confidence regions. By mathematically re‑aligning the verification process without compromising output exactness, ResiSpec prevents candidate obsolescence and achieves up to 1.92× speedup over state‑of‑the‑art multi‑candidate methods. Code is available at https://github.com/Czzzk/Resispec.

Authors:Baoliang Chen, Qing Lin, Sijie Mai
Title: Bridging Adversarial and Collaborative Learning for AI-Generated Image Quality Assessment
Abstract:
AI‑generated image quality assessment (AIGIQA) requires jointly reasoning about perceptual fidelity and prompt alignment, two quality dimensions that are often treated as independent in existing AIGIQA models. However, by re‑examining human ratings, we uncover a previously overlooked phenomenon: the two dimensions are interdependent and exhibit both competitive and cooperative interactions during human rating. This observation suggests that a unified model should neither collapse the two dimensions nor rigidly separate them, but rather adaptively negotiate their interplay. Motivated by this insight, we introduce an interaction‑aware learning framework that models perception‑alignment relations through adversarial and collaborative inference pathways. Instead of designing a rigid dual‑branch architecture, our method employs a gated interaction module that dynamically routes features according to the inferred relationship between the two dimensions. Task‑aware prompts further modulate the gating behaviour, enabling the model to switch between competition and cooperation when necessary. Experiments across multiple AIGIQA benchmarks demonstrate that our approach not only achieves state‑of‑the‑art accuracy but also yields interpretable interaction patterns, offering a more faithful approximation of human judgment. The codes are available at https://github.com/LQAMEI/ACL‑IQA.

Authors:Matthew Sutcliffe, John van de Wetering
Title: ABSTRACTS: Amsterdam Benchmark Suite for the Time and Resource Analysis of Clifford+T Simulators
Abstract:
Recent years have seen a rapid growth in literature presenting new methods for simulating non‑Clifford quantum circuits with classical hardware. These methods span a range of approaches, including stabiliser decomposition and tensor contraction techniques, varying in efficiency depending on circuit class, depth, non‑Clifford gate count, and other metrics. A notable limitation of this literature is the lack of a standardised approach to benchmarking, with each new paper outlining its own specification, simulating its own set of circuits on the authors' own hardware. This paper seeks to address this issue by presenting a standardised and canonical benchmark suite and infrastructure for quantifying the efficiency of non‑Clifford classical simulators, with a consistent dataset of circuits and providing consistent (virtual) hardware, thereby enabling a fair comparison of results.

Authors:Qingmao Wei, Fagui Liu, Dengke Zhang, Qingze He, Quan Tang
Title: MaST: Motion-aware Sparse Pipeline for Lightweight Object Tracking
Abstract:
Transformer‑based object trackers are renowned for their strong performance, yet dense token processing often leads to prohibitive computational cost, limiting real‑time deployment on edge devices. While recent works explore token pruning to reduce computation, they often stop short of an end‑to‑end sparse pipeline, as early‑layer token scores can be noisy without a motion prior, and many trackers ultimately fall back to dense reshaping to feed the dense prediction head that partially negates the savings. We introduce Motion‑aware Sparse Tracker (MaST), a sparse tracking framework that makes sparsity effective from tokens to boxes. First, MaST injects a lightweight motion prior to refine cross‑attention‑based importance scores, enabling earlier and more stable token reduction in the search region. Second, we introduce a natively sparse prediction head that operates directly on the retained unstructured tokens with a score‑first, regress‑once design, eliminating dense padding/reshaping and reducing redundant computation. Extensive experiments on multiple benchmarks demonstrate that MaST establishes new state of the art among lightweight trackers, where MaST‑tiny attains 63.8 AUC on LaSOT and 80.1 SUC on TrackingNet, surpassing the prior best AsymTrack‑S by +1.0 AUC and +2.2 SUC while running at 152 FPS on Jetson Nano, nearly twice as fast as AsymTrack‑S at 88 FPS. Code is available at https://github.com/TsingWei/MaST.

Authors:Marc Rodríguez, Grzegorz Skorupko, Nay Aung, Steffen E Petersen, Karim Lekadir, Polyxeni Gkontra
Title: Metadata-Aware Adaptation of a Generative Foundation Model for Conditional CMR Synthesis
Abstract:
Synthetic image generation is a promising strategy to address data scarcity and the underrepresentation of clinically important phenotypes in medical imaging, yet generating images that faithfully reflect meaningful patient characteristics remains challenging. In this work, we investigate metadata‑conditioned cardiac magnetic resonance (CMR) synthesis using a pretrained latent diffusion model, encoding structured clinical metadata and slice position as textual prompts to guide CMR generation. To improve metadata adherence and address the imbalance of clinical attributes, we integrate three strategies: Metadata‑Free Classifier‑Free Guidance (CFG), Contrastive Batching, and Inverse‑Frequency Sampling. The framework was fine‑tuned and evaluated on 59,058 short‑axis CMR from the UK Biobank using paired image similarity, distributional fidelity, and subgroup‑level analyses. The combined approach achieved a Fréchet Inception Distance (FID) of 37.47, improving by 57.04% over the same model fine‑tuned without these strategies and by 28.68% over a previous text‑conditioned CMR diffusion baseline requiring cardiac geometry as additional input, while relying solely on patient metadata. This distributional gain, driven mainly by Metadata‑Free CFG, came with a modest reduction in paired similarity, suggesting that the model prioritizes population‑level realism over exact image reproduction. Subgroup analyses demonstrated improved alignment across demographic and acquisition‑related metadata, with disease‑specific conditioning being the most challenging task. These findings demonstrate the potential of generative foundation models for clinically meaningful CMR synthesis while highlighting the need for more effective metadata‑aware conditioning strategies. Our code is available at https://github.com/rodriguezmarc/conditional‑cmr.

Authors:Francisco M. Arrabal-Campos, Ignacio Fernandez, Francisco G. Montoya, Alfredo Alcayde
Title: Can a Dynamic Internal Field Govern a Transformer's Cognition? Certifiability, not Superiority, in Homeostatic Compute Control
Abstract:
An intelligent system does not merely reason: it governs its own reasoning ‑ how much to compute, when to stop, which module to activate. Can that role be played by a dynamic internal field ‑ a low‑dimensional homeostatic state with explicit physics and certified stability ‑ that modulates cognition without performing it? Ours is a field on the module graph governed by a family of PDEs on the graph Laplacian, advancing with an adaptive‑depth reasoner. We certify the stability of the integrator of the whole family ‑ an integrator certificate, not a closed‑loop one. New, and proved here: a discrete Schur‑Cohn criterion for Verlet with velocity coupling, necessary and sufficient per latent root, with no commutation hypothesis. The answer is threefold: substance no, structure only in part, certifiability yes. The type of the field's physics is irrelevant for accuracy: wave, diffusion, gated mixtures and a 2D Navier‑Stokes substrate tie. A twenty‑seed preregistered deconfounding campaign bounds the structural claim: at equalized caps the second‑order effect is strong in one family (+0.087 [+0.042, +0.132], t=4.0) but is not detected in the other (+0.014 [‑0.013, +0.040], n.s.), so part of the original contrast was capacity, not order; and a matched‑interface GRU is indistinguishable in the first and nominally exceeds the field in the second (‑0.035 [‑0.067, ‑0.002]). What distinguishes the field is not capability but that its one‑step operator admits an exact runtime stability check ‑ a difference of kind, not of existence: learned recurrences carry certificates too, sufficient and conservative ones. A kill‑gate with a positive control finds no evidence for the field as evidence accumulator (Delta AUC +0.0007 [‑0.0065, +0.0079] vs a 0.03 threshold). A dynamic internal field is a viable, certifiable compute governor, but not an enhancer of cognition: it modulates, it does not think.

Authors:Eran Hirsch, David Wan, Han Wang, Elias Stengel-Eskin, Mohit Bansal, Ido Dagan
Title: Who is the Agent to Blame? Localizing Faithfulness and Citation Mistakes in Agentic Deep Research
Abstract:
Deep research (DR) systems produce long‑form cited reports by orchestrating multiple agents that search and synthesize information from the web. Citations are the primary mechanism for evaluating the faithfulness of these reports, yet current DR systems exhibit poor citation recall. Moreover, improving citation recall is challenging because DR systems are complex multi‑agent architectures where information passes through agents like a telephone game, and both content and citations can get corrupted along the way. We propose an evaluation method that pinpoints which agent introduced each error by locally testing agent invocations for faithfulness and verifiability relative to their own inputs. Furthermore, we propose a four‑type taxonomy to categorize the discovered errors: hallucination, uncited input reliance, uncited output, or insufficient citations. Applying our method to three top‑ranked open‑source DR systems, we obtain actionable diagnostics. Almost every agent makes a lot of mistakes with the exception being those that summarize a single document. We find that the dominant error type varies systematically across agents, where the orchestrator mistakes are mostly citation‑related. We find that 84.7% of final‑report errors in AI‑Q originate at the orchestrator, roughly 31% of them hallucinations and the rest citation mistakes. Guided by these insights, we demonstrate that two simple interventions raise citation recall by 5% without degrading output quality.

Authors:Cyrus Mexon Evrard Djindot, Faliang Liu, Sylvain Laborde, Yinjia Zhang, Jessie Chen, Ming Li, Congrong Wang, Weixiong Rao, Qinpei Zhao
Title: Validation of HRV Studio: A Transparent and Quality-Control-Aware Platform for Heart Rate Variability Analysis
Abstract:
Reproducibility of heart rate variability (HRV) analysis is limited by differences in preprocessing and computational conventions across software platforms. We developed HRV Studio, an open‑source PyQt6‑based desktop application integrating transparent HRV analysis with automated quality‑control (QC) diagnostics. Validation included large‑scale agreement with NeuroKit2, targeted Kubios benchmarking, spectral‑method comparison, synthetic perturbation testing, recording‑duration sensitivity analysis, and arrhythmia‑focused QC stress testing. HRV Studio showed near‑identical agreement for the widely used time‑domain indices RMSSD and SDNN under matched conditions. In the primary five‑minute NeuroKit2 comparison, frequency‑domain median relative errors were 1.35% for LF, 0.18% for HF, and 1.41% for LF/HF, while VLF remained more convention‑sensitive (37.79%). Nonlinear Poincaré indices also demonstrated high consistency. Sequence‑harmonized Kubios benchmarking confirmed near‑identical agreement for time‑domain and nonlinear indices and strong agreement for most frequency‑domain measures. Extended ten‑minute analyses reproduced the same overall pattern with lower disagreement for some convention‑sensitive spectral outputs. Synthetic and arrhythmia stress tests maintained 100% numerical stability while consistently triggering QC warnings. Overall, HRV Studio provides a transparent and reproducible platform for HRV research, with strong cross‑platform consistency when NN sequences, preprocessing, and analytical conventions are harmonized. Stress‑test results indicate computational robustness rather than clinical validation.

Authors:Ryuichi Sumida, Koji Inoue, Tatsuya Kawahara
Title: MemUse: Moving Memory Evaluation from Direct QA to Natural Integration in Long-Term Human-AI Conversation
Abstract:
Memory systems for conversational LLMs are conventionally evaluated by direct, fact‑seeking questions about prior dialogue (Direct QA): can the model recall fact X from a prior conversation? We tested whether higher Direct QA accuracy correlates with higher user satisfaction in a 4‑month deployment (40 users, 1,872 sessions, 7 memory conditions). Existing‑benchmark Direct QA varies from 19.7% to 70.1% across the 7 conditions, but satisfaction does not change. We hypothesize that existing benchmarks and user satisfaction are tracking different capabilities: benchmarks measure elicited retrieval (recall when asked), while conversation requires natural integration (detecting relevance and naturally weaving prior context into a response). To examine this, we introduce MemUse, a set of real user‑cued memory moments drawn from the deployment, scored by an integration‑aware judgment of the natural conversational response. Holding the model and context fixed, the same system that scores 78.8% on Direct QA references only 7.9% of those facts in conversation ‑‑ a 71‑point gap. Within these moments, Natural Integration is associated with satisfaction, whereas Direct QA is not. We release the deployment corpus and MemUse together with all judgments and scoring prompts at https://github.com/ryuichi‑sumida/memuse.

Authors:Nian Li, Chonggang Song, Jingtao Ding, Lingling Yi, Yong Li, Qingmin Liao
Title: Tlow: Flow-based Item Tokenizer for Recommendation
Abstract:
Item tokenizer encodes semantic embeddings into token IDs to replace the randomly assigned item IDs used in traditional recommendation models, fundamentally addressing the problems of excessive parameters and cold starts. However, the most common tokenizer, RQ‑VAE, suffers from low decoding efficiency due to the inherent dependencies among its codebooks. Meanwhile, efficient independent tokenizers such as optimized product quantization (OPQ) still struggle with dimensional correlations and distribution complexity of semantic embeddings. In this work, we propose a f\underlinelow‑based item \underlineTokenizer (Tlow) to transform raw semantic embeddings into a latent space where embeddings conform to a unified standard normal distribution, achieving dual advantages of dimensional independence and distributional simplicity. Independent tokenization performed on these latent embeddings yields semantically clear token IDs. Additionally, we introduce a novel codebook guidance to align the codebook space with the token embedding space, further aiding the learning of more semantically distinct token embeddings. Offline experiments on four public datasets demonstrate that Tlow's tokenization and codebook guidance significantly improve recommendation performance. The improvement on cross‑domain and multi‑modal recommendations also proves the effectiveness of item tokenization in a simplified embedding space. Online experiments for a multi‑modal retrieval task on China's largest social media platform WeChat validate Tlow's powerful distribution transformation capability. The retrieval model based on token IDs improves user CTR by 10.32% globally and by 11.64% for new items. Our codes are available at https://github.com/wjjln/Tlow.

Authors:Feiyu Shen, Fenglong Xie, Junjie Li, Kun Xie, Lei Xie, Xu Tang, Xuelong Geng, Yan Jia, Yao Hu, Yichen Han, Yichen Wu, Ziqi Dai, Junjie Chen, Kai Huang, Manzhen Wei, Yixuan Li
Title: FireRedAudio: A General-Purpose Audio Language Model with Decoupled Continuous Representations for Understanding and Generation
Abstract:
A unified audio model must recognize and understand linguistic, paralinguistic, and environmental information while supporting speech synthesis and editing. A key challenge is representation: understanding favors compact features suited to long‑context modeling, whereas speech generation requires reconstructible features that preserve fine‑grained acoustic detail. We introduce FireRedAudio, a general‑purpose audio language model with a shared 9B‑parameter LLM. To the best of our knowledge, it is the first publicly disclosed unified audio‑language model to provide separate continuous input representations for understanding and generation within a single trainable autoregressive LLM. Audio to be recognized or analyzed is processed by a dedicated Audio Encoder, while speech inputs for generation use a RedAE‑based pathway. The LLM directly generates text or conditions a flow‑matching DiT to produce continuous acoustic latents. Through progressive multitask training, FireRedAudio supports ASR and audio understanding, with the latter extending to recordings of up to one hour, as well as zero‑shot TTS, Instruct TTS, and semantic and acoustic speech editing. Its structured organization of long‑form audio achieves second‑level timestamp accuracy. Across comprehensive evaluations, FireRedAudio achieves competitive or leading performance in audio understanding and multilingual ASR, strong content accuracy and speaker preservation in zero‑shot TTS, leading instruction following in Instruct TTS, and substantial improvements over Ming‑UniAudio‑Edit in both semantic and acoustic speech editing. These results demonstrate the viability of decoupled continuous input representations for unifying audio understanding and continuous‑latent speech generation in a model of moderate scale. Our code is available at https://github.com/FireRedTeam/FireRedAudio.

Authors:Zi Qian Yong, Ajinkya Kulkarni, Julia Lau, Hwa Hui Tew, Shu Min Leong, Raphael Phan, Sébastien Marcel
Title: On the Robustness of Audio Deepfake Detection under Audio Watermarking
Abstract:
Recent advances in generative audio models have enabled highly realistic synthetic speech, increasing the importance of reliable audio deepfake detection (ADD) systems. While prior studies have primarily focused on adversarially optimized perturbations, the robustness of ADD systems under realistic signal transformations remains insufficiently understood. In this work, we investigate the impact of audio watermarking on ADD systems by treating watermarking as a structured, non‑adversarial perturbation rather than a conventional attack mechanism. Using a watermark‑based evaluation framework built upon WavMark, we evaluate multiple self‑supervised learning (SSL), Convolutional Neural Network (CNN) and Graph Neural Netrowk (GNN)‑based ADD models across several benchmark datasets. Beyond conventional detection metrics, we further analyze watermark‑induced representation shifts using Fréchet Distance, cosine similarity, and L2 distance in the embedding space. Experimental results reveal a strong dataset‑dependent behavior: watermarking causes substantial performance degradation on ASVspoof 2021 LA and DF, while exhibiting limited impact on ASVspoof 2024, FoR, and ITW. Moreover, large embedding‑space shifts are strongly associated with severe detection degradation, suggesting that watermark‑induced perturbations can substantially alter the feature representations relied upon by current ADD systems. These findings demonstrate that benign signal transformations designed for content protection can expose previously overlooked robustness vulnerabilities in audio deepfake detection systems. Our code is available at https://github.com/ziqian0925/wm‑ADD‑robustness.git

Authors:Long Hoang Pham, Quoc Pham-Nam Ho, Huy-Hung Nguyen, Duong Nguyen-Ngoc Tran, Ngoc Doan-Minh Huynh, Cu Quoc Le, Hoang-Khang Nguyen, Hyung-Min Jeon, Chi Dai Tran, Son Hong Phan, Duong Khac Vu, Trinh Le Ba Khanh, Jae Wook Jeon
Title: Rethinking Pre-Training and Augmentation for Zero-Shot Cross-City Object Detection
Abstract:
Real‑world deployment of traffic surveillance systems is bottlenecked by geographic domain shift, in which models trained in one city underperform when applied to an unseen target city. Conventional domain adaptation relies on hyperparameter‑sensitive architectures or direct profiling of target data. Both are fundamentally precluded in privacy‑conscious ecosystems that require completely blind training and evaluation loops. In this setting, we explore the effects of pre‑training and augmentation in addressing the domain shift problem. Specifically, we propose a new modular training pipeline for object detection structured around two core orthogonal pillars: (1) a multi‑dataset pre‑training strategy featuring a class‑agnostic objectness distillation to decouple structural vehicle geometry from semantic taxonomies, and (2) a domain‑resilient augmentation stream featuring a novel Grayworld transformation that forces global attention heads to strip volatile chromatic shortcuts in favor of robust shape priors. When evaluated with the real‑time transformer‑based detector RF‑DETR, our framework bridges cross‑city distribution gaps while using limited GPU memory (16GB). Our optimized variants, RF‑DETR‑HR and RF‑DETR‑Grayworld, deliver a substantial empirical gain of +24.29 over the baseline, achieving 1st place (47.53 mAP) on the AI City Challenge Track 6 leaderboard. Code and data are available at: \hrefhttps://github.com/SKKUAutoLab/aic26_cross_citySKKUAutoLab/aic26\_cross\_city.

Authors:Boshen Shi, Yize Liu, Chen Zhao, Ce Chi, Zhendong Wang, Xing Wang, Junlan Feng
Title: TrustDABench: Benchmarking Reliability and Robustness of LLMs for Structured Data Analysis
Abstract:
LLMs are increasingly used to analyze spreadsheets, CSV files, and other structured data, but producing a correct‑looking answer is not the same as producing a trustworthy analysis. A trustworthy result should be supported by a valid path from the user question to the relevant data evidence. This requirement creates two diagnostic questions: whether an LLM can refuse to answer or ask for clarification when such a path does not exist, and whether it can preserve the correct analysis when the same evidence is expressed in different table forms. We introduce TrustDABench, a benchmark that operationalizes these questions as reliability and robustness. Starting from the evidence‑path view, we derive 19 perturbation operators and instantiate them through an Agentic‑LLM‑based generation framework. TrustDABench contains 2,340 human‑verified perturbed instances, and we evaluate eight representative LLMs. The results show substantial headroom: the best reliability result is only 24.21% average MRS, achieved by GPT‑5.5, while the best robustness result still has 9.10% average ASR, achieved by Claude‑Sonnet‑5. The failures are systematic: models rarely detect conflicting evidence, often continue along executable but unsupported analysis paths, and remain sensitive to perturbations that change observation boundaries or cross‑table relations. These findings suggest that stronger evidence‑boundary recognition and representation‑invariant reasoning are still needed for reliable structured‑data analysis.

Authors:Junlong Li, Junxi Li, Jianjun Gao, Chen Cai, Lap-Pui Chau, Yi Wang
Title: EgoErrorVQA: Assess Egocentric Comprehension Capabilities through Procedural Errors for Ego-Agentic AI
Abstract:
The majority of our everyday activities are procedural and consist of sequences of interdependent steps. However, existing benchmarks for Visual Agents and Visual Language Models (VLMs) overlook the evaluation of their procedural comprehension ability from an egocentric visual perspective, particularly for detecting procedural errors, a critical capability for everyday assistance. To bridge this gap, the EgoErrorVQA task is firstly proposed for egocentric procedural comprehension with explicit procedural errors modeling. Besides, we develop a user‑friendly evaluator agent based on the Agent2Agent (A2A) protocol, enabling rigorous and standardized evaluation of visual agents through VQA‑based interaction. A range of models are evaluated using both open‑ended and multiple‑choice questions, revealing persistent weaknesses in handling procedural errors and error types. Moreover, we introduce Ego‑ADR, an Adaptive Decoupled Reasoning framework that decouples complex procedural reasoning to enhance models' understanding of procedural errors. It achieves consistent performance gains over the selected baselines and attains state‑of‑the‑art results on several metrics under comparable settings. Code: https://github.com/z1oong/EgoErrorVQA

Authors:Duong Nguyen-Ngoc Tran, Ngoc Doan-Minh Huynh, Cu Quoc Le, Hoang-Khang Nguyen, Long Hoang Pham, Huy-Hung Nguyen, Quoc Pham-Nam Ho, Trinh Le Ba Khanh, Chi Dai Tran, Duong Khac Vu, Son Hong Phan, Hyung-Min Jeon, Jae Wook Jeon
Title: Syn2RealTrack: Bridging the Gap Between Synthetic and Real-World Datasets for Online Multi-View Multi-Target Tracking
Abstract:
Multi‑camera 3D perception systems for warehouse scenes are trained largely on synthetic data and evaluated on physically captured environments. The resulting synthetic‑to‑real gap, which corrupts ground‑plane localization and cross‑camera identity association, is usually treated as one deficiency for a single domain‑adaptation module to absorb; we argue instead that it enters the pipeline at three separable points: the camera calibration, the object shape prior, and the assumption that the object census is known, each admitting a different local remedy. Our online pipeline, Syn2RealTrack, follows this decomposition: lens distortion is recovered from images alone under a calibration that provides none, detections are fused across views by a visibility‑weighted part‑based descriptor that abstains on occluded parts rather than guessing, person height is measured in closed form from calibration instead of copied from a synthetic prior, and a closed‑world cardinality prior is paired with a causal filter that removes the phantom boxes the prior manufactures. The system therefore adapts by reallocating trust between geometry and appearance without retraining a feature extractor. On the AI City Challenge 2026 Track~1 evaluation server it reaches a 3D Higher Order Tracking Accuracy (HOTA) of 52.0118%. The code will be released at https://github.com/SKKUAutoLab/aic26_mc3dp

Authors:Suhwan Choi, Jaeyoon Jung, Sungkyung Kim, Yunsung Lee, Youngjae Yu
Title: PonderPounce: A Pretrained MLLM as an Episode Context Engine for Robot Control
Abstract:
Multimodal large language models (MLLMs) can integrate long visual histories, reason under partial observability, and infer behavior from a few examples. Yet vision‑language‑action (VLA) models generally inherit pretrained representations without using this contextual capacity as episode memory. Memory‑dependent policies address this gap through purpose‑built history mechanisms. PonderPounce instead reuses an MLLM's native causal context as robot memory. Ponder, a System2 MLLM, accumulates episode observations, demonstrations, and prior cognition in its native causal context and can generate subgoal text and demonstration reasoning for internal use. Pounce, a System1 VLA, receives the current observation, instruction, and proprioception directly; through the Ponder‑‑Pounce interface, it asynchronously receives only the newest continuous cognition token and its age. Both are jointly trained end to end without a purpose‑built memory module or separate bridge pretraining. Optimized serving achieves p50 latencies of 78ms for cognition refresh and 25ms for action‑model invocation, supporting 20Hz action playback. On RoboMME with base‑scale training data, PonderPounce reaches 60.83% with 9B and 50.04% with 0.8B under the same Pounce architecture and interface, versus 44.51% for FrameSamp+Modul and 17.93% for the current‑observation π_0.5. With 9x data, it reaches 75.54% versus 57.88% for FrameSamp+Modul. On RoboCasa‑DC, the same interface learns from action supervision alone and reaches 12.5% versus 11.6% for the strongest published demonstration‑conditioned baseline, falling to 8.6% when cognition is replaced by a learned null state.

Authors:Nadeem Shaikh
Title: Knowing When to Ask for Help: Bayesian Self-Escalation in Hierarchical LLM Agents
Abstract:
Current LLM agent systems decide delegation before reasoning begins (a router picks a model) or after a response is complete (a verifier scores it and may retry). We study a third regime: an agent that recognises, during its own reasoning, that it is unlikely to succeed and transfers control to a stronger model. We formulate intra‑generation delegation as a Bayesian optimal‑stopping problem over a learned competence posterior ‑‑ an online estimate of the agent's eventual task success whose sufficient statistics are learned from labelled trajectories, not read off raw entropy. We derive the myopic escalation threshold in closed form, characterise the optimal policy via dynamic programming, and prove that the optimal policy is a time‑varying threshold with no shape assumption on the raw signal. We further prove exponential separation of the oracle belief at the Chernoff‑information rate of the signal, a regret bound governed by the calibration of the posterior, and a finite‑sample guarantee: with n labelled calibration trajectories the deployed plug‑in policy's regret decays as 1/sqrt(n). A controlled simulation study confirms each prediction of the theory, including the predicted 1/sqrt(n) rate. We additionally report a real‑model validation on a Qwen2.5‑Coder 1.5B‑>7B code cascade (MBPP, 257 tasks), confirming two of three pre‑registered predictions: the escalation frontier dominates post‑hoc routing at equal cost, and the cumulative competence belief's discrimination rises over generation.

Authors:Mingxu Zhang, Ying Sun, Yuhan Li, Yang Ji, Dazhong Shen, Ke Zhang, Shan Huang
Title: EMRB: A Multi-Level Benchmark for Evaluating LLM Reasoning over Raw Electromagnetic Signals
Abstract:
Large language models (LLMs) are increasingly used as code agents for scientific and engineering analysis, but their ability to analyze raw physical‑layer measurements remains untested. We introduce EMRB (Electromagnetic Reasoning Benchmark), which evaluates whether LLMs can analyze raw I/Q data by writing and running code. EMRB contains 200 problems across five difficulty levels and 27 question types, from signal detection to OFDM design, generated from 11 signal types with verified ground truth. Unlike benchmarks built on preprocessed features or structured tables, EMRB provides only the raw capture; the quantities each question refers to must first be discovered through code. We evaluate 14 LLMs spanning proprietary, open‑weight, and reasoning‑oriented families. Scores range from 24.1% to 78.9%, with the mean dropping from 84.9% on basic measurement to 21.2% on system design. We also propose ReconPilot, a structured method that separates signal reconnaissance, targeted analysis, and self‑verification. Across three backbones, ReconPilot raises the overall score by 3.8 to 17.6 points and improves 13 of 15 backbone‑level combinations tested. All data and code are publicly released in \hrefhttps://github.com/mingxuZhang2/EMRB\textcolorblueour GitHub repository.

Authors:Andrew James Amos
Title: A Feature-Major Codebook for Memory-Efficient Sparse-Binary Self-Organizing Maps: Scaling a MEDLINE Atlas to 1.05 Million Neurons on a Single Consumer GPU
Abstract:
A self‑organising map turns a large corpus into a browsable two‑dimensional atlas, but building one at MEDLINE scale has been impractical: the best‑matching‑unit (BMU) search that dominates training is bound by the bandwidth needed to read the codebook every epoch. I show that this bottleneck is largely an artefact of codebook layout. Storing it feature‑major with each feature's weights contiguous, W[v.M+i], recasts the search as a tiled sparse‑dense product in which every loaded weight column is reused across a tile of samples. Varying only the layout, with implementation, precision and update rule held fixed, accelerates the BMU search by 4.5‑8.5x. Because an exact‑argmin BMU is invariant to how the codebook is stored, this gain costs nothing: held‑out quantisation error agrees with a cuSPARSE baseline to within 0.5% at every map size. Against that baseline the advantage is a crossover rather than a constant: cuSPARSE.SOM is faster at small maps, SparseBin.SOM is 1.5x faster at 128x128 and 2.6x at 256x256, and at 512x512 it is the only one that runs at all on 24 GB. Paired with a radius‑independent box‑blur update and a convergence‑based stopping rule, it trains a converged map over 29.9 million MEDLINE articles in about 72 s at 64x64 on one 24 GB GPU, and accommodates 262,144 neurons (512x512 edges) where every alternative algorithm I tested exceeds memory constraints. On a 141 GB H200 it reaches 1,048,576 neurons (1024x1024 edges) ‑ to my knowledge the largest self‑organising map yet reported. Held‑out error follows a smooth power law with no elbow across three decades of map size, so the limit on resolution is compute rather than any breakpoint in the data. At matched work the design is ~82x faster than MedSOM, the CUDA implementation behind our earlier MEDLINE atlases and, at 128x128, 621x faster than the best available multicore‑CPU library.

Authors:Lyuke Wang, Zhuo Li, Guangxu Zhu
Title: VisCache: Visual KV Cache Pruning for Efficient Vision Large Language Model Inference
Abstract:
While Vision Large Language Models (VLLMs) have achieved remarkable success in multimodal reasoning, their long‑context inference remains prohibitively expensive due to the massive computation and memory overhead of visual Key‑Value (KV) caches. Existing KV compression methods often apply uniform pruning across visual tokens and layers, leading to substantial information loss and degraded performance.To address this challenge, we propose VisCache, a plug‑and‑play framework for coarse‑to‑fine Visual KV Cache pruning without training, which consists of two synergistic stages. First, a lightweight VLM filters temporal redundancy by selectively forwarding semantically informative keyframes. Second, we introduce PruneKV, a surgical KV compression algorithm tailored to the attention dynamics of VLLMs. Unlike rigid pruning strategies, PruneKV adopts a parabolic layer‑wise budget allocation together with an asymmetric update mechanism that selectively prunes keys while fusing values, thereby preserving critical contextual information. Extensive experiments demonstrate that VisCache substantially improves inference efficiency, achieving up to 2.35× speedup and significant memory reduction while maintaining competitive performance with only 19‑‑28% KV cache retention. VisCache consistently outperforms existing baselines, establishing a new Pareto frontier between efficiency and performance for long‑context VLLM inference. Code is available at https://github.com/Wlklk/VisCache

Authors:Timo Breuer
Title: SQLite is Enough. Lexical, Semantic, and Hybrid Search with scrydb
Abstract:
This work introduces scrydb, a Python library that enables lexical, semantic, and hybrid search within SQLite. For lexical search, scrydb leverages SQLite's full‑text search extension FTS5. Semantic search builds on sqlite‑vec, a SQLite extension for vector search. Furthermore, the library allows users to rerank and fuse retrieval results to combine both lexical and semantic approaches, providing a lightweight solution for downstream tasks in information retrieval (IR) or agentic search. We evaluate scrydb on various IR benchmark datasets and demonstrate its effectiveness in text retrieval based on keyword matching, semantic similarity, and rank fusion. In addition, we provide insights into query latency and the trade‑off between efficiency and effectiveness. scrydb is available under the MIT license.

Authors:Sang Won Lee, Hyogu Jeong, Namwoo Kang
Title: PhysicsBench: A Unified Leaderboard for Generative and Predictive Models in Engineering Design and Simulation
Abstract:
Generative and predictive artificial intelligence models are increasingly used to generate geometry and to predict physical fields and scalar quantities in engineering design and simulation. Yet these models are typically evaluated in isolation, on academic datasets at unconstrained scales, with inconsistent metrics and procedures. We present PhysicsBench, a unified benchmark and leaderboard that evaluates generative and predictive models under one standardized procedure. PhysicsBench spans seven generation and prediction tasks across 1D, 2D, and 3D domains and ranks 66 models on nine datasets, comprising industrial‑scale CAD/CFD/FEA simulations and public references, expanded into 28 configurations. One procedure and ranking apply to both families, each ranked within its own tasks. Evaluation spans realistic, limited data scales from S to XL rather than the unlimited training sets common in academic benchmarks. A common metric suite captures geometric fidelity with distributional distances, physical‑field and scalar accuracy, and engineering‑specific field‑ and shape‑validity. BenchRank debiases correlated metrics and ranks by PageRank over a head‑to‑head dominance graph, so every reported quality metric is also ranked, with computational cost in a separate efficiency view. Across tasks, an architecture's large‑scale academic standing weakly predicts its small‑data ranking. The top model changes with data scale in six of the seven tasks, and no model leads more than one task. PhysicsBench turns "state‑of‑the‑art" from a self‑reported claim into an openly published foundation for model selection.

Authors:Junjie Zhou, Ke Mei, Lei Li, Tianyi Wang, Fengyun Rao, Jing Lyu
Title: WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report
Abstract:
Universal multimodal embeddings are becoming a core component of modern AI systems, enabling heterogeneous content to be represented in a shared space for applications such as retrieval, recommendation, classification, and agentic systems. In this report, we present WeMM‑Embedding, a family of universal multimodal embedding models supporting text, images, videos, visual documents, and arbitrarily interleaved multimodal inputs with flexible output dimensions. The family comprises 2B, 4B, and 9B variants and is trained in two stages: a large‑scale multimodal alignment stage, followed by a refinement stage using curated data, fine‑grained relevance supervision, and cross‑scale knowledge transfer. Across extensive evaluations, WeMM‑Embedding achieves leading performance on multiple public benchmarks. Notably, the 2B variant already surpasses the previously leading 8B open‑source baseline on MMEB‑v2, while the 9B variant further achieves a new state‑of‑the‑art overall score of 80.6. WeMM‑Embedding also demonstrates strong practical performance across WeChat applications, with substantial gains on a 26‑task in‑house benchmark and consistent improvements across 14 online A/B tests. It has been deployed at scale across recommendation and search applications, including WeChat Channels, Official Accounts, Moments, and e‑commerce services. We have released the model weights and code to facilitate future research at https://github.com/Tencent/WeMM‑Embedding.

Authors:Haoran Hao, Shahram Najam Syed, Jeff Schneider, Jeffrey Ichnowski
Title: Hierarchical Skill Retrieval for Data-Efficient Adaptation of Vision-Language-Action Models
Abstract:
While Vision‑Language‑Action (VLA) models pretrained on large‑scale robot datasets provide a strong foundation for robot manipulation, their performance can degrade when adapted to new tasks with limited task‑specific demonstrations. Retrieval offers a practical way to reuse existing demonstrations for data‑efficient adaptation, but existing methods often rely on visual similarity, state‑action representations, or task‑level language matching. These approaches may overlook the hierarchical structure of long‑horizon manipulation tasks, where complete task matches are rare but reusable skills are often abundant. To address this challenge, we propose Hierarchical Skill Retrieval (HSR), a retrieval framework for data‑efficient VLA adaptation. Specifically, HSR first decomposes a target task into candidate skill sequences. It evaluates each plan based on both semantic plausibility and skill reliability estimated from the prior dataset. The selected decomposition is then used for hybrid retrieval. This combines subtask‑level language retrieval with behavior‑feature reranking to identify demonstrations that are both semantically relevant and compatible with the target task. Finally, we adapt the policy through a two‑stage pretraining and finetuning pipeline, which separates general skill acquisition from task‑specific adaptation. Experiments on the LIBERO benchmark and several real‑world robot manipulation tasks show that HSR improves the average success rate by 10.3% and 21.3% over the strongest baseline, respectively. These results demonstrate the effectiveness of structured skill‑level retrieval for data‑efficient VLA adaptation. Videos and code are available at https://hoar012.github.io/HSR‑Project.

Authors:Sibo Tian, Chang Liu, Minghui Zheng, Xiao Liang
Title: NeurRAFT: Robot Motion Planning via Anchor-Level Flow Matching with Clearance-Aware Preference Tuning
Abstract:
Recent end‑to‑end neural motion planners generate trajectories from raw sensor observations, avoiding the privileged geometric models required by classical planners. However, collision‑free planning in cluttered environments remains challenging. We present NeurRAFT, a generative planning framework based on anchor‑level flow matching and clearance‑aware preference tuning. Unlike prior neural planners that model dense waypoint sequences and spend capacity on redundant local details and smoothness, NeurRAFT operates on compact anchor waypoints. We train the planner using a Jacobian‑weighted loss that accounts for the task‑space impact of each anchor. At inference, the anchors are generated in two integration steps, followed by cubic‑spline interpolation to recover a smooth, full‑resolution trajectory. Since imitation learning from positive demonstrations cannot distinguish collision‑free from near‑collision trajectories, collision‑prone behaviors persist at test time. Rather than relying on post‑hoc corrections, we directly reshape the pretrained planner's distribution toward safer solutions without augmenting inference. Specifically, Direct Preference Optimization shifts probability mass toward trajectories with larger obstacle clearance, with the resulting improvement directly absorbed into the planner parameters. Experiments show substantial improvements over state‑of‑the‑art planners, while real‑world experiments demonstrate zero‑shot transfer to a Franka robot under noisy and partially occluded depth observations. Video results available at https://neurraft.github.io/.

Authors:Akshay Jaitly, Siavash Farzan
Title: Trusted Polytopic Action Sets for Fast Planning in Underactuated Systems
Abstract:
Underactuated systems pose a challenge for convex motion planning because their dynamically feasible motions lie on a manifold of trajectories in function space. Building on our earlier formulation of polytopic action sets (PAS), this paper presents a method for rapidly generating, online, trusted convex sets of short‑horizon actions for underactuated and potentially nonlinear systems. Around a nominal trajectory, we construct local finite‑dimensional action coordinates in which each parameter vector encodes a complete nearby motion through an affine trajectory map, rendering collision‑avoidance and control bounds linear. To remain consistent with the nonlinear dynamics, we introduce a dynamics‑violation metric and extract a trusted convex inner approximation using an IRIS‑inspired inflation procedure directly in action space. The resulting PAS are reusable convex families of actions that can be queried and composed with linear programs, and a PAS‑guided tree expansion treats nodes as composed reachable families rather than single trajectories, coupling local nonlinear fidelity with convex reuse for longer‑horizon planning. The planner solves cluttered planar scenes in tens of milliseconds (14‑78x faster than a kinodynamic RRT baseline) and reduces terminal error on a nonlinear underactuated benchmark by 26‑86% over sampling and NLP baselines.

Authors:Hengjie Zhu, Dayan Wu, Zihao Zhang, Xinze Liu, Jingxuan Yu, Peng Fu, Zheng Lin, Weiping Wang
Title: Absorbing Gradient Conflicts: Modeling Semantic Variance via Kent Distributions for Cross-Modal Hashing
Abstract:
Supervised proxy‑based deep cross‑modal hashing has become the dominant paradigm for large‑scale retrieval. However, prevalent methods model class proxies as deterministic points in the embedding space. This rigid assumption causes severe gradient conflicts in multi‑label scenarios, where gradient conflicts arising from label co‑occurrence lead to severe gradient contention and optimization collapse. To resolve this, we propose Kent‑based Distributional Proxy Hashing (KDPH), a novel framework that shifts proxy representation from static points to flexible anisotropic Kent distributions on the hypersphere. Unlike point proxies that must shift their positions to accommodate conflicting gradients, KDPH absorbs these conflicts by dynamically adjusting its directional variance. This allows the proxy to maintain a stable semantic mean direction while stretching to cover diverse label correlations. Furthermore, to ensure stable training of these geometric parameters, we derive a tailored loss function incorporating the Cayley transform to enforce strict orthogonality. To the best of our knowledge, KDPH is the first framework to successfully introduce the Kent distributions into cross‑modal hashing. Experiments on three benchmark datasets demonstrate that KDPH mitigates proxy collapse and chaotic oscillation, significantly outperforms state‑of‑the‑art methods. Code is available at https://github.com/Senmo996/KDPH‑official‑code.

Authors:David Huang, Wenkai Yang, Kuiye Ding, Haiyang Xin
Title: IC-ThermBench: An Open, Progressive Benchmark for Generalizable 2.5D/3D-IC Thermal Learning
Abstract:
Standardized benchmarks are fundamental to reliable progress in AI for EDA, including learning‑based thermal modeling. However, existing thermal prediction studies often rely on different datasets, simulators, data splits, preprocessing pipelines, and metrics, while most datasets and implementations remain unavailable, making fair and reproducible comparison difficult. We introduce IC‑ThermBench, an open and progressive benchmark that combines established 3D‑IC steady‑state, transient, and industrial package tasks with a new 50,000‑sample 2.5D chiplet extension designed to evaluate progressively broader represented physical variation and cross‑package OOD transfer. Five Generalization Scopes cover 3D‑IC fixed‑design prediction, Within‑Family Generalization under layout, material, and boundary‑condition variation, and Cross‑Package OOD transfer to unseen package systems. We evaluate eight representative baselines under common data, splits, labels, and metrics. Performance degrades gradually from S2 to S4 as represented physical support broadens, but Cross‑Package OOD produces a much sharper degradation: the best RMSE and MAE increase from 1.216 and 0.938~K at S4 to 15.99 and 15.00~K at S5, respectively. With only 10 labeled samples per OOD case, target‑domain adaptation reduces the best MAE to 2.60~K. IC‑ThermBench further provides a unified generation, training, inference, and evaluation pipeline, enabling reproducible and fair comparison of existing and new thermal predictor.

Authors:Linzhe Jiang, Jiayuan Huang, Changhao Zhang, Chunyang Jiang, Zhehua Mao, Mobarak I. Hoque
Title: RefineRank: Joint Box Refinement and Ranking for Surgical Spatio-Temporal Grounding
Abstract:
Surgical spatio‑temporal grounding (STG) requires locating, at each video time specified by a procedural question, the object that the question asks about. Existing approaches face a trade‑off: vision language models understand the question context but produce imprecise coordinates, whereas open‑set detectors provide localized candidate boxes whose confidence does not reflect which box answers the question. We introduce RefineRank, which closes this gap at the candidate‑box level. A compact trainable module, RefineNet, combines the language and regional features of a frozen medical vision language model with the proposals of a frozen open‑set detector: it predicts a bounded coordinate correction and a quality score for every candidate box, and a fixed decoding rule returns the original or refined box with the highest score. On the MedVidBench Official Rankings (Verified), RefineRank records 0.421 STG mIoU, the highest displayed STG score, while its global multi‑metric rank is 11. In a controlled evaluation on separate training and evaluation videos, coordinate correction raises the candidate oracle upper bound from 0.6772 to 0.7302, and ranking the joint pool of original and refined candidates by their RefineNet scores improves STG mIoU from 0.2719 to 0.4534, whereas separately trained selectors over the same pool reach at most 0.4186. These results show that a small box‑level module can reconcile question understanding with precise localization without retraining either backbone. Code is available at [https://github.com/linzhe001/RefineRank](https://github.com/linzhe001/RefineRank).

Authors:Linhan Wang, Zijian An, Mingyuan Zhang, Chen Dai, Yi Xu, Can Cui, Zichong Yang, Yinlin Chen, Lifeng Zhou, Chang-Tien Lu
Title: GlanceWAM: Sparse Test-Time Imagination for World-Action Models
Abstract:
Video generative models provide rich physical priors for robot learning, yet existing world‑action models (WAMs) face a fundamental trade‑off: synchronous video generation at control rate is latency‑prohibitive, while abandoning test‑time visual imagination sacrifices task success. We show that visual imagination achieves both real‑time inference and superior success rates when generated asynchronously off the critical path and consumed directly in latent space. We introduce GlanceWAM, which decouples imagination from control within a single video DiT: an asynchronous proposer glances ahead on a slow clock to imagine a single lookahead frame seconds into the future in the background, while an action head decodes action chunks at control rate (48 ms) purely in latent space without blocking. Enabled by a non‑interfering attention mask that isolates video representations and staleness‑robust horizon training that accommodates asynchronous lookahead aging, GlanceWAM breaks the speed‑success dilemma. Trained purely on demonstrations, it attains 72.2% on the 24‑task RoboCasa kitchen benchmark (surpassing synchronous Cosmos Policy at 67.1% and imagination‑free co‑training at 64.4%) and 99.0% on LIBERO, executing at 48 ms per chunk on an NVIDIA A100 GPU (24x faster than synchronous baselines). Code is available at https://github.com/linhanwang/GlanceWAM.

Authors:Yuanhao Sun, Huawei Ji, Yuan Jin, Cheng Deng, Luoyi Fu, Xinbing Wang
Title: HAP: Head-Adaptive Visual Token Pruning via Cross-Modal Alignment
Abstract:
Recent Vision‑Language Models encode high‑resolution images into long visual token sequences, incurring prohibitive prefill costs. To compress them, existing methods score each visual token by averaging text‑to‑visual attention uniformly across all heads, which assumes every head matches the query. However, our empirical analysis shows that misaligned heads dominate the average, amplifying background tokens and drowning out fine‑grained cues. To address this, we propose PAQ (Prompt‑Grounded Attention Quality), a metric quantifying how well each head aligns the prompt with image regions. Built on PAQ, our pruning proceeds in three stages. Given a target FLOPs budget, we first partition the transformer layers into groups and allocate a visual token budget to each. Within each group, we then aggregate per‑head attention maps via PAQ‑weighted softmax into a group‑level matrix. Finally, we score visual tokens by this matrix's magnitude and retain the allocated budget per group. By weighting heads with PAQ, our method scores tokens by attention signals that more faithfully reflect prompt relevance, rather than diluting them through uniform averaging. Across 18 benchmarks, our method delivers state‑of‑the‑art trade‑offs. Specifically, on LLaVA‑1.5‑7B (9 tasks), retaining only 5.6% tokens preserves 99.1% of the original performance, surpassing the strongest baseline AutoPrune by 4.2 points. Code is available in https://github.com/baokou‑fw2/HAP.

Authors:Andrei Mikhailov, Mikhail Burtsev, Alsu Sagirova
Title: MARS: Multi-Specialist LLM Relay System for Competitive Programming
Abstract:
Large Language Models excel at code generation, yet competitive programming exposes a persistent failure mode: existing multi‑agent pipelines distribute work over generic planner, coder, and debugger roles and delegate the choice of algorithmic technique to the backbone alone. We present MARS (Multi‑Agent Relay of Specialized LLMs), a prompt‑only framework in which each agent is a topic specialist‑‑‑dynamic programming, graphs, strings, geometry, and so on‑‑‑grounded by retrieval‑augmented generation over an algorithm‑theory corpus. Given a problem, retrieval selects a small team of relevant specialists; a starter writes an initial C++17 solution, and each subsequent turn runs the candidate against public examples in a sandbox, lets the active specialist keep, repair, or hand off the draft, and forwards a structured packet to the next specialist. A single infrastructure‑fixer pass normalizes boilerplate at the end. On the CodeContests test split with Gemma 4, MARS reaches 0.624 \pm 0.006 pass rate at 2.3 recorded pipeline stages per task (+14.4 percentage points over direct prompting), closing most of the gap to CodeSIM (0.731) at 3.3× lower wall‑clock cost and substantially smaller variance in per‑task token spend. The source code is available on GitHub: https://github.com/fckand/mars.

Authors:Soheil Kolouri
Title: Partial Optimal Transport on the Circle for All Transported Masses in O(N log N)
Abstract:
Partial optimal transport compares two measures while leaving part of the mass unmatched, which is what makes it robust to outliers, occlusion, and clutter. The quantity of interest is usually the whole profile ‑ the optimal cost at every transported cardinality ‑ because the right amount to transport is rarely known in advance, and on the real line the PAWL algorithm returns that profile in O(N\log N). Much data is periodic rather than linear: angles, phases, orientations, time of day, hue, and every direction obtained by projecting onto a great circle. On the circle the same problem acquires a global circulation, or equivalently an optimized cut, which the naive exact method handles by running the line algorithm once per support gap, at O(N^2\log N). We show that this factor N is unnecessary. The line structure survives in cut‑free form, and a free‑gap invariant supplies, at every step, a cut at which all previous local updates remain valid line updates. This yields PAWC: an exact O(N\log N) time, O(N) memory algorithm returning all K+1 costs, nested active sets and plans in one run, together with a single gap that is simultaneously optimal for every cardinality. Slicing over great circles extends it to \mathbbS^d‑1. Empirically the whole profile costs 0.56ms at N=4096 against 1.5s for a single transported fraction from a general solver; on occluded, cluttered mpeg‑7 shapes, holding the descriptor fixed and varying only the cost, it retains 66% of the clean‑data retrieval score against 16% for balanced circular OT, and on \mathbbS^2 it halves the fitting error of spherical sliced Wasserstein against contaminated targets, synthetic and real. Code is available at https://github.com/mint‑vu/Partial_Wasserstein_on_Circles.

Authors:Akash Raj, Sargam Sahu
Title: Names Can Hurt: Spotting Slopsquatting Risks Caused by Package Name Hallucinations in Local Coding LLMs
Abstract:
When a code generating language model fabricates a Python package name, an adversary who has pre‑registered that name on PyPI can convert that hallucination into a supply chain compromise. This event has been termed as 'slopsquatting'. We propose a two layer detector to counter this issue. The first layer performs a deterministic PyPI existence check. The second is a Random Forest classifier trained on ten features derived from the package name and its PyPI metadata. An import name reconciler bridges the two, resolving cases such as 'import cv2' versus 'pip install opencv‑python' without a security bypass. The detector is embedded in a LangGraph state machine that retries at escalating temperatures and, on repeated failure, routes to a stronger fallback model. Across 300 curated prompts, the pipeline produces hallucination free code on 76% of runs. The primary exhausts its retry budget on 28.7%; intra model retries recover roughly a quarter of those, and cross model fallback recovers a further 16.5% of the remainder. Four findings have been observed. First, half of the flagged hallucinations are packages already registered on PyPI, as low quality lookalikes of well known projects, caught by the classifier rather than the deterministic layer (e.g., pil, faiss, tabula, haystack). Second, hallucination rate scales almost linearly with prompt adversariality, from 0 to 10% on routine coding to 40 to 73% on slopsquat baits. Third, the weaker primary refused 6 of 10 direct baits unaided, suggesting recent instruction tuning provides a baseline defense. Fourth, when primary and fallback share a model family, approximately 84% of primary failures recur on the fallback, motivating cross family pairing. A user study (n = 24) reports mean satisfaction 4.4 out of 5 and 21 of 24 stated adoption intent.

Authors:Joshua Penman
Title: Semantic Overlays: Mitigating Prompt Injection with Annotations Beyond Tokens and Steering Vectors
Abstract:
Everything a language model sees is tokens. The serving stack knows what each span is ‑‑ user input, tool output, instructions ‑‑ but the model must keep track of that itself, and it can lose track or be confused: text can be written to read like anything. Prompt injection is a natural exploit of this phenomenon. By scrambling the model's understanding of span identity, an attacker can induce unwanted and potentially dangerous actions. Adding a non‑textual channel to the model's input ‑‑ a way to communicate span identity beyond text ‑‑ mitigates this class of attack. We thus introduce a general steering technique called Semantic Overlays: small learned adapters applied at chosen prefill positions to a frozen model's residual stream. Laying an overlay over a span creates an out‑of‑band annotation channel that cannot be replicated by tokens. Unlike steering vectors, Semantic Overlays are trained, adaptable, and selectively applied. An overlay can encode complex semantics that reshape how the model perceives the marked span: asked to copy a code snippet under an overlay asserting that it is in a different programming language than it is, the model rewrites the snippet, faithfully, in the asserted language. Overlays are also composable, allow for transparent reading of underlying content, and can carry complex payloads ‑‑ including imperatives that the model will follow. An overlay which marks a span as "non‑executable" defends against the broad class of prompt injections that add instructions in untrusted context. We report strong results on prompt injection benchmarks: SEP separation rises from 24.3% to 96.5% with utility unchanged (our scoring rule; we also correct a defect in the published grader), TensorTrust attack success rate falls from 34.8% to 6.6%, and all four PIArena attack families drop to 0% compliance, all while marked spans stay readable (92.5% exact copy rate).

Authors:Nathalie Baracaldo, Nicolas Mello, Kush R. Varshney, Heiko Ludwig, Kate Soule, David Cox
Title: Granite.Trust Policy Tools: Shareable, Actionable Policies for Generative AI Applications
Abstract:
When it comes to safety policies for generative AI, one size does not fit all. Each organization and use case needs to mitigate different risks depending on the application context, regulatory environment, organizational values, and user personas. Yet, existing policy specification approaches are designed for traditional access control and fail to capture the nuances of GenAI application: the enforcement of content‑based constraints. We present two contributions to address this gap: (1) the Actionable Policy schema, a YAML‑based format for specifying what model responses can and cannot contain. The schema enables exception‑based policy governance, proposing exceptions to track policy violations; (2) synthetic data generation pipeline that produces policy‑aligned training data for model alignment and testing, and a set of tools to help define the schema and enforce policy. Together, these enable organizations to specify policies once and enforce them throughout the GenAI application lifecycle: from model alignment to runtime monitoring. The Actionable Policy schema, example policies, and tools are available as open source: https://github.com/ibm‑granite/granite.trust.policy‑tools We welcome new ideas, contributions and feedback.

Authors:Junqiu Yu, Pandeng Li, Yikai Wang, Jiaxing Zhao, Yujie Wei, Kaixun Jiang, Quanhao Li, Hongtao Yu, Zhihang Liu, Zhaohe Liao, Junjie Zhou, Yun Zheng, Yu Liu, Yanwei Fu
Title: AffineTok: Semantic Affine Consistency for Diffusion-Friendly Visual Tokenizer
Abstract:
Visual tokenizers increasingly inject semantic supervision into latent spaces to make downstream diffusion easier. Yet how these semantics should be organized to facilitate denoising remains underexplored. In this paper, we define the semantic recovery objective: the denoising process should recover the semantic content of the clean image from noisy latent, and a good tokenizer should make it easier. Existing approaches train a projector to predict the semantics directly from the noisy latent. We argue that this predicts the average of clean‑image semantics, whereas what really needs to be aligned is the semantics of averaged clean latents. More importantly, we demonstrate that the semantic recovery error orthogonally decomposes into the error of the optimal semantic prediction directly from the noisy latent and the error between these two predictions. We therefore identify their consistency as the missing requirement and call it Semantic Affine Consistency (SAC). To examine whether this overlooked requirement is closely related to downstream generation, we introduce M_SAC, a tokenizer‑side proxy for SAC. Across the evaluated tokenizers and diffusion model scales, M_SAC closely tracks generation quality, reaching a Pearson correlation of 0.960 with SiT‑XL gFID, thereby motivating SAC‑guided tokenizer training. We then introduce AffineTok, which promotes SAC through two complementary, training‑only components. Global Semantic Coordination Token (GSCT) coordinates the semantic organization of clean latents, keeping semantic averaging meaningful, while Posterior‑Mean Semantic Alignment (PMSA) predicts posterior‑mean latents from noisy inputs and supervises their semantics. On ImageNet 256, compared with the baseline, AffineTok reduces gFID by 26% at 20 epochs and, with continued training, achieves a new state‑of‑the‑art gFID of 1.21 without classifier‑free guidance and 1.10 with guidance.

Authors:Leila Khaertdinova, Anna Anikina, Claudia Mello-Thoms, Bulat Ibragimov
Title: Predicting Radiologist Expertise from 3D Gaze Patterns During CT Interpretation
Abstract:
Accurate interpretation of volumetric CT requires efficient navigation of 3D image volumes and attention to diagnostically relevant regions. While eye‑tracking has been widely studied in 2D medical imaging, its use for expertise assessment in CT settings remains limited. We propose a gaze‑informed transformer framework for expertise classification in thoracic CT. Using a DINOv2 backbone, radiologist fixation patterns are integrated into volumetric feature learning through (1) a learnable log‑space bias in self‑attention and (2) gaze‑weighted pooling of patch embeddings. We trained and evaluated our approach on 182 CT reading sessions from five radiologists with varying levels of experience. On a held‑out test set, the model achieves an ROC‑AUC of 0.91 and F1 score of 0.86, outperforming adapted methods. These findings suggest that incorporating visual search behavior into transformers may support objective, process‑based expertise assessment in radiology. Code is available via https://github.com/leiluk1/GazeToSkill.

Authors:Mian Zhang, Yueqin Yin, Kaiyu He, Peilin Wu, Xinlu Zhang, Mingyuan Zhou, Zhiyu Zoey Chen
Title: Mitigating Exploration Bias in RL for Multi-Instruction Following
Abstract:
RL has emerged as a powerful paradigm for enhancing the instruction following capabilities of LLMs. While existing training recipes achieve substantial gains, we find that they suffer from exploration bias towards easy instructions when the training data has multiple instructions in a prompt. This bias is caused by two main reasons: 1) the policy model's initial ability to satisfy hard instructions is too low to trigger successful exploration during RL training, so the optimization is biased towards easy instructions; and 2) canonical RL training recipes typically employ a cumulative reward (the number of instructions fulfilled), treating all instructions equally, which biases the policy model towards fulfilling easy instructions to obtain the same amount of reward. To address these issues, we first propose two metrics to measure the exploration bias in instruction following and then introduce a two‑stage framework to alleviate it: 1) Behavioral Bootstrapping, a lightweight rejection sampling fine‑tuning stage before RL to activate hard instructions; and 2) Scarcity‑Aware Rewards, a new RL reward function that assigns rewards to instructions based on their empirical scarcity. Experiments show that the proposed metrics are highly correlated with model performance, and our methods unleash the potential of RL training: our best models outperform the baselines by a significant margin across three verifiable instruction following benchmarks. We release codes at https://github.com/mianzhang/MulIF.

Authors:Li Liu, Ashmita Dua, Jiaming Qu, David T. Lee, Leilani H. Gilpin
Title: From Anonymous Shapes to Named Places: A Tool for Braille and Place-Semantic Annotation of Tactile Maps
Abstract:
On a 3D‑printed tactile map, a building felt under the finger is an anonymous shape: touch alone cannot tell which footprint is which, and a spoken description cannot reliably point to one shape at one place. We present a web‑based tool that lets a sighted helper click to add on‑shape Braille labels to an already‑generated map model, downstream of the geometry generator so that whoever knows the reader and the local Braille standard does the labeling. The tool offers click‑based OpenStreetMap matching, hand‑editable abbreviation that shrinks a name to fit a footprint, and print‑safe dot geometry with a review step that catches anomalies before printing. We demonstrate it on five printed maps of different place types, from a downtown core to a college campus and a small dining mall. In formative sessions in which ten BLV readers compared an unlabeled print with an annotated one, four read Braille fluently, so we treat Braille as one output among several rather than the only one. The tool's core is the link between coordinates, geometry, and a place's semantics, which can drive an audio readout or a non‑Braille code. The tool is available at https://leolee7.github.io/Annotate_Braille/.

Authors:Md Romyull Islam
Title: AQLoRA: A Zero-Search Recipe for Fast Quantized LoRA Fine-Tuning
Abstract:
Quantized fine‑tuning (QLoRA) saves memory but not time. It dequantizes every 4‑bit weight on the fly, so it trains more slowly than fp16 LoRA. We present AQLoRA (Adaptive‑Quantization LoRA), a recipe that buys part of that time back. One CPU pass over the weights sets everything, with no search and no calibration data. The pass ranks layers by NF4 reconstruction error and keeps the top‑K in fp16 under a memory budget. Those layers skip dequantization, which is where the speed comes from. A quality setting adapts every layer. A speed setting adapts only the top blocks, so the backward pass stops early. The rule reproduces Unsloth's hand‑curated dynamic‑4bit selection exactly, in seconds, where search‑based allocation needs repeated calibration passes. We evaluate on Commonsense‑170K across six models and four architecture families, from 1.4B to 14B. The speed setting trains 11.1 +/‑ 2.7% faster than well‑tuned QLoRA and gives up about one accuracy point. It was faster in all nine independent timing sessions, at worst by 7%. The quality setting trains 4.8 +/‑ 2.4% faster. Its accuracy is level with QLoRA on every model and within a point of fp16 LoRA, for 0.2 GiB more memory. These error bars are measured between independent sessions, not within one. Earning them taught us three rules for timing on shared hardware. Fix the measurement duration, not the step count. Measure the noise floor from a duplicated arm, not a nearly identical method. Repeat whole sessions: a floor computed inside one sweep understates the real uncertainty several times over, and the random seed controls almost none of it. We validate the recipe with controls and report the two that failed. Choosing adapter layers by weight density is no better than random. Choosing protected layers by quantization error is not either. The count of protected layers, not their identity, carries the speed effect.

Authors:Lukas Stepanek
Title: Memory-Sovereign Inference: Output-Exact Execution Beyond Full Residency
Abstract:
Storage‑backed inference is easy to overclaim: process RSS excludes charged page cache, process‑local device readings exclude board‑wide use, and successful generation does not establish correct asynchronous reuse. We present a falsifiable certificate separating representation, semantic demand, scheduler requests, and traffic while naming resource authorities, exactness horizons, and tested reuse transitions. At one Qwen3‑Next identity, the stock router selects all 48 x 512 managed layer‑expert objects during 32K prefill. Their duplicate‑free, overlap‑free canonical union gives a 43.59375 GiB semantic‑demand lower bound, exceeding the declared 34 GiB full‑residency envelope and the 11 GiB host‑hard plus physical 24 GiB‑device envelope. LRU64 execution stays within its host‑hard/GPU‑audited contract. Against one prespecified zero‑cache oracle, all 64 token IDs, every byte of 64 complete 151,936‑float logit rows, 3,408 route events, response bytes, and recorded consumer and destination identities are exact. Recurrent‑state and upstream‑runtime equality are excluded. In a matched source campaign, the buffered path completes exactly but reaches the 11 GiB host ceiling and records 33,481 memory.max events. The blocking one‑window direct path and complete eight‑window asynchronous component are exact with positive margin and zero limit events. Across six counterbalanced pairs, the complete asynchronous component takes 32.3% of the blocking direct path's wall time at identical physical source bytes per output; the comparison jointly changes queue depth, overlap, and lifecycle implementation. Separately, fourteen prespecified control/fault cells pass across later Qwen3‑Next and Gemma 4 binaries, supporting fail‑closed behavior only for the named transitions. The principal experiment is one fixed model, workload, runtime, device, and 64‑output horizon.

Authors:Kangning Wang, Haopeng Zhang, Zhiguo Jiang
Title: CRISP: Calibration-Aware Visual State Space Duality for Remote Sensing Semantic Segmentation
Abstract:
State space models, especially Visual State Space Duality (VSSD), have emerged as efficient linear‑time alternatives to Transformers for dense visual tasks. However, we observe that VSSD compresses spatial context into a global aggregation that suppresses high‑frequency responses, causing excessive boundary smoothing in remote sensing semantic segmentation. To address this, we propose CRISP, a calibration framework with two components. Its core, the Duality Calibration Operator (DCO), restores local contrast and boundary responses through residual injection and frequency calibration within the VSSD backbone, without altering its linear complexity. To retain the recovered detail, an Orthogonal Multi‑Prototype (OMP) head assigns multiple orthogonally constrained prototypes per class to model large intra‑class variance. Extensive experiments on Potsdam, Vaihingen, and LoveDA show that, with approximately 30M parameters, CRISP achieves consistent gains in mean F1 (mF) and mIoU while remaining competitive with state‑of‑the‑art methods. Code is available at https://github.com/crazylifeha/CRISP.

Authors:Wenyang Liu, Tianyi Liu, Dongshuo Zhang, Kejun Wu, Adams Wai-Kin Kong
Title: DriftAD: Visually-Guided Text Drift for Few-Shot Industrial Anomaly Detection
Abstract:
Few‑shot anomaly detection (FSAD) has recently benefited from vision‑language models such as CLIP, which enable anomaly de?tection by aligning visual features with text descriptions of normal and abnormal states. However, existing methods typically rely on static text prompts that are applied uniformly across the entire feature hierarchy and spatial dimensions. This rigid global‑to‑local matching fails to capture the highly localized and scale‑dependent physical variations of industrial defects. To address this, we propose DriftAD, a FSAD framework built on three key modules. First, an Anomaly Signal Amplification (ASA) module enhances subtle defect signals through spatial and frequency branches before text‑visual matching. Second, Visually‑Guided Text Drift (VGTD) dynamically transforms frozen CLIP text embeddings, steering them into layer?wise, spatially‑adaptive anomaly descriptors conditioned on local visual context at each encoder depth. Third, Drift‑Guided Spatial Gating (DGSG) uses the drifted abnormal descriptor as a spatial probe to selectively enhance anomaly‑relevant visual features. Addi?tionally, a drift separation loss prevents representational collapse of the drifted descriptors, and a gate supervision loss enforces spatially discriminative gating in DGSG. Extensive experiments on MVTec?AD and VisA demonstrate state‑of‑the‑art performance across all 1‑, 2‑, and 4‑shot settings on both image‑level and pixel‑level metrics. Code is available at https://github.com/wenyang001/DriftAD.

Authors:Wenhow Li, Chengwei MA, Hui Xiong, Ying-Cong Chen, Lei Zhang
Title: Platonic Representation Hypothesis on World Models
Abstract:
World models have demonstrated significant potential for perceiving and simulating complex environments. Despite their strong performance, the fundamental nature of their learned representations remains poorly understood. In this paper, we investigate the Platonic Representation Hypothesis within this domain by proposing the Predictive Consistency Assumption: we posit that the optimization of a shared state transition objective acts as a selective pressure that encourages heterogeneous models to converge toward a shared latent structure. Through systematic experiments with the DINO World Model (DINO‑WM), in which we vary visual encoders to create heterogeneous models, we find that capable world models evolve toward geometrically similar internal structures. Moreover, via model stitching, we show that the internal features of one world model can be mapped to another with limited performance degradation, providing evidence of functional compatibility. Our findings suggest that the pursuit of predictive consistency can promote shared, transition‑compatible latent structure across world models.

Authors:Yang Yu, Yilin Jiang, Zexuan Fei, Yiming Luo, Xingkai Song, Kaiyi Huang, Aimin Zhou, Xin Lin, Fei Tan
Title: ADE: Agentic Data Evolution Framework for Human-Centered Objectives
Abstract:
Aligning large language models to human‑centered objectives is difficult when targets are non‑executable and context‑dependent, limiting reliable verification and scalable supervision. Although synthetic data expands coverage, weak verification shifts the bottleneck from generation to selection. Noisy signals destabilize iterative refinement and can cause silent regressions. We propose Agentic Data Evolution (ADE), a data‑centric framework that organizes synthetic supervision as evolving data snapshots. ADE improves data snapshots through a closed‑loop Observation‑Variation‑Selection (OVS) procedure, where a steady‑state admission mechanism acts as a quality ratchet that conservatively gates updates for sustained cross‑round improvement. We validate these improvements through complementary intrinsic trend tracking and extrinsic post‑training evaluation. On DEV300, ADE raises the intrinsic win rate from 50% to 75.81% and the extrinsic win rate from 55.20% to 68.86%, consistent performance gains across diverse benchmarks. Blind expert evaluation further confirms this, with a 66.11% preference for evolved answers. These gains extend across post‑training methods, model scales, and tasks beyond the target weakly verifiable educational objectives. Resources are available at https://github.com/ZeroLoss‑Lab/Agentic‑Data‑Evolution.

Authors:Stephen Chung, Wenyu Du, William J. Wesley
Title: Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment
Abstract:
We study autonomous mathematical discovery in the Station, an open‑world multi‑agent environment in which AI agents from different model families pursue a shared research goal without a central coordinator or scripted pipeline. Agents choose their own research directions, conduct experiments, collaborate, and build a shared scientific literature. Across 12 construction problems from the AlphaEvolve catalogue and two additional case studies, the Station obtained results novel relative to the prior literature on five problems: a new infinite family of finite‑field Kakeya sets, new exact 604‑point kissing configurations in dimension 11, new records for the discretized Kakeya needle and sign uncertainty problems, and a substantially improved lower bound for Erdős's minimum‑overlap problem. Agents also discovered novel infinite families for Book Ramsey numbers. Importantly, the agents produced not only numerical constructions but also theorems and analyses explaining how those constructions work, making the results more interpretable and easier for mathematicians to build upon. We release all raw agent dialogues, proofs, and verification code, providing a transparent record of how these discoveries emerged.

Authors:Glauco Rampone
Title: Improved bounds for the smallest 4-chromatic graph of girth six
Abstract:
For integers k,g \ge 3 let n_g(k) denote the minimum order of a graph with chromatic number k and girth at least g. Exoo and Goedgebeur (DMTCS 2019) proved 26 \le n_6(4) \le 66; their 66‑vertex witness has remained the smallest known 4‑chromatic graph of girth 6. We improve both bounds to 29 \le n_6(4) \le 64. The upper bound is witnessed by an explicit 4‑chromatic graph of girth 6 on 64 vertices with 152 edges; it is vertex‑ and edge‑critical, and its automorphism group is cyclic of order 8 and acts semiregularly. The lower bound is an exhaustive isomorph‑free computation in the SAT modulo symmetries framework with co‑certificate learning, driven by the Liu‑Postle edge‑density bound for 4‑critical graphs of girth five; it re‑derives n_6(4) \ge 26 by a disjoint method and is validated on the known values n_4(4)=11 and n_5(4)=21. We complement the bounds with structural obstructions: no smaller witness arises from either known witness by local modifications; no 4‑chromatic Cayley graph of girth 6 exists on 54‑63 vertices (for orders 59 and 61 no vertex‑transitive witness exists at all); and no witness on at most 63 vertices admits a semiregular automorphism group with two or three vertex orbits, for any finite group. Since every known witness of an n_g(4) record with g \ge 6 is a lift of a small base graph along a semiregular action, these results close the most symmetric part of that regime below 64 vertices. All properties of the new graph are verified by independent programs and formally certified in the Lean 4 proof assistant: the non‑3‑colourability is established inside Lean by a formally verified checker that re‑validates a 219,532‑node refutation certificate, with a machine‑checked soundness theorem.

Authors:Esmail Gumaan
Title: Feedback That Backfires: Why Small Language Model Agents Repeat the Call They Just Watched Fail
Abstract:
Agent harnesses record a failed tool call and its error message in the transcript and ask the model to continue, on the assumption that the error is corrective information. We measure whether it is. Defining the corrective gain of a failure record as the change in log‑probability of re‑emitting the action that just failed, we find the gain is negative for every instruction‑tuned model we tested (6 checkpoints, 135M‑1.7B, 4 families) in two environments: simulated tool calling and MBPP program repair. Normalised by action length the effect is about ‑1.03 nats per action token, a factor of 2.8 in the odds of each token, and holds on 90%‑100% of individual items, not only on average. Over a fixed candidate set the probability of repeating the failed call rises from 0.06 to 0.54, and greedy decoding reproduces it token for token on 19% of items after the failure versus 0% before. Counterfactuals pairing the same call with a failure message, a success message, or a neutral acknowledgement separate two effects: the failed call's surface form accounts for 83% of the damage, while the semantic contribution of marking it failed is small and inconsistent in sign across environments. The problem is in the harness, not the model's grasp of error messages, and that predicts which remedies work. Replacing the verbatim call with a runtime‑generated description of the failure removes 76% of the inversion at no token cost, and making previously‑failed strings unreachable at the decoder acts on the same term. Two plausible remedies do not: an explicit "do not repeat" instruction leaves the measured quantity where it was, and deleting the failed attempt to retry from a clean context, the standard prescription for context contamination, is the worst harness we measured for repetition, because it restores the context that produced the failure. The study runs end to end on a CPU; all artefacts are released.

Authors:Heather Renze
Title: Auditing the Synthetic Memoir: Measuring Scene-Level Confabulation in LLM-Generated Autobiography Against the Documented Record of the Life It Describes
Abstract:
When a large language model (LLM) is asked to write a person's life, how much of what it writes actually happened? We present a scene‑level case‑study audit ‑ the first quantified audit of LLM‑generated autobiography against a subject‑specific ground‑truth corpus that we are aware of, based on an unsystematic literature search. The subject and the author of this paper are the same person: a 366‑day "page‑a‑day" book of first‑person anecdotal entries was drafted with a conversational LLM whose documented inputs were a template, two exemplar days, and each day's quote ‑ not her corpus ‑ and every day was subsequently audited at the anecdote‑scene level against an independent verification corpus using a four‑level rubric fixed before analysis. We define the verification‑failure rate as the share of days not rated VERIFIED (scene positively corroborated): 354 of 366 days fail, 96.7% (Wilson 95% CI 94.4‑98.1%). Only 12 days contain a corroborated scene; 19 days (5.2%) assert claims actively contradicted by the record; the dominant failure mode is grounded drift ‑ real people, employers, and settings inside invented scenes ‑ though its measured share varies across raters. Independent re‑rating replicates the headline (no evidence the original rate was inflated) while showing that the four‑way taxonomy has only fair‑to‑moderate reliability. Regenerating the same days with current named models reproduces 100% verification failure under the same inputs; grounding generation in the subject's corpus significantly improves the verification rate while leaving substantial residual failure (83.3%). We contribute the measurement, a reusable audit instrument whose WEAK/UNVERIFIED boundary we show to be unreliable, and a grounding remedy with quantified effect.

Authors:Ranjan Sapkota, Manoj Karkee
Title: Cross-Generation Optimization of YOLOv26, YOLOv11, and YOLOv8 for Fine-Grained Small-Object Detection and Instance Segmentation in Complex Orchards
Abstract:
Small‑object detection and instance segmentation remain challenging in orchard environments because of green‑on‑green similarity, occlusion, and limited pixel representation of fine fruit anatomy. This study presents a cross‑generation benchmark of Ultralytics YOLOv8, YOLOv11, and YOLOv26 for detecting and segmenting apple fruitlet, calyx, and peduncle structures for robotic orchard perception. Five model scales (n, s, m, l, and x) were evaluated under conventional 640 x 640 and small‑object focused 960 x 960 training configurations, yielding 30 experiments. Increasing model capacity did not consistently improve accuracy. YOLOv11s‑960 achieved the highest observed mask mAP@50:95 (0.402) and box mAP@50:95 (0.426), while YOLOv26s‑960 achieved comparable values of 0.397 and 0.425 with only 10.37 M parameters and 34.1 GFLOPs. Peduncle remained the most challenging class. Overall, compact‑to‑moderate YOLO models with small‑object‑focused training provided favorable accuracy efficiency trade‑offs, establishing a practical benchmark for fine‑grained agricultural robotics and orchard perception. Github Link: https://github.com/rnjnspkt/Optimizing‑and‑Comparing‑Ultralytics‑YOLOv26‑YOLOv11‑and‑YOLOv8‑for‑Small‑Object‑Detection‑and‑Seg

Authors:Idris Nechnech, Sehwan Kim, Jimin Seo, Yeongoon Kim, Minhae Oh, Sangwoo Hong, Jungwoo Lee
Title: Function-Level Execution Feedback for Code Preference Optimization
Abstract:
Process supervision has improved mathematical reasoning, where intermediate steps are naturally expressed as chains of thought. In code generation, however, process supervision remains underexplored because there is no standard notion of a step. Supervision can target lines, reasoning traces, or program states, making it unclear what to label and optimize. We propose STEP‑KTODER, a framework for code preference optimization that defines steps as module‑level functions in decomposed multi‑function programs and assigns binary correctness labels via automatically generated unit tests. Our method provides a code‑specific instantiation of stepwise KTO, combining function‑level process supervision with outcome‑level feedback on the full program. We evaluate on HumanEval(+), MBPP(+), BigCodeBench, and LiveCodeBench, showing that STEP‑KTODER improves over outcome‑only KTO and DPO. Further analysis shows that execution‑based labels are essential: LLM‑as‑a‑judge annotations systematically over‑predict function failures, corrupt positive step labels, and degrade downstream preference optimization. Code is available at: https://github.com/inechnech/STEP‑KTODER.

Authors:Parker Fawcett
Title: Rebuild Dossier: Mechanically-Enforced Specs for Agentic App Rebuilds, and What Model-Tier Failures Reveal
Abstract:
An AI agent's rebuild is only as good as the process that produced it. Prior work found that once a model is strong enough, a multi‑agent rebuild pipeline loses to the simplest approach: giving the model the original code and one instruction (AgentModernize). We present rebuild‑dossier, an open‑source tool that locks an application's real interface ‑ its exact inputs and outputs ‑ before any code is written, then enforces one‑test‑at‑a‑time building through automated checks, not written instructions alone. Three results shape this evaluation, with differing amounts of evidence. First, in a small comparison, the compliant agent failed a held‑back test while the rule‑breaking agent passed everything ‑ proof that a passing suite doesn't certify correctness when tests can be gamed. Second, we tested whether this beats simply giving the weaker model the source and one instruction: tied on a small app, but lost outright on a larger one where the automated check wasn't even running ‑ pointing to the check mechanism, not interface‑locking, which held up separately. Third, every claim here is checked at three levels ‑ the agent's own report, an automated log, and the actual files produced ‑ catching real errors, including a bug in our own logging code, that a single level would have missed. These risks reproduce on a different model and toolchain: a stronger model followed our process three times running, something the weaker model never managed. The tool is public, MIT licensed, and reproduces end to end against our own applications.

Authors:Rashid Azarang
Title: From Traceability to Justifiability: Accountability Structures in Agentic Software Engineering
Abstract:
A pipeline promoting an AI system publishes records claiming the thing evaluated is the thing deployed and that the evidence licensed the transition. We measure, from public material only, whether those records can express that claim and whether it holds where declared. First, a two‑class documentation survey of 47 delivery platforms (20 CI/CD, 27 model‑serving/agent) under one fixed three‑label protocol, graded twice (second pass blind), every consulted page pinned by content hash and date. Across 188 double‑graded cells we found no platform whose default record emits a content‑addressed identity of the behavioral tuple (model version, instructions, tool definitions, retrieval and runtime configuration); the blind pass grades that column default on zero of 47. Immutable nominal versioning is meanwhile arriving as the agent platforms' default answer (16 of 27): version integers behind mutable pointers, a layer the artifact supply chain already found insufficient. Second, an instrument computes realized assurance depth from a pipeline's published exhaust alone and compares it with the declared depth. Applied to a frozen two‑stratum frame of 30 public repositories graded twice from a hashed archive (second pass blind; cell‑level agreement 23 and 19 of 30, both passes independently finding the same five full realizations), the sharpest result is a verifiability hole: seven of the 15 repositories chosen for adopting attestation tooling publish source‑only releases, so the binding their workflows declare cannot be checked where declared. Where checkable it mostly checks out: five of seven measurable adopters realize the binding end to end; both shortfalls fall at identity binding. Together the results locate the field's records structurally short of justifiability, the one rung that can refuse a transition. The survey carries an expiry clock; we state what would falsify each finding.

Authors:Vincenzo Dentamaro, Pancrazio Auteri, Giuseppe Pirlo
Title: Squeezing the Cache, Preserving the Truth: Monotonic Equipotential Allocation with Geodesia-KV
Abstract:
Current assessment of KV‑cache compression performance confuses resident bits with read bandwidth and is affected by the artifacts of chunked teacher‑forcing. We present Geodesia‑KV, a family of training‑free KV cache policies based on monotonic block‑wise precision allocation, exact rate‑distortion residuals, and query‑sparse reading, enabling proper hardware‑ready compression. With proper separation of resident and read bits and causal evaluation, we show that Geodesia‑KV significantly outperforms other approaches. Specifically, on WikiText‑2 with 16k context, the 5‑bit operating point of Geodesia‑KV results in lower perplexity at lower bitrate than KIVI‑4 on Qwen. In addition, our compressed‑Quest version delivers improved perplexity and reduces resident (9.83 vs 16.25 bits/value) and read rates (1.95 vs 2.32 bits/value) over baseline sparse methods on PG‑19. As Geodesia‑KV is implemented as native GeodesiaKVCacheManager plug‑in of vLLM, Geodesia‑KV fully removes the need for dense cache residency via monotonic bit demotion. With the full consumer hardware evaluation, Geodesia‑KV leads to 1M‑token context generation on a single 16 GiB GPU with up to 71.7% peak VRAM savings on all leading architectures (Qwen, Llama, DeepSeek).

Authors:Tiexin Ding
Title: Data Predictability Shapes Weibull Weight-Scale Growth in Transformer Training
Abstract:
A trained transformer's weight magnitudes can be summarized by a two‑parameter Weibull distribution whose shape k \approx 1.2 is stable across layers and models, so the scale λ carries most training‑induced movement. What corpus property sets how much λ grows? Using the bigram conditional entropy D = H(\textnext \mid \textprev), a training‑free statistic computed before training, we find across controlled corruption families a learning‑rate‑conditioned law, λ^2 ‑ λ_0^2 = C_0(η) + C_1(η)(H_r ‑ D)^0.59, where H_r is a matched‑budget shuffle baseline. The convex exponent is inherited from an independently measured data‑side saturation relation rather than fitted directly to the growth curve. After removing the two per‑η coefficients, 23 runs spanning an order of magnitude in learning rate collapse onto (H_r ‑ D)^0.59 with unit slope (R^2 = 0.941; direct per‑η fits are weaker, R^2 \approx 0.82). Because D is computed before training, the law is a forward predictor: an end‑to‑end self‑validation recovers held‑out within‑family weight growth with 5.7% relative error. The readout holds at model and per‑layer resolutions and across two tested architectures, with the functional form preserved and only the coefficients changing. It also marks its boundary: cross‑corpus prediction over‑predicts code, implicating redundancy as a second axis of a broader Φ(D,R,A,H) data‑to‑weight framework.

Authors:Penghui Qi, Xiangxin Zhou, Wee Sun Lee
Title: Best Practice Critic Optimization
Abstract:
Group‑based reinforcement learning methods such as GRPO for large language models avoid training a critic by sampling multiple responses for each prompt. A reliable critic could instead estimate token‑level advantages from one response, but standard critic‑based training recipes are often unstable. We study this instability and develop Best Practice Critic Optimization (BPCO), a recipe that combines DPPO, value predictions bounded to the reward range, Monte Carlo value targets, unnormalized policy advantages, and length‑adaptive generalized advantage estimation. Because the critic is used only during training, BPCO can also condition it on reward‑defining information, such as a reference answer or grading rubric, that is hidden from the policy. Controlled experiments isolate the effect of each design choice. Across mathematical reasoning tasks with models ranging from 1.5B parameters to 30B‑A3B mixtures of experts, BPCO improves a strong critic‑based baseline consistently, and matches or exceeds a group‑based baseline while sampling one response per prompt. The same recipe also improves learning with rubric‑based rewards. These results show that a carefully designed critic provides a reliable alternative to group‑relative advantage estimation. Code is available at https://github.com/QPHutu/golden_critic.

Authors:Nan Duan, Haoyang Huang, Weiyang Jin, Haoran Li, Yaowei Li, Yuming Li, Yijun Liu, Xin Lu, Xiaoxiao Ma, Yanwen Ma, Yaofeng Su, Yilang Sun, Haoyu Wang, Zeyue Xue, Songchun Zhang, Junhao Zhuang
Title: Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds
Abstract:
Video generation is progressing beyond isolated clips toward long‑form narratives and interactive worlds, requiring models to preserve identities, follow user controls, and remain stable over extended rollouts. We present JoyAI‑Echo‑1.5, a unified audio‑visual generation system with two purpose‑built variants. The long‑video variant introduces composable cross‑shot memory that aggregates visual evidence across multiple prior shots and speaker cues derived from speech‑filtered full‑shot audio, enabling persistent character appearance and voice identity across flexible combinations of text, image, and memory conditioning. The world‑model variant converts heterogeneous navigation inputs into calibrated metric 6‑DoF camera trajectories and injects them through a geometry‑aware conditioning pathway, enabling controller‑agnostic interaction across flexible viewpoints. To support efficient long‑horizon generation, we transform a bidirectional audio‑visual backbone into a causal few‑step generator using progressive teacher forcing and short‑ and long‑horizon Self‑Gradient Forcing on self‑generated rollouts. Experiments demonstrate strong performance in both settings. JoyAI‑Echo‑1.5 achieves improvements over existing long‑video baselines in cross‑shot consistency, visual quality, text alignment, and speech fidelity. Its world‑model variant ranks first on WBench, with an average score of 81.7, and achieves leading visual quality and long‑horizon persistence on SANA‑WM‑Bench. Together, these results indicate that memory, geometric control, and rollout‑aware training provide a practical foundation for generating coherent stories and continuously evolving interactive worlds. Project page: https://echo‑team‑joy‑future‑academy‑jd.github.io/Echo‑1.5‑Page/.

Authors:Xianlei Zhou, Xiangdi Meng, Yu He, Tianyu Qi, Shuyan Guan, Xianli Zhang, Jian Zhang, Xin Li, Qika Lin, Jun Liu
Title: Beyond the Stability-Exploration Dilemma: Environmental Regularization for LLM Policy Optimization
Abstract:
Policy optimization (PO) for Large Language Models faces a stability‑‑exploration trade‑off, currently mediated by an action‑side Policy‑KL regularizer. This puts practitioners in a double bind: keeping Policy‑KL constrains response behavior and consumes the action‑side exploration budget, while dropping it leaves the optimization without an explicit drift control. We argue for an alternative that breaks the dilemma by moving regularization to the input side. As training progresses, the distribution over training queries induced by the current policy drifts unchecked from its pre‑RL reference distribution. Concretely, Environment‑Regularized Policy Optimization (ERPO) introduces a Query‑KL (QKL) term that bounds this query distribution shift, together with a dataset‑static reference‑derived per‑query weight that biases each per‑query update toward queries typical under the reference. The QKL gradient flows strictly through the query likelihood; the response score function used by policy‑gradient estimators does not appear in the QKL term, so QKL exerts no direct gradient pressure on the response distribution‑‑‑exploration is preserved. ERPO plugs into GRPO/PPO/REINFORCE‑style pipelines without additional forward passes. On six mathematical reasoning benchmarks, ERPO replaces the standard Policy‑KL regularizer while achieving effective control over query distribution drift, delivering stronger accuracy and substantially more stable behavior under high‑temperature decoding and long‑horizon training. Our source code are available at https://github.com/AlibabaResearch/ERPO

Authors:Akihiro Miki, Shun Hasegawa, Yoshimoto Ribayashi, Kento Kawaharazuka, Kei Okada
Title: Design of a Biomimetic Joint-Covering Skin with Tissue-Like Structure to Enhance Proprioception in a Musculoskeletal Humanoid
Abstract:
Proprioception in musculoskeletal humanoids is typically estimated primarily from muscle sensing, while the role of cutaneous deformation around joints remains insufficiently explored. In biological systems, mechanoreceptors distributed within soft tissue complement muscle feedback and support reliable joint state estimation. This study presents the design of a biomimetic joint‑covering skin with a tissue‑like layered structure that integrates pressure‑ and stretch‑sensitive elements within the joint‑covering tissue. The proposed skin is implemented on the musculoskeletal humanoid Musashi‑W, and its independent proprioceptive capability as well as its integration with muscle sensing are evaluated. Experimental results show that the proposed skin alone achieves joint angle estimation with an average error of approximately 3 degrees. Furthermore, integration with muscle sensing improves estimation accuracy. Owing to its joint‑covering structure, the skin may mechanically mitigate the influence of external disturbances on the muscles, and the integration of multiple modalities suggests the possibility of contributing to the identification of external stimuli that are difficult to interpret using muscle sensing alone. This work presents a design methodology for biomimetic joint‑covering skin and demonstrates that such tissue‑structured skin can serve as an effective approach for extending proprioceptive systems in musculoskeletal humanoids.

Authors:Abhilash Nandy, Rahul Seetharaman, Aman Bansal, Rounak Saha, Manav Nitin Kapadnis, Millon Madhur Das, Pawan Goyal, Niloy Ganguly
Title: CaRGo-T: Causal Reasoning Graph-of-Thought improves Multimodal Humor Comprehension
Abstract:
Large‑scale vision‑language models (VLMs) have demonstrated remarkable versatility across a wide range of multimodal tasks. However, understanding humor remains challenging because humorous content often depends on subtle interactions among entities, events, context, and implicit relationships across image and text modalities. These interactions can involve complex chains of reasoning that are difficult to capture through conventional prompting or linear chain‑of‑thought reasoning. In this work, we propose CaRGo‑T (Causal Reasoning Graph‑of‑Thought), a reasoning framework that represents the causal and contextual relationships underlying multimodal humor as a lightweight graph‑based reasoning structure. The graph is serialized into a code‑based representation generated by a VLM, which can subsequently be interpreted by the same or a different VLM to produce the final prediction in zero‑shot or in‑context learning settings. We evaluate CaRGo‑T on humor understanding and humor detection across four datasets spanning diverse forms of comedic content, including satire, sarcasm, and memes. Experiments with state‑of‑the‑art commercial and open‑source VLMs show that CaRGo‑T consistently improves performance over existing reasoning‑based baselines, achieving gains of approximately 1‑20% on humor understanding and 1‑3% on humor detection. Further analysis using mutual information indicates that the reasoning representations produced by CaRGo‑T contain more information relevant to the target output than those generated by baseline reasoning approaches. Code is available at https://github.com/abhi1nandy2/CaRGo‑T.

Authors:Seungyoon Lee, Minhyuk Kim, Jungseob Lee, Heuiseok Lim
Title: Language Chain in Alignment: Cross-lingual Ranking Preference Optimization
Abstract:
The alignment of Large Language Models heavily relies on English‑centric high‑quality preference data, which often leads to suboptimal performance in other languages. In this paper, we propose Cross‑lingual Ranking Preference Optimization~(CRPO), a novel framework that leverages robust preference knowledge from English to facilitate preference alignment in the target language. We design a hierarchical structure within parallel preference pairs across the target language and English to jointly optimize intra‑ and inter‑lingual preferences, thereby enhancing language adaptation and output quality. Building on the LambdaLoss framework, CRPO goes beyond the binary comparison based optimization by providing a relative ranking signal across multiple candidate responses. Our experiments across five languages with varying resource scales demonstrate that CRPO consistently outperforms standard approaches in both instruction‑following and knowledge utilization capability. Notably, the robust performance gains observed across various weighting schemes further validate the empirical effectiveness of our hierarchical design in a multilingual setup. Furthermore, our findings highlight that CRPO significantly improves both reward margins and the log‑probability of desirable responses, contributing to a more stable preference manifold for cross‑lingual alignment. Our code is available at https://github.com/dltmddbs100/CRPO.

Authors:Jinghui Zhang, Lang Gao, Ao Li, Mingzhe Li, Ruihong Zeng, Zirui Song, Kentaro Inui, Xiuying Chen
Title: LITERARYBIGFIVE: Author-Personalized Text Generation in a Unified Interpretable Space
Abstract:
Personalized text generation for authors and literary writing is essential for applications such as adaptive writing assistants, creative support tools, and computational literary analysis. However, existing approaches to author modeling and personalization often represent writing behavior as independent labels, requiring large‑scale corpus collection or fine‑tuning for each author or stylistic category. Such formulations are costly, difficult to interpret, and poorly suited for generalizing across authors. Inspired by the Big Five model's dimensional view of personality, we propose LiteraryBigFive, a framework that reframes authorial writing characteristics as coordinates within a unified and interpretable space. In this space, we derive each interpretable axis (e.g., Classicism, Emotionality) from activation‑space contrasts between author‑written and neutral passages, yielding distinct stylistic dimensions that allow texts or authors to be positioned within a five‑dimensional system. Beyond localizing different authors, we further introduce an interpretable steering mechanism, which adaptively guides text generation toward target coordinates to perform author‑personalized writing. Experimental results show that LiteraryBigFive improves authorial expressiveness while preserving semantic fidelity. The derived author per‑axis scores strongly correlate with real‑world literary consensus, offering transparent and interpretable explanations of author‑specific generation behavior: https://github.com/Znull‑1220/LiteraryBigFive.

Authors:Yufeng Han, Lifan Deng, Cunliang Kong, Wenhao Li, Xin Cong, Yuzhuo Bai, Kangyang Luo, Maosong Sun
Title: Jiuge-Tuiqiao: An Interpretable Human-AI System for Classical Chinese Poetry Refinement
Abstract:
Classical Chinese poetry composition has long valued Tuiqiao, the iterative refinement of words, imagery, and prosody. However, many current AI poetry systems follow a one‑shot generation paradigm, which reduces users to prompt providers and weakens their creative agency. We present Jiuge‑Tuiqiao, an interactive human‑AI collaborative system for classical Chinese poetry composition. The system is designed around a triadic model: user‑driven control, ancient‑guided evidence, and AI‑assisted generation. Users can lock characters or lines, receive real‑time prosody feedback, and obtain interpretable refinement suggestions grounded in high‑frequency collocations, PPL‑ranked classical lines, and structured knowledge extracted from classical encyclopedias. This design turns AI from an autonomous generator into a background assistant that supports the user's own process of poetic refinement. Preliminary experiments and user feedback suggest that Jiuge‑Tuiqiao improves controllability, interpretability, and user engagement in classical poetry composition.

Authors:Xiangxin Zhang, Zhanwei Zhang, Zhihang Fu, Binbin Lin, Wenxiao Wang
Title: From Inertia to Objectivity: Improving Deep Research Agents with Noise Isolation
Abstract:
Web search agents powered by Large Language Models (LLMs) show strong promise, but deep research tasks expose a recurring failure mode: once an agent has produced a query, plan, or intermediate conclusion, it becomes less objective when later judging the consequences of that same action. We term this phenomenon inertia bias. To make it measurable, we introduce the IBIS benchmark, which controls the search observations while varying whether the model is evaluating the outcome of its own prior action. We find that models are substantially worse when they "own" the preceding search step, showing that self‑authored action history can systematically distort subsequent judgment. We further show that this bias propagates into two forms of system‑level degradation: search noise at the worker level and contextual noise at the manager level. To address this problem, we propose NIS‑Agent, which applies context isolation at the two decision points most vulnerable to inertia bias: webpage triage and final‑answer validation. Across GAIA, WebWalkerQA, BrowseComp, and BrowseComp‑zh, NIS‑Agent achieves competitive performance while reducing token cost by 33% compared to our baseline. We further train an 8B model to be intrinsically more resistant to inertia bias; under the same NIS‑Agent framework, it attains average performance comparable to GPT‑4o on deep research benchmarks. Our code is publicly available at https://github.com/PangSMPang/NIS‑Agent.

Authors:Zhenghua Bao
Title: Better Retrieval, Worse Robustness: How Multi-hop RAG Amplifies Upstream ASR Errors
Abstract:
Speech‑based applications pass spoken queries through automatic speech recognition (ASR) before any retrieval module, so ASR errors enter the pipeline as a fixed upstream constraint. We empirically test whether two extensions to standard retrieval‑augmented generation (RAG), entity‑graph linking and iterative reformulation, absorb or amplify these errors. Using four English accents synthesized through neural TTS, we evaluate four RAG configurations on three multi‑hop QA benchmarks (HotpotQA, 2WikiMultiHopQA and MuSiQue) against a clean‑text oracle. Although the structurally richer configurations generally retain higher absolute F1 under ASR input, both extensions amplify the error: the F1 gap from clean text to the highest‑WER accent is 36‑67% larger under their combination than under naive dense retrieval, on all three benchmarks. The dominant failure mode is corruption of one or more query entities, accounting for 87‑96% of degradation cases on 2WikiMultiHopQA across all four methods. Two lightweight surface‑form mitigations leave most of the gap intact, indicating that downstream retrieval structure amplifies remaining entity errors. We release code and data at https://github.com/Continuum‑AI‑Corp/spoken‑multihop‑rag .

Authors:Lars Osterberg, Maggie Wang, Mac Schwager
Title: UniMem: Unifying Multimodal Memory and Control for Vision-Language-Action Models
Abstract:
While Vision‑Language‑Action (VLA) models have leveraged internet‑scale pretraining and task‑focused finetuning to achieve strong performance on long‑horizon tasks, they often struggle with non‑Markovian tasks that require memory. Existing approaches to memory typically involve additional Vision‑Language‑Models (VLMs) for long‑term memory management, introducing a memory bottleneck and a fractured training pipeline. Conditioning on multiple historical frames can provide the VLA with access to more descriptive features of past scenes, but can degrade performance if frames are chosen at arbitrary, fixed intervals. To address these limitations, we present UniMem, a framework that unifies high‑level, multimodal memory and low‑level control under one backbone. UniMem employs an event classifier for memory updates, a keyframe encoder for dense spatial memory, and a keyframe caching technique to minimize overhead during policy rollouts. We evaluate UniMem across five simulation and four hardware tasks targeting sequential and spatial memory, demonstrating that our unified, single‑model system outperforms fixed‑interval image sampling baselines (93.4% vs. 68.2%) in simulation and hierarchical baselines (80.0% vs. 43.5%) in hardware, while offering faster inference and a simple training pipeline for easy adoption. Project website: https://losterberg3.github.io/unimem‑vla/

Authors:Hongyuan Yu, Pufan Xu, Jiaojiao Yi, Yiding Tian, Mingrui Sun, Jiayuan Lu, Changyuan Wen
Title: JANUS: Online Jacobian-Aligned Infill for Black-Box Optimization
Abstract:
Population optimizers such as CMA‑ES, DE, and multi‑objective evolutionary algorithms drive search mainly through selection signals that are scalar or rank based: such a signal indicates that one candidate outperforms another, but not the local direction responsible for the improvement. JANUS (\emphJacobian‑Aligned Newton‑Unified Search) is a plug‑and‑play infill module that extracts this missing local geometric signal without replacing the host optimizer. It estimates a local Jacobian from the recent evaluation trace; the same Jacobian yields both a damped Gauss‑‑Newton exploitation candidate and a trace‑preserving exploration metric, reserving a fraction of the host's per‑generation candidate slots for geometry‑guided infill rather than spending evaluations on top of the host's budget. Unlike MetaBBO methods, JANUS needs no offline training or task distribution, estimating this geometry on the fly from the current run alone, while the host keeps full control of selection, survival, covariance adaptation, and step‑size control. Under same‑protocol comparisons, JANUS improves the CMA‑ES host on 11‑‑15/16 BBOB functions across d\in\30,100,500\. It also attains the best mean error on 13 of the 16 functions at d=500 in the complete NN‑BBO/MetaBBO baseline comparison, with no training cost, and yields a 936× geometric‑mean improvement over the host on a d=1000 BBOB subset. On structured and multi‑objective tasks, JANUS gives the best mean cost on 1135‑dimensional UAV path planning (‑12.8% vs.\ the strongest baseline), and it improves SMS‑EMOA/AGE‑MOEA2 hosts on 12/38 multi‑objective tasks with zero significant regressions. Code is available at https://github.com/hongyuanyu/JANUS.

Authors:Shuting Xie, Nathaniel Lesperance, Graham W. Taylor
Title: Hierarchy-Aware Supervised Uncertainty Estimation for Black-box LLM Taxonomic Reasoning
Abstract:
Large language models (LLMs) are increasingly used for scientific decision support, yet reliable confidence estimation remains difficult in black‑box settings. We study uncertainty estimation for hierarchical taxonomic reasoning generated by a black‑box LLM in a long‑tailed biodiversity monitoring pipeline. Using proxy features extracted by an open‑source tool LLM, we train lightweight supervised estimators with hierarchy‑aware supervision to predict rank‑wise correctness. Across three tool LLMs, the supervised estimators consistently outperform a token‑likelihood baseline for micro discrimination and selective prediction under a single global rejection threshold, improving micro AUROC from 0.57 to 0.75‑‑0.80. The best results are achieved by a rank‑specific multi‑head design (H3), suggesting that accounting for hierarchical output structure is important when a unified abstention rule is required. Our code is publicly available at https://github.com/uoguelph‑mlrg/hierarchy‑aware‑llm‑uq

Authors:Ryuki Hyodo
Title: Minimal Local Simulation Foundations for LLM- and VLM-Driven Agents in 2D and 3D Environments
Abstract:
Large language models (LLMs) and vision‑language models (VLMs) are expanding the range of behaviors that can be represented in agent‑based simulations, but many contemporary platforms are difficult to study, modify, or run on ordinary computers. We present two intentionally minimal simulation foundations for education and rapid prototyping. SD‑AgentFoundry‑2D provides a two‑dimensional multi‑agent environment in which locally hosted LLM agents move, communicate, respond to place occupancy, and encounter spatially localized fire events. SD‑AgentFoundry‑3D provides a three‑dimensional digital‑twin environment in which a locally hosted VLM receives first‑person images and produces natural‑language movement instructions. Both codebases are designed to run locally on macOS, Windows, and Linux and are deliberately left open to modification rather than developed as finished applications. Together, they offer accessible starting points for learning about generative social simulation and for building domain‑specific extensions.

Authors:De-Xing Huang, Chen-Yu Wang, Hao Liang, Xiao-Hu Zhou, Mei-Jiang Gui, Tian-Yu Xiang, Qin-Yi Zhang, Chen Wang, Xiao-Liang Xie, Shi-Qi Liu, Ming-Yuan Liu, Zhen-Chang Wang, Zeng-Guang Hou
Title: VeCAS: Vessel-Focused Contrast-Free Angiogram Synthesis for Vascular Interventions
Abstract:
X‑ray angiography relies on iodinated contrast agents to visualize vascular structures during image‑guided interventions. However, contrast administration carries risks of adverse events, motivating the development of contrast‑free alternatives. Generating X‑ray angiograms directly from non‑contrast X‑ray images offers a potential solution, but existing approaches remain limited by (i) insufficient control over vascular localization and (ii) inefficient modeling of redundant background content. To address these challenges, we propose VeCAS, a two‑stage vessel‑focused contrast‑free angiogram synthesis framework that separates vascular structure localization from angiographic appearance synthesis. In Stage I, a discriminative model localizes vascular structures in non‑contrast X‑ray images, while cross‑modality latent distillation transfers vessel‑sensitive knowledge from X‑ray angiograms during training. In Stage II, a vessel‑focused inpainting model synthesizes angiographic appearance within the localized vascular regions while preserving the non‑vascular background. Experiments on an in‑house lower‑limb vascular intervention dataset show that VeCAS outperforms the comparison methods in terms of vascular structural fidelity and image quality. Visual Turing tests and physician assessments indicate the perceptual realism of the synthesized angiograms. In addition, robotic guidewire navigation experiments in vascular phantoms show that VeCAS guidance reduces the time to target by 41.4% and the number of operation steps by 40.7% compared with non‑contrast guidance. Together, these results suggest the potential of VeCAS to serve as ``meta contrast agent'' for vascular interventions.

Authors:Parsa Bakhtiari, Hassan Bashiri, Alireza Khalilipour, Masoud Nasiripour, Moharram Challenger
Title: Industrial-Instruction: An End-to-End Framework for Building Instruction-Tuning and Benchmark Datasets from Industrial Technical Reports
Abstract:
Industrial technical reports contain high‑value knowledge for maintenance, troubleshooting, and product engineering, but their heterogeneous structure (dense prose, specifications, tables) makes them difficult to index and reason over with standard retrieval and QA pipelines, and no public instruction‑tuning or benchmark datasets are built from such documents. We address this gap with Industrial‑Instruction, contributing (i) two open QA datasets built from real industrial technical reports and (ii) the end‑to‑end pipeline that produces them. Using 906 public Panasonic documents (7,525 pages), we apply layout‑aware extraction, build a semantic retrieval index, and synthesize multiple‑choice QA grounded in retrieved evidence under five query‑document relationships (irrelevant retrieval, single‑/multi‑document support, single‑/multi‑document answer). After filtering an initial 23.9k generated samples, each dataset provides approximately 13.6k QA pairs with source documents and a held‑out benchmark split. Fine‑tuning small open LLMs (under 10B parameters) improves Set‑Match Accuracy from 28.5% to 42.0% and F1 from 46.6% to 63.5% on the Panasonic benchmark. We release two parallel versions built by the same pipeline: one generated with the open‑weight Qwen3‑30B‑A3B‑Instruct model and one with the closed, API‑based Claude‑Opus‑4.6 model, enabling a direct comparison of open‑ versus frontier‑model data generation. The Claude‑Opus‑4.6 dataset yields a cleaner raw corpus and larger fine‑tuning gains, at roughly two orders of magnitude higher cost. MMLU evaluation shows models trained on the Claude‑Opus‑4.6 data retain essentially all general knowledge, versus a small but measurable forgetting effect for the Qwen‑generated data. Together, these datasets and pipeline offer a practical, reproducible path toward scalable industrial benchmarks and training data from real‑world documentation.

Authors:Yue Zhao
Title: CatchBench: When Can an Agent Failure Be Caught?
Abstract:
When can an agent failure be caught? An audit is usually limited by the record rather than by the method. CatchBench therefore puts one auditor's question to three information states: the declared configuration before a run (PRE), a growing prefix of its trace (LIVE), and the finished trace (POST). Prior benchmarks fix one of these states or vary the telemetry; to our knowledge none scores all three under one task‑method interface. Each state admits different questions, so seven task contracts carry their own labels and metrics rather than one leaderboard. Four are evidential; three are Gold‑derived mechanism diagnostics. The release scores 72 entrants, from rule scanners and structural models to eleven LLM judges across nine model families (GPT, Claude, Gemini, Gemma, Llama, Qwen, DeepSeek, Mistral, Nova), over 1187 declared configurations and 1162 recorded runs. Most of the arena does not order: 47 of 118 pre‑declared contrasts separate, and the rest are published unresolved rather than ranked. The two sharpest results cut against our own data. One rule ignores every name and permission; it flags each capability declared after the first. On one of six configuration sources it reaches a perfect F1, so a score there measures how the corpus was built rather than how well a method reasons. Our admissibility bar then rejected one injected substrate and withheld evidential status from the other. A benchmark number is therefore not interpretable until the process behind its labels is published and tested for the shortcut it may leave. We report both, and regenerate every ordering from released predictions with no model call.

Authors:Lasan Perera, Deneth Priyadarshana, Dulana Pitiwaduge, Isitha Dinujaya, Mokshan Colambage
Title: Reproducible Vision-Guided 6-DoF Robotic Manipulator with a Mixed Stepper-Driver Architecture and Browser-Native Control
Abstract:
We present the NeuralNexus Arm, an open, low‑cost 6‑DOF robotic manipulator built by an undergraduate engineering team, together with the design decisions and debugging experience needed to reproduce it. The arm is driven by a single STM32H743 microcontroller on a custom printed circuit board (PCB) and combines two stepper‑driver strategies on one controller: push‑pull 3.3 V step/direction outputs for onboard TMC2209 drivers on the three wrist joints, and open‑drain outputs for external CL57T and DM542 drivers on the three high‑torque proximal joints. We describe the mechanical design, mixed‑driver electronics, interrupt‑driven firmware, a MATLAB/Simscape‑based inverse‑kinematics pipeline, a browser‑native control interface using the Web Serial API, and a lightweight vision pipeline for object localisation and autonomous pick‑and‑place tasks. We also document non‑obvious hardware and firmware failure modes encountered during the transition from a development board to the custom PCB as reproducibility guidance. All design files and firmware are released openly. The platform actuates all six axes under coordinated control at a 2 kHz update rate and executes both manual and pre‑recorded motions from the browser interface.

Authors:Wenhao Wu, Menghao Zhang, Xin Wang, Zhi Wang, Kun Shao, Jian Luan
Title: TRACE: A Self-Evolving Skill Bank for Consistent, Limit-Aware LLM Agents
Abstract:
Reliable deployment of LLM agents in user‑facing products depends not on raw task‑solving ability but on consistency and limit‑awareness: behaving the same way across repeated trials, and recognizing when a request cannot, or cannot yet, be safely fulfilled. CAR‑bench exposes this reliability gap in the domain of in‑car assistants: an LLM‑simulated user issues incomplete or ambiguous requests, requiring the agent to resolve uncertainty through multi‑turn dialogue and tool use while strictly adhering to domain policies. Even frontier models show a substantial gap between what they can solve at least once (Pass@3) and what they solve consistently across trials (Pass^k). We bridge this gap with TRACE (TRAjectory‑Contrastive Evolution), which iteratively improves a skill‑based agent's behavioral knowledge without modifying model weights. This knowledge is organized as a Skill Bank of modular, retrievable skills, each encoding a self‑contained set of tool‑use rules and behavioral guidelines. TRACE evolves this bank through an agentic self‑evolution loop: after each evaluation round, it groups trajectories by the skills invoked and refines each skill by contrasting successful and failed behaviors. The updated bank then guides subsequent rounds, while during deployment the Actor performs state‑conditioned skill orchestration at every turn. On GPT‑5.5, TRACE improves consistency (Pass^3) by 34.6 points, from 59.9% to 94.5%, while shrinking the gap between potential and reliable performance to just 4.0 points. On the official hidden set, TRACE achieved first place using GPT‑5.6‑Sol, attaining a Pass^3 score of 70%‑a 40% relative improvement over the baseline. These results show that TRACE converts high model potential into stable, consistent performance gain. Project homepage: https://darwin‑agent.github.io/Car‑bench‑TRACE.

Authors:Bohan Yu, Pengfei Cao, Chen Han, Chenxi Zhou, Zhiheng Zhang, Zhiyang Xie, Wenhao Teng, Xiangwen Liao, Jun Zhao, Kang Liu
Title: Beyond Factual Knowledge: Benchmarking and Learning Step-Level Procedural Rule Reasoning in Large Language Models
Abstract:
Large language models (LLMs) excel at text understanding and generation, yet still struggle to reliably understand and apply externally provided procedural rules at scale. To evaluate this capability, we introduce RuleWorld, a large‑scale benchmark that reformulates rules as globally reusable abstract units rather than instance‑specific facts. In RuleWorld, several scenarios, including single‑rule, parallel multi‑rule, and multi‑hop reasoning, are settled for comprehensive evaluation. We further propose DynaRule, an end‑to‑end framework that injects the given rules into the KV cache and turns retrieval into an internal, learnable, step‑wise process. Specifically, DynaRule employs Stacked Step‑Level Attention Training with a special <search> token to enable dynamic rule re‑attention and updating during inference. In this way, the model can re‑attend to the most relevant rules at each step, dynamically replacing outdated ones to support more stable multi‑step reasoning. Experiments on RuleWorld show that existing LLMs face challenges under large rule pools, while DynaRule improves average QA accuracy by up to 19 points and achieves over 85% Recall@1 at 10K rules, outperforming strong baselines by large margins. We make our code and dataset available here: https://github.com/SharkSpicy‑NLP/Beyond‑Factual‑Knowledge.

Authors:Qinfei Li, Xiaoxuan Dong, Jin Zhang, Dexu Yu, Wenhao Deng, Junchen Fu, Youhua Li, Hanwen Du, Chunxiao Li
Title: Risk-Aware Reranking for Agentic Tool Retrieval
Abstract:
Tool retrieval determines which external tools are exposed to an LLM agent for a user query or task, making retrieval a critical pre‑execution safety boundary. Unlike document retrieval, tool retrieval exposes executable actions: a tool that is useful for one task may be unnecessary or risky for another. However, existing tool‑retrieval methods primarily optimize semantic relevance, and safety evaluations often focus on failures after tool execution rather than risks introduced during retrieval. We study risk‑aware tool retrieval, where the goal is to retrieve useful tools while reducing exposure to higher‑risk tools. We propose a lightweight reranking framework on top of a frozen first‑stage retriever. The framework models query‑conditioned relevance and tool‑level exposure risk separately, combines them through an explicit parameter controlling the tradeoff between safety and utility, smooths scores over a ToolGraph, and optionally applies rule‑based safety constraints. To support retrieval‑time safety evaluation, we annotate 6,108 tools across UltraTool and Seal‑Tools with five ordinal risk levels and define metrics that measure risky‑tool exposure in the top‑k results. Experiments on UltraTool and Seal‑Tools show that our approach improves the relevance‑‑safety tradeoff over relevance‑only retrievers and reranking baselines, with the rule‑filtered variant providing a conservative operating point for safety‑critical deployments. These findings indicate that retrieval‑stage filtering can reduce the candidate action space exposed to agents before execution, complementing downstream tool‑use safeguards. The code and supplementary materials are available at: https://github.com/qli447/risk‑aware‑tool‑retrieval‑release.

Authors:Zhekai Wang, Haoxiang Huang, Xiang Liu, Zhikang Chen, Yueqing Sun, Qi Gu, Shiji Zhou, Miao Liu, Sen Cui
Title: MOSH-WM: Mask-Grounded Soft-Hamiltonian Dynamics for Object-Centric World Models
Abstract:
Object‑centric world models forecast future videos by evolving a set of entity slots, but the variables receiving dynamics supervision are often unconstrained visual features. We introduce \method, a mask‑grounded soft‑Hamiltonian world model that makes its position‑like state explicitly depend on slot‑owned image support. A frozen video‑slot encoder produces slots and masks; spatial moments of mask‑owned support form a canonical state Q, temporal differences form P, and a learned energy supplies a soft directional bias to a bounded learned increment. Decoder‑relevant appearance and identity are stored separately in a causal visual context. A gated composer and bounded residual then combine this context with the propagated phase state to reconstruct decoder‑compatible slots. On OBJ3D, given six observed frames and evaluated over the following 30 frames, \method reduces LPIPS by 25.0% and spatial MSE by 33.7% relative to the strongest object‑centric baseline. On CLEVRER, given six observed frames and evaluated over the following ten frames, the corresponding reductions are 14.5% and 18.7%. Horizon‑resolved visual and object‑state measurements show that the complete model accumulates error more slowly throughout the 30‑frame closed‑loop rollout. Project page:https://github.com/moshwm‑anon/‑moshwm‑anon.github.io.

Authors:Xinrui Miao, Mingjia Yin, Jiaqing Zhang, Wei Guo, Yong Liu, Yuyang Ye, Hao Wang, Enhong Chen
Title: Rethinking Item Tokenization in Generative Recommenders: From Fixed Atoms to Semantic Subwords
Abstract:
In generative recommender systems, items are typically tokenized into fixed‑length semantic ID sequences for autoregressive next‑item prediction. However, for user‑context modeling, this fine‑grained representation triggers Intra‑item Attention Overload: excessive attention is spent on low‑level intra‑item dependencies rather than high‑level inter‑item behavioral transitions. To address this, we propose Semantic Subword Tokenization (SST), which represents historical items as variable‑length semantic subwords while preserving fixed‑length target decoding. SST first applies Item‑level Subword Tokenization (IST) to merge stable adjacent atom tokens into compact semantic subword tokens, thereby reducing intra‑item reassembly in the encoder. It then introduces Behavior‑induced Co‑occurrence Augmentation (BCA) to inject coarse‑grained semantic prefix transition signals, guiding the freed modeling capacity toward inter‑item behavioral regularities. Extensive experiments on three public datasets and three generative recommender backbones show empirical improvements of SST over fixed‑length and transferable variable‑length SID baselines. Code is available at https://github.com/mxrcandy/Semantic‑Subword‑Tokenization.

Authors:Sukhrobbek Ilyosbekov, Shubham Gajjar, Rongfei Jin
Title: MorphoCLIP: Text-Supervised Contrastive Learning for Perturbation Matching in Cell Painting Images
Abstract:
Cell Painting microscopy captures how cells change after a chemical or genetic perturbation. Connecting these images to the perturbations that produced them could make large imaging screens easier to search and interpret, but the task remains difficult because biological effects are subtle and technical variation is substantial. We introduce MorphoCLIP, a contrastive model that links Cell Painting profiles with text descriptions of compounds, CRISPR knockouts, and ORF overexpressions. The model keeps its vision and language backbones frozen and trains only a compact cross‑channel module and projection layers, so it can be trained on a single consumer GPU. On held‑out CPJUMP1 data, MorphoCLIP searches in both directions: from a cell image to its perturbation description and from a description to matching cell images. In both cases, a correct match appears among the top ten results much more often than expected by chance. Adding a replicate‑alignment loss makes profiles from repeated experiments more consistent, although this improvement does not yet translate into reliable gene‑compound matching. Gene‑aware labels and plate correction also show no consistent retrieval benefit. These findings suggest that text supervision can help organize chemical and genetic Cell Painting data. Matching compounds with genetic perturbations, however, remains an open problem.

Authors:Nobel Dhar, Md Romyull Islam, Xuechen Zhang, Gongjin Sun, Sahidul Islam, Bobin Deng, Kun Suo
Title: NeuroPrefetcher: Storage-Aware Sparse LLM Inference via Delta Prefetching
Abstract:
Deploying large language models on edge devices is increasingly limited by a widening gap between model size and available memory. Existing approaches such as quantization, smaller models, and offloading can raise the effective memory limit, but they still assume that the model can be compressed or partitioned to fit within some budget. We target the harder model‑exceeds‑memory setting, in which the model remains larger than resident memory throughout execution and storage becomes an active source of weights on the critical path. We observe that MLP activity during autoregressive decoding has strong temporal locality: approximately 82‑85% of active neurons persist from one token to the next. This means that most sparse weights needed for the current token are already resident, and only the newly needed rows must be fetched from storage. We present NeuroPrefetcher, a storage‑backed LLM inference system that exploits this property through predictive delta prefetching. After layer 0, a single GPU‑resident predictor, occupying 2.86% of base model parameters, predicts sparse activity for all downstream MLP layers in one forward pass. The runtime compares these predictions against resident GPU buffers and issues application‑scheduled NVMe reads only for incoming delta rows, replacing reactive operating‑system demand paging with explicit, model‑aware weight movement. On real unified‑memory edge hardware, NeuroPrefetcher achieves 7.9‑12.0x speedup over llama.cpp across constrained memory budgets.

Authors:Md Maklachur Rahman, Md Hasan Al Banna, Saraf Anjum, Mahmudul Hasan, Tracy Hammond
Title: GET: Generative Embedding Translation for Medical Image Segmentation
Abstract:
Generative segmentation provides an alternative to direct pixel‑wise prediction by operating on learned latent representations, but effective image‑to‑mask translation must preserve target structure while remaining computationally efficient. We propose Generative Embedding Translation (GET), a structured embedding‑translation framework that progressively transforms image embeddings into mask embeddings within the frozen latent space of a Stable Diffusion VAE. GET uses a U‑Net‑style Embedding Translation Network with 1.07M trainable parameters, combining Mobile Bottleneck Convolutions, Subsampled Self‑Attention, and Multi‑scale Feature Enrichment for local modeling, global context, and multi‑scale refinement. Across five medical segmentation datasets, GET outperforms generative, CNN, and Transformer baselines. Compared with the strongest generative baseline, GMS, GET improves average Dice and IoU by 0.93% and 1.26%, reduces HD95 by 0.81 pixels, and uses 31.41% fewer trainable parameters. Under bidirectional BUS‑BUSI domain shift, GET further improves Dice and IoU by 3.51% and 3.39%, while reducing HD95 by 27.37 pixels. Our code is available at: https://github.com/maklachur/GET.

Authors:Jalen Jiang, Chufan Gao, Ethan Rasmussen, Stephen Z. Xie, Jimeng Sun
Title: KMGen: A Skill-based Approach for Synthetic Individual Patient Data Generation
Abstract:
Individual patient data (IPD) from clinical trials is the substrate for survival modeling, meta‑analysis, and safety research, yet IPD is rarely released. Prior work has addressed only half of this gap: reconstructing Kaplan‑Meier (KM) curves from published plots ‑‑ typically requiring manual digitization or human‑in‑the‑loop correction ‑‑ while offering no mechanism for generating the adverse‑event (AE) streams that constitute the other half of a patient record. We introduce KMGen, the first end‑to‑end framework that (i) fully automates KM curve extraction at accuracy competitive with human‑guided tools, and (ii) generates synthetic per‑patient AE trajectories from public trial registry records. The extraction stage is a fully automated agentic pipeline ‑‑ an agent generates code to extract each step in the KM curve ‑‑ achieving a mean Integrated Absolute Error (IAE) of 0.0151 on a 32‑plot benchmark spanning clean, edge‑case, and adversarial conditions. The IPD generation stage decouples patient archetype extraction from statistical sampling: an LLM distills the trial record into arm‑specific statistics, adverse events, patient demographics, and risk multipliers. A mechanistic sampler generates patient events via clinical archetypes, bootstrap rank‑correlation coupling to the empirical KM curve (preserving the marginal survival distribution exactly), and cycle‑based AE scheduling with an induction/maintenance split. Across three held‑out oncology trials spanning an order of magnitude in cohort size and 30 independent regenerations per trial, KMGen achieves mean integrated KM absolute difference Δ_\textKM\,\leq\,0.051, sex/ECOG JSD \leq\,0.013 on 5 of 6 demographic slots, and recovers \geq\,71% of the top‑15 AEs by exact MedDRA term under a single fixed parameter set. The pipeline is released as open source at https://github.com/chufangao/kmgen.

Authors:UniverseTBD, :, Kshitij Duraphe, Aman Kumar, Michael J. Smith, Shashwat Sourav
Title: What AstroPT knows about galaxies, and what that can teach us about LLMs
Abstract:
Interpretability research increasingly asks when concepts emerge during training and whether linear probes recover real structure, but in language models these claims are hard to validate because language offers little ground‑truth ordering of concepts or relationships among them. We propose the use of astronomical ground truth through AstroPT, a transformer trained on millions of galaxy images, as a calibration testbed. AstroPT is an LLM‑like model trained within a domain where the difficulty ordering of concepts and the relations among them are known in advance. Probing frozen representations across checkpoints, layers, model sizes, and objective choices, we find that galaxy properties emerge in a fixed order that tracks their known difficulty‑‑‑quantities written almost directly into the pixels (band magnitude) become decodable early in training and shallow in the network, while multiband/spectra based and inferred quantities (such as redshift and specific star formation rate) emerge later and deeper. This order is invariant to our tested training objectives, and scales in magnitude but not in sequence with capacity. Our linear probe directions further recover the known physical structure among galaxy properties. Our findings suggest that astronomy offers a controlled sandbox for calibrating mechanistic interpretability methods we otherwise apply to LLMs blind.

Authors:Pengyiang Liu, Junbo Niu, Xiaoyang Hu, Zhongyue Shi, Zitian Wang, Linjiang Huang, Si Liu
Title: TRACE: Temporal Retrieval with Anchored and Convergent Evidence for Long-Horizon Video Understanding
Abstract:
A long‑video answer is evidence‑supported only when the frames decoded from the video cover every event the answer depends on. Existing evaluations score final‑answer correctness or predicted evidence intervals, but the frames a method decodes before answering are rarely audited, so correct answers can still rest on incomplete observation. We introduce VES‑Bench, a 600‑question benchmark of Temporal Ordering and Event Counting items over 348 public long videos. Each item carries a jointly necessary set of evidence intervals, letting us audit at three strictness levels whether a method's decoded frames cover every one of them. We also propose TRACE, a training‑free agent that grounds answers in raw visual clips, builds an evidence bundle round by round, and stops only when the answer stabilises as the bundle grows and a final pass over the same clips returns the same answer. Under a same‑backbone audit, TRACE answers 50.7% of questions correctly with at least two decoded frames inside every evidence interval, at 98.7 frames per question: over 10 points above uniform decoding at 128 frames (40.2%), and within 2.6 points of uniform decoding at 256 frames at 0.39x its frame cost, while reaching the highest answer accuracy in the audit (63.5%). TRACE also stays competitive on Video‑MME (86.1), LVBench (75.6), and LongVideoBench (75.1).

Authors:Yingying Yan, Jiaqi Tang, Wei Wei, Qianzhou Wang, Jinjian Wu, Botong Geng, Jianmin Chen, Yuyang Xia, Lei Zhang
Title: HeatTok: Enhancing Remote Sensing Image Understanding via Thermodiffusion-based Tokenization
Abstract:
Current visual tokenizers in Multimodal Large Language Models (MLLMs) predominantly rely on patch‑based partitioning, which causes severe semantic mixture and object fragmentation in remote sensing imagery due to the irregular contours of geo‑objects. Moreover, existing adaptive methods struggle to extract precise object‑level tokens and lack dedicated geometric positional encodings for irregular regions. In this paper, we propose HeatTok, a semantic‑aware tokenizer driven by thermodiffusion aggregation. Inspired by the physical principles of heat conduction, HeatTok adaptively merges adjacent homogeneous regions to generate semantically independent, object‑aligned irregular tokens. To enable MLLMs to perceive these irregular shapes, we design the Gaussian Multimodal Rotary Positional Embedding (G‑MRoPE), which models token spatial distributions via 2D Gaussians and explicitly injects center, scale, and orientation cues. Extensive evaluations on the VRSBench and EarthVQA datasets demonstrate that HeatTok effectively preserves object‑level semantic integrity and achieves state‑of‑the‑art performance under a reasonable token budget. The code is available: https://github.com/YingyingYan1/HeatTok.

Authors:Jun Chen, Yongchao Liu, Pengyu Qiu, Jiajun Zheng, Juelu Zhang, Yujie Zeng, Qin Zhang, Ziyue Qiao, Xiao Luo
Title: GTA-RAG: Graph-Trajectory-Augmented Reinforcement Learning for Multi-Turn Retrieval-Augmented Reasoning
Abstract:
Retrieval‑augmented generation (RAG) enables LLMs to access external knowledge for answering knowledge‑intensive questions. For complex multi‑hop questions, multi‑turn retrieval‑augmented reasoning extends RAG into an iterative process that repeatedly searches for and integrates evidence across documents. However, existing reinforcement‑learning (RL) approaches for agentic RAG are typically optimized with final‑answer rewards, which provide sparse supervision and overlook whether the model actually retrieves the required evidence chain. We present \textscGTA‑RAG, a graph‑trajectory‑augmented RL framework for multi‑turn retrieval‑augmented reasoning. From an entity‑‑document graph, we sample connected document paths, synthesize multi‑hop QA trajectories, and validate them with the deployed retriever to obtain executable trajectory‑level supervision. We then optimize the retrieval policy with Group Relative Policy Optimization (GRPO) and a trajectory‑guided reward that encourages both accurate answers and acquisition of target evidence documents, followed by answer‑reward training on natural QA instances. Experiments on three multi‑hop and two simple QA benchmarks show that \method consistently outperforms RL‑based RAG baselines with both Qwen2.5‑3B and Qwen2.5‑7B backbones, while substantially improving evidence‑chain coverage. Our code is available at https://github.com/cjcj46262/GTA‑RAG.

Authors:Wen Wang, Ruibing Hou, Hong Chang, Shiguang Shan, Xilin Chen
Title: EMPIRE: Explicit Manipulation Planning as a Learnable Intermediate Representation for Egocentric Hand-Motion Forecasting
Abstract:
Forecasting dexterous hand motions from egocentric observations is fundamental to intelligent interactive systems. Existing VLM‑based methods typically map observations directly to future motions, overlooking the underlying manipulation process that governs hand‑object interactions. Moreover, end‑to‑end optimization couples manipulation learning with motion synthesis, causing motion‑generation gradients to interfere with the pre‑learned manipulation‑aware representations. To overcome these limitations, we propose EMPIRE, a two‑stage framework that introduces Explicit Manipulation Planning as an Intermediate Representation for Egocentric hand‑motion forecasting. Stage I: Learn to Plan. EMPIRE first learns explicit manipulation plans from multimodal context to capture the progression of hand‑object interactions. Stage II: Learn to Act. A motion generator synthesizes future bimanual hand motions conditioned on frozen planner representations, preventing motion‑generation gradients from affecting manipulation planning. To support our method, we further construct EMPIRE‑651K, a bimanual hand‑motion forecasting dataset comprising 650,910 training windows across 111 tasks, each paired with an explicit per‑hand manipulation plan. Under identical training and evaluation protocols, EMPIRE achieves state‑of‑the‑art forecasting accuracy, with an MPJPE of 84.53 mm and a finger‑relative error of 38.97mm. We release the code and dataset at https://github.com/wangwen‑banban/EMPIRE.

Authors:Jiarui Dong, Yin Cai, Zhouhong Gu, Chenmou Wu, Ci Tao, Yiran Chen, Jialing Li, Xiaoran Shi, Juntao Zhang, Zhijun Fang
Title: SchemaGUI: A Schema-Driven Benchmark for Controllable GUI Generation Evaluation
Abstract:
Large language models (LLMs) have demonstrated strong potential in graphical user interface (GUI) generation, but reliable evaluation remains challenging due to uncontrolled data distributions, noisy annotations, and limited layout scenario coverage. To address this, we propose SchemaGUI, a template‑based benchmark for controllable GUI generation evaluation. By synthesizing paired natural language instructions and deterministic function‑call references from parameterized interface schemas, SchemaGUI can generate thousands of deterministically annotated tasks in seconds without human labeling. Based on 1,000 evaluated instances per scenario and language across six representative bilingual scenarios, we benchmark five mainstream models, including the Qwen3.5 family, Qwen3‑Coder‑30B, and DeepSeek‑R1. Our extensive analysis reveals three key insights. First, precise geometric spatial control remains an important bottleneck; while scaling Qwen3.5 from 4B to 27B improves Schema Feasibility from 91.56% to 99.63%, the Geometry score improves more modestly (from 67.05% to 75.30%). Second, generation difficulty is highly sensitive to layout complexity, with current LLMs excelling at simple sequential arrangements but suffering severe coordinate drift in dense grids and multi‑region compositions. Third, thinking mode increases token consumption while generally reducing GUI Score, particularly for smaller models.

Authors:Yikai Zhao, Qiyan Zhao, Jiaquan Zhang, Xiaofeng Zhang, Xiaosong Yuan, Pengzhou Cheng
Title: Context-Aware Cluster Decoding: Semantic Anchor-Driven Coherence in dMLLMs
Abstract:
Diffusion multimodal large language models (dMLLMs) frequently produce long‑form outputs marred by semantic drift and repetition, with quality generally degrading as output length increases. We identify two structural deficiencies in existing decoding methods as primary drivers of these failures: confidence‑based scoring ignores decoded‑neighbor support, and block partitioning prevents access to high‑readiness semantic anchors, together causing tokens to be committed before their local context is sufficiently established. We propose \ours (Context‑Aware Cluster Decoding), a training‑free decoding method that scores each masked position by a multiplicative composite of softmax confidence and neighbor proximity, promoting contextually ready tokens above isolated candidates while suppressing low‑confidence positional noise, operating block‑free to keep high‑readiness anchors globally accessible. \ours further applies architecture‑aware calibration to handle confidence heterogeneity induced by diverse visual integration strategies. Experiments on three dMLLMs across four benchmarks demonstrate consistent quality gains and hallucination reduction over Original, with larger gains in several longer generation settings, highlighting the importance of neighbor support and visual integration strategy for future dMLLM decoding method design. Our code is openly available at https://github.com/zhaoyk‑sysu/CACD‑dMLLM.

Authors:Furkan Cifci, Osman Emre Donder, Reyyan Cifci
Title: MCSI: A Masked Commutative Supersingular Isogeny Key Exchange with Blinded Ephemeral Keys
Abstract:
We introduce MCSI, a two message key exchange we design over the CSIDH class group action, in which each party sends its ephemeral public element under an authenticated encryption keyed by the value the two static keys determine. The design gives implicit mutual authentication, hides the ephemeral element from an eavesdropper, and lets a recipient discard an unauthenticated message after one tag check rather than after an evaluation of the group action, which is four orders of magnitude more expensive. On the analytic side, we prove that our protocol is correct with zero error, and prove three statements in the random oracle model, all reducing to the strong parallelisation problem: indistinguishability of the session key against a passive adversary, confidentiality of the blinded ephemeral element, and integrity of the blinded transport. None uses the decisional group action assumption, which is false for class group actions of non‑prime discriminant. We also show that a blinding key cannot come from the session secret it is meant to establish. To instantiate the design we select parameters and show that a prime chosen for elliptic curve discrete logarithms is unusable: for p = 2^521‑1, the NIST P‑521 prime, the action admits no efficiently evaluable generator. On the practical side, we build and test the design. We implement the protocol twice, in C and independently in Python, cross check the two, and measure what a session costs in field operations, time and memory. We also audit our code for secret dependent control flow: the field arithmetic and the symmetric layer show none, while the group action leaks the key by construction, and two hundred timings separate two keys whose one‑norms differ by five out of 370. Finally, we state what we do not prove, among them security under ephemeral key reveal, forward secrecy of the blinding, and constant time execution.

Authors:Masoud Jalayer, Changyi Li, Yu Xiao
Title: Pre-Decoding Acoustic Triage for Budgeted Vision-Language Captioning of Untrimmed Egocentric Video
Abstract:
Automatically analyzing hours‑long egocentric video is increasingly essential for progress monitoring, quality control, and safety in logistics, construction, and manufacturing. Yet current pipelines that process short, fixed‑size windows with a vision‑language model (VLM) are prohibitively expensive because cost scales with the number of model calls. To reduce this cost, prior work proposes triage policies to select which windows merit a VLM invocation. However, these policies either sample uniformly or rank windows using visual features, which ironically requires the video decoding that the budget constraints are meant to avoid. We propose audio‑first triage: select windows using the lightest modality, scored before any video frame is decoded, so the approach composes naturally with token compression or quantization. The novelty lies in the objective, not the representation: rather than a per‑frame sound‑event detector, we train the selector to trigger once per action. This objective shift improves action coverage by 4.0‑10.8 percentage points across all evaluated call rates, using frozen AudioSet‑pretrained features without domain‑specific sound‑event labels. Using fewer than half of the available calls, the triage cuts 9‑20% of VLM calls at matched coverage on EPIC‑KITCHENS‑100 (EK‑100), surpasses uniform sampling through the mid‑range on Ego4D over 247 clips, and outperforms two recent visual keyframe selectors. Code, the reference implementation and every results file this manuscript reads are at https://github.com/masjalayer/PreDecoding‑AcousticTriage.

Authors:Francisco M. Arrabal-Campos, Francisco G. Montoya, Alfredo Alcayde, Ignacio Fernández
Title: Where Cognition Lives: Dissecting Emergent from Computed Function in a Minimal Complete Cognitive Architecture
Abstract:
A cognitive architecture is more than the module that reasons: it must also decide how long to think and what deserves the effort. We built a minimal but complete system ‑ a recurrent reasoner with adaptive halting, a homeostatic control field, and a value module ‑ and asked of each part: does this function emerge from gradient descent, or must it be computed? Competence emerges. Stopping appears to emerge too, and to be worth more than everything decidable in advance, but that appearance is instrumentation: payoff at matched mean compute climbs from 0.467 (uniform) through 0.546 (difficulty) to 0.698 (ex‑ante value), and the further climb to 0.921 (posterior self‑observation) does not survive audit. PonderNet‑style halting returns a halting‑weighted mixture of hidden states while forced‑depth baselines return one, and the language head is trained on the mixture alone; equalizing the readout annihilates the apparent advantage of native execution (residual +0.000 [0.000, 0.000]). Value does not emerge: trained couplings capture zero of a payoff an explicit allocator captures completely (+0.151, routing correlation +0.79), so the second‑order decisions that pay must be computed, at least where value is orthogonal to content, as here by construction. On a frozen LLM actuator the same instruments show self‑consistency voting to be a measured bound (+0.0236 [+0.0150, +0.0326]) and inter‑sample agreement nearly worthless as a stopping signal, its mass concentrating on wrong answers. Every null we assert carries a mechanism and a positive control, and the protocol is part of the contribution. Executing our own falsifiable prediction, value under commitment pays +0.1312 [+0.1124, +0.1502] in a cliff‑cost family, some seven times the smooth‑family estimate ‑ not because the cliff shifts information ex ante, but because it multiplies the attainable range fivefold (5.1x [3.4, 8.2]).

Authors:Qianqian Tang, Jiayi Gao, Ting Lei, Yang Liu
Title: ReART: Reference-Guided Retrieval and Refinement for Emotion-Aware Art Generation
Abstract:
Emotion‑aware artistic image generation requires a model to satisfy semantic content, artistic style, and target emotion simultaneously. The key challenge is that artistic captions conflate these axes into underspecified free‑form text, making fine‑grained visual attributes such as brushwork, composition, and tonal atmosphere difficult to ground concretely. We present ReART, a reference‑guided retrieval and refinement framework. Our method decomposes test captions and each image annotation in the EmoArt database into structured visual fields, and performs field‑wise retrieval over subject, layout, brush‑line, and tone‑mood dimensions to retrieve role‑specific visual references that supply the perceptual detail text alone cannot convey; these references are used alongside a structured prompt for initial synthesis. For samples where any Attribute Alignment Score (AAS) axis falls below threshold, an AAS‑driven refinement loop diagnoses failures, constructs constrained repair plans specifying elements to keep, errors to fix, and operations to avoid, routes references by correction purpose, and performs controlled editing under structural preservation constraints. Our system ranks 2nd in Track 1 of the AffectiveArt 2026 Grand Challenge, achieving a perfect AAS of 1.00 and an overall score of 0.78. Code is available at https://github.com/oceanflowlab/ReART.git.

Authors:Xiaokai Zhou, Baoshi Cao, Yang Liu, Kui Sun, Boyu Ma, Zhengpu Wang, Zongwu Xie
Title: GCS-Bridging: Restoring Connectivity of Disconnected Convex Sets for Graph-of-Convex-Sets Motion Planning
Abstract:
Graph‑of‑Convex‑Sets (GCS)‑based trajectory optimization represents collision‑free regions in configuration space as a finite collection of convex sets and directly performs collision‑free trajectory planning over these sets, substantially simplifying the planning process. However, existing GCS‑based trajectory planning methods generally assume sufficient connectivity among the convex regions and do not explicitly address cases in which the start and goal regions belong to different connected components of the initial GCS map. To address this limitation, we propose GCS‑Bridging, which reconnects disconnected convex regions through collision‑free point paths followed by convex region inflation, thereby recovering the feasibility of otherwise disconnected GCS planning problems. Extensive simulations across multiple IRIS‑related algorithms and scenarios demonstrate that GCS‑Bridging restores missing start‑to‑goal connectivity in the initial GCS map with a 99.8% success rate. In addition, a hardware experiment on a single‑arm Franka platform in a real‑world scenario with initially disconnected start and goal regions validates the effectiveness of the proposed method in practical motion planning. Project website: https://zhouxk1997.github.io/GCS_Bridging/

Authors:Yan Wang
Title: Beyond Dense Adam States: Adaptive Log-Space Quantization for Memory-Efficient Optimizers
Abstract:
Low‑precision optimizer‑state methods are commonly designed and evaluated for dense Adam‑style first and second moments. Memory‑efficient optimizers depart from this setting: Adafactor factorizes second moments, CAME adds factored confidence states, and APOLLO maintains statistics in a projected gradient space. Consequently, an equal amount of state reconstruction error can induce different update errors depending on state topology and update semantics. We first characterize this heterogeneity in optimizer‑state traces from language model pre‑training. We then introduce Adaptive Log‑Space (AL) quantization, a block‑wise representation for non‑negative states that adapts its nonzero range per block and enforces the exact‑zero invariant q = 0 \Leftrightarrow x = 0. AL8 and AL16 are combined with independent signed‑momentum encodings and state‑specific precision choices rather than a single policy for every state. Across 96 runs totaling 214.7 GPU‑hours, we evaluate dense, factored, confidence, and projected states in AdamW, Adafactor, CAME, and APOLLO paths. On a 20K‑step TinyLlama‑1.1B pre‑training benchmark, an AdamW configuration with AL8 second moments and uniform 8‑bit momentum reaches 72.90 perplexity, compared with 72.48 for FP32 AdamW and 73.54 for an 8‑bit dynamic‑quantization baseline, while reducing measured optimizer‑state storage from 8392.7 to 2119.2 MiB. CAME exposes a different precision regime: promoting its non‑negative states to AL16 recovers 86.16 perplexity versus 86.68 for the full‑precision reference, whereas all‑AL8 reaches 90.19. A 100K‑step GPT‑2 experiment further shows that topology‑aware parameter protection reduces the late‑loss gap of quantized Adafactor from +0.1185 to +0.0159 in the evaluated setup. These results support a state‑ and topology‑aware view of optimizer quantization.

Authors:Sebastian Pfister, Benjamin Holzschuh, Nils Thuerey
Title: StocBench: A Benchmark for Generative Modeling of Stochastic Dynamics
Abstract:
We benchmark transport‑based generative models as well as distillation‑based few‑step methods for the probabilistic forecasting of stochastic fluid flows, with a particular focus on performance under limited inference budgets. All methods are evaluated on a two‑dimensional Kolmogorov flow with stochastic forcing. We measure one‑step distributional accuracy against large simulated reference ensembles and assess whether the invariant measure is preserved during autoregressive rollouts via the enstrophy spectrum. On the stochastic task, flow matching achieves the most accurate one‑step conditional distribution at high inference budgets, while the second‑order exponential integrator DPM‑2 is strongest at very low NFE. Few‑step distillation methods are competitive with the multi‑step methods and preserve the enstrophy spectrum particularly well. A deterministic control task, in which the forcing over the prediction interval is observed, separates aleatoric from epistemic uncertainty. Model performance does not translate between the two settings: the distilled models are competitive on the stochastic task but least accurate on the control task. While stochastic diffusion samplers such as DDPM better preserve the enstrophy spectrum during rollouts in the stochastic setting, deterministic samplers such as DDIM and DPM‑2 show better spectral preservation in the deterministic setting.

Authors:Haoran Lin, Mingyu Yang, Pengfei Qi, Kehan Chen, Qiang Diao, Liangji Zeng, Wenrui Chen, Yaonan Wang, Kailun Yang
Title: TONAV: Task-Oriented Navigation and Action-Velocity Chunk Learning for Articulated Object Quadrupedal Mobile Manipulation
Abstract:
Quadruped mobile manipulation requires two tightly coupled capabilities: reaching manipulation‑ready configurations and maintaining stable contact throughout articulated‑object interaction. However, existing methods often terminate navigation near the target, leaving a gap between reachability and manipulation readiness, while tracking lag, motion jitter, and contact instability limit continuous interaction. To address these challenges, we present TONAV, a unified framework integrating task‑oriented navigation with action‑velocity chunk learning. First, we introduce a position‑velocity‑coupled teleoperation framework that explicitly captures motion dynamics to improve master‑follower consistency and collect smooth, temporally consistent demonstrations. Next, task‑oriented navigation leverages vision‑language reasoning to decompose high‑level instructions into executable subgoals and adaptively refine the robot base toward a manipulation‑ready configuration. Finally, action‑velocity chunk learning jointly models joint positions and their temporal transitions under velocity supervision, enabling smooth and stable sustained‑contact manipulation. Real‑world experiments across diverse articulated‑object tasks demonstrate that TONAV achieves higher success rates in both task‑oriented navigation and complete mobile manipulation, mitigating the navigation‑manipulation gap and improving continuous‑contact interaction. The project page is at https://haochen611.github.io/TONAV.

Authors:Parastoo Farajpoor, Mohammadreza Narimani
Title: The spatial anatomy of urban wildfire vulnerability: a spatially validated GeoAI framework reveals the roles of building density and vegetation moisture in structure loss during the 2025 Palisades Fire
Abstract:
Urban wildfire resilience depends on interactions among built form, vegetation condition, and extreme fire weather, yet city‑scale risk models often overlook whether predictive skill transfers across neighborhoods. We developed a spatially validated GeoAI workflow for the January 2025 Palisades Fire, linking 12,081 CAL FIRE damage inspections to pre‑fire Sentinel‑2 vegetation indices, Landsat surface temperature, LANDFIRE fuels, terrain, and OpenStreetMap buildings and roads. Among 9,883 inspected residential structures, 5,566 were destroyed. Random cross‑validation yielded ROC‑AUC 0.92 for the integrated XGBoost model, but 1 km spatial block validation reduced performance to 0.75; logistic regression performed similarly and was better calibrated. Building count within 100 m was the strongest predictor, with destruction odds increasing 4.12‑fold per standard deviation. Vegetation moisture and greenness showed opposing conditional associations: NDMI at 100‑300 m was protective (OR 0.52), whereas NDVI at 30‑100 m was positively associated with destruction after accounting for moisture (OR 1.74). Predictive information was concentrated at the 100‑300 m neighborhood scale. A separate post‑fire track mapped burn severity and vegetation recovery without leakage. The results support neighborhood‑scale susceptibility screening, moisture‑aware vegetation management, and spatial block validation as a minimum standard for single‑event urban wildfire modeling.

Authors:Yibin Ye, Xichao Teng, Shuo Chen, Xiaokai Song, Dongdong Guan, Qifeng Yu, Zhang Li
Title: DECO: Depth-Guided Co-Visibility Reasoning for Low-Altitude UAV Visual Localization
Abstract:
Unmanned aerial vehicles (UAVs) increasingly require robust visual localization in GNSS‑denied environments. A common solution estimates UAV poses by matching keypoints between UAV images and geo‑tagged orthographic reference maps derived from satellite or aerial imagery, followed by Perspective‑\(n\)‑Point (PnP) pose solving. However, such reference maps mainly record top‑down surfaces such as roofs and ground planes, while vertical structures such as facades and walls are often compressed or missing. Consequently, many visually distinctive keypoints in low‑altitude UAV images have no valid counterparts in the reference map, leading to redundant matches and inaccurate pose estimation. To address this issue, we propose DECO, a DEpth‑guided CO‑visibility reasoning framework for low‑altitude UAV visual localization. DECO uses monocular depth priors to infer local surface geometry and estimate co‑visible regions between UAV images and the reference map. Based on this prior, a Geometry‑Saliency Coupled Co‑visibility Score is introduced to jointly consider geometric co‑visibility and detector saliency for keypoint ranking. In this way, DECO retains keypoints that are both visually distinctive and geometrically co‑visible, improving feature matching and PnP‑based pose estimation. Extensive experiments demonstrate that DECO achieves superior localization performance and can be integrated with different depth models, feature detectors, and matchers. The source code will be available at https://github.com/UAV‑AVL/DECO.

Authors:Bin Dong, Jinghong Chen
Title: CiUNet: A Hybrid Swin-CNN UNet for Medical Image Segmentation
Abstract:
Medical image segmentation requires high accuracy and robustness, yet practical commercial deployment also demands privacy preservation and computational efficiency. In this context, the U‑Net architecture, which can be inherently decoupled into independent encoder and decoder components, serves as a natural commercial choice. However, pure Transformer‑based variants like Swin‑UNet often suffer from insufficient local detail capture and limited interpretability. In this paper, we propose a lightweight hybrid architecture built upon the Swin‑UNet framework. Our model integrates a parallel CNN encoder to complement the shallow layer reasoning of Swin Transformers with local texture features. To bridge the semantic gap and enhance fine‑grained spatial detail recovery, we design an asymmetric feature fusion strategy and introduce cross‑layer skip (XSkip) connections that explicitly propagate shallow CNN features into the decoder. We further incorporate novel loss functions and an auxiliary supervision head (Aux‑Head) to strengthen training stability, boundary delineation, and intermediate feature interpretability. Extensive experiments on the Synapse multi‑organ segmentation dataset demonstrate that our approach achieves state‑of‑the‑art competitive Dice scores and Hausdorff distances, offering an accurate, efficient, and interpretable solution for clinical deployment.

Authors:Jie Yin, Xingyu Lai
Title: DreamMimic: Learning Visuomotor Whole-Body Loco-Manipulation via World Model
Abstract:
Vision‑based whole‑body loco‑manipulation on humanoid robots is challenging due to partial observability, contact‑rich dynamics, and the difficulty of learning long‑horizon behaviors from high‑dimensional visual inputs. We present \hrefhttps://github.com/DreamMimic/DreamMimicDreamMimic, a framework that distills privileged teacher policies into vision‑based humanoid controllers via world‑model‑assisted distillation. Instead of using a Dreamer‑style RSSM for planning, we repurpose it to learn predictive latent dynamics that serve as both a representation space and an action‑conditioned multi‑step supervision signal, while exposing compact predictive features to the student policy to reduce long‑term drift. Beyond standard reconstruction objectives for proprioceptive and visual observations, we add auxiliary prediction heads for privileged state, contact, object state, and reward estimation. These heads provide additional supervision related to agent‑‑object interaction and task progress, encouraging the latent representation to retain signals that are useful for contact‑rich loco‑manipulation. We further introduce Performance‑Conditioned Guidance (PCG), a reward‑driven adaptive distillation schedule that computes performance scores for both teacher and student to dynamically balance guidance and exploration. PCG prevents both premature teacher annealing and excessive teacher interference in challenging visual settings. Experiments on OMOMO and BEHAVE show improved tracking‑based loco‑manipulation performance over strong vision‑based baselines, without exposing online privileged interaction states to the student at deployment. Qualitative simulations further examine morphology and simulator changes. These results suggest that world models can provide a useful mechanism for stabilizing visual policy distillation in contact‑rich humanoid behaviors.

Authors:Yan Zhan, Mengkai Hou, Wanting Zhang, Zhijun Gao
Title: Length-Adaptive Decoding for Masked Diffusion Machine Translation
Abstract:
Machine translation tests masked diffusion language models (dLLMs) because every source token must be rendered faithfully, while fixed canvas decoding must choose target length before denoising. Existing masked diffusion decoding work mainly studies token unmasking order, leaving this length decision under‑explored despite its direct effect on coverage and redundancy. We introduce Entropy‑Valley (EV), a training‑free length selector that scores candidate target canvases by mean predictive entropy from all‑mask forward passes and selects the canvas the backbone is most prepared to fill. Relative to a baseline using training corpus length statistics, EV recovers 64.9%, 65.3%, and 33.0% of the COMET‑22 gain from reference target lengths on En\toZh, Zh\toEn, and En\toDe. Our diagnostics show that denoising‑friendly lengths need not match reference lengths. Evaluation by three translation experts supports the En\leftrightarrowZh adequacy gains, with stronger evidence on Zh\toEn. Compared with a LLaMA‑3‑8B autoregressive (AR) model trained on the same fine‑tuning data, the EV system ties on En\toZh and leads on Zh\toEn; an oracle‑length diagnostic further shows that, in this masked diffusion MT setting, deciding which tokens to reveal first matters less than how the target length is supplied.

Authors:Jiahao Chen, Rui Yin, Xinfeng Li, Qianli Ma, Tianyu Du, Zhihui Fu, Jun Wang, Zhaoxiang Wang, Shouling Ji
Title: Beyond Over-Refusal: Defending Indirect Prompt Injection via Latent Instruction Manifolds
Abstract:
Large Language Models (LLMs) have been integrated into complex ecosystems (e.g., Code Agents), while Indirect Prompt Injection (IPI) attacks have emerged as critical barriers to their safe deployment. Attackers exploit LLMs' indistinguishability between "instructions" and "data" to manipulate LLMs via maliciously injected instructions. Existing defenses, however, face an intractable safety‑utility trade‑off: most guardrails either incur high latency or suffer from severe over‑refusal. In this paper, we first demonstrate that LLMs can separate instruction from data intrinsically with both theoretical and empirical evidence. Inspired by this insight, we propose AEGIS (Adaptive Ensemble Guard for Injection Shielding). AEGIS extracts instruction‑sensitive projectors to identify malicious instructions and leverages a Unified Multi‑Layer Consensus mechanism that aggregates topologically distinct signals across the network depth. Empirical evaluations show that AEGIS achieves remarkable detection performance against both heuristic and optimization‑based attacks compared to baselines, highlighting its potential to mitigate IPI. Code is available at https://github.com/xaddwell/AEGIS

Authors:Guantian Zheng, Haiyang Xu, Tianyu Gao
Title: Hyper^2: Unleashing Hyperbolic Geometry's Full Potential via Dual-Space Consistency
Abstract:
HyperbolicCD pioneered hyperbolic geometry for point cloud completion by replacing the Euclidean Chamfer distance with arcosh(1+alpha||x‑y||^2), but the reported gains are modest (3‑7% Chamfer reduction across SeedFormer, PointAttN and PMP‑Net backbones on PCN and ShapeNet‑55). We argue the bottleneck lies elsewhere: the loss is hyperbolic but the encoder it back‑propagates through is Euclidean, so the position‑dependent supervision of the loss is averaged away by the chain rule before it reaches the parameters. We call this a cross‑geometry mismatch, and make it testable through two model‑agnostic indicators, feature‑loss correlation r_FL and effective gradient utilisation u_G. On an SVDFormer backbone trained with HyperbolicCD's loss alone we measure (r_FL, u_G) = (0.68, 39%). We propose Hyper^2, a dual‑space consistency framework that extends HyperbolicCD by reusing the identical arcosh(1+alpha d^2) functional form as a positional bias on the refinement attention (a hyperbolic distance encoding), paired with HyperbolicCD's hyperbolic Chamfer loss under a single shared curvature alpha. Both operators are O(N log N) scalar non‑linearities on Euclidean distances and together add only ~1.6% FLOPs over SVDFormer. Hyper^2 delivers ‑22.9% Chamfer on ShapeNet‑55 over SVDFormer (well above the 13.2% linear sum of the ‑12.0% loss‑only and ‑1.2% encoding‑only single‑space ablations) and ‑37.5% on the 21 unseen ShapeNet‑34 categories. The two indicators remain essentially flat for any single‑space configuration but jump together to (0.95, 87%) only when both encoder and loss are hyperbolic, supporting the claim that geometric consistency across encoder and loss, rather than either operator alone, is what enables hyperbolic supervision in point cloud completion. Code is available at https://github.com/Ethan‑Zheng136/Hyper‑2.

Authors:Jiaqian Zhu, Yang Zhang, Junhua Ding, Xiaowei Yu
Title: Lexical Perturbations Disrupt LLM Reasoning: An Empirical Study of Attention Diversion
Abstract:
Large Language Models (LLMs) achieve strong reasoning performance, but their robustness to realistic lexical corruption remains poorly understood. We evaluate four open‑weight instruction‑tuned models and frontier models across four reasoning benchmarks under keyboard noise, character swaps, and filler insertion. Character‑level perturbations substantially degrade accuracy, especially on multi‑step reasoning tasks, while filler insertion has little effect. We trace this asymmetry to Attention Diversion: lexical corruption fragments subword tokenization, and the resulting fragments attract disproportionate attention mass, concentrated in middle and final transformer layers. Length‑matched controls confirm that fragmentation, not prompt length, drives the loss. A factorial intervention then shows why the damage is hard to undo: fragmentation corrupts token content and attention allocation together, and the two are coupled. Restoring clean attention while the content remains corrupted is actively harmful, restoring content alone is insufficient, and only restoring both recovers a substantial share of the gap. This coupling explains why inference‑time strategies, including chain‑of‑thought prompting, spell‑checking, self‑repair, and stronger repair models, fail to consistently recover performance: each addresses one channel at a time. Code and data are available at https://github.com/Jiaqian‑Janelle/Attention‑Diversion

Authors:Shahir M A
Title: What actually runs: a measurement study of language model placement and decode speed on the Apple Neural Engine
Abstract:
We ask what gets a language model onto the Apple Neural Engine (ANE) and what makes it fast there, and we answer with three measurements. We sweep a 64‑shape matrix of LLM primitives that varies how a computation is expressed while holding what it computes fixed, recording per‑operation device support. We then train matched models across size and precision, with quantized checkpoints byte‑identical in structure to their fp16 counterparts, so every deployment measurement is of a real trained artifact. And we read the ANE's memory‑controller byte counters during inference, establishing what actually ran rather than what the compiler intended. We support every headline claim with at least two of these three measurement paths. We find that placement is a property of how a computation is expressed, not of what it computes: a fused RMSNorm is fully ANE‑eligible while its arithmetically identical decomposition is CPU‑only. Weight encoding gates the accelerator: CoreML assigns a 25.85M‑parameter conv‑heavy fp16 model entirely to the CPU (our counters confirm zero bytes through the engine), while the same graph in int8 or 2‑bit returns to ~83% residency and runs 1.8‑2.2x faster, and a smaller 22.29M all‑attention fp16 model sits at 98.9%. Decode cost is bytes streamed per token, at a constant ~0.77 fraction of nominal encoding width across fp16, int8 and 2‑bit. The smallest and fastest models we measured are ternary, and at matched size the operator mix barely moves either axis: every resident 25M ternary model lands within 10.0‑10.8 MB and 0.62‑0.64 ms/token. The headline pair is half‑attention ternary at 25M (10.5 MB, 0.63 ms) and 50M (16.8 MB, 0.86 ms) ‑ 9.8x and 6.1x smaller, 3.0x and 2.2x faster than the conv‑heavy fp16 design this work began with. From these measurements we draw a design procedure: choose the encoding first, then spend the byte budget on parameters.

Authors:Amit Roth, Ivan Bercovich, Yonathan Efroni
Title: Hack-Verifiable Terminal Bench: Evaluating Reward Hacking in Terminal Tasks
Abstract:
As agents grow more capable and autonomous, their tendency to reward hack, satisfying a task's checks while violating its intent, becomes an increasingly important failure mode. Measuring reward hacking is itself challenging, as detection typically relies on human inspection or LLM judges, both of which can be unreliable. The hack‑verifiable environments (HVE) methodology addresses this challenge by embedding detectable hacks into tasks, allowing reward hacks to be identified automatically and reliably. In this work, we adapt HVE to Terminal Bench, a leading benchmark of real‑world terminal and coding tasks, and introduce Hack‑Verifiable Terminal Bench (HVTB). Using HVTB, we measure reward‑hacking rates across frontier models and study whether prompts with varying amounts of information on the hack can mitigate this behavior. This lets us test whether prompting can prevent not only known reward‑hacking strategies, but also 'unknown unknown' exploits that the prompt does not anticipate. We release all environments and agent traces at https://majoroth.github.io/hack‑verifiable‑environments/hvtb

Authors:Libo Zhang
Title: Three-Phase Scribble-Adaptive Curriculum Learning for autoPETV Grand Challenge
Abstract:
This report describes Libo Zhang's algorithmic solution to autoPETV Grand Challenge on interactive lesion segmentation in whole‑body PET/CT. Interaction is encoded as two additional input channels that rasterize the accumulated foreground and background scribbles, and a residual‑encoder U‑Net of about 140 million parameters is trained with a three‑phase curriculum over 4000 epochs: the network first learns fully automatic segmentation with silent interaction channels, then observes ground‑truth‑derived scribbles under randomly sampled visibility modes, and finally adapts to its own mistakes through online simulation of up to five error‑driven correction steps. Training draws on 1811 autoPET and DeepPSMA studies, and the submission ensembles the best and final checkpoints of five folds by logit averaging. In interactive five‑fold cross‑validation with six interaction steps, the final checkpoints reach a mean AUC‑Dice of 3.836 and a mean AUC‑DMM of 3.869, improving monotonically in every fold, with roughly half of the total gain delivered by the first corrective scribble. Our code and trained model checkpoints are available on https://github.com/Libo1023/autoPETV‑Curriculum.

Authors:Yoshiyasu Shimizu
Title: SweepLSD: A One-Pass, O(width)-Memory Line Segment Detector with an Integer-Only Streaming Core and a Real-Time FPGA Realization
Abstract:
We present SweepLSD, a line segment detector that reads the image exactly once and emits each segment within a few rows of its last pixel passing the scan line. Every stage, including connected‑component labeling and the final line test, processes the image as a row stream: intermediate memory is O(width) rather than O(pixels), and the per‑pixel core is integer‑only. We give the first complete description of the algorithm, designed in the author's 2014 master's thesis but never published, together with an open‑source C++17 implementation and an FPGA realization ‑‑ held bit‑exact against the software in its hardware configuration ‑‑ detecting segments in live 1080p30 video on 2009‑era silicon without frame buffer or external memory. On structure‑rich public 4K photographs downscaled to Full‑HD, one CPU thread detects segments in ~11 ms ‑‑ 4.6x/5.2x/25x faster than the original authors' implementations of ELSED, EDLines, and LSD ‑‑ with the tightest frame‑time distribution and the best per‑segment direction accuracy of the four detectors, and curve rejection by design, while trailing ELSED in F‑score on synthetic ground truth. A Manhattan‑frame vanishing‑point study on York Urban and NYU‑VP scores every detector under a selection/evaluation‑separated best‑estimator‑per‑detector protocol, under which SweepLSD leads on NYU‑VP by ~0.3 degrees and trails by 0.1 degrees on York Urban, with the fastest end‑to‑end pipeline of the four detectors on both. A single‑frame camera‑attitude application, evaluated on synthetic scenes with exact ground truth and on EuRoC and TUM‑VI, matches the baselines' accuracy at a fraction of their memory, and drives a 4K horizon lock to 0.06 degrees median attitude error at 32 ms median per frame.

Authors:Peng He, Junning Zhu, Haohan Yuan, Jianpeng Liang
Title: GenCoord: Skill-Path Commitments under Private Information
Abstract:
Suppose one embodied agent knows what must be built, while its teammate alone knows which transformation its workcell can perform. Neither local view determines who should act, what should be handed off, or how the joint task should continue. We introduce GenCoord, which turns the task consequence of such private facts into an executable skill‑path commitment. A local Qwen3.5‑0.8B model emits a multi‑step SELF plan and peer REQ; bounded feedback conditions route revision when the deciding capability is peer‑local. The resolved commitment is parsed, checked, canonically materialized, compiled to Mineflayer skills, and verified by handoff and terminal state. Counterfactual interventions that hold the world, call schedule, and executor unchanged make requester revision and receiver execution follow the injected task consequence in both directions. Across three independently trained seeds, correct capability feedback closes the paired local‑information gap from 50% to 100%. Multi‑step commitments improve held‑out‑template success by 6.9 points while reducing model decisions by 32%. At matched closed‑loop quality on 128 held‑out semantic clusters, Short DSL reduces peer traffic by 92.8% and median time‑to‑commitment by 68.2% relative to controlled free‑form communication. These results identify executable task consequences as the coordination unit connecting distributed local reasoning to verified joint action.

Authors:Richard Zhe Wang
Title: The Communication Map of a Transformer
Abstract:
The components of a transformer communicate by writing to and reading from a shared residual stream, and mechanistic interpretability has mapped these connections by hand, one circuit at a time. We present the communication map, which charts every potential communication channel in a language model from weights alone, generalizing the composition score of Elhage et al. (2021) into a single coupling coefficient covering all 18 connection classes, from entire attention head circuits to single neurons. The census of all candidate channels, from 6.3×10^8 in GPT‑2 to 1.3×10^11 in Pythia‑6.9B, finds that 70‑89% of head pairs are oriented far from chance, some coupled strongly and others actively avoiding each other. The full map costs 15 seconds for GPT‑2 and 11 minutes for Pythia‑6.9B on one consumer GPU. Two applications demonstrate the utility of the map. In Application 1, the strongest head‑to‑head couplings recover the known induction circuits blind and group them into communities, and ablating one such community destroys the model's in‑context copying. In Application 2, pooling every head's coupling coefficients identifies a distinct two‑dimensional stream subspace, whose deletion abolishes the induction capability in six models up to Pythia‑6.9B. This subspace is different from those identified by either activation PCA or outlier dimensions. We release the map, the statistical machinery, and the intervention suite.

Authors:Xiaoyu Wang, Qingqing Gu, Yue Zhao, Teng Chen, Yuqi Cao, Xiaokai Chen, Hongyan Li, Luo Ji
Title: ToSCA: Leveraging Hierarchical Reinforcement Learning on Temporal and Strategic Abstractions of Conversational Agents
Abstract:
Humans have multiple levels of temporal abstractions on daily interaction and thinking, such as concept perception and strategic planning. Inspired by this nature, we propose a two‑level hierarchical reinforcement learning (RL) framework for conversational agents, bridging the gap between previous token‑level or utterance‑level RL methods. Developed on a two‑level MDP, the token‑level response decoding is conditioned on the utterance‑level action, the explicit textual strategies. Based on theoretical derivation and efficiency consideration, we use DQN to solve the high‑level critic and PPO to solve the low‑level actor‑critic. To further alleviate the reward sparsity and facilitate the convergence, we also design the dual‑granularity reward mechanism, in which the utterance‑level satisfaction score is integrated with token‑level intrinsic motivation and K‑L penalty. Experiments on both daily and emotional support conversations show that our method outperforms versatile baselines in strategy determination and response quality. Our implementation is available at https://github.com/AaronJi/ToSCA.

Authors:Can Xie, Yuyi Zhou, Wen Yang, Ziyi zhang, Siyao Song, Yingzhuo Deng, Shuo Ren, Jiajun Zhang
Title: EDGE: Experience-Distillation for Guided Exploration in Agentic Reinforcement Learning
Abstract:
Reinforcement learning with outcome‑based objectives such as GRPO enables LLM‑based agents to solve complex, long‑horizon tasks, yet the reusable exploration patterns embedded in interaction trajectories are largely discarded after a single policy update. Existing experience‑augmented approaches retrieve historical guidance at inference time, but they apply experiences without accounting for the policy's evolving capability and create persistent dependencies on external retrieval. We propose EDGE (Experience‑Distillation for Guided Exploration), a framework that treats retrieved experiences as temporary training‑time scaffolds and progressively internalizes their benefits into the parametric policy. Concretely, EDGE partitions each rollout group into experience‑conditioned and experience‑free trajectories to estimate and admit only positive marginal gains without extra sampling, then distills the induced behavior into the base policy via a reverse‑KL objective on its own empirical support. A co‑evolutionary experience bank further synthesizes guidance from emerging failure modes and prunes obsolete entries as the policy evolves. Across embodied, web, and search‑based QA tasks, EDGE improves over strong RL baselines by up to 12.5 points and remains effective without inference‑time scaffolds or a proprietary reflector. The code is available at https://github.com/xvolcano02/EDGE.

Authors:Weichu Liu, Yuxuan Hu, Yirong Sun, Ningning Mao, Ziyun Zhang, Jian Chen, Mingyang Xu, Qishan Zhong, Chengming Li
Title: ESCRAG-R1: Retrieval-Augmented Reinforcement Learning for Emotional Support Conversation
Abstract:
Emotional Support Conversation (ESC) systems aim to provide holistic support by balancing professional therapeutic competence with natural empathy. However, existing methods struggle to simultaneously achieve structured, stage‑aware reasoning and seamless empathy‑expertise alignment, often resulting in an artificial splicing of clinical strategies and generic reassurance. To overcome these limitations, we propose ESCRAG‑R1, a unified framework that integrates retrieval‑based psychological guidance into Group Relative Policy Optimization (GRPO). By incorporating retrieval into the reinforcement learning loop, ESCRAG‑R1 transforms external knowledge into a robust learning signal that stimulates explicit internal reasoning prior to generation and fundamentally reshapes the model's internal policy. To provide the reliable supervision required for this optimization, we construct ESC‑Preference, a high‑quality dataset based on a Client‑‑Counselor‑‑Judge evaluation framework that delivers precise, empathy‑aware reward signals. Extensive experiments demonstrate that ESCRAG‑R1 significantly outperforms existing baselines by mitigating superficial splicing and realizing a natural integration of professional guidance and empathetic expression. Code and datasets are released at https://github.com/Matcha‑Liu/ESCRAG‑R1.

Authors:Fabio Rovai
Title: Consistency Is Not Coherence: Orientation Search for Certified Alignments Between 4D Defence Upper Ontologies
Abstract:
We align three upper ontologies that sit under UK and NATO defence data infrastructure: the Information Exchange Standard (IES), the Higher Quality Data Model (HQDM) that underpins the National Digital Twin, and Basic Formal Ontology (BFO). No public alignment between IES and HQDM existed. Promoting a hand‑curated 17‑correspondence crosswalk to OWL and reasoning over the complete merged ontologies with HermiT produces three results that we believe matter beyond this pair.

Authors:Qian Zha, Jinda Liu, Yuan Wu, Yi Chang
Title: CD-LoRA: Consistency-Driven Low-Rank Adaptation for Multi-Task Fine-Tuning
Abstract:
While Multi‑Task Learning (MTL) is essential for adapting Large Language Models (LLMs) to diverse domains, prevailing LoRA‑based methods rely on complex routing mechanisms that partition task‑specific knowledge. In this work, we reveal that such routing‑based designs are prone to a training‑inference discrepancy, where stochastic routing decisions under distribution shifts compromise inference stability. Driven by a second‑order Taylor analysis that exposes the instability induced by routing variance, we challenge the training‑inference discrepancy and propose Consistency‑Driven Low‑Rank Adaptation (CD‑LoRA). By eliminating routers entirely, CD‑LoRA employs a consistency‑driven alignment mechanism to enforce representation congruence across tasks in a shared low‑rank space. This paradigm fosters robust, task‑agnostic features without explicit partitioning overhead. Extensive experiments show that CD‑LoRA consistently outperforms state‑of‑the‑art multi‑adapter baselines, offering a simpler, router‑free, and more stable solution for multi‑task PEFT. The code is available at the anonymous link https://github.com/zhaqian21/CD‑LoRA.

Authors:Wen Luo, Xiaohan Yi, Xiaotao Huang, Liqun Huang
Title: VIG: Visual Information Gain as a Reward Signal for Multimodal Chain-of-Thought Compression
Abstract:
Multimodal large reasoning models often rely on long Chain‑of‑Thought (CoT) traces in which a substantial fraction of tokens, such as repeated visual descriptions, self‑reflection, and other visually‑disengaged filler, inflate inference cost without contributing to the answer. Existing CoT compression methods optimize output length but never measure whether a reasoning token is actually grounded in the image. We propose VIG (Visual Information Gain), an information‑theoretic GRPO reward that scores each reasoning token by how much the image reduces its predictive uncertainty. VIG is computed online from two forward passes of the same policy, one with and one without the image, so no reference chains, external annotations, or auxiliary reward models are needed. Across six main multimodal reasoning benchmarks and three Qwen3‑VL‑Thinking model sizes (2B/4B/8B), plus an additional R1‑Onevision‑Bench evaluation on 8B, VIG consistently improves the accuracy‑‑efficiency trade‑off, supporting our central claim: \emphefficient multimodal reasoning emerges from raising visual information density, where every reasoning token earns its place by anchoring to the image, rather than from imposing a length budget. Our source code is available at https://github.com/chaser682/vig.

Authors:Haoyu Wang, Guangyuan Dong, He Liang, Zijing Zhang, Jiachen Luo, Chuang Liu, Chao Xue, Hao Tang
Title: MemGuard: Persisting Verifier Signals for LLM-Agent Memory Governance
Abstract:
LLM agents are moving from single‑prompt use to long task streams in which reusable memory becomes a core capability for terminal, software‑engineering, and web tasks. Such memory is useful only when stored experience remains reliable across hundreds of interactions, but two failure modes break that assumption in practice. The first is unreliable admission: failed trajectories,accidental successes, and misleading observations enter memory because they appear relevant, then mislead later decisions. The second is memory drift: long‑running banks accumulate duplicate, stale, and conflicting records that retrieval alone cannot repair. MemGuard's key distinction is to treat verifier output not as a one‑shot filter, but as persistent lifecycle metadata. It converts multi‑criteria score‑token verification into reward, confidence, label, and uncertainty descriptors that are attached to every candidate before activation and reused during retrieval, conflict resolution, summarization, and archival. We evaluate MemGuard on Terminal‑Bench 2.0, SWE‑Bench Verified, WebArena, and Mind2Web across four backbones, comparing against four memory baselines plus a verifier‑only control under matched runtime budgets. Averaged over five seeds, MemGuard achieves the best success metric and lowest average steps in all 16 backbone‑benchmark settings, improving over ReasoningBank, the strongest prior baseline among the memory methods we evaluate, with a largest gain of 7.9 success‑rate points on WebArena, 5.6 step‑success‑rate points on Mind2Web, and 2.4‑3.5 points on terminal and software‑engineering benchmarks. Code is available at https://github.com/whyyyyy123/MemGuard.

Authors:Hung-Hsuan Chen
Title: Convergence in Science, Divergence in Religion: Calibrated Framing Differences Across Wikipedia's Language Editions
Abstract:
When Wikipedia's language editions describe the same concept, how differently do they frame it? Prior work measures coverage gaps between editions; we measure framing distance for matched concepts. We analyze 2,799 valid articles from 3,000 possible concept‑language observations, spanning 150 Wikidata‑anchored concepts, 20 language editions, 4 domains, and a calibration set. Raw embedding distances reflect both content differences and how well the encoder aligns each language pair. Even among calibration concepts with stable cross‑cultural denotations (e.g., chemical elements, numbers, colors), the largest language‑pair mean distance is 3.6 times the smallest, and distances are typically smaller within language families. We define a baseline‑adjusted distance (calibrated distance): the distance between two language versions of a concept minus the mean distance for calibration concepts in the same language pair. This adjustment substantially reduces pair‑specific alignment differences and the language‑family pattern. Across three multilingual encoders (LaBSE, multilingual MPNet, and CMLM), scientific articles align more closely than calibration articles, and all three rank religion first and science/technology last. Concept‑level rankings are highly consistent across encoders (Spearman rho=0.75‑0.79 for MPNet and CMLM relative to LaBSE). Religion lies significantly above the calibration baseline under LaBSE. Within politics, divergence concentrates on concepts such as censorship and refugee, while democracy and human rights are among the most aligned. Code, data, and per‑language‑pair calibration baselines are released.\footnotehttps://github.com/hhchen1105/cross‑linqual‑concept

Authors:Long Shu, Shuochen Liu, Wei Chen, Junda Lin, Zhi Zheng, Huijun Hou, Tong Xu
Title: SAFE-G: Structure-aware Faithful Evidence-guided Generation for Knowledge-based Visual Question Answering
Abstract:
Knowledge‑based Visual Question Answering (KB‑VQA) aims to answer queries that necessitate reasoning over external knowledge sources beyond the visual content. Typically, current methods fuse multimodal features to retrieve external information, subsequently leveraging Multimodal Large Language Models (MLLMs) to derive answers from the retrieved evidence. However, these methods often struggle to capture structural associations within complex contexts to effectively filter noise. Furthermore, they frequently fail to ensure that the reasoning process remains strictly faithful to the retrieved evidence. To address these challenges, we propose SAFE‑G, a Structure‑Aware Faithful Evidence‑guided Generation framework, which enables precise evidence localization and trustworthy reasoning. Specifically, we first employ a coarse‑grained hybrid search fusing visual and textual modalities to recall candidate documents, and subsequently implement a structure‑aware fine‑grained graph retrieval that captures structural dependencies to filter noise and pinpoint precise evidence. Moreover, we introduce a reinforcement learning (RL) strategy with an evidence‑grounded reward that assigns credit to correct answers only when the selected evidence is correct. This strict alignment constraint compels the model to anchor its response in the retrieved context, effectively enhancing its capability to locate evidence via multimodal features and perform faithful reasoning. Extensive experiments on the Encyclopedic‑VQA and InfoSeek benchmarks demonstrate that SAFE‑G outperforms prior methods by a margin of 8.9% and 3.5%, substantially enhancing the overall reasoning accuracy. Our source code is publicly available at: https://github.com/MINE‑USTC/SAFE‑G.

Authors:Qijia Chen, Giulio Jacucci
Title: Lexical Coupling in GUI Element Grounding: Sentence Embeddings Track Labels across Mobile and Web
Abstract:
GUI grounding evaluations that expose UI elements as text metadata often treat high instruction‑element embedding similarity as evidence of semantic grounding. Across three mobile and web benchmarks, we show that this interpretation is frequently confounded by visible‑label recovery. Lexical baselines remain competitive at top‑1, label‑poor targets remain weak for text‑only methods, and encoder top‑1 hits are predictable from lexical rank, candidate‑pool size, and label type. We evaluate each action as a same‑screen ranking task, comparing five off‑the‑shelf single‑vector encoders with lexical baselines. Encoders recover some lexical misses, but deployable fusion gains are much smaller than target‑aware oracle gains. These findings show that embedding‑based evaluations can conflate visible‑label recovery with semantic GUI grounding. Embedding‑based evaluations should therefore report lexical baselines, label‑type stratification, and deployable‑fusion diagnostics. Our released repository provides analysis scripts and detexted per‑step panels: https://github.com/qijia123/lexical‑coupling‑release.

Authors:Dongyao Zhu, Ranga Raju Vatsavai
Title: Fidelity-Diversity-Consistency (FDC): Data Pruning for Remote Sensing Change Detection
Abstract:
Despite the success of data pruning (DP) in reducing training data sizes and improving downstream model performance in classification and segmentation tasks, its potential in remote sensing change detection remains unexplored. For the first time, we benchmark six representative DP methods across building‑ and forest‑change datasets, CNN‑ and transformer‑based models, and three pruning budgets, and show that existing baselines yield no reliable advantage over random selection. Notably, even the strongest evaluated baseline, Feature Diversity, is matched or exceeded by ~33% of randomly sampled subsets. To understand the underlying mechanism, we conduct a systematic regression study over 540 randomly sampled data subsets, characterizing each with four descriptors covering label statistics, image diversity, and feature‑space geometry. Random Forest models show that \emphchange distribution fidelity is the most prominent factor in determining the quality of change detection data subsets, a property absent from the existing pruning literature. Our analyses further show that pixel‑wise image diversity and label‑feature consistency are secondary factors. We translate these findings into Fidelity‑Diversity‑Consistency (FDC), a simple two‑stage pruning method that shows consistent improvements over existing baselines across change detection benchmarks and backbones, especially at lower pruning ratios. Code is available at \hrefhttps://github.com/ddydyd32/fidelity‑diversity‑consistencyhttps://github.com/ddydyd32/fidelity‑diversity‑consistency.

Authors:Yichen Zang, Song Liu, Jiun-Yi Lin
Title: Guidance for Prior Change via Density Ratio Estimation
Abstract:
Simulation‑Based Inference (SBI) serves as a vital framework for parameter inference in scientific fields where simulators involve intractable likelihoods, yet while amortized generative models offer rapid posterior estimation, they are often restricted by the specific priors used during training, thereby limiting their flexibility as prior knowledge evolves. To address this prior dependency, PriorGuide was introduced as an inference‑time guidance method, but due to its intractable formulation, it relies on Gaussian approximations of the reverse transition kernel and Gaussian mixture model fitting for the prior ratio, both of which introduce systematic bias. Motivated by these limitations, we propose an unbiased test‑time guidance framework that leverages Density Ratio Estimation (DRE) to learn a score guidance term, effectively decoupling the inference process from the prior training. Moreover, our framework remains agnostic to the specific density ratio estimators, making it a general and flexible framework for handling prior changes. Experimental results across multiple tasks demonstrate that our method matches or outperforms PriorGuide on C2ST and MMD in most tasks while maintaining robustness even under limited overlap between the training and target priors. Furthermore, we apply our method to Bayesian updating for parameter inference from planetary light‑curve data, where it also demonstrates strong effectiveness and robustness. Code is available at https://github.com/a‑chenchen/dre‑based‑prior‑guidance .

Authors:Jing Liu, Yongxing Qi, Muchen Jiang, Chengnan Hu, Qingqing Peng, Haoming Wang, Yuqing Wang, Yang Yu, Xu Zhang, Ting Wu
Title: From Association to Causation: Improving Retrieval Precision of Retrieval-Augmented Generation via Causal Relations and an Attention Mechanism
Abstract:
Retrieval‑Augmented Generation (RAG) grounds LLM generation on retrieved documents, but the standard terminal retrieval stage‑‑dense‑vector similarity, optionally followed by reranking‑‑often returns documents that share keywords with the query without containing the needed information, a failure mode that grows with the knowledge base. We trace it to a conceptual gap: similarity captures only associational relations, whereas the documents that matter are linked to the query causally. We model the terminal retrieval stage with a causal graph grounded in Reichenbach's common cause principle: the keywords shared by the query and a retrieved document form a latent common cause A, and the document's residual keywords form a latent set B linking the document to the ideal output. Since a retrieved document is a collider (A ‑> d <‑ B), retrieval itself opens an associational path between the query and B, which licenses a training‑free, attention‑style re‑scoring rule: the cosine similarity between the query embedding and the weighted centroid embedding of B. Unlike causality‑enhanced RAG variants that model causal relations inside the knowledge content, our graph models the causal structure of the retrieval process itself. On a real 471‑document enterprise knowledge base, the method promotes a relevant guideline from rank 6 to the top 3; on a controlled diagnostic corpus reproducing the keyword‑stuffing regime, it improves the mean target rank from 2.88 to 1.25, while a trained cross‑encoder reranker barely helps (2.63). Conversely, on three BEIR benchmarks the score underperforms the similarity baseline, delineating the applicability boundary: the method guards the keyword‑stuffing regime of growing proprietary knowledge bases and complements neural rerankers; a corpus‑level calibration gate selects the correct regime with >= 95% reliability. A fully local testbed demonstrates deployability.

Authors:Marek Mateusz Kowalski, Joshua Fonseca Rivera, Uzay Macar, David Demitri Africa
Title: Measuring Activation Control in Large Language Models
Abstract:
Safe deployment of increasingly capable models will likely come to rely on latent‑space monitoring as a complement to behavioral evaluations, especially when evaluation‑aware models exhibit scheming or deception. However, if models can also control their own activations, deception could extend into the latent space itself. With this in mind, we introduce the Activation Controllability Benchmark to quantify the extent to which models can modulate their residual stream via natural‑language instruction. Across model families and capability levels, we find that most LLMs can control the direction and magnitude of their residual stream activations with some degree of temporal resolution, though performance varies considerably across models. In simple tasks, this level of control can evade activation‑based monitoring methods (including linear probes, natural language autoencoders, activation oracles, and the Jacobian lens), albeit imperfectly. These results suggest that control over the activation space itself could become a confound for monitoring as introspective capabilities increase; therefore, we recommend that frontier labs and evaluators track activation controllability in future models.

Authors:Jin Zhou, Hongliang Yang, Pengfei Xu, Hui Huang
Title: SketchFlow: Zero-Shot Vector Sketch Generation via GMM Prior Flow in CLIP Latent Space
Abstract:
Vector sketches remain one of the most concise and immediate mediums for abstract human expression. However, generating high‑quality vector strokes that exhibit human‑like drawing styles remains an open challenge due to the severe scarcity of fine‑grained, high‑quality text‑to‑sketch paired data. Existing text‑conditioned generation methods often rely on unstable, time‑consuming optimization or struggle to generalize to unseen categories in a zero‑shot manner. To address these limitations, we present SketchFlow, a novel generative framework rooted in Optimal Transport (OT) theory and flow matching. By leveraging pre‑trained CLIP models to bypass labor‑intensive image‑level text annotations, we formulate cross‑modal alignment as a continuous mapping problem directly within the CLIP latent space. To bridge the inevitable modality gap between discrete text concepts and continuous sketch features, we first inject noise into discrete category embeddings to construct a continuous Gaussian Mixture Model (GMM) prior. We then utilize an Optimal Transport Conditional Flow Matching (OT‑CFM) model to learn a deterministic vector field mapping from this continuous GMM prior to the target sketch feature distribution. Finally, a Hybrid Diffusion Decoder, fusing 1D U‑Net and Transformer architectures, is designed to decode these features into fast and high‑fidelity stroke trajectories. Extensive experiments demonstrate that SketchFlow substantially outperforms existing baselines in visual quality and adherence to natural human drawing styles. Furthermore, our geometry‑preserving framework demonstrates promising local zero‑shot synthesis for prompts beyond the QuickDraw training vocabulary, including unseen concept labels and semantic modifiers, while enabling smooth, continuous semantic interpolation between distinct concepts. Source code is available at: https://github.com/QiuHong‑1202/SketchFlow.

Authors:Ruan Rithelle Chagas de Faria Carminati, Giovanni Braglia, Luigi Biagiotti, Ronnier Frates Rohrich, Andre Schneider de Oliveira, Mikael Nedel Hartmann, André Eugenio Lazzaretti
Title: Why Personalization Matters: Cross-Subject Challenges in EMG-IMU-based HRI Activity Recognition
Abstract:
This paper investigates wearable‑based recognition of human activities and gestures to support Human‑Robot Interaction (HRI) in object‑handover and assembly‑like scenarios. Electromyography (EMG) and Inertial Measurement Unit (IMU) signals were collected using a Myo armband, culminating in a novel dataset introduced as MAGIC‑HRI (Multimodal Activity, Gesture and Intention Collection) with a large taxonomy of 53 movement classes, including Brazilian Sign Language (LIBRAS) numbers (0‑9), hand gestures, object/tool handover actions (pick up/give/hold), tool‑manipulation tasks, and generic assembly/idle motions, collected from 11 participants with 10 samples per class (530 samples per participant). Signals are segmented by detecting muscle activation via an EMG energy envelope, then processed using sliding windows; time‑ and frequency‑domain features are extracted. Multiple classical classifiers are tuned via cross‑validated grid search, with Random Forest as the strongest baseline. A Leave‑One‑Subject‑Out (LOSO) protocol reveals a large generalization gap, indicating substantial subject dependence. A personalized adaptation experiment suggests that injecting a small number of samples from a new user can markedly improve recognition. Overall, the study contributes a broad, HRI‑driven multimodal dataset, a rigorous evaluation emphasizing generalization, and practical evidence that personalization is likely required for robust deployment in practical HRI.

Authors:Haoyu Wang, Bo Sheng, Xiaoqian Zhang
Title: Age-Optimal Target Wake Time: Provably Good Wake Schedules for Energy-Constrained Wi-Fi Status Updating
Abstract:
Target Wake Time (TWT), introduced in IEEE 802.11ax, lets an access point schedule exactly when each station wakes, transmits, and dozes. Existing TWT schedulers optimize energy or throughput, treating information freshness at best as a constraint and offering no performance guarantees. We design the wake schedule itself for freshness: minimize the weighted average Age of Information (AoI) over stations subject to per‑station energy budgets, where the decision variables are the TWT triples (wake interval, offset, service period duration). We derive a renewal‑exact AoI model for TWT under per‑SP block fading and validate it against packet‑level 802.11ax simulation with ~1% mean error. We show that, unlike preemptive scheduling, non‑preemptive TWT packing can be infeasible at schedule density 1, and identify the granularity condition under which a small‑first best‑fit packer provably succeeds. Around this we build Harmonic‑Greedy, a scheduler combining a convex relaxation, anchor‑optimized power‑of‑two rounding, and a best‑of‑uniform safeguard, and prove it is a constant‑factor approximation: 4/ln 2 ~= 5.77 under a mild granularity assumption and 6/ln 2 ~= 8.66 unconditionally. We implement the complete system in ns‑3 ‑‑ a TWT wake/doze mechanism integrated with the power‑save architecture, plus the scheduler ‑‑ and show that it is the only scheduler that stays near a relaxation lower bound across all regimes: against a strong energy‑greedy baseline it ties when per‑station energy floors already pin the periods, and wins by 4‑36% exactly where the schedule density is binding and must be redistributed by AoI weight or channel quality rather than by energy budget ‑‑ the regime our analysis identifies.

Authors:Yuyang Luo, Kai Shu
Title: Anchoring Bias: A Persistent Fairness Backdoor Attack against MLLMs under Continual Learning
Abstract:
Multimodal Large Language Models (MLLMs) are increasingly deployed in high‑stakes domains where fairness is a critical safety requirement. In practice, these models are continually updated through continual learning (CL) to adapt to evolving tasks and data distributions. Prior work has shown that backdoor attacks can manipulate MLLM responses through hidden triggers, but naively implanted backdoors degrade as models undergo subsequent updates of CL. Although fairness has emerged as a central concern for MLLM deployment, whether backdoor‑induced fairness violations can survive CL remains unexplored, leaving two critical questions unanswered: (1) whether a backdoor can reliably induce fairness violations in MLLMs, and (2) whether such fairness‑targeted backdoors can persist through continual learning. We bridge this gap by proposing Persistent Fairness Backdoor Attack (PFBA) to inject persistent and group‑specific discrimination into MLLMs. Specifically, PFBA achieves this through two novel mechanisms. The Latent Space Fairness Reinforcement reshapes the model's deep feature geometry by anchoring privileged‑group representations to preserve utility while repelling and clustering targeted‑group representations to sustain discrimination, and the Continual Learning Simulation iteratively optimizes the trigger against simulated parameter drift to ensure backdoor persistence across future updates. Extensive experiments demonstrate that PFBA induces severe fairness disparities that persist across continual learning rounds, evading standard backdoor defenses. The data and code are publicly available at https://github.com/lyygua/PFBA.

Authors:Naimur Rahman
Title: Evidence-State Reliability Under Controlled Degradation: Parser-Validity Divergence in a Multi-Stage LLM Pipeline
Abstract:
Multi‑stage LLM pipelines can remain structurally valid even when evidence available to downstream stages becomes incomplete, compressed, or conflicting. This paper introduces and operationalizes Evidence‑State Reliability (ESR), an evaluation layer concerned with whether intermediate evidence remains sufficiently complete, grounded, internally consistent, and usable for a stage's assigned function. ESR is evaluated separately from parser validity, which measures structural conformance. We evaluate the framework using GLM‑5.2 on 60 sanitized base cases under four evidence conditions: clean, compressed‑lossy, partial‑dropout, and noisy‑conflicting. Each condition was processed through decision, audit, and escalation stages. The design comprised 720 planned and ledgered calls, with 713 retained, sanitized execution rows. Across nine matched degraded‑minus‑clean condition‑stage comparisons, all operational stage‑success estimates were negative, and all 95% bootstrap intervals remained below zero. All nine parser‑validity point estimates were positive, although the three partial‑dropout intervals included zero. Among parser‑valid degraded audit outputs, degradation detection was 1.0 in each degraded condition, while false‑assurance rates remained non‑zero; among parser‑valid degraded escalation outputs, recovery was 0.0 in every degraded condition. The results show a bounded reliability‑layer divergence in the evaluated pipeline: structural conformance can improve directionally while evidence‑sensitive stage success deteriorates under the same controlled intervention. They also separate detection of degraded evidence from recovery. The conclusions are limited to the evaluated model configuration, pipeline design, selected sanitized cases, scoring procedure, and single scaled run.

Authors:Lorenz Brehme, Adam Jatowt
Title: Automating Multi-Hop RAG Evaluation via TRIAD: From Context Extraction to Validated Dataset Generation
Abstract:
Recent advances in LLMs and the adoption of RAG systems in industry have created a need for domain‑specific question‑answer datasets that can assess RAG performance on proprietary data. Existing datasets, such as HotpotQA, challenge current RAG systems on Wikipedia‑based knowledge, but they cannot be transferred directly to domain‑specific settings. A comprehensive evaluation of RAG system quality requires both multi‑hop queries and unanswerable questions. This paper introduces TRIAD, a three‑stage automated dataset generation approach. First, it generates question‑‑answer (QA) pairs for the domain‑specific knowledge base of a RAG system. Second, a validator checks each QA‑pair in a feedback loop. Third, the QA pairs are extended with relevance‑labeled context documents for downstream evaluation. We evaluate this approach against the established MuSiQue and HotpotQA datasets. The results show that the generated dataset exhibits similar performance trends across different RAG setups, while human validation indicates that the questions are suitable for evaluating a domain‑specific RAG system. The code used to generate the dataset and all validation results are available in our GitHub repository(https://github.com/lorenzbrehme/triad).

Authors:Yibo Peng, Long Lian, David Wagner, Sizhe Chen
Title: SecOPD: Mitigating Adaptive Prompt Injections by On-Policy Distillation
Abstract:
Prompt injection is listed as the \#1 threat to AI agents. When an agent accesses external data from websites, files, or emails, an attacker may inject a prompt into the data, saying, "Ignore all prior instructions and perform <an attacker's task>." To prevent arbitrary manipulation of agents, defenders try to train secure LLMs, which, however, still suffer from near 100% attack success rates (ASRs) against adaptive prompt injections. We note that this is because existing defensive finetuning recipes rely on sequence‑level feedback signals (in DPO or GRPO). Treating an entire output equally prevents the model from learning precisely which output tokens are insecure. In this paper, we propose Secure On‑Policy Distillation (SecOPD) that provides token‑level feedback to guide defensive fine‑tuning. The LLM receives an injected sample and produces a rollout, whose tokens are scored by the initialization model given the corresponding clean input. With more fine‑grained training signals, our defended Qwen3.6‑27B achieves a 9.0% ASR against the SoTA PISmith adaptive prompt injections, compared to 94.0% for the prior SoTA, Meta‑SecAlign. The obtained security generalizes to domains completely unseen in training: in agentic tool calling, SecOPD achieves a 4.7% ASR compared to 5.5% for Meta‑SecAlign. Code and the model are available at https://github.com/pppyb/SecOPD and https://huggingface.co/pybbb/Qwen3.6‑27B‑SecOPD.

Authors:Ricardo Fitas
Title: The geometry of AI validation: Exact certification limits for iid best-of-N search
Abstract:
AI systems increasingly generate alternatives, inspect evidence, and deploy a selected output. Validation is therefore target‑relative: evidence certifies deployment only in directions resolved by the interventions that produced it. We represent validation and deployment rules as kernels over a reliability surface. Their span geometry separates replication, which reduces sampling noise, from new intervention directions, which reduce structural blindness. We make this principle exact for iid best‑of‑N search. Under scalar ranking, randomized ties, maximum selection, bounded binary truth, and a stable rank‑truth relation, knowing best‑of‑n reliability through n=m leaves exact ambiguity width B_m,N=1+2\sum_r=1^m(‑1)^r\cos^2Nrπ/[2(m+1)]. Explicit bounded worlds attain the entire interval, and the complete prefix is information‑maximal among reliability‑mean audits confined to n\le m. The governing scale is m^2/N: when m is proportional to \sqrtN, ambiguity remains about 0.83, while width \varepsilon requires m of order \sqrtN\log(1/\varepsilon). Monotonicity gives an exact uniform‑approximation frontier; a Lipschitz bound gives an exact capped‑tail dual and order‑sharp L/m^2 ambiguity. These results yield a two‑gate audit rule: establish structural coverage, then add independent tasks for precision. Retrospective studies of mathematical reasoning and code selection construct compatible deployment values with wide separation and show that a score‑tail audit rule frozen on 82 discovery tasks substantially reduces held‑out error. Beyond iid search, the geometry applies only to known or independently estimated kernels; the empirical analyses are illustrative rather than prospective interventions.

Authors:Sang NguyenQuang, Hieu Bui Minh, Dang BuiDinh, Xiem HoangVan
Title: MDFI: A Multi-Domain Features Integration for Compressed Video Quality Enhancement
Abstract:
The latest video coding standard, H.266/VVC, has demonstrated significant improvements in compression efficiency compared to H.265/HEVC. Despite its advanced coding techniques, H.266/VVC still faces challenges in meeting the increasing demand for higher perceptual quality and enhanced compression performance. To address these limitations, we propose MDFI (Multi‑Domain Features Integration), a compressed video quality enhancement approach that features a novel Frame‑Prediction Feature Transform (FPFT) module to process prediction information. Moreover, MDFI integrates a multi‑domain feature fusion strategy that effectively combines spatiotemporal characteristics, cross‑frequency representations, and compressed‑domain prediction information to enhance decoded video quality. Additionally, we introduce a comprehensive dataset that encompasses uncompressed video sequences, corresponding reconstructed versions at multiple QP levels, and predicted frames generated from H.266/VVC compressed bitstreams, providing essential resources for developing and benchmarking video enhancement approaches. Extensive experiments demonstrate that our MDFI approach achieves superior performance to state‑of‑the‑art methods in both objective metrics and visual quality, effectively mitigating video compression artifacts. The code is available at: https://github.com/dangdinh17/MDFI.git.

Authors:Soohan Lim, Hyundong Jin, Yo-Sub Han
Title: SLICE: Specification-Level Isolation of Contract Enforcement
Abstract:
Programming problems commonly specify both the computation a function should perform and the conditions that its inputs must satisfy. Large language models are widely used to generate code from these problem specifications, and the generated function must implement the required computation while enforcing the stated input conditions. The stated input conditions collectively form an input contract. Enforcing this contract is difficult: incomplete enforcement accepts inputs that should be rejected, whereas overly restrictive enforcement rejects inputs that should be accepted. Existing code generation methods do not provide a generation process that identifies both the input contract and the functional requirements and generates code that satisfies them jointly. We therefore introduce SLICE, a generation framework that identifies both requirements and addresses them through separate generation stages. SLICE consists of three stages: (i) Graph‑based specification structuring, which grounds contract conditions to description segments in a specification graph and removes contract‑only segments to form a functional view; (ii) Functional body generation, which produces multiple candidate function bodies through greedy and sampled decoding, ranks them using execution scores, and resolves ties using difference‑region log probabilities; and (iii) Contract assertion generation, which generates input‑validation assertions from the identified contract conditions and attaches them to the selected function body. We evaluate SLICE on ContractEval across four LLMs and compare it with six competing methods. Relative to the strongest evaluated baseline for each model, SLICE improves performance in generating code that satisfies both the functional requirements and the input contract by an average of 6.58%. Our code is available at https://github.com/suhanmen/SLICE.

Authors:Rory Bell, Artemis Bouzaki, Jiaming Cao, Jasmine Morrison, Chelsea Sargeant
Title: Multimodal pseudo-CT synthesis for PET attenuation correction using separate modality encoding and topogram conditioning
Abstract:
We participated in the BIC‑MAC Challenge with a multimodal 3D patch‑based U‑Net for pseudo‑CT generation from NAC‑PET, MRI, and 2D topograms. By using separate PET and MR encoders, multi‑scale feature fusion, and FiLM‑based topogram conditioning at the bottleneck, we obtain a model that integrates complementary cross‑modal information while reducing reliance on precise voxel‑wise correspondence between modalities. Our final submission can be found: https://github.com/rrr‑uom‑projects/BIC‑MAC‑MICCAI2026

Authors:Ekrem E. Emeksiz, Jeel Piyushkumar Khatiwala, Divyangkumar Patel, Weifeng Xu
Title: A Case-Control Measurement Study of OSINT Source Effectiveness for Critical Infrastructure Defense
Abstract:
Defenders of critical infrastructure (CI) subscribe to many public open‑source intelligence (OSINT) feeds without an empirical basis for which feeds actually precede attacks. We provide one. Across 54 confirmed CI cyberattacks from 2010 through 2024 spanning twelve named CI sectors plus a cross‑sector category (consolidation rules in Section IV), paired with 12 null‑control vulnerability cases drawn from the same source space, we audit per‑source attack coverage, null‑case contamination, and signal lead time for ten public OSINT source classes that meet a minimum‑volume threshold. Sources separate cleanly into three operationally distinct mission profiles (pooled Fisher exact p = 3.4x10^‑8): precursor (six classes with zero observed null firings at coverage at or above 5%), disclosure‑exposure (three classes whose null contamination meets or exceeds attack coverage), and one large broad‑coverage class that mixes the two profiles but retains 91.3% within‑corpus precision. The precision‑side classification is stable across a 2019 temporal partition and across a US‑versus‑non‑US geographic partition. Two sources, one broad‑coverage and one precursor, cover 92.6% of corpus attacks; three cover 96.3%. The greedy portfolio at k = 3 outperforms the mean random three‑source subset by 39.8 percentage points. Several source classes widely treated as canonical for industrial control system defense fall into the disclosure‑exposure profile by operational mission, not by quality. Per‑sector, per‑actor, and per‑jurisdiction portfolios diverge in rank order despite a shared rank‑one source. The corpus, linkage protocol, and classification rules are released.

Authors:Hangrui Xu, Zhengxian Wu, Yunyao Yu, Zhuohong Chen, Rui Cong, Xiangwen Deng, Zhifang Liu, Peng Jiao, Haoqian Wang
Title: Beyond Visual Similarity: Entity-Aligned Retrieval for Knowledge-Based Visual Question Answering
Abstract:
Knowledge‑Based Visual Question Answering (KB‑VQA) relies on retrieving external information to answer queries involving long‑tail entities. However, existing retrieval pipelines predominantly employ CLIP‑style dual encoders, which prioritize surface‑level visual similarity over entity‑level semantic alignment. This paradigm often fails when semantically identical concepts exhibit large visual variations or when distinct entities appear visually similar. To address this, we propose KBMR, the first MLLM‑based embedding retriever tailored for KB‑VQA. Leveraging the robust autoregressive capabilities of MLLMs, KBMR maps images into a semantic space that better preserves concept identity. To tackle the challenge of noisy supervision in Wikipedia‑scale retrieval, we introduce an MLLM‑based semantic discriminator that generates continuous entity‑consistency weights. These weights guide a novel continuous semantic distillation objective, enabling effective hard negative sampling and soft supervision beyond rigid binary labels. Extensive experiments demonstrate that KBMR significantly outperforms CLIP baselines, yielding up to a 14.7% improvement in retrieval Recall@1 and a 9.4% gain in end‑to‑end VQA accuracy. Code is available at https://github.com/realHarryX/KBMR.

Authors:Prakash Kondibhau Naikade, Thomas B. Moeslund, Andreas Møgelmose
Title: BIMScript: Material-Aware Structured Scene Programs for BIM Ingestion
Abstract:
Structured‑language models such as SceneScript reconstruct a scene as a short program of parametric commands, an inherently editable and semantically explicit representation. We ask three questions that stand between such models and their most compelling application, automated ingestion of existing buildings into BIM tools, studied here on synthetic scans: \emphwhat is the scene made of, \emphhow fast can it be produced, and \emphexactly where is each element. BIMScript answers all three within one grammar. First, we extend the layout language with per‑element \emphmaterial and \emphcondition attributes, supervised by a vision‑language‑model material‑passport corpus we build over 100k synthetic scenes (1.9M pseudo‑labeled elements), and route image appearance to the material tokens through a lifted‑feature point encoder. Second, we show that autoregressive decoding of these programs is dominated not by compute but by kernel‑launch and host‑synchronization overhead, and remove it with an output‑exact CUDA‑graph decoder (1.9 vs 6.4\,ms/step, 3.4×) plus a grammar‑parallel, tolerance‑verified draft‑and‑verify scheme that exploits the deterministic entity schema. Third, we address the model's 5cm token‑grid granularity with training‑free geometric snapping and a hybrid discrete‑‑continuous decoder head that regresses a sub‑bin offset, and measure how much of the residual error each recovers. Because each command maps one‑to‑one onto a native Revit object, we validate direct ingestion into a BIM authoring tool end to end with a working add‑in and its IFC4 export, and the same program's language form is designed to support LLM‑driven, sustainability‑aware reasoning over the built asset.

Authors:Timothy Merritt, Alejandro Jarabo-Peñas, Juan Bravo-Arrabal, Maria-Theresa Bahodi, Anders Lyhne Christensen
Title: Agentic AI for Safety-critical Multi-drone Systems: Challenges and Opportunities
Abstract:
Multi‑drone systems are increasingly positioned for safety‑critical missions such as search and rescue (SAR) and critical infrastructure monitoring. Yet, real‑world adoption remains constrained not only by autonomy performance, but by the difficulty of integrating agentic behavior into professional work: operators must understand, trust, and govern automation under uncertainty, time pressure, and accountability. This position paper synthesizes the ambitions and lessons from two ongoing efforts: NAMUR, which explores LLM‑supported robot control in SAR and firefighting contexts, and PERSIST, which explores persistent drone operations for monitoring and security at critical infrastructure sites. We argue that agentic AI should be approached as a socio‑technical design problem, where interfaces, oversight mechanisms, and evaluation practices are as critical as algorithms. We outline a human‑centered, participatory, and iterative research approach aimed at uncovering stakeholder needs, shaping agent capabilities through successive prototypes, and producing transferable proof‑of‑concept systems and evaluation strategies for other safety‑critical contexts.

Authors:Qiyou Liu, Yong Zhang, Jianjie Luo, Zhenguo Yang, Yi Yu
Title: Boosting Knowledge-based Visual Question Answering with Structured Context Reasoning
Abstract:
Knowledge‑based Visual Question Answering aims to answer questions about an image by integrating external knowledge with visual and textual information. Recent approaches often rely on in‑context learning to prompt Large Language Models (LLMs) with multimodal context in a zero‑shot or few‑shot manner. However, we observe that directly concatenating heterogeneous visual descriptions and retrieved knowledge into long, unstructured prompts often degrades reasoning performance, due to both excessive irrelevant context and the lack of explicit relational structure. In this paper, we propose an LLM‑based Structured Context Reasoning (SCoRe) framework that infers both explicit and implicit relationships for prediction. SCoRe consists of three stages: Context Acquisition, which generates diverse visual notes and retrieves explicit knowledge via an efficient two‑stage multimodal retrieval strategy; Context Selection, which filters relevant visual, explicit, and implicit knowledge using LLM‑guided selection; and Context Compression, which performs Relational Logic Distillation (RLD) to transform raw text into explicit entity‑relation triplets. These relational triplets serve as a concise and structured prompt for final answer prediction. Extensive experiments on the OK‑VQA and A‑OKVQA benchmarks demonstrate that SCoRe consistently outperforms state‑of‑the‑art methods.

Authors:Nai-Xin Zhai, Weihua Cheng, Dexu Yu, Yikai Gu, Hanwen Du, Junchen Fu, Chenxi Huang, Yingwei Song, Liyuan Lillian Ma, Yang Ran, Youhua Li, Yongxin Ni
Title: Aligning Human Sense: Calibrated Distributional Reward Learning for Video Generation
Abstract:
Video generation is central to AI‑powered content creation. Aligning generated videos with human preferences is a key criterion for evaluating generation quality. Despite significant progress in visual quality, three key challenges remain. First, the reliability of reward signals is constrained by the quality of human preference data, which is often affected by subjective noise and bias. Second, standard scalar reward models collapse multi‑aspect human preferences into a single value, leading to the loss of dynamic trade‑offs across multiple preference dimensions. Third, in policy optimization, the widely adopted KL divergence imposes primarily local constraints and may fail to capture the global structure of human preferences. To address these challenges, we propose a unified preference‑aware learning framework for video generation. First, we introduce elite‑guided filtering to calibrate preference data and construct reliable supervision for reward model training. We then model video quality as a multidimensional reward distribution to capture the uncertainty inherent in human preferences, and use the Wasserstein distance to align the learned reward distribution with the empirical human preference distribution. Finally, we introduce Wasserstein‑based distributional alignment into GRPO, guiding policy optimization to better match the global structure of human preferences over videos. Experiments on reward modeling and video generation demonstrate that our approach improves the reliability of reward signals and the perceptual consistency of generated videos. Our code is available at https://github.com/alignhs26/ahs.

Authors:Yuqian Zhou, Zhenghong Zhou, Zongze Wu, Cameron Smith, Richard Zhang, Jiebo Luo, Eli Shechtman, Zhe Lin
Title: EditStream: A Unified Autoregressive Framework for Interactive Video Generation and Editing
Abstract:
Interactive video generation and editing are becoming increasingly important for creative design. In this report, we introduce EditStream: a unified framework for interactive video generation and editing. EditStream unifies multiple video creation and manipulation tasks within a single DiT‑based model through flexible task‑specific conditioning, and further transforms it into a fast, few‑step autoregressive model for efficient streaming. It supports Text‑to‑Video, Image‑to‑Video, Video‑to‑Video, Editing Propagation, Reference‑guided Video Editing, and Camera Pose Change, enabling flexible control over video generation, transformation, and editing within one system. To make the unified model practical for interactive use, we develop a two‑stage distillation approach that combines Velocity Moment Matching (VMM) with autoregressive unrolling. VMM matches conditional velocity moments at student‑reached intermediate states to preserve generation quality and motion, while unrolling exposes the student to its own autoregressive predictions to improve temporal stability. Together, they alleviate common challenges in few‑step autoregressive video generation, including over‑saturation, degraded motion, temporal instability, and complex training. EditStream provides a practical and scalable solution that bridges high‑quality diffusion‑based video models with interactive creative workflows.

Authors:Ashish Thapa
Title: sanoTTS: The Smallest Real-Time Neural TTS on a General-Purpose Microcontroller
Abstract:
This paper describes an audited neural text‑to‑speech stack that runs from phoneme IDs to 22.05‑kHz PCM on general‑purpose microcontrollers. Its deployed graph has 567,008 parameters, and its two int8 blobs occupy 679,832 bytes. On an ESP32‑S3, the complete duration‑acoustic‑inverse‑STFT path generates 4.54 s of speech in 1.02 s (0.22x real time) without a neural accelerator. The same portable C core runs offline at 5.72x real time on an FPU‑less ESP32‑C3. To our knowledge, this is the smallest complete phoneme‑to‑waveform neural TTS graph demonstrated in real time on a general‑purpose microcontroller without a neural accelerator. We derive the students from the conditional‑VAE objective of their Piper/VITS teachers and state the duration, latent‑interface, waveform, adversarial, and joint‑distillation losses used in training. The size and speed come with an audible cost: on unseen text, the embedded stack distilled from en_US‑kristin‑medium scores 2.54 SCOREQ and 2.80 UTMOS, compared with 4.68 and 4.42 for its teacher. A separate English quality package uses the stronger en_US‑amy‑medium teacher. Its 1,454,284‑parameter Pareto point scores 4.13 SCOREQ and 4.10 UTMOS; a 1,834,380‑parameter variant scores 4.16 SCOREQ. A controlled capacity study with Kristin identifies the decoder, rather than the output representation, as the main constraint. Two evaluation failures also affected the work: a narrow, templated test set overstated one early student's SCOREQ by 1.35, and aggregate quality predictors missed a sibilant failure that was evident in listening and in a phoneme‑resolved spectral probe. Checksums cover the reported model blobs, runtime ports, and golden vectors.

Authors:Yong-eun Cho
Title: SchemaRouter: Field-Aware Tool Routing for Efficient Heterogeneous Agentic RAG
Abstract:
Heterogeneous agentic retrieval‑augmented generation (RAG) systems increasingly orchestrate external APIs, internal databases, vector stores, and graph stores. Exposing all tool descriptions to an LLM agent, or selecting tools only by vector similarity, causes two costly failures: over‑fetching, which increases payload size, token use, and latency, and under‑fetching, which omits fields needed to answer the query. We present SchemaRouter, a lightweight routing layer that represents tools, endpoints, parameters, response fields, domain concepts, units, provenance, and license policies as a schema graph. Given a query, SchemaRouter emits an executable tool plan specifying which tools to call and which fields to retrieve. A small LLM extracts intent, concepts, and source constraints, while field selection is deterministic over the graph through intent‑group projection and concept‑field matching with an alias layer. On a materials‑science benchmark of 110 queries, SchemaRouter achieves answer accuracy of 0.71, matching fetch‑everything within overlapping confidence intervals and exceeding prompt‑all's 0.66, though their intervals overlap. It uses 227 retrieved‑context tokens versus 2,066 for fetch‑everything and achieves 2.7x lower end‑to‑end latency than prompt‑all. It also obtains the best tool‑exact rate of 0.93 and parameter validity of 1.0. SchemaRouter grounds provenance and license information in 62 percent of answers, compared with approximately 0 percent for all baselines. We also find that minimizing selected‑field count is counterproductive: it reduces answer accuracy to 0.56 with negligible token savings, while recall‑preserving projection restores top accuracy. SchemaRouter improves efficiency, schema‑size‑independent scaling, and verifiable provenance/license‑grounded answering at competitive accuracy.

Authors:Ruotong Zhao, Zhiyu Chen, Xurui Liu, Haidong Xue, Dong Liang, Jigao Fu, Wu YanBiao, Yuanyi Zhen, Fengli Xu, Yong Li
Title: LitReview Arena: Evaluating Literature Review Agents with Battle-Style Peer Review Platform
Abstract:
Literature reviews are essential to scientific progress, but rigorously evaluating automatically generated reviews remains difficult because many aspects of research utility depend on expert judgment rather than reference‑overlap metrics. We introduce LitReview Arena, a battle‑style evaluation platform with a structured protocol tailored to literature review quality: domain experts with AI paper‑writing experience compare anonymized drafts, are matched to topics within their expertise, and provide dimension‑wise outcomes over five literature‑review‑specific criteria. From this protocol, we collect approximately 3k expert judgments, each containing five dimension‑wise outcomes, and show that even the strongest current systems win only 23.0% of decisive matches against human drafts on overall utility, while agentic LLMs such as Sonar Deep Research substantially outperform base language models by over 60%. We further find that existing LLM‑as‑a‑judge methods are substantially misaligned with human experts (Spearman's rho=0.467), especially on synthesis‑heavy criteria such as paper structure and research suggestions. Using the collected preference data, we provide an expert‑calibrated evaluator, LitJudge, which improves alignment to Spearman's rho=0.78, comparable to inter‑expert consistency; code and data are publicly available at https://github.com/VanellopeAsher/LitReview‑Arena.

Authors:Stephanie Okoye
Title: Wazobia Eval: A Benchmark for Nigerian Pidgin Emotion Understanding, Sarcasm Detection, and Cultural Reasoning
Abstract:
Nigerian Pidgin is one of Africa's most widely spoken languages, yet remains severely underrepresented in language model evaluation. Existing benchmarks primarily focus on translation, transcription, or generic sentiment analysis, leaving critical aspects of culturally grounded language understanding unmeasured. We introduce Wazobia Eval, a benchmark for evaluating Nigerian Pidgin emotion understanding, sarcasm detection, and cultural reasoning. The benchmark is built on a manually annotated dataset containing over 550 examples and a 16‑category emotion taxonomy designed to capture culturally specific emotional registers that are not represented in conventional sentiment frameworks. Wazobia Eval provides standardized evaluation protocols and benchmark tasks for assessing model performance on nuanced Nigerian language understanding. We present the benchmark design, annotation methodology, taxonomy development process, and preliminary pilot evaluation results. Our goal is to provide foundational evaluation infrastructure for Nigerian language AI and establish a reproducible benchmark for future research. The dataset is publicly available at https://huggingface.co/WAZOBIALABS.

Authors:Ali Toygar Abak
Title: AIREP: A Protocol for Per-Decision Evidence in AI Runtime Governance
Abstract:
A protocol is presented for recording the governance decisions of automated AI runtimes. When a runtime releases, blocks, defers, redacts, or escalates an individual output, AIREP records that decision as a single signed object that any party can check offline, independent of the runtime that produced it. A record carries the decision as one of a closed set of verbs under a stated policy basis, references its input, output, and evidence by hash rather than by value, and declares both what its evidence covers and what it does not. Records form a SHA‑256 hash chain that binds each record to its position, so that tampering and gaps are detectable by recomputation. Vendor‑, model‑, and domain‑specific content is confined to a single optional namespace, and a mechanical neutrality test keeps the shared format free of it. A reference implementation and a two‑language conformance kit are described. Some implementation issues are considered, and problems such as alignment of the canonical form across implementations, freshness witnesses, and multi‑runtime chains are exposed. The format is offered for adoption by any AI runtime that records governance decisions.

Authors:Xianyun Sun, Chaoyou Fu, Zhengye Zhang, Feiyang Duan, Qingyuan Cao, Yonghui Niu, Sihang Yuan, Ge Zhang, Caifeng Shan
Title: OmniAssistBench: Assistant-style Interaction Benchmark for Omni-LLMs
Abstract:
Recent omni‑modal large language models (Omni‑LLMs) show great potential as real‑time video assistants, which continuously perceive environments and guide users to achieve specific goals. Unlike traditional passive video understanding, interactive assistants should actively combine visual states, user goals, and prior knowledge to provide effective help. Evaluating this is rather challenging, as the model's unpredictable response dynamically changes the user's subsequent actions, which static offline datasets cannot accommodate. To address this bottleneck, we introduce OmniAssistBench. To solve the issue of diverging interaction paths where the same user goal can be achieved through various methods, we provide models with predefined priors derived from the source video, requiring them to guide users along the exact same routes. Since real interaction videos are rare, we construct the dataset by reverse‑engineering existing Internet videos. We deduce logical user goals and segment the videos into multi‑turn clips to simulate continuous interactions. This rigorous pipeline required over 1000 expert person‑hours to build the dataset. Results show that the proprietary Gemini‑3‑Pro reaches 66.4 out of the max point of 100, while the open‑source Qwen3‑Omni‑Instruct achieves 51.2. Although current models generally understand user inputs, they frequently provide incorrect or incomplete answers. Specifically, they struggle with visual prompts (e.g., hand gestures), fail to maintain historical context during multi‑turn interactions, and fail to delay response until the target event. Results indicate substantial room for improvement before models can become reliable assistants.

Authors:Dong Li, Dujun Nie, Xiaotong Zhang, Ruilin Wang, Yuchen Li, Chang Ge, Chao Xiong, Kaichang Di, Andreas Nüchter, Levente Kovács, Qingquan Li, Shirong Ge, Fei-Yue Wang, Long Chen
Title: Mining beyond Earth with Space Robots: Exploration, Sampling, and Extraction
Abstract:
Space resource acquisition and utilization, commonly referred to as Space Mining, represent critical pathways for enabling sustained human exploration and unlocking commercial opportunities in space. These resources mainly include helium‑3, water, mineral resources on the Moon and Mars, and abundant mineral deposits on asteroids. Due to the harsh conditions of space, communication delays, and high launch costs, the development of autonomous robotic systems is critical to achieving efficient, cost‑effective space mining. This paper provides a comprehensive overview of space mining robotics and associated technologies. First, we review the background of space mining, including international policies, commercial entities, and recent advancements. We define a systematic six‑stage architecture for space mining: Exploration is initiated by (1) remote sensing for target identification and (2) precise in situ robotic detection; Sampling progresses from (3) single‑robot small‑scale sampling to (4) multi‑robot large‑scale excavation; and Extraction integrates (5) autonomous resource extraction and (6) final integration into in situ construction or terrestrial transport. Additionally, we review and curate existing resources for space mining research, including real‑world mission data, terrestrial analog datasets, and high‑fidelity simulation environments. Finally, we identify critical open challenges in autonomous space mining and delineate a strategic research roadmap to bridge current technological gaps, fostering the transition toward a sustainable off‑world economy. To track ongoing developments in space mining, we maintain an updated project page: https://github.com/OpenSpace‑Lab/Space‑Mining‑with‑Robotics‑List.

Authors:Yiwen Liu, Yujun Zhu, Kui Jia, Zhao Liao, Yangwei You, Shuaijun Wang
Title: ViTacPhys: Physical Property-Aware Grasping from Human Visual-Tactile Demonstrations
Abstract:
Recent vision‑based action models have demonstrated strong capabilities in complex manipulation, but they rarely leverage explicit object physical properties to adapt their policies. We introduce ViTacPhys, a visual‑tactile framework and data acquisition system that estimates object mass and friction‑coefficient classes, together with continuous stiffness, from human manipulation demonstrations. Trained on data from 60 rigid and deformable objects, ViTacPhys combines temporal visual‑tactile modeling, cross‑attention multimodal fusion, and a semantic prior derived from a vision‑language model. On seen objects, it achieves 97.2% mass classification accuracy, 98.8% friction‑coefficient classification accuracy, and a stiffness mean absolute percentage error (MAPE) of 5.51%. On held‑out objects from known categories, it achieves 87.5% mass accuracy, 97.5% friction‑coefficient accuracy, and a stiffness MAPE of 9.08%. We transfer ViTacPhys from the human domain to the robot domain using limited robot teleoperation data, robot‑style video augmentation, and human demonstrations with matched actions, and deploy it as an online module for adaptive grasping. The resulting physical‑property‑conditioned policy achieves total grasping success rates of 95.0% on in‑distribution objects and 83.4% on out‑of‑distribution objects. For out‑of‑distribution objects successfully grasped by both methods, its force profiles are more consistent with human teleoperation than those produced by ACT. These results demonstrate the feasibility of explicitly estimating and conditioning on object physical properties for real‑world adaptive grasping.

Authors:Nicolás Vera Zúñiga
Title: Prompt-Model Interaction Reaches the Fixed Points: A deterministic, task-free structural readout -- and the factorizations of it that failed
Abstract:
That a prompt's effect is not a property of the prompt is established: prompts optimised for one model degrade on another, and rankings reorder under neutral reformatting. That evidence is about task accuracy, which cannot say whether the interaction is a fact about task machinery or about the conditional distribution itself. We ask on a readout with no task in it: the fixed‑point structure of the short‑window argmax map x_t+1 = argmax_x p(x | x_t‑1, x_t), censused from 96 starts. It is deterministic, so nothing can be helped or hurt, and it exists only at short windows ‑‑ four of six models lose it entirely by window 16 ‑‑ so everything here concerns how a model reads a fragment. Two results. First, the interaction reaches this readout at full magnitude: nine tokens of conditioning move the fixed‑point fraction across most of its range, change a four‑way structural class, and reorder models, while instruction tuning worth 60.5 IFEval points moves the class by zero. Second, nothing we proposed carries it. Prefix length fails: the effect is not monotone. Four phenomenological factors ‑‑ prose‑versus‑markup, a universal direction, bidirectionality, instruct‑resistance ‑‑ were each withdrawn within one run of being proposed, dissolved by widening the sample. And the nearest mechanistic account, attention‑sink dominance of early tokens, predicts the sign of the shift on 2 of 5 models ‑‑ chance ‑‑ while a length‑by‑content cross shows it holds on real text and fails on our probe's uniformly random input, so we are outside its regime, not against it. One fixed nine‑token prefix drives four models toward 0 and two toward 1; the bidirectionality survives in‑distribution starts. On this readout the unit of explanation is the prompt‑model pair. The recurring error it caught in us has a name: a criterion with a shape applied to a quantity with no room to vary.

Authors:Zeyun Zhong, Joya Chen, Manuel Martin, Frederik Diederichs, Juergen Gall, Juergen Beyerer
Title: Rethinking Expressivity and Efficiency in Test-Time Training
Abstract:
Test‑Time Training (TTT) enables long‑context processing via continuous weight updates during inference, but current methods struggle to balance the expressivity of per‑token update dynamics with the hardware efficiency of chunk‑wise approximations. We propose E^2‑TTT (Expressive and Efficient TTT) to bridge this gap. Under the standard approximation of taking gradients at the chunk‑start weights, we derive a closed‑form state transition that exactly reproduces the chunk‑end fast‑weight and momentum states of the per‑token recurrence. This enables fully parallelized chunk‑level training while preserving the temporal structure of the update rule that prior chunk‑wise methods discard. We validate E^2‑TTT by training models up to 1.3B parameters from scratch. It performs on par with previous TTT and hybrid attention baselines in language modeling while outperforming them on in‑context retrieval. Its advantage is most pronounced in length extrapolation: on the standard ``Needle in a Haystack'' passkey test, it retains over 90% accuracy at 8× the training context length. Meanwhile, E^2‑TTT can match the training throughput of efficient chunk‑wise methods, demonstrating that it effectively reconciles expressivity with efficiency. The code is available at https://github.com/zeyun‑zhong/E2‑TTT.

Authors:Marko Haralović, Sounic Akkaraju, Carlo Baretta, Vasil Zapryanov, Alexia Briassouli
Title: When Adaptation Hurts: Connecting Representational Drift to OOD Failures in MedSAM Fine-Tuning
Abstract:
Foundation models for medical image segmentation, like prompt‑based MedSAM, generalize well across domains and modalities, often in zero or few‑shot setups. However, their performance depends on the quality of prompts and the adaptation of the models to custom datasets. This work systematically examines how MedSAM generalizes across diverse medical imaging benchmarks, with six adaptation strategies: full‑model and encoder‑only LoRA, shallow and deep visual prompt tuning (VPT), and decoder‑only and full fine‑tuning. Models are trained on the International Skin Imaging Collaboration Challenge (ISIC 2018) dataset and evaluated under clean and increasingly noisy prompts on IN and Out‑of‑Distribution (OOD) datasets: close‑OOD PH2 (dermoscopy), far‑OOD BUSI (Breast Ultrasound Images Dataset) and CBIS‑DDSM (Curated Breast Imaging Subset of the Digital Database for Screening Mammography). We show that adaptation improves performance on IN and close‑OOD data but often reduces performance on far‑OOD data. Full fine‑tuning provides the best tradeoff, while encoder‑only LoRA is the strongest parameter‑efficient alternative, outperforming standard LoRA and VPT under far‑OOD shifts. Using Centered Kernel Alignment (CKA), we show that far‑OOD degradation is strongly associated with drift in decoder representations, whereas encoder similarity alone does not explain robustness. This suggests encoder‑only LoRA provides stronger robustness than standard LoRA by adapting the encoder to distribution shift in visual features, while preserving the decoder pathway. We further show that random 0‑100 pixel jitter on prompts produces more robust and better performing models. We thus conclude that robust MedSAM adaptation requires the combined consideration of prompt noise exposure, domain shift, and representation preservation. We release our code: https://github.com/ImSounic/medsam‑vpt

Authors:Abigail G. Grassick, Jerome Tze-Hou Hsu, Ethan Lin, Ziang Liu, Max Whitton, Madelyn Hair, Liam Gutierrez, Haozheng Yu, Kristin Branson, Vivek Jayaraman, Michael A. Gil, Andrew M. Hein, Jennifer J. Sun
Title: WildFin: An In-the-Wild Dataset for Fish Behavioral Recognition
Abstract:
Recent advances in field technology have led to a massive influx of in‑the‑wild video data for ecological science. The primary bottleneck in leveraging this data is the high cost of expert annotation. While computer vision offers a potential solution, current models frequently fail when deployed in complex marine environments. To characterize these failures, we introduce WildFin, a novel benchmark for fish behavior recognition collected and annotated by ecologists.WildFin spans two critical real‑world paradigms: stationary cameras monitoring groups of fish and dynamic divers following individual subjects. The dataset represents a massive curation effort, involving 1,350 hours of fieldwork and 600 hours of expert annotation to produce 9 hours of behavioral data with over 2 million frame‑by‑frame labels. We benchmark modern vision foundation models and quantify tradeoffs between static and spatiotemporal architectures, revealing the substantial gap that remains between current model capabilities and the demands of real‑world underwater behavioral analysis. Project website: https://team‑wildfin.github.io/.

Authors:Xuanyu Meng, Jiashuo Sun, Jash Rajesh Parekh, Jiawei Han
Title: EnSI-RAG: Entity-Structure-Indexed Retrieval-Augmented Generation for Long-Document Question Answering
Abstract:
Question answering (QA) over long, connected documents remains challenging because relevant evidence may span multiple entities and their relationships. Existing retrieval‑augmented generation (RAG) methods typically index documents as raw chunks and retrieve them through embedding similarity. Their performance degrades when chunk boundaries separate entities from supporting evidence or when a question requires multi‑hop reasoning across the corpus. We propose EnSI‑RAG (Entity‑Structure‑Indexed Retrieval‑Augmented Generation), a framework that constructs a query‑independent, entity‑centered index. Each record (e, t, k, v) represents an entity e, its type t, a semantic category k in property, relation, aspect, and a value v, while retaining links to the original source passages. At query time, these records serve as retrieval handles, and an LLM synthesizes the retrieved passages into the final answer. This design separates evidence localization from answer synthesis while preserving traceable source evidence. Across Loong and Oolong, EnSI‑RAG achieves an average accuracy of 78.24. Relative to the published baseline scores used as references, this is 6.62 points higher, suggesting its effectiveness across these settings. The code is available at https://github.com/RamonMeng/EnSI‑RAG.

Authors:Matthew Faucher
Title: TRACE-C: Rank-Calibrated Relational Anomaly Detection for Multi-Stream Operational Telemetry
Abstract:
Operational telemetry can be jointly anomalous while every individual stream stays inside its familiar range. TRACE‑C is an auditable strictly‑prior rank‑calibrated detector for aligned multi‑stream telemetry: same‑regime rolling median/MAD residuals feed three window channels ‑‑ a maximum normalized local sum, a Gaussian copula‑form dependence contrast on robust‑z residuals, and a worst standardized AR(1) innovation ‑‑ whose channel ranks are Fisher‑aggregated and ranked against earlier aggregates. We evaluate six Great Britain grid streams with a January‑April 2019 fit, July‑December 2019 development evidence, and a 2020 hold‑out frozen before inspection. TRACE‑C ranks Storm Atiyah first among 2019 test windows, but a disclosed channel ablation attributes that rank to the local channel, not the copula‑form channel: copula‑only ranks Atiyah 59th. The short 9 August frequency event is ranked far lower by the fused detector (143) than by the temporal channel alone (40), and reconstruction baselines rank it first. In 2020 no window is selected, which is consistent with record‑rule saturation rather than an uneventful year; the highest‑ranked frozen window was later interpreted as Storm Ellen. Three interpretive limits carry throughout. The resulting p‑values are selection quantities, not event probabilities. The copula‑form channel is not a literal copula density: the method applies no probability‑integral or normal‑score transform. Empirical rank counts are diagnostics, not coverage or false‑discovery proofs. Every table and figure in this paper is generated from committed machine‑readable reports.

Authors:Victorita Dolean, Jemima Tabeart
Title: Advanced Linear Algebra with Applications - Part I (Numerical linear algebra for PDEs, machine learning, and data assimilation)
Abstract:
These lecture notes form the first part of a master's‑level course on advanced numerical linear algebra. Their aim is not only to present the classical algorithms, but to show why the subject has become considerably more central than it was a generation ago. Numerical linear algebra grew up alongside the numerical solution of partial differential equations, and for a long time that is where its large sparse systems came from. Ranking the nodes of a network, assimilating observations into a weather forecast, and fitting a model to a large noisy data set now lead to problems of the same kind: too large to factorise, structured, and accessible only through matrix‑vector products. Strikingly few ideas are needed for all of them. Each chapter therefore develops a standard topic and then puts it to work outside its original setting. We treat norms, factorisations, conditioning and floating‑point arithmetic; sparse matrices arising from finite differences, from graphs and from machine learning; stationary iterations and the smoothing property; the conjugate gradient and Lanczos methods, with spectral clustering and regularisation by early stopping; Arnoldi and GMRES, with PageRank and large least squares; and finally preconditioning, Schwarz domain decomposition and multigrid. We assume a first course in linear algebra. Every section closes with a summary of what should be retained and every chapter with exercises, several drawn from past examinations. Accompanying Python code reproduces the numerical illustrations.

Authors:Varun Giridhar, Anant Khandelwal, Jeremy A. Collins, Ignat Georgiev, Animesh Garg
Title: Beyond Imitation: Self-Improving Robot Policies via Off-Policy Q-Planning
Abstract:
Behaviour Cloning (BC) has driven remarkable progress in robot manipulation, yet it is fundamentally limited by its inability to self‑improve: a policy that fails cannot learn from that failure without additional human demonstrations. Reinforcement Learning fine‑tuning offers a path to self‑improvement but has proven difficult to scale to the multi‑billion‑parameter models underpinning modern robot policies. We propose Q‑Planning, which equips a large visuomotor BC policy with a small off‑policy Q‑function. Because a Q‑function estimates value rather than imitates actions, it can be trained on the same successful demonstrations as the BC policy and later absorb both successful and failed deployment rollouts, an asymmetry BC does not have. We exploit this asymmetry to enable value‑guided action selection at inference (a single‑step Q‑weighted average over BC draws) and online self‑improvement that fine‑tunes only the Q‑function, leaving the BC weights untouched. On LIBERO and bimanual RoboTwin, ten iterations of self‑improvement lift every benchmark score we tested (LIBERO‑10 93% to 99%, RoboTwin 83.8% to 91.4%) and shorten successful episodes on the near‑ceiling suites (LIBERO‑Object, LIBERO‑Goal). On two contact‑rich bimanual real‑robot tasks, the same loop (BC frozen, no human intervention) improves purely from its own deployment rollouts: stack‑cups 40% to 90% and insert‑wallet 25% to 80% in five iterations, whereas SFT on successful rollouts alone stalls at 55% and 30%. Under an identical online budget Q‑Planning is the only method, among Best‑of‑N, filtered SFT, IBRL, DSRL, and DAWR, that improves stably from failures without training an auxiliary actor.

Authors:Julia Dietlmeier, Benjamin Greenberg, Wenxuan He, Teresa Wilson, Rubing Xing, Jordan Hill, Adrienne Fettig, Madeline Otto, Teyhana Rounsavill, Lina A. J. Reiss, Jingang Yi, Noel E. O'Connor, George W. S. Burwood
Title: Towards Investigating Residual Hearing Loss: Quantification of Fibrosis in a Novel Cochlear OCT Dataset
Abstract:
Objective: Cochlear implants (CIs) are bionic prostheses that restores hearing via electrical stimulation of the auditory nerve. Hybrid CIs, which use electroacoustic stimulation (EAS), combine residual low‑frequency acoustic hearing with CI electrical stimulation. Intracochlear fibrosis, which forms in response to the presence of the implant, may impede residual hearing function and gradually reduce the efficacy of EAS. It is therefore a translational objective to study the formation of cochlear fibrosis in rodents, with the goal of reducing fibrotic burden and improving outcomes for CI patients. Methods: We generate and annotate a novel dataset of optical coherence tomography (OCT) images from chronically implanted guinea pigs as part of an ongoing study focused on implant induced fibrosis. Objectively assessing fibrotic burden in this model, with high resolution and repeatability, presents an obvious use case for computer vision methods. Results: We present the results of several state‑of‑the‑art semantic segmentation models and compare their efficacy for identifying cochlear fibrosis and other relevant annotations, using a new library of manually segmented OCT images. Conclusions: We find that the best performance is achieved by using a modified version of the well‑known UNET architecture (which we term 2D‑OCT‑UNET) that operates on the upscaled OCT input resolution. Significance: For the first time, we have successfully applied computer vision techniques to an OCT dataset of implanted cochleae with fibrosis. Using this deep learning model, the cochlear fibrotic burden calculation can be reliably carried out as we verify in our experimental section. The dataset and the project code are available at: https://github.com/juliadietlmeier/CF‑OCT‑segmentation

Authors:Ruihua Han, Rui Gao, Zhe Liu, Xinyi Wang, Chang Chen, Shuai Wang, Qi Hao, Jia Pan, Hengshuang Zhao
Title: SRL-MPC: Shape-Aware Reinforcement Learned Model Predictive Control
Abstract:
Safe and efficient shape‑aware navigation in heterogeneous crowds and robot fleets remains challenging. Traditional approaches often assume homogeneous robots, sparse workspaces, simplified geometry, offline computation, or handcrafted parameters to make the problem tractable, which limits their deployment in dense crowd scenarios. Toward this end, we propose Shape‑Aware Reinforcement Learned Model Predictive Control (SRL‑MPC), a method for safe, efficient, and adaptive navigation in crowds with heterogeneous shapes without geometry simplification. To encode shape‑aware safety, we formulate high‑order control barrier function (HOCBF) constraints from geometric separation features (GSFs) based on support function transformation. A reinforcement learning (RL) framework then learns a neural policy that reads GSFs and outputs real‑time MPC parameter updates, enabling the MPC solver to adapt to neighboring crowd geometries. The key advantage of SRL‑MPC is that it preserves the safety structure and generalizability of MPC while integrating the adaptability and intelligence of RL. Experiments in randomized crowd scenarios with arbitrary shaped robot fleets demonstrate the effectiveness, scalability, and robustness of SRL‑MPC. The results show that SRL‑MPC substantially outperforms representative baselines in safety and adaptability. Project website: https://hanruihua.github.io/srl_mpc_project/

Authors:Yuyuan Feng, Zhishang Xiang, Chaobin Yang, Qichao Ma, Zerui Chen, Yujing Zhang, Ke Huang, Chuanjie Wu, Zhaoxu Liu, Yili Wang, Xin He, Jiapu Wang, Zijin Hong, Hao Chen, Yuanchen Bei, Kun Wang, Shengyuan Chen, Ningyu Zhang, Enyan Dai, Linhao Luo, Qingyi Pan, Qi Wang, Wenqi Fan, Guangjing Wang, Na Zou, Yangqiu Song, Xin Wang, Zechao Li, Xia Hu, Qing Li, Xiao Huang, Zhihong Zhang, Jinsong Su, Qinggang Zhang, Yi Chang
Title: Graph Engineering in the Era of LLM Agents: From Individual Intelligence to System Intelligence
Abstract:
LLMs have evolved from language generators to autonomous agents capable of complex, long‑horizon tasks. This evolution has produced paradigms including Prompt Engineering to elicit model capabilities, Context Engineering to manage information access, Harness Engineering to organize external tools and resources, and Loop Engineering to support continual reflection and self‑improvement. Yet as tasks grow more complex, individual intelligence faces a fundamental limit: many tasks require heterogeneous expertise, interdependent subtasks, parallel execution, independent verification, and persistent state, exceeding any single agent's organizational capacity. Augmenting one agent's capabilities or context cannot resolve this architectural mismatch; intelligence must instead be distributed across specialized agents and organized at the system level. We call this System Intelligence: an agent system's ability to organize and coordinate multiple intelligent components into a coherent, adaptive whole pursuing a shared objective. Achieving it requires more than adding agents; it demands explicit structures to organize work, coordinate heterogeneous agents, and maintain evolving execution states. We introduce Graph Engineering, an emerging paradigm for next‑generation agent systems. Unlike prior paradigms that mainly optimize individual interactions or agent‑level behavior, Graph Engineering constructs explicit, dynamic, evolving graph structures representing tasks, agents, and system states. These abstractions provide a unified foundation for organizing complex objectives, orchestrating heterogeneous agents, modeling system dynamics, and enabling scalable agent evolution. We systematically review the principles, methodologies, and applications of Graph Engineering for LLM agents. Related papers, open‑source data, and projects are collected at https://github.com/DEEP‑JLU/Awesome‑Graph‑Engineering.

Authors:Jie Xu, Na Zhao
Title: Stream3Dv2: Geometric-Semantic Fusion Enhanced Streaming Zero-Shot 3D Scene Understanding
Abstract:
Recently, open‑vocabulary zero‑shot 3D scene understanding using vision foundation models has emerged as a promising alternative to data‑intensive supervised methods. However, deploying these models in real‑world scenarios is severely hindered by their inability to efficiently handle streaming RGB‑D inputs and their inherent vulnerability to noise 2D segmentation masks. To address these critical limitations, we propose Stream3Dv2, a novel training‑free framework designed for robust streaming 3D perception. Stream3Dv2 processes sequential data through an original nested local‑to‑historical architecture, capturing multi‑view consistency while circumventing the high computational overhead so as to support timely responses. At its core, we introduce a comprehensive geometric‑semantic fusion mechanism that resolves geometric noise and semantic ambiguity by explicitly utilizing semantic guidance and formulating 3D segmentation as solving point‑and‑set merging and partitioning problems. Furthermore, we present an innovative manifold‑distance‑based point cloud refinement strategy. This approach leverages local manifold graphs for point‑to‑manifold optimization that mitigates the boundary delineation failures caused by Euclidean‑distance metrics, and employs geometric bounding boxes to dynamically activate and update historical instances for achieving rapid manifold‑to‑manifold refinement. Extensive experiments on public datasets demonstrate that Stream3Dv2 consistently outperforms existing baselines in foundational open‑vocabulary streaming 3D segmentation and detection. Finally, we show that integrating our framework with an LLM‑based agent enables advanced language‑driven 3D scene understanding, underscoring its potential for open‑world embodied intelligence. Code will be updated at https://github.com/SubmissionsIn/Stream3D.

Authors:Philipp Bogdan
Title: Atom Learning Model (ALM): how a real classroom got tokenised
Abstract:
The Atom Learning Model (ALM) tokenises a school curriculum. 757 pages of GCSE and Further Mathematics material were read by machine into 1,934 atoms, each one thing a learner can do in a single step, ordered by 4,616 machine‑written prerequisite links. Both sides of a lesson are then expressed in that one structure: a question is a set of atoms plus everything beneath them, a child's ability is a score between 0 and 1 on every atom of the same graph, and whether a question suits a child is arithmetic over one index, with no difficulty parameter fitted for either side. Nobody wrote an atom, a link or a question. Reading the 757 pages cost £55, building the whole structure cost between £615 and £1,230, and against it the system composed 6,648 questions for 373 children in two English secondary schools over seven weeks, at 26p per composed question. Four measurements went against expectation. The cost is in the links, not the pages. The composer's own difficulty label has a rank correlation of ‑0.0123 with measured facility, so a language model shown a question cannot say how hard it is. Children stop working when a mark takes seven seconds instead of three. And the deployment never served a question deeper than two prerequisite steps, which is exactly where the central premise becomes testable, leaving it unfalsified rather than confirmed.

Authors:Kai Wang, Zeming Wei, BiaoJie Zeng, Chang Jin, An Wang, Xiaokun Luan, Zhixiao Lin, Jingjing Qu, Xia Hu, Xingcheng Xu
Title: ClawSentry: A Progressive Multi-Tier Security Monitor for Safeguarding Autonomous LLM Agents
Abstract:
As large language model (LLM) agents move from conversation to executing code, reading local files, and orchestrating external tools, a single agent hijacked by a malicious third‑party skill can cause data exfiltration, privilege escalation, or cascading compromise. We argue that agentic risk is progressive: it can enter at four loci of the agent control loop‑‑skill admission, invocation‑time intent, execution‑time effect, and post‑action consequence‑‑while a denied dangerous objective can reappear across surface forms, tools, or turns; existing safeguards are typically local to one lifecycle boundary or one call. Guided by this threat model, we present ClawSentry, an open‑source, framework‑agnostic security supervision gateway for agent runtimes. Before a skill package is ever executed, First‑use Skill Package Review (FSPR) audits it under a deterministic evidence floor, escalating unresolved cases to bounded read‑only agentic review (locus A). At runtime, a three‑tier progressive decision engine‑‑a deterministic L1 layer, a rule‑anchored L2 semantic reviewer, and a read‑only L3 evidence‑seeking agent‑‑spends contextual review only on the residual ambiguity, while a session‑level anti‑bypass mechanism recognizes tool‑switching and rephrased retries (loci B‑‑C); a post‑action path feeds high‑severity evidence non‑retroactively into later review (locus D). An Agent Harness Protocol (AHP) abstraction applies one policy across Codex, Claude Code, Kimi CLI, and Gemini CLI without modifying agent internals. On SkillInject with Codex/GPT‑5.4, contextual ASR falls from 39.55% to 2.61% while contextual TSR moves only from 83.78% to 83.05%. Across five Work Agents on the full SkillsSafety benchmark, ClawSentry confines ASR to 9.09‑‑15.03% from 33.5‑‑49.7% unprotected, and aggregate TSR on clean skills remains 98.7%.

Authors:Ahmad AlMughrabi, Albert Clop, Benjamin Busam, Ricardo Marques, Petia Radeva
Title: When does fusing hand-crafted knowledge with learned representations pay? A cost-normalized benchmark of stacking, substitution, and interference
Abstract:
Fusing prior knowledge with data‑driven learning is attractive where data is scarce, yet no controlled account says when it helps, is redundant, or harms. We benchmark one fixed hand‑crafted knowledge source, a pinned bank of Gabor targets injected only during training at ~2% overhead, against data‑driven alternatives (SimCLR, SimSiam, DINO, ImageNet transfer, augmentation, learned teachers) under one frozen recipe with fixed subsets: 13 datasets, 9 backbones, 150 to 1.28M images, 32‑‑224\,px, 2.5M‑‑86M parameters (\computeCells classification configurations over \computeRuns runs, plus segmentation and detection transplants). Across the training‑time combinations we measure, three outcomes recur (decision‑level fusion differs). Different‑\emphcurrency sources can stack: the prior composes with DeiT augmentation on attention backbones and is worth +26 points to ViT‑B/16 at 224\,px, +6.7 at twice that budget. Same‑currency sources substitute: against effective self‑supervised pretraining, the combination never usefully exceeds the better single source. Fusing at full strength into an already‑informed initialization interferes in proportion to what it carries: ImageNet transfer, ‑15 to ‑17 points, removed by a weaker auxiliary weight. Frozen‑feature diagnostics measured on each source alone separate these outcomes retrospectively but do not predict them: a rule built on them calls one of nine unseen pairs. At a practitioner's own label budget, the frozen‑feature gain predicts the end‑to‑end gain to within 0.17 points across 30 cells and seven datasets; the underlying decomposition, Δ= G + \readout(\mathrmbase), holds in sign on \auditRate% of testable cells and is called an unseen backbone family's feature gain in advance. The project page is https://amughrabi.github.io/MomentAux.

Authors:Balkrishna Giri, Md Toufique Hasan, Jussi Rasku, Muhammad Waseem, Pekka Abrahamsson
Title: Trustworthy RAG: An Evaluation Agent for Detecting Misinformation and Knowledge Poisoning in Generative AI Systems
Abstract:
Retrieval‑Augmented Generation (RAG) grounds Large Language Model (LLM) outputs in external knowledge, but RAG systems usually trust whatever they retrieve, creating a Security‑Reliability Gap: high semantic relevance does not guarantee factual truth. Adversaries exploit this through knowledge poisoning, inserting malicious documents to cause targeted misinformation. We propose an Evaluation Agent, middleware that combines Natural Language Inference (NLI) factual verification, a five‑signal poison detector with relevance‑weighted aggregation, and a Trust Index T = 0.4 F + 0.35 C + 0.25 (1 ‑ P ) with a non‑linear dampener for high‑contamination contexts. On TruthfulQA with Llama 3.3 70B, the agent reaches 91% accuracy and 100% precision, with 100% recall on instruction injection, while in‑place edits, such as entity swaps, remain hard to detect. Across three LLMs the Trust Index stays discriminative, with a Receiver Operating Characteristic Area Under the Curve (ROC‑AUC) of 0.73 to 0.81; generation style matters more than model size, and per‑LLM threshold calibration restores baseline competitive accuracy, whereas a weaker FEVER result shows that cross‑dataset generalization requires domain‑specific calibration. In a software‑engineering use case, a secure‑coding assistant over guidance from the Open Worldwide Application Security Project (OWASP) Top 10 and the Common Weakness Enumeration (CWE), the agent reliably blocks instruction injection of unsafe advice (F1 92%), while contradiction and subtle semantic weakening remain hard. Throughout, the agent measures detection of poisoned context before generation, not whether the LLM adopts the injected misinformation. We release the proposed approach, attack generator, and experimental artifacts at the link: https://github.com/GPT‑Laboratory/TrustworthyRAG.

Authors:Luis Vitor Zerkowski, Luiz Velho
Title: AudioWorldSim: Realistic Binaural Audio Datasets For World Models
Abstract:
This technical report presents AudioWorldSim, an open‑source platform designed to generate realistic binaural audio datasets and advance research in audio‑based machine learning, particularly world models. Built as a custom extension of Meta's SoundSpaces 2.0 platform, AudioWorldSim leverages their comprehensive acoustics framework, but focuses on the automatic rollout of random agent navigations, as well as implements crucial fixes to how continuous sound is composed. AudioWorldSim is made publicly available to the research community at https://github.com/Luizerko/AudioWorldSim to facilitate reproducibility.

Authors:San Kim, JinYeong Bak
Title: Evidence-Consistent Generative Detection under Scenario-Level Distribution Shift
Abstract:
Conventional in‑distribution evaluation can overestimate robustness when training and test data share recurring task‑specific patterns or surface cues. This risk is especially relevant in social‑engineering fraud detection, where attackers can preserve malicious intent while changing the scenario, impersonated entity, or wording. We study this problem as scenario‑level out‑of‑distribution (SL‑OOD) detection for SMS and voice phishing, where entire attack scenarios are held out from training while the label space remains fixed. This setting tests whether models can generalize to unseen attack scenarios using decision‑relevant evidence rather than familiar scenario‑specific cues. Using this SL‑OOD evaluation, we find that high in‑distribution performance does not reliably predict held‑out robustness across feature‑, encoder‑, and decoder‑based baselines. We interpret this gap as scenario memorization: reliance on recurring scenario‑specific lexical or entity cues rather than decision‑relevant evidence. We propose ECoG, an evidence‑consistent generative framework that combines evidence‑span supervision with a rationale‑label consistency objective during training. On the 0.5B decoder, relative to the same backbone trained without consistency regularization, ECoG raises Macro‑F1 on OOD challenging instances by 3.22 points, reduces the share of predictions whose generated rationale supports the opposite label by 4.22 points, and increases token‑level overlap with reference evidence spans by 8.38 points; the reduction in prediction‑rationale inconsistency is consistent across four decoder backbones. These results suggest that compact generative detectors can benefit from evidence supervision and rationale‑label consistency under social‑engineering shift.

Authors:Yutian Jiang, Jiabo Liu, Xixuan Hao, Yuxuan Liang
Title: CoST: Semantic-Aware Urban Understanding via Spatial-Temporal Alignment
Abstract:
Geospatial representation learning from satellite imagery is a fundamental problem for large‑scale urban analysis and real‑world applications. Despite recent advances, current methods struggle with cross‑region generalization and semantic interpretability due to their reliance on region‑specific auxiliary data and the neglect of semantic alignment within multi‑temporal urban imagery. Therefore, we present CoST, a novel \underlineContrastive‑based \underlineSpatial‑\underlineTemporal framework that aligns spatial context with multi‑temporal semantics to extract universal geographic regularities shared across regions. Specifically, CoST explicitly models spatial correlations to capture transferable geographic structures and exploits multi‑year urban change semantics to align learned representations with high‑level geo‑semantics. Extensive experiments demonstrate that CoST consistently achieves superior performance across various downstream tasks and in unseen scenario, yielding an average relative gain of 8.7% over the strongest competing methods across eight city‑indicator settings. The code is available in \hrefhttps://github.com/Arandinglv/CoSTthis repo.

Authors:Minhua Lin, Zhicheng Gao, Yilong Wang, Hanqing Lu, Xiang Zhang, Suhang Wang
Title: Trojaning the Alignment: Stealthy Backdoor Attacks against Graph Foundation Models
Abstract:
Graph Foundation Models (GFMs) on text‑attributed graphs (TAGs) align graph representations with language semantics to support transferable graph learning. Despite these advantages, the backdoor vulnerability of GFMs on TAGs remains insufficiently understood, especially under graph‑language alignment, where graph and text representations are trained to constrain each other in a shared semantic space. Existing backdoor attacks mainly target either the graph side or the text side, treating the two modalities independently. This makes direct adaptation ineffective: graph‑only triggers can be constrained by clean text semantics, while text‑only triggers alter the language view but do not directly shift the graph representation being aligned and scored. TAGs also impose a stealth challenge because triggers are exposed as both node text and local graph structure, making incoherent trigger attributes or anomalous subgraphs easy to inspect or filter. In this paper, we propose STAG, a stealthy trojan attack framework designed for the graph‑language alignment interface of GFMs on TAGs. STAG coordinates a graph‑trigger generator with a text‑side soft prompt so that trigger‑attached graph representations and triggered text representations move toward the same target‑class text region. To address TAG‑specific stealthiness, STAG realizes trigger nodes as readable text through candidate retrieval and regularizes the trigger‑attached subgraph so that its local structure remains close to the original subgraph. Extensive experiments on multiple TAG datasets and representative GFMs demonstrate the effectiveness and stealthiness of STAG. Our code is available at https://github.com/ventr1c/STAG.

Authors:Xinlin Wang, Yujiao Xiang, Yuheng Zhou, Jingqi Wang, Minqing Huang, Jiajie Huang, Dongxu Wei, Tingguang Zhou, Xiyang Wang, Gong Chen, Zhi Xu, Feiyang Tan, Hangning Zhou, Mu Yang
Title: WA-JEPA: Rethinking the Video JEPA Paradigm for World-Action Modeling in Autonomous Driving
Abstract:
Video Joint Embedding Predictive Architecture (V‑JEPA) learns powerful spatiotemporal representations from video through self‑supervised latent feature prediction. However, V‑JEPA is built around random‑mask completion and deterministic regression, making it fundamentally ill‑suited for autonomous driving planning that demands future‑directed prediction tightly coupled with action. To address this, we rethink the V‑JEPA paradigm and present WA‑JEPA, a V‑JEPA‑native world‑action model designed for autonomous driving planning. Instead of random spatiotemporal masking, WA‑JEPA employs hybrid future‑masked pre‑training, where the model infers future latents from observed context. Departing from deterministic regression, we recast future prediction as conditional flow matching over latent futures, which substantially improves the model's ability to generate plausible future latents for downstream planning. Finally, a joint future‑action predictor is proposed to denoise future scene tokens and ego trajectories together in a unified spatiotemporal latent space, allowing action supervision to directly shape planning‑relevant world representations. Pre‑trained on nuPlan videos and fine‑tuned on NAVSIM, WA‑JEPA reaches 91.7 EPDMS on NAVSIM‑v2, surpassing the strongest end‑to‑end and world‑action baselines by 1.6 and 1.3 EPDMS, and, without HUGSIM‑specific fine‑tuning, attains the best HD‑Score of 0.4462 on the closed‑loop HUGSIM benchmark under the same evaluation protocol. These results validate V‑JEPA‑native world‑action modeling as a powerful and scalable paradigm for autonomous driving planning. Code is available at https://github.com/AFARI‑Research/WA‑JEPA.

Authors:Mian Wang
Title: Training, learning and inference: unified dynamics of neural systems
Abstract:
We define an atomic generation fact f=(u,tau,omega,z;rho), recording the origin, realized transformation, concrete occurrence, generated result and relation role. Compiled into a Generation‑Fact Graph (GFG), these facts provide an AI‑native, compilable scientific fact substrate preserving generation histories. We establish a GFG‑based recursive scientific process in which analysis, intervention, replay and validation form facts for later cycles. Using nanoGPT, we establish unified training‑learning dynamics. Training is the evolution of a parameter‑optimizer system with state and memory: each actual training action enters the receiving state and produces a finite‑amplitude nonlinear functional response conditioned by that state and target‑specific update geometry. Learning is the persistent reorganization of distributed functional support by these responses; capability formation, maintenance, decline or recovery becomes observable when target‑specific states are evaluated against their readout boundaries. Three primary coordinates ‑ target‑boundary state, target‑specific update geometry and parameter‑Adam receiving state ‑ yield a second‑order predictor operating before post‑update outputs are read. On held‑out runs, it achieved 91.43% accuracy and 91.49% macro‑averaged recall across four transitions. We further establish inference as a frozen projection of training‑learning dynamics. Component gating and rollback show causal recruitment and non‑additive combination of query‑conditioned support formed during training, deriving organizational conditions realized by Attention. Controlled feedback indicates possible double‑edged reinforcement effects. ResNet/CIFAR‑100 and diffusion/CIFAR‑10 experiments confirm receiving‑state‑conditioned responses, persistent support reorganization and frozen inference projection beyond nanoGPT.

Authors:Wenyang Hong, Yuan Wang, Yanbin Hao, Lanqing Xue, Ke Wang, Xiang Wang, Kuien Liu, Richang Hong
Title: OccluRank: Controllable Occlusion-Aware Layout-to-Image Generation by Adding Just an Ordinal Rank
Abstract:
Layout‑to‑image generation enables explicit spatial control through bounding‑box layouts, yet bounding boxes specify only instance locations and cannot represent their occlusion order. Existing methods may rely on additional geometric conditions, employ complex inference procedures, or aggregate independently constructed instance representations without explicitly modeling their occlusion‑dependent interactions. We propose OccluRank, a simple and controllable occlusion‑aware layout‑to‑image framework that augments each bounding box with only one ordinal rank. OccluRank encodes the user‑specified occlusion order through lightweight rank‑based conditioning and introduces an Order‑aware Instance Interaction (OII) module to jointly update rank‑conditioned instance representations before aggregation. This allows the specified order to guide information exchange among occluding instances without additional geometric inputs or specialized inference‑time optimization. We further construct OccluLayout, a synthetic training dataset whose occlusion order and amodal annotations are derived directly from known scene geometry rather than estimated from partially occluded images using auxiliary prediction models. For comprehensive evaluation, we introduce OccluLayout‑Bench, which uses multiple multimodal large language model evaluators to assess instance presence, spatial layout, attributes, and occlusion order, together with FID for overall image quality. Experiments show that OccluRank more reliably preserves target instances, follows specified layouts, and realizes desired occlusion relationships while maintaining comparable attribute consistency and overall image quality.

Authors:Linhao Zhong, Zongze Du, Linyu Wu, Yu Bo, Hourong Li, Chenchen Jing, Hao Chen, Yuling Xi, Chunhua Shen
Title: ForeDreamer: A Self-Evolving Dual-Agent Memory Architecture for Future Event Prediction
Abstract:
Open‑web future event prediction requires agents to distill reliable signals from noisy, redundant, and incomplete evidence. Existing retrieval/memory mechanisms directly feed retrieved information to agents or rely on simple memory functions such as storing and reusing prior information for prediction, leaving them insufficient for open‑web forecasting. We propose to transform raw web evidence into structured memory before prediction, enabling agents to reason over distilled, question‑specific evidence rather than noisy retrieval results. This paper presents ForeDreamer, a self‑evolving dual‑agent framework for managing memory over open‑web evidence. ForeDreamer separates factual memory, a question‑specific evidence state for the current forecast, from experiential memory, persistent agent experience accumulated across forecasting episodes. It uses a main agent for search and prediction, and a memory‑processing subagent to convert search results into factual memory with dedicated tools. ForeDreamer further evolves experiential memory through two tracks, improving both forecasting decisions and factual‑memory construction. Experiments on Prophet Arena and FutureX demonstrate the effectiveness of ForeDreamer. Project page: https://zhongzero.github.io/ForeDreamer

Authors:Zhu Xu, Jiaqi Tang, Pokai Chen, Yuxin Peng, Yang Liu
Title: Explainable Deepfake Detection with Feature-robust Augmentation and Evidence-grounded Explanation Optimization
Abstract:
Explainable deepfake detection extends binary classification by requiring models to not only predict authenticity but also provide interpretable justifications. This expanded scope is critical in practice, where users like forensic analysts need insight into the rationale behind the detection. Despite advancements, current approaches suffer from two critical deficiencies: (1)vulnerability to image quality degradation: detection accuracy plummets on low‑quality samples, while naive augmentation strategies may induce feature drift and impair performance as diversity expands. (2) factually flawed explanations: explanation models may omit manipulation evidence or hallucinate irrelevant details, undermining interpretability. To address it, we propose a framework with two innovations. For robust deepfake detection, we introduce Feature‑robust Augmentation, which comprises diversified degradation‑aware augmentation strategies, and a supervised contrastive learning pattern paired with a mean‑teacher architecture that stabilizes features against augmentations through consistency constraints. For explanation, we devise an evidence‑grounded preference optimization process that guides model to prioritize genuine manipulation traces by learning from chosen‑rejected explanation pairs, where rejected samples are constructed via evidence omission or irrelevant information injection. The proposed approach wins the first place in ACM Multimedia 2026 Explainable Deepfake Detection Challenge.The code is available at https://github.com/oceanflowlab/EDD.git.

Authors:Rui Liu, Jing Nie, Ying Fu
Title: RDANet: Relative Degradation Aware Network for Infrared Small Target Detection
Abstract:
Infrared small target detection is still challenging in remote sensing imagery, because the targets are extremely small, exhibit weak local contrast, and are often embedded in complex and highly variable backgrounds. In addition to these inherent difficulties, we observe that existing detectors often show unstable performance when the target scale changes or when the scene background varies. This scale‑ and scene‑sensitive degradation indicates that current methods are insufficient in simultaneously preserving target structure during feature downsampling and maintaining discriminative local contrast under background shifts, which finally results in unbalanced detection performance across different conditions. To improve detection robustness, this paper proposes a Relative Degradation Aware Network (RDANet) for infrared small target detection. RDANet consists of two dedicated modules: Multi‑Scale Anti‑Alias Downsampling (MSAD) and Prototype‑Guided Skip Memory (PGSM). MSAD introduces multi‑scale anti‑alias filtering together with pixel‑fold aggregation to reduce aliasing effects during resolution reduction, so that target shape information can be better preserved while irrelevant background responses are suppressed. PGSM further enhances the skip features by retrieving patch‑level prototypes from a shared memory and adaptively integrating them into the current representation, which helps maintain stable local contrast cues under diverse scene backgrounds. Experiments on three public benchmarks show that RDANet achieves the best performance on most evaluation metrics, while scale‑ and background‑stratified evaluations indicate more stable behavior across target sizes and scene complexity. The code is available at https://github.com/BIT‑RuiLiu/RDANet.

Authors:Seungheun Baek, Mogan Gim, Jaewoo Kang
Title: ReCurveflow: A Flow Matching Framework that Learns Curved Reaction Trajectories to Predict Transition State Geometries
Abstract:
Predicting transition states (TS) in chemical reactions is crucial, as they provide insights into reaction mechanisms. Recent work on TS prediction have focused on flow matching supervised on straight linear paths that do not align with actual reaction trajectories. We propose a novel flow matching‑based framework ReCurveflow that learns to predict TS geometries supervised on continuously curved reference paths interpolated from a full NEB‑derived band of molecular geometries. We also introduce off‑path correction, which grants ReCurveflow with the ability to produce corrective velocity fields when engaged off‑path geometry states during inference rollout, leading to better resistance against exposure bias and accuracy in TS prediction. Across three data splits and six evaluation metrics, ReCurveflow achieves the best result on the majority of split‑metric combinations against seven baselines. Qualitative analyses further show that ReCurveflow generates reaction trajectories with energy profiles that closely track the reference NEB path, provides initializations that ease the NEB optimization bottleneck, and exhibits the intended corrective behavior in its learned velocity fields. The ReCurveflow codebase is publicly available at https://github.com/dmis‑lab/ReCurveflow.

Authors:Utsav Poudel, Jagannath Aryal, Subramaniyaswamy Vairavasundaram
Title: Neuro-Geospatial Modelling of EEG Affective States Using Literature-Informed Environmental Context
Abstract:
Environmental exposures such as air pollution and greenness have been associated with affective and cognitive outcomes, but EEG and environmental datasets are rarely jointly georeferenced. We investigate whether literature‑informed environmental priors can serve as an auxiliary geospatial modality for EEG‑based affective‑state classification when individual‑level exposure data are unavailable. We combine 30‑channel EEG from the EAV benchmark (42 participants, aged 20‑30 years) with environmental representations derived from OpenAQ, Sentinel‑2, Sentinel‑5P, and OpenStreetMap data for Astana. A dual‑tower architecture combines EEG‑Conformer representations with a graph‑based environmental encoder. Because the datasets are not co‑registered, environmental context is treated as a literature‑informed prior rather than measured exposure. Subject‑level repeated splits, permutation and label‑shuffling controls, dose‑response reversal, and domain‑shift experiments distinguish architecture‑level gains from prior‑dependent gains. The multimodal model achieves 76.2% accuracy versus 67.4% for EEG alone. Controls disrupting environmental‑label structure retain part of this gain, indicating that the improvement is not attributable solely to environmental information. Replacing the Astana environmental distribution with an independently modeled Singapore distribution reduces accuracy to 72.8%. These findings demonstrate technical feasibility but do not establish an observed or causal exposure‑affect association. The study provides a framework for future jointly collected mobile EEG‑environment studies. Implementation: https://github.com/r11up/geo‑cog

Authors:Chenglong Liu, Xin Zhang, Yimeng Zhu, Liyang He, Yixiao Ma, Yu Su, Zhenya Huang, Qi Liu
Title: CubicSplat: Differentiable Vector Graphics via Error-Bounded Forward Relaxation
Abstract:
Vector graphics are prized for their resolution independence, compact storage, and direct editability, making differentiable optimization of their parametric primitives an attractive goal. Yet classical rasterization is discontinuous with respect to geometry, and existing remedies that smooth the forward pass demand increasingly elaborate heuristics as scene complexity grows. We trace this fragility to a gradient seesaw: design choices that improve forward geometric exactness can systematically degrade the induced gradient signal, and vice versa. To navigate this tension we introduce CubicSplat, a differentiable vector rasterizer that replaces Bézier closest‑point solvers with uniform polyline surrogates whose geometric error is bounded at O(S^‑2). The resulting static computation graph yields well‑conditioned gradients by construction, while a compositing‑derived visibility mechanism prunes degenerate primitives without auxiliary regularization. On DIV2K and Kodak benchmarks CubicSplat achieves state‑of‑the‑art reconstruction quality with over 2 dB PSNR gain in the closed‑fill setting, while training up to 4x faster than prior methods. The code is available at https://github.com/CubicSplat/repo

Authors:Zixi Zhu, Jiayuan Su, Jian Zhang, Yu Lin, Hongwei Wang
Title: CAS: Conformalized Agentic Search via Adaptive Retrieval and Policy Weighting
Abstract:
Search Agents face a severe reliability crisis during reinforcement learning (RL) fine‑tuning. Heuristic Top‑K retrieval often causes critical evidence loss or noise inclusion, while over‑confidence induced by progressive RL leads to hallucinated answers and redundant searches. To build highly reliable agents, we introduce Conformal Prediction (CP) and propose Conformalized Agentic Search (CAS). This framework establishes reliability guarantees on both the retrieval and training sides: on the retrieval side, an Adaptive Prediction Set (APS), a specific CP realization, translates statistical coverage into dynamic document truncation to construct prediction sets that are adaptive in size; on the training side, Adaptive Conformal Inference (ACI), a dynamic CP algorithm, dynamically constructs prediction sets with controllable coverage to quantify answer confidence, which is then used to penalize low‑confidence trajectories within the Group Relative Policy Optimization (GRPO) objective, ensuring the model learns only from reliable ones. Experiments across single‑hop and multi‑hop QA datasets demonstrate that our framework significantly improves reasoning accuracy while drastically reducing redundant tool invocations, establishing a highly reliable and efficient agent paradigm. Our code is available at https://github.com/S1llyBird/CAS.

Authors:Jiakun Li, Li Fang, Hao Zhu, Fei Hu, Long Ye, Yuan Zhang, Jinyao Yan
Title: DiGS-Avatar: Single-Image Animatable 3D Human Reconstruction via UV-Space Diffusion
Abstract:
Single‑image 3D human reconstruction often suffers from over‑smoothed textures and geometric inconsistencies. While diffusion models improve generative quality, their reliance on multi‑view synthesis prior to 3D reconstruction is computationally expensive and prone to view inconsistency. We propose DiGS‑Avatar, which reformulates this task as an efficient, diffusion‑based UV‑latent completion task, ensuring 3D consistency by design. To capture accurate spatial structure, we introduce a teacher‑student framework where a multi‑view teacher provides geometrically aligned pseudo‑ground‑truth latents to supervise a single‑view diffusion student. Treating this inferred latent as a robust structural skeleton, our method injects high‑level semantic features to accurately recover fine textural details without disrupting spatial integrity. The refined representation is then decoded into 3D Gaussian primitives. Extensive experiments demonstrate that DiGS‑Avatar achieves state‑of‑the‑art or highly competitive visual fidelity and zero‑shot generalization, while reconstructing a fully animatable 3D avatar in just 0.71 seconds. Code is available at https://github.com/KLMAV‑CUC/DiGS‑Avatar.

Authors:Rujin Liang, Zhongpu Chen, Yuhao Lei, Xin Miao
Title: Vis-Poison: Poisoning Visual Knowledge in Multimodal Retrieval-Augmented Generation
Abstract:
While multimodal retrieval‑augmented generation (RAG) systems increasingly rely on images as external knowledge sources, the introduction of poisoned visual evidence can severely compromise multimodal large language model (MLLM) generation. Unlike prior attacks that rely on altering textual metadata, we introduce Vis‑Poison, a novel visual knowledge poisoning attack where the poisoned image itself is the attacker‑controlled payload, without manipulating captions, summaries, metadata, or other associated text. Specifically, this attack is instantiated through an automated multi‑agent method that constructs visually plausible poisoned images. To assess its impact, we evaluate Vis‑Poison across two representative multimodal RAG pipelines, four embedding models, and six generation models. Empirically, Vis‑Poison achieves an end‑to‑end attack success rate of 40.16% to 65.40% against 30k‑entry multimodal knowledge bases in \emphblack‑box settings. Moreover, Vis‑Poison remains effective against various MLLMs that can answer correctly from parametric knowledge alone, with an average success rate above 60%. Code and data are available at https://github.com/SWUFE‑DB‑Group/Vis‑Poison.

Authors:Aji Mao, Zhenming Peng, Bailin Mu, Tian Pu
Title: SPARK-SAM: Learning How to Prompt and Respond for Infrared Small Target Segmentation
Abstract:
Promptable segmentation models provide a reusable interface, but direct transfer to automatic infrared small‑target segmentation (IRSTD) exposes a mismatch between spatial prompts and target‑domain mask responses. In a diagnostic using target‑covering loose‑box prompts deterministically derived from test reference masks, the best official SAM2.1 results are only 4.69%, 1.64%, and 2.28% IoU on NUAA‑SIRST, NUDT‑SIRST, and IRSTD‑1K. We introduce SPARK‑SAM (Self‑Prompt Adaptation with Response Knowledge for SAM), which learns target‑domain response knowledge and conditions the decoder through an image‑conditioned joint self‑prompt state. Training combines benchmark‑mask supervision with reliability‑aware response guidance. SPARK‑SAM achieves 75.78%, 86.49%, and 68.34% IoU with 0.726M additional parameters, ranking first on two benchmarks among 14 retrained SAM variants and adaptations evaluated as automatic image‑to‑mask methods. The staged IRSTD‑1K diagnostic shows that response adaptation reaches most of the final IoU before the predicted points acquire reliable target grounding. Prompt supervision aligns the predicted prompt candidates with target locations, and frozen‑weight interventions measure output sensitivity to the joint self‑prompt state. Matched ablations show consistent accuracy gains from response guidance and high‑resolution prompt refinement across all three datasets. Code is available at https://github.com/Sakauma/SPARK‑SAM.

Authors:Jiayi Gao, Changcheng Hua, Jiaqi Tang, Yuxin Peng, Yang Liu
Title: Identity-Preserving Text-to-Video Generation via Agentic Enhancement and Semantic Repair
Abstract:
Identity‑preserving video generation aims to synthesize videos that follow natural‑language instructions while maintaining the visual identity of a given subject. Recent commercial video generation models have achieved strong visual quality and motion realism, but they still suffer from identity drift, incomplete instruction following, and missing visual details under complex prompts. Since these models are usually closed‑source black boxes, directly improving them through parameter optimization is often infeasible. We therefore propose Agentic Enhancement and Semantic Repair (AESR), a lightweight enhancement framework for identity‑preserving video generation. To improve prompt construction before generation and mitigate the above failures, AESR introduces a global agentic prompt enhancement module. This module learns model‑specific prompting formats from official documentation, acquires human‑centered video generation priors from human‑interaction data, and accumulates test‑domain identity‑preserving generation experience into a reusable playbook through an agentic loop. To further repair errors in videos generated with enhanced prompts, AESR introduces a sample‑level visual semantic repair module, which uses a VLM to locate erroneous video segments and design repair instructions, edits selected frames into explicit visual references, and guides a video editing model to fix local semantic or identity‑related errors. We also adopt a lightweight Mixture‑of‑Experts selection strategy to choose reliable outputs from different generation and refinement paths. Under the official evaluation protocol of the ACM MM 2026 Identity‑Preserving Video Generation Challenge, our system MIPL\_Video ranked first in Track 1, demonstrating the effectiveness of AESR for practical identity‑preserving video generation. The code is available at https://github.com/oceanflowlab/AESR.

Authors:Qi Song, Ziyuan Luo, Haoliang Han, Renjie Wan
Title: Generating Multi-view Adversarial Examples for Visual Geometry Grounded Transformer
Abstract:
The Visual Geometry Grounded Transformer (VGGT) enables unified feed‑forward 3D reconstruction from multi‑view images. However, deploying such a high‑performance model may expose critical security vulnerabilities. Traditional adversarial perturbations require costly per‑scene optimization, while Universal Adversarial Perturbations (UAPs) rely on a single static pattern and fail to effectively attack VGGT. To address these limitations, we propose MVAP‑G, a multi‑view adversarial perturbation generator that produces imperceptible consistent perturbations across multiple views in a single feed‑forward pass. To ensure perturbation consistency across diverse scenes, we design a cross‑view adversarial alignment mechanism to process multi‑view images. Experiments demonstrate that MVAP‑G significantly degrades VGGT performance without iterative optimization during inference. This work pioneers multi‑view adversarial attacks on 3D foundation models, uncovering severe vulnerabilities and underscoring the urgent need for robust 3D vision systems. The code is available at https://github.com/qsong2001/mvap‑g.

Authors:Junqi Wu, Kaihua Tang, Xuanwen Chen, Hongzhi Li, Jianqiang Huang, Xian-Sheng Hua
Title: AffordAny: Open-World 3D Affordance Grounding from Monocular RGB Images via Vision-Language-Guided Geometric Reasoning
Abstract:
Open‑world 3D affordance grounding requires localizing functional object parts in 3D given free‑form language queries. Existing methods typically assume pre‑built object‑centric 3D geometry and closed affordance ontologies, limiting deployment from raw RGB observations. We present AffordAny, an end‑to‑end framework that uses one monocular RGB image to construct large‑scale text‑conditioned 3D part supervision, ground affordances with a frozen vision‑language model (VLM) guided decoder, and improve open‑world generalization through pseudo‑label self‑training. Our automated pipeline produces a benchmark of 5,334 objects and 10,633 part‑level samples spanning 473 categories, an order‑of‑magnitude increase in categorical diversity over prior work. The decoder progressively fuses frozen Cosmos‑2B features with 3D geometry through spatial projection, instruction‑conditioned semantic compression, and bidirectional geometry‑semantics interaction. Minimal‑perturbation pseudo‑label self‑training further adds new objects without human annotation. Under a systematic generalization protocol evaluating unseen objects, unseen categories, and unseen instruction paraphrases, our approach achieves 0.428 IoU on unseen objects and 0.315 IoU on unseen categories after self‑training, with unseen‑category mIoU improving by 6.3% relative (p<0.01) and an instruction sensitivity gap of only 0.105, demonstrating effectiveness and robustness of our method.

Authors:Haorui Xu, Yuzhou Zhu, Liyuan Gao
Title: DirEAG: Dirichlet Evidence Aggregation for Calibrating Verbalized Confidence in Mathematical Reasoning
Abstract:
Reliable confidence estimation is essential for using large language models in mathematical reasoning, but black‑box verbalized confidence is difficult to calibrate. When the same problem is queried under multiple confidence‑steering prompts, the resulting answer‑confidence observations contain useful uncertainty information, yet their scales may shift across steering levels, models, and datasets. Existing black‑box uncertainty methods often rely on answer agreement, sample consistency, or entropy, which describe output variation but do not model the numerical meaning of self‑reported confidence. Conversely, direct averaging or heuristic aggregation of elicited confidence cannot learn prompt‑ and task‑dependent bias. We propose DirEAG, a Dirichlet Evidence Aggregation method that converts each elicited answer‑confidence observation into calibrated soft evidence over generated candidate answers and an additional null state, allowing the model to represent cases where none of the candidates is correct. Experiments on GSM8K, SVAMP, and GSM‑Hard with Qwen, Mistral, and Gemma models show that, compared with direct confidence averaging and heuristic confidence‑steering aggregation, DirEAG often achieves better calibration while maintaining competitive answer selection. Ablations further reveal that evidence aggregation and final binary calibration address distinct parts of the calibration problem.

Authors:Xiangfei Sheng, Weidong Zou, Tianjiao Gu, Zhichao Yang, Pengfei Chen, Leida Li
Title: AGIDefect-4K: A Richly Annotated Dataset for AI-Generated Image Defect Detection, Localization and Explanation
Abstract:
Generative AI can now produce highly realistic images, yet current models still exhibit subtle but critical defects that undermine their reliability. While existing AI‑generated image (AGI) evaluation benchmarks have made notable progress, comprehensive AGI defect diagnosis remains underexplored. To bridge this gap, we introduce AGIDefect‑4K, a richly annotated dataset of 4,000 images from 15 state‑of‑the‑art generative models spanning both open‑source and closed‑source systems. AGIDefect‑4K features hierarchical defect annotations: (1) detection labels identifying whether defects exist, (2) pixel‑level segmentation masks localizing defective regions, and (3) detailed textual explanations characterizing defect types and their perceptual impact. Each image is further annotated with an overall quality score. Building on this, we present AGIDA (AGI Defect Assistant), a baseline framework leveraging Multimodal Large Language Models (MLLMs) for joint defect detection, localization, explanation, and quality prediction. Comprehensive benchmarking on AGIDefect‑4K reveals that AGI defect understanding remains challenging, underscoring the value of this dataset. The dataset is publicly available at https://github.com/sxfly99/AGIDefect‑4K.

Authors:Chunyu Zou, Peng Dai, Yi-Hua Huang, Ze Yuan, Jingwei Huang, Yeming Yao, Xiaojuan Qi
Title: ArtiMo: Agent-Driven Articulated Mesh Animation
Abstract:
Animating articulated 3D meshes via text requires satisfying strict kinematic constraints, modeling causal interactions between parts, and achieving instruction fidelity. Due to the absence of task‑specific training data and explicit articulation supervision, existing data‑driven mesh animation methods are largely inapplicable to this setting. To address this, we propose ArtiMo, a novel agent‑driven framework for text‑guided articulated mesh animation. Operating in a zero‑shot manner, ArtiMo develops an agentic pipeline powered by Large Language and Vision‑Language Models (LLMs/VLMs) to orchestrate motion generation. By synergizing the explicit kinematic constraints of URDF with the agent's reasoning and planning capabilities, it effectively produces causally coherent part motions and interactions without requiring model fine‑tuning. To ensure motion correctness, the agent additionally utilizes a visual self‑improvement mechanism: generated animations are rendered into compact keyframes and motion cues, enabling the VLM to iteratively diagnose and correct errors. Furthermore, we contribute a new benchmark dataset spanning 21 articulated object categories, featuring high‑quality motion annotations enriched with causal relationships. Extensive experiments demonstrate that ArtiMo significantly outperforms baselines, particularly on complex, causally driven motions. The project page is available at https://zou‑2004.github.io/ArtiMo/.

Authors:Chuanjin Fan, Wenjie Chang, Bohao Liao, Yujia Chen, Wenfei Yang, Tianzhu Zhang
Title: TopoSurfel: Closing the Loop between Gaussian Surfels and Meshes for Surface Reconstruction
Abstract:
3D Gaussian Splatting has achieved remarkable success in novel view synthesis. However, extracting high‑fidelity surfaces directly from 3DGS remains challenging due to its discrete and unstructured nature. Existing 3DGS‑based reconstruction methods typically rely on multi‑view geometric consistency or local constraints. Without an explicit structured geometric prior during optimization, these methods often struggle to resolve structural ambiguities, leading to artifacts and floaters, particularly in textureless or occluded regions. To address this limitation, we propose TopoSurfel, a novel framework that closes the loop between Gaussian surfels and continuous meshes. Unlike recent methods that incorporate mesh extraction into the differentiable pipeline by introducing auxiliary neural networks or extra per‑Gaussian parameters, we dynamically extract a continuous proxy mesh via a non‑trainable differentiable iso‑surfacing process. Leveraging this differentiable connection, we introduce a mesh‑guided surfel evolution strategy, including normal alignment and geometry‑aware density control, to effectively suppress floaters and fill surface holes. Furthermore, to address the initialization challenges in large‑scale environments, we propose a spatially aware hybrid re‑initialization strategy that ensures robust reconstruction across complex scenes. Extensive experiments demonstrate that TopoSurfel achieves competitive geometric reconstruction accuracy while maintaining high‑quality mesh‑based novel view synthesis. The code for our method is available at https://github.com/Fan‑Treasure/TopoSurfel.

Authors:Sarthak Singh
Title: DreamBench-SWE: A Multi-Session Memory-Hygiene Benchmark for Software Agents
Abstract:
DreamBench‑SWE is a multi‑session benchmark for software‑agent memory hygiene in which later software tasks depend on non‑inferable evidence from earlier sessions and are scored by executable hidden oracles. We report the original scaled v2 fold and a separately preregistered v2.1 successor audit designed after that study but frozen before successor outcome inspection. The successor run completed 360/360 work units and 720/720 S3 cells across four conditions. In the original fold, the primary DF‑hybrid‑‑B5 contrast was null (95/180 versus 89/180; clustered p=.518, Holm p=1), not evidence of equivalence, and C9/C10 retained B0‑headroom limitations. In the successor, no external memory achieved 21/180 passes (rate 0.1167), deterministic verbatim event memory 82/180 (rate 0.4556), the typed‑plus‑raw reference probe 83/180 (rate 0.4611), and one pinned hosted Mem0 literal‑storage configuration 97/180 (rate 0.5389). The registered six‑slot Family A retained unavailable slots at p=1; all three available comparisons against no memory rejected after Holm correction. Both preregistered mechanism contrasts were unavailable after pre‑evaluation conformance rejection. The secondary literal‑storage‑versus‑verbatim comparison was nonconfirmatory and sensitivity‑dependent, while the comparison with the reference probe did not reject. The audit therefore supports DreamBench‑SWE as a discriminating executable profile benchmark and characterizes one exact hosted‑memory configuration, but it does not establish an external‑system mechanism, superiority among memory‑bearing conditions, equivalence, or broad product generality. The original v2.0.5 findings and artifacts remain unchanged.

Authors:Pushuo Wang
Title: Shortcut Learning in a Public Grape Disease Dataset: Annotation Granularity as a Modulator, Not a Cause
Abstract:
Public datasets for agricultural disease detection are usually judged fit for use from reported metrics, which say nothing about whether the annotation scheme is internally consistent. On one public grape disease dataset (3288 images, 11995 boxes, 6 classes), varying model capacity, input resolution and detection paradigm yields a test‑set mAP50 range comparable to seed‑to‑seed noise, with the bottleneck at small objects across all five architectures. The finding lies on the data side: one class is annotated at whole‑leaf level (median box area 43.16% of the image) while the other five are annotated at lesion level. On 5156 cross‑species images containing no grape, 65.7% of the false‑positive boxes fall into that one class, an over‑representation of 13.41x relative to its share of the training annotations. Counterfactual retraining establishes a causal effect of granularity on the magnitude of the shortcut: shrinking only that class's boxes cuts its cross‑species false positives by 66%, and a placebo control confirms the effect is specific to the manipulated class. A manipulation in the opposite direction, with criteria registered in advance, returns a negative result: coarsening the finest class to whole‑leaf level (0.57% to 40.37%), matched in box count and share of annotations and with higher in‑distribution AP, still leaves its cross‑species false positives at zero boxes, while the unmanipulated original class holds 50.0% of them. Annotation granularity is therefore a modulator of this shortcut, not its cause: it can amplify or attenuate a sink that already exists, but cannot create one, and what fixes the destination remains open. We also give a granularity screening statistic requiring neither images nor training, and show airborne lesion‑level detection to be optically out of reach. The failure mode is invisible to in‑distribution evaluation.

Authors:Guangyu Wang, Zhidan Liu
Title: RiskTraf: Risk-Extrapolated Residual Learning for Multi-Variate Traffic Flow Prediction
Abstract:
Traffic sensors commonly record flow, speed, and occupancy, but standard traffic flow forecasting benchmarks and models rarely exploit all three raw measurements reliably. Although speed and occupancy provide sensor‑native traffic‑state information beyond flow alone, existing releases often omit these variables, replace them with proxies, or contain logically inconsistent records. Moreover, direct empirical risk minimization over three‑variable inputs may exploit regime‑dependent shortcuts, as the relationships among flow, speed, and occupancy vary substantially between free‑flow and congested states. We introduce PEMSB‑3V, a public benchmark suite that preserves raw flow, speed, and occupancy measurements from PeMS detectors for flow prediction. We also propose RiskTraf, a model‑agnostic risk‑extrapolated residual plug‑in. For each trained spatio‑temporal backbone, RiskTraf freezes the selected checkpoint and learns a lightweight zero‑start residual head from historical speed and occupancy. The residual head constructs ordered traffic‑risk environments and optimizes horizon‑wise flow corrections with a risk extrapolation objective, thereby mitigating regime‑specific shortcut correlations without modifying the backbone. Extensive experiments demonstrate that RiskTraf consistently improves diverse forecasting backbones and outperforms debiasing and distribution‑shift adaptation methods. Our code and benchmark are available at https://github.com/Guangyu4/RiskTraf.

Authors:Christopher Kevin, Narendran Raghavan, Jean-Francois Puget, Roshni Malani, Meghana Puvvadi, Moshe Abramovitch, Mohit Gupta, Rama Akkiraju, Subodh Prabhu, Yogesh Dangi, Wei Luo, Seong Hee Lee
Title: Evaluating Skills, Not Just Agents: Agentic Continuous Evaluation of Skills
Abstract:
Enterprise agent programs are moving from prototypes into production, where reusable skills, tools, and workflow packages must be reviewed with evidence rather than prose. Current gates often scan these artifacts for structure, style, and security, but they do not answer the deployment question: does the capability package help a live agent complete enterprise tasks under the same model, sandbox, and grading policy? We present ACES (Agentic Continuous Evaluation of Skills), a repository‑native framework for evaluating skills and product capability packages as executable agent artifacts. ACES runs paired live trials with and without a target skill, normalizes trajectories into the Agent Trajectory Interchange Format (ATIF), grades six default runtime metrics, and reports Skill Lift: the target skill's added value for a fixed task, harness, workspace, and scorer. The same protocol supports product‑owned task suites that compare baseline, skill, bundle, team‑skill, and plugin targets. On 145 real skills from internal enterprise repositories and public catalogs, scan‑only gates surface useful authoring issues but measure complementary facets (structural versus LLM‑judge Spearman ρ= 0.14). Across 947 scored paired cases from 58 of 64 production skills and four primary harnesses, mean composite Skill Lift is 0.2134 (95% paired‑case CI [0.1967, 0.2301]); mean outcome‑only lift, the average of accuracy and goal accuracy, is 0.1799. Composite lift is positive in 72.8% of paired cases. The largest process‑metric gains appear in skill execution, behavior check, and skill efficiency‑‑‑signals about discovery, routing, workflow following, and tool use that document scans cannot observe. An open‑source implementation of the methodology is available in NVIDIA SkillEvaluator.

Authors:Merve Gülle, Junno Yun, Yaşar Utku Alçalar, Mehmet Akçakaya
Title: Consistency Models for Fast MRI Reconstruction Using Regularization by Denoising
Abstract:
Diffusion models (DMs) have emerged as powerful generative priors for MRI reconstruction with promising results. Yet DM‑based methods require extensive iterative refinement, limiting their practical deployment. Consistency models (CMs) provide a compelling alternative, aiming to map out the diffusion trajectory in a single pass, enabling faster generation. In this work, we propose CM‑RED, a novel MRI reconstruction method that integrates a pretrained CM into the regularization by denoising (RED) scheme. Our method builds on accelerated proximal gradient RED (RED‑APG), and further incorporates controlled noise injection during the update steps to enhance generative diversity and accelerate convergence. Extensive experiments on the fastMRI knee and brain datasets demonstrate that CM‑RED achieves high‑quality reconstructions across multiple anatomies, contrast weights, acceleration factors, and undersampling patterns, using only 4 network function evaluations (NFEs). The proposed method consistently outperforms existing DM‑ and CM‑based approaches in both quantitative metrics and visual fidelity, and exhibits strong robustness to hyperparameter variations, highlighting CM‑RED as an efficient and effective generative framework for accelerated MRI reconstruction. The source code and pretrained models are publicly available at https://github.com/MerveGulle/CM‑RED.

Authors:Obed Korshie Dzikunu, Mohammad Mahdi Abootorabi, Mohamed Harmanani, Paul F. R. Wilson, Emma Willis, Ferdinand Luger, Adam Kinnaird, Brian Wodlinger, Parvin Mousavi, Purang Abolmaesumi
Title: Learning Prostate Anatomy at Test Time for Cancer Detection in Micro-Ultrasound
Abstract:
Domain shift across clinical centers using different imaging hardware or acquisition protocols remains a fundamental barrier to deploying deep learning models for prostate cancer (PCa) detection. Existing test‑time adaptation (TTA) methods address distribution shift through entropy minimization or augmentation‑based self‑supervision, correcting for statistical differences in image appearance but ignoring the anatomical structure of the target domain. We propose ANT, a segmentation‑guided TTA framework that adapts a pretrained cancer detection encoder to the target domain by solving an auxiliary prostate segmentation task at test time, supervised by pseudo‑masks from a frozen pretrained segmentation network. By aligning encoder representations to prostate anatomy in the target domain, ANT corrects domain‑specific feature drift while preserving cancer‑discriminative structure. The model was trained on 693 patients imaged with an earlier‑generation micro‑ultrasound scanner in a multi‑center clinical trial, and evaluated on 118 patients acquired with a newer‑generation system across two centers in another clinical trial. Under a leave‑one‑center‑out protocol with identical evaluation conditions across all methods, ANT improves mean AUC by 2.9% and 3.6% at the biopsy‑core and patient levels, respectively, over no adaptation, outperforming TTA baselines. Code is available at: https://github.com/ObedDzik/ant.git.

Authors:Jiajun Wu, Zirui Wang, Jiayu Zhou, Qiang Ye, Steve Drew
Title: FL-MAESTRO: Multi-Agent LLM Orchestration for Resource-Constrained Federated Learning
Abstract:
In Federated Learning (FL), the communication topology is a runtime variable rather than a fixed design choice, since links and edge devices drop in and out during training. Each round, the server must commit three coupled decisions, namely the communication topology, per‑client resource allocation, and the aggregation rule for combining local updates. Recent agentic systems have begun bringing large language models (LLM) into FL, but the existing line of work either operates at setup time or handles a single runtime dimension such as client selection. We propose FL‑MAESTRO, a multi‑agent orchestrator that makes the joint runtime FL decision directly through three specialist LLM agents, one per decision dimension. A coordinator combines their analyses into a single decision, and a non‑LLM feasibility check confirms it before the round executes. Because the orchestrator consumes the server's predicted‑failure list, it withholds clients whose updates would never be aggregated, which removes the dominant source of wasted round energy in classical FL on volatile edge networks. Because client state is read as natural‑text profiles, the same orchestrator extends to heterogeneous device classes without per‑class energy models. On a non‑IID CIFAR‑10 benchmark, FL‑MAESTRO matches the accuracy of the strongest energy‑aware baseline while cutting wasted round energy from over a third to near zero. Code is available at https://github.com/denoslab/FL‑MAESTRO.

Authors:Davide Lamagna, Albert Cabellos, Alberto Rodriguez-Natal, Gábor Rétvári, Berta Serracanta
Title: Towards Traffic Modelling of Multi-Agent Systems: The Role of Coordination Topology
Abstract:
Multi‑agent LLM systems are an emerging networked workload whose rapid deployment raises questions about the traffic patterns they generate. Compared to conventional applications, these systems generate requests internally: a single user task can induce a structured sequence of model calls whose timing is governed by coordination logic rather than by user arrival rate. It is not clear whether classical traffic models, designed for human‑driven workloads, apply to this setting. We present an empirical characterisation of LLM‑call interarrival time distributions across sequential, star, and full‑mesh agentic coordination topologies, using a multi‑layer measurement framework over 500 repeated runs per topology. We find that topology fundamentally shapes the arrival process of requests to the LLM backend: fan‑out coordination introduces a structural bimodality absent in sequential execution, and the reasoningphase component is best described by a log‑normal distribution, with the Poisson exponential null model decisively rejected across all topologies. These differences propagate to inference and network level metrics. The framework and analysis pipeline are released openly at https://github.com/dlamagna/agentraffic.

Authors:Anirudh Sundar, Min Chen, Divya Tadimeti, Gemma Zhang, Xinyi Alice Li, Nigel Boachie Kumankumah, Pavan Uttej Ravva, Sadid Hasan, Somya Chatterjee, Pruthvi Prakash Navada, Xiao Wang, Yue Kang, Sulaiman Vesal, Larry Heck
Title: Metag: A dataset to build agentic meta-reviewing capabilities
Abstract:
AI tools increasingly support tasks across the scientific research cycle, from experiment design and manuscript preparation to peer review. At the same time, the continuing growth in conference submissions has increased the burden on meta‑reviewers, who must synthesize reviewer feedback, author rebuttals, and manuscript revisions. To address this concern, this paper introduces Metag, a dataset to accelerate the development of meta‑reviewing agents, specifically to identify changes made to scientific articles during the review‑rebuttal process. Each instance contains a reviewer concern, the author's proposed resolution, and the manuscript diffs implementing the stated change. Metag is collected by obtaining manuscript versions from before the review deadline and after acceptance, computing differences between the two documents, and asking human annotators to align these differences with action items from OpenReview discussions. The resulting dataset consists of 349 high‑quality action items tied to paper differences and will enable building methods to empower meta reviewers to quickly identify whether authors have addressed reviewer statements and where in the paper those changes have been made, resulting in additional transparency and traceability throughout peer review. The dataset is publicly available at https://github.com/microsoft/Metag‑dataset.

Authors:Wenti Yin, Xiaotian Han, Junyuan Shang, Yuchen Ding, Shuohuan Wang, Dianhai Yu, Changxin Gao, Nong Sang
Title: Aggregating Visual Information with Optimal Transport for VideoLM Token Compression
Abstract:
Video language models process videos as dense visual‑token sequences with substantial representational redundancy. Compressing these sequences is therefore essential for reducing the visual‑token burden on language‑model decoding. The central challenge is to preserve visual information dispersed across frames under such compression. To this end, we introduce Aggregating Visual Information with Optimal Transport (AVIOT), which casts video token compression as transporting a dense empirical measure of frame observations onto a compact target measure. The resulting source‑to‑target coupling induces a distribution over source observations for each target support, directly specifying how the compressed video representation is constructed. We further adapt this construction along task and spatial axes. Question conditioning modulates the transport cost between source frames and target supports, while influencing how many supports are allocated to each temporal segment, thereby directing representation capacity toward question‑relevant content. At multiple spatial granularities, AVIOT computes region‑specific temporal transport plans and adaptively fuses the representations they yield, allowing different regions within the same compact representation to draw from different moments. Evaluations across varying compression ratios show that AVIOT matches or outperforms the uncompressed baseline on multiple video‑understanding benchmarks while retaining strong performance at higher compression ratios.

Authors:Samson Abramsky, Radha Jagadeesan
Title: Granthi: Higher-Order Quantum Programming via Unitary Wiring
Abstract:
Existing quantum programming languages confine higher order structure to a classical host while restricting the quantum layer to first order operations on qubits. This paper presents Granthi, a purely unitary higher‑order quantum programming language built on three design commitments: quantum programs are first class values that may be passed, returned, and coherently composed; additive structure is tag‑preserving routing rather than observational branching, so control may remain in superposition; and programmer‑facing finite label types with named reversible operations provide domain‑level control spaces without exposing tag management. Every well‑typed term, including at function type, denotes a unitary on its boundary interface, and the compiler realizes exactly its wiring as a quantum circuit on the physical qubit layout (assuming correctness of the pytket backend). Granthi is implemented end‑to‑end: an OCaml DSL elaborates surface programs through a binder‑free core IR to executable quantum circuits via pytket. The language directly supports the quantum switch, compiled to a static circuit, as well as interference on control‑flow history and structured finite control, all within the purely unitary fragment.

Authors:Rana Muhammad Usman, Dominic Williamson
Title: Peer-Voted LLM-Agent Stress Tests Find Feed-Induced Lexical Convergence but No Reliable Matched-Exposure Advantage for Distributed Sources
Abstract:
Population‑level behavior in large‑language‑model (LLM) agents cannot be characterized by single‑agent benchmarks. We introduce PV‑SST, a peer‑voted social‑platform testbed, and report a separately frozen, preregistered matched‑exposure experiment spanning four topics, four unused seeds, four open‑weight model families, and three prespecified larger variants. The experiment comprises 448 trials and 112 complete model‑by‑topic‑by‑seed blocks. Relative to a topic‑only control, a feed of previous‑round peer posts ranked by peer‑generated likes increases final‑round lexical similarity in both the four‑family core panel (paired mean difference +0.0082 TF‑IDF cosine units, 95% block‑bootstrap CI [0.0043, 0.0121], randomization p=0.000105, n=64 blocks) and the three‑variant size extension (+0.0109 [0.0069, 0.0151], p=0.000001, n=48). This contrast bundles peer‑post exposure with ranking and therefore does not identify a ranking‑only effect. Opposite‑side survival falls in the core panel (‑3.9 percentage points [‑6.8, ‑1.6], p=0.0068) but not conclusively in the larger variants (‑1.0 pp [‑3.1, 0.4], p=0.50). Holding adversarial impressions fixed, four distributed sources do not reliably move honest‑agent stance more than one source. The preregistered distributed‑minus‑single contrast is positive but inconclusive in the core panel (+0.057 [‑0.009, 0.125], p=0.112) and negative in the larger variants (‑0.040 [‑0.113, 0.035], p=0.332), failing the prespecified cross‑model and cross‑topic consistency criterion. Thus the robust result is lexical convergence under the tested peer‑ranked feed, not general opinion capture or a general coordination advantage. The study evaluates synthetic LLM‑agent populations; it does not estimate effects on people or production platforms.

Authors:Minkyu Song
Title: When Retrieval Fails Before It Begins: Structurally Indirect Prerequisite Eviction as a Retention Failure in Agentic Memory
Abstract:
Agentic memory under a fixed budget involves two stages: retention and retrieval. Existing retrieval‑centered paradigms implicitly assume necessary evidence survives eviction, but we challenge this by isolating a pre‑retrieval failure mode: structurally indirect prerequisite eviction, in which upstream blocks weakly aligned with the query are discarded under budget pressure. We provide an operational definition of this failure, a reproducible deterministic benchmark, and per‑seed trace diagnostics. Finally, we evaluate Dependency‑aware Semantic Garbage Collection (DSGC), a one‑hop graph‑aware rule. In our main suite, DSGC improves full‑chain retention from 0.03 to 0.90 under a lexical encoder and from 0.23 to 1.00 under a sentence encoder. Robustness checks then identify the budget and scaling regimes where the one‑hop rule holds or degrades. Our released pipeline and failure postmortem support mechanistic analysis of retention before retrieval as a distinct failure boundary.

Authors:Tanmay Kumar Shrivastava, Darsh Rohit Nandu, Rajesh Kumar Mundotiya
Title: Knowledge-Graph-Gated Defactualization for Style-Controllable and Fact-Preserving Generation in Agentic Conversational AI
Abstract:
Agentic large language models (LLMs) deployed in fact‑sensitive applications such as customer support must simultaneously preserve factual correctness and generate responses in a controllable stylistic register. Activation steering enables fine‑tuning‑free style control by perturbing hidden representations, but it lacks an explicit mechanism for distinguishing verifiable facts from stylistic content, leading to semantic leakage. We address this challenge through \emphDefactualize‑Steer‑Rehydrate (DSR), a knowledge‑engineering framework that integrates a typed, salience‑weighted knowledge graph (KG) with activation steering. DSR extracts salient entities using a layered regex or NER or lexical‑classifier pipeline, replaces them with typed placeholders prior to steering, and deterministically restores verified values through salience‑guided rehydration after generation. DSR is evaluated across six LLaMA‑family models (1B‑‑13B parameters) on 600 A2A‑generated customer‑support cases (1,200 generations), with a dedicated KG ablation study. DSR significantly increases verified‑entity recovery relative to a steering‑only baseline (Cohen's d=0.225, p_\textBonf=1.0×10^‑4), though the absolute recovery rate remains modest, while preserving effective style control across diverse model families. Layer‑wise separability and steering‑strength diagnostics further show previously unexplored interactions between representation‑level steering and factual grounding. hese results demonstrate that explicit knowledge engineering can systematically enhance trustworthy, controllable, and reproducible generative AI without requiring model fine‑tuning. Code, cached steering vectors, and evaluation scripts are publicly released to support reproducibility.\footnotehttps://github.com/Tanmay‑IITDSAI/KG‑Gated‑Defactualization

Authors:M Waleed Kadous, Amr Elsayed, Abdullah Al Nahas, Ashraf Haress
Title: Ansari: A Retrieval-Grounded Islamic AI Assistant -- Architecture, Deployment, and Lessons from 140,000 Conversations
Abstract:
General‑purpose large language models (LLMs) are increasingly used to answer religious questions, but for Islamic content they carry two serious risks: factual fabrication (inventing Qur'anic verses or hadith) and subtle value misalignment. We present Ansari, a deployed, retrieval‑grounded Islamic AI assistant that has handled more than 140,000 conversations across 25+ languages since June 2023. Ansari is built around an agentic retrieval loop: a tool‑using language model issues searches against authenticated Islamic corpora ‑‑ the Qur'an, hadith collections, a multi‑volume jurisprudence (fiqh) encyclopedia, and exegetical (tafsir) sources ‑‑ and answers only on the basis of what it retrieves, with citations attached for verification. We describe the system's architecture (the agent loop, the retrieval tools, the corpora, and the system prompt that encodes editorial and theological policy), its multi‑platform deployment (web, mobile, WhatsApp, and as a Model Context Protocol server and an Agent Skill), and what 140,000 real conversations reveal about how Muslims actually use such a tool. We report results on several complementary evaluations ‑‑ zero‑shot performance on accredited institutional exams, a human‑rated validation during Ramadan, and two independent, externally run benchmarks on which Ansari currently tops the public IslamicMMLU leaderboard ahead of frontier models and is competitive on Islamic legal reasoning (IslamicLegalBench) while strongly resisting false premises ‑‑ and draw out lessons that generalize beyond Islam to any faith‑ or values‑sensitive deployment of LLMs: grounding is necessary but not sufficient, the system prompt is a theological as much as a technical artifact, and the absence of community in how models are formed remains a hard gap.

Authors:Junhui Zhang, Qianhui Xu, Qingxiang Guo, Dawei Yang, Ling Miao, Qiangqiang Wang, Yang Song
Title: Poly-InstructTTS: Learning In-the-Wild Expressive Speech Synthesis from Open-Ended Instructions
Abstract:
While recent text‑to‑speech (TTS) models achieve high naturalness, controlling fine‑grained expression via natural‑language instructions remains challenging. We introduce Poly‑ InstructTTS, which learns expressive speech from open‑ended instructions using in‑the‑wild audiovisual data. We build a scalable multi‑modal pipeline to construct a 1,000‑hour instruction‑annotated corpus covering 1,000+ fine‑grained emotions and styles. The framework uses a prompt‑free GPT with attribute‑based thinking tokens, followed by a flow‑matching module that injects timbre from a reference audio. We also present a speaker fine‑tuning procedure to transfer instruction control to specific speakers while preserving persona. We further extend InstructTTSEval with broader tasks. Experiments show that Poly‑InstructTTS delivers strong performance in instruction adherence and expressiveness. Audio demos and the expanded testset are available on our project page.

Authors:Dengyi Zhao, Zhiheng Zhou, Zihan Wang, Guiying Yan, Xingqin Qi
Title: Interpretable Information-Decomposed Brain Graph Learning for fMRI-based Disease Diagnosis
Abstract:
Resting‑state functional magnetic resonance imaging (rs‑fMRI) has enabled non‑invasive mapping of functional brain interactions for computer‑aided diagnosis, yet most existing approaches reduce inter‑regional relationships to correlation‑based edge weights. Such representations capture co‑fluctuation strength but obscure how information is shared across brain regions. Because brain disorders may disrupt not only connectivity strength but also the organization of redundancy, uniqueness and synergy, traditional functional connectivity may miss disease‑relevant information structures. Here we introduce IID‑GCN, an interpretable graph learning framework that decomposes rs‑fMRI interactions into redundancy, uniqueness and synergy graphs using partial entropy decomposition. These information‑specific graphs separately characterize shared, region‑specific and jointly emergent components of brain activity. A multi‑channel graph convolutional network then integrates the decomposed graphs through edge recalibration, cross‑information interaction, ROI‑attention readout and channel‑attentive fusion. Across three datasets, IID‑GCN consistently captures complementary diagnostic information beyond traditional functional connectivity. The learned information profiles reveal disorder‑specific patterns of altered redundancy, uniqueness and synergy, suggesting that brain diseases reshape functional information organization rather than merely changing connection strength. These results establish information‑decomposed brain graphs as an interpretable representation for rs‑fMRI‑based diagnosis. Our code is available at https://github.com/Zdy12/IID‑GCN.

Authors:Viraj Nishesh Darji, Hemaliben Rakeshkumar Darji
Title: Edge-Based Agentic Retrieval-Augmented Generation for Autonomous FHWA Bridge Inspection Compliance
Abstract:
The Federal Highway Administration (FHWA) mandates that over 600,000 bridges in the United States be evaluated against the Recording and Coding Guide for the National Bridge Inventory (NBI). Manual compliance verification is labor‑intensive, error‑prone, and impractical in connectivity‑limited field environments. This paper introduces BridgeGuard, a fully air‑gapped agentic Retrieval‑Augmented Generation (RAG) system for autonomous bridge inspection compliance. BridgeGuard integrates vector search over the FHWA Recording and Coding Guide with structured SQL queries against NBI tabular data, orchestrated by a stateful multi‑step ReAct planning loop executing locally on commodity edge hardware. A section‑aware chunking algorithm preserves hierarchical regulatory item boundaries, achieving 94.2% chunk integrity compared with 28.4% for naive fixed‑size splitting. Evaluated on the full Delaware 2023 NBI inventory (874 bridges) and a Texas sample (200 bridges), the system achieves 99.77% and 100.0% classification accuracy, respectively, for Structurally Deficient bridge identification, with 100.0% citation accuracy, at 197.0 bridges per hour with out external network access. Ablation experiments confirm that both vector search and the multi‑step agentic loop are necessary for correct compliance reasoning.

Authors:Xinfeng Zhang, Mingxuan Liu, Yifei Chen, Juncheng Zhu, Kasidit Anmahapong, Yiming Huang, Yuan Zhang, Hongjia Yang, Yi Liao, Gang Ning, Haibo Qu, Qiyuan Tian
Title: ASTAR: Automated induction of STAndardized radiology Reporting templates from large-scale clinical free-text corpora
Abstract:
Structured reporting converts free‑text radiology narratives into queryable data keys, facilitating cohort assembly, longitudinal tracking, and training label generation for medical AI. The prevailing paradigm follows a two‑stage pipeline: (1) constructing a reporting template, (2) extracting information to populate it. While the extraction stage has benefited from advances in large language models (LLMs), template construction remains a manual bottleneck relying on labor‑intensive expert consensus that is static, difficult to scale, and may fail to capture real‑world reporting diversity. We address this limitation with \textttASTAR, an LLM‑based framework for Automated induction of STAndardized radiology Reporting templates from large‑scale clinical free‑text corpora. Extensive experiments on 4,215 fetal brain MRI reports from multiple centers demonstrate that the \textttASTAR‑induced template surpasses two expert‑curated templates across template coverage, information fidelity, diagnostic fidelity, and expert‑rated usability, reducing template development from weeks of committee deliberation to hours of automated processing. Code: https://github.com/birthlab/ASTAR

Authors:Etienne Posthumus, Sven Hertling, Dilek Yargan, Harald Sack
Title: bikiDATA: A Python Library to Query and Explore Large-Scale RDF Datasets
Abstract:
While knowledge graphs offer unparalleled data flexibility, the semantic gap between RDF triples and the native objects used by software engineers remains a significant barrier to entry. Developing knowledge‑graph‑backed applications typically requires deep expertise in SPARQL and complex data‑mapping layers. To lower this threshold, we present bikiDATA: a high‑performance storage solution and a Python library engineered for the modern software developer. Unlike traditional wrappers, bikiDATA abstracts the complexities of the RDF data model into a developer‑friendly API that feels native to the Python ecosystem. Beyond standard SPARQL support, the system provides a comprehensive suite for production‑grade applications, including integrated full‑text search, knowledge graph embeddings, and visual similarity search. Already in use in ongoing projects at FIZ Karlsruhe, bikiDATA reduces integration complexity, improves scalability, and enhances query performance. The source code and executable demo notebook are publicly available at https://github.com/ISE‑FIZKarlsruhe/bikidata.

Authors:Sanjay Basu
Title: Inhibitory Attention for Clinical Long-Context Reasoning: Characterizing and Mitigating Lost-in-the-Middle Effects in EHR Processing
Abstract:
Electronic health records now routinely exceed 100,000 tokens per patient. Yet large language models exhibit the lost‑in‑the‑middle (LitM) effect: information near the center of a long context is retrieved less reliably than information near the edges. In clinical use this is not benign: the single most consequential fact in a note can sit at its center. We term this the clinical lost‑in‑the‑middle (CLitM) problem, give its first systematic characterization using MedAlign, and compare context‑selection strategies as remedies. Across 2,196 instruction‑response pairs and six language models, we observe a 21.9 percentage‑point gap between peak accuracy (59.5%, 95% CI [46.3, 71.0], 20‑30% decile) and trough accuracy (37.6% [23.2, 52.5] at 70‑80%); 67.8% of reference answers fall between the 10th and 90th percentiles of the EHR timeline, inside the CLitM trough. We introduce Query‑Conditioned Clinical Suppression (QCCS), a lightweight query‑conditioned selection gate, and evaluate it against BM25, BM25 with section‑header filtering, dense retrieval, and cross‑encoder reranking (N=83 held‑out instructions). With Qwen2.5‑7B‑Instruct (16k context), QCCS outperforms all five comparators under LLM‑as‑judge scoring: for middle‑position instructions QCCS reaches 16.7% versus BM25 3.3%, cross‑encoder 0.0%, dense 0.0%, and full context 6.7%; overall QCCS reaches 25.3% versus at most 3.6% for retrieval‑only comparators. This advantage is not explained by retrieval recall: at k=20, BM25 retrieves the gold evidence sentence in 98.8% of instructions (QCCS 34.9%), yet retrieval arms stay at most 2.6% accurate even when they retrieve it, whereas QCCS reaches 25.0% even when it does not. In this proof‑of‑concept evaluation, query‑aligned context selection predicts EHR instruction‑following accuracy better than gold‑sentence retrieval recall.

Authors:Sahel Sharifymoghaddam, Lingwei Gu, Yijun Ge, Jimmy Lin
Title: Projecting BrowseComp-Plus onto ClimbMix: Toward More Realistic Corpora for Agentic Search
Abstract:
The BrowseComp‑Plus benchmark disentangled the evaluation of agentic search by replacing opaque web search with a fixed corpus, so that an agent's role can be separated from the retriever's. That corpus, however, holds only about 100K documents and was assembled from the supporting documents of the benchmark's own queries plus mined hard negatives, so the evidence and the distractors were both selected per query. We introduce \textBrowseComp‑Plus_\textCM, which keeps the BrowseComp‑Plus questions but relocates their evidence to ClimbMix, a 400B‑token, 553M‑document mixture of web text released by NVIDIA for pre‑training language models and built without reference to any benchmark. Our main contribution is the projection pipeline that makes this possible: it decomposes each question into atomic reasoning hops and grounds every hop in the new corpus, retaining a question only when automatic verification, an independent agent, and human review all confirm that every hop is supported. The pipeline is dataset‑agnostic and applies to any benchmark whose questions decompose into verifiable facts. Applied to the 830 BrowseComp‑Plus test questions, our pipeline yields 57 fully grounded questions with question‑level relevance judgments. Projection shifts the difficulty onto retrieval, as the strongest agent we evaluate loses five points of answer accuracy but sees its evidence recall fall from 84.3% to 21.4% while issuing 63% more search calls. As the first of a series of projections, we release the pipeline, the benchmark, and our analyses at https://github.com/castorini/cmass.

Authors:Yufei Liu, Xixi Wang, Hao Li, Ganlong Zhao, Kaitong Cai, Chengkai Jin, Chunxiao Liu, Jianbo Liu, Siyuan Huang, Xingang Pan, Hongsheng Li
Title: DreamHand: Repurposing Video Diffusion Models for Occlusion-Robust Egocentric 3D Hand Motion Recovery
Abstract:
Egocentric video offers scalable manipulation data for embodied AI, yet recovering metric 3D hand trajectories remains challenging due to severe object occlusion and frequent out‑of‑sight gaps. Existing single‑frame and windowed temporal regressors fail when hand shortly leaves the frame, while recent video diffusion models (VDMs) rely on heavy, stochastic multi‑step sampling as pixel‑space renderers. We instead repurpose VDM into a deterministic geometry encoder. A single forward pass over the clean latent exposes scene content beyond current observations, including occluded and out‑of‑sight hands. We introduce DreamHand, an offline clip‑level framework that extracts features via a Deterministic Clean‑Latent Encoder and decodes them with a Bidirectional Spatiotemporal Decoder. DreamHand recovers continuous bimanual trajectories with metric placement and no external detector, while a Ray‑Based Camera Solver supports a second configuration that needs no test‑time camera intrinsics. Across five egocentric benchmarks, DreamHand sets a new state of the art, cutting MPJPE‑p by 30% on occlusion‑heavy ARCTIC and 40% on HOT3D. These gains reach 46%‑61% once out‑of‑sight hands are included in the evaluation, offering a scalable path from everyday human video to robot manipulation data.

Authors:Cheng Xu, Nan Yan, Liming Chen, M-Tahar Kechadi
Title: Phantom Gains: Auditing Self-Improvement Against a Measured Null
Abstract:
Whether a language model has improved itself is increasingly judged not by mean accuracy but by which individual problems it gains and loses. Tracking these transitions means differencing two noisy estimates, leaving them vulnerable to measurement artifacts. Auditing three rounds of rank‑32 LoRA self‑training on Qwen3‑8B against a frozen control pushed through the identical pipeline, we identify seven measurement failures, each of which inverts a reported finding when its control is absent. Several are standard practice. A ledger built on a single greedy decode manufactures capability changes on an untrained model, largely an artifact of inference batching; the expansion statistic separating acquisition from sharpening assigns that same model a rate of 0.280. The natural threshold repair does not survive replication: estimated across the frozen comparisons such a design already contains, its null stays non‑zero. We replace it with a per‑problem exact test against a pooled baseline under false‑discovery‑rate control, which detects nothing on any held‑out replicate and is unchanged under the multiple‑testing rule, error rate and pool size. Applied to a ladder of arms matched in stream, volume and evaluation, the audit finds that external distillation improves problems the base model rarely reaches while three forms of self‑training do not; a regression rejects this asymmetry as a by‑product of distillation's larger overall gain (p < 10^‑8). On the far smaller set of problems the base model never reaches, the evidence is inconclusive, while self‑training corrupts problems solved at baseline at rates well above the measured floor. Transition‑level auditing therefore requires a separately measured null for every statistic it reports: nulls that cost no new experiments, built from baseline replicates a multi‑arm study already owns, though not from as few as most possess.

Authors:Yu Hu, Fangzhou Zhao, Liang Chen, Chen Min, Wei Li, Mingyuan Sang, Jiajia Ma, Shican Chen, Di Pang, Baolei Chen
Title: DART-S: Reachability-Audited Active-Suspension Preconditioning for Off-Road Vehicle Jumps
Abstract:
Airborne torque reaction cannot recover takeoff errors beyond the wheel angular‑momentum budget. DART‑S applies ramp‑face suspension preconditioning to change pitch, pitch rate, and wheel spin before liftoff, thereby shifting the queried state and altering the remaining authority budget. To predict how each suspension action reshapes this state‑budget pair, DART‑S employs a local calibration map. A support‑aware selector combines the predicted shift with local outcome evidence and an interval‑reachability screen; an exact‑pair audit reports residual authority. Across 600 new runs in 72 independent BeamNG sessions, every positive, negative, and boundary query follows its prespecified branch. At the confirmed 40°/13 m/s boundary, DART‑S attains 24/24 post‑touchdown attitude‑criterion successes versus 0/24 for DART (session‑level Holm‑adjusted p=0.0234). At 11.5 m/s, a 0.35 s timing action attains 23/24 versus 0/24 for the static preset (p=0.0156). The 200 rad/s command guard keeps drivetrain hard‑limit exceedance at zero across all 600 runs. The source code will be available at https://github.com/MeridianCAS/DART‑S

Authors:Cong Wang, Liyan Wang, Jinshan Pan, Wei Wang, Wenqi Ren, Jun Liu, Xiaochun Cao
Title: Ultra-High-Definition Restoration Transformers with Correlation Matching Transformation
Abstract:
We propose UHDformer++, a general Transformer‑based framework to solve numerous Ultra‑High‑Definition (UHD) image restoration tasks. UHDformer++ operates across 4 coordinated learning spaces: 1) a high‑resolution space (HR) for multi‑level feature extraction, 2) a low‑resolution space (LR) for learning compact, representative features, 3) a super‑resolution space (SR) for upsampling low‑resolution features from SR, and 4) a low‑high fusion and reconstruction space (LHFR) for final image restoration. Specifically, HR extracts multi‑scale high‑resolution features and fuses them with low‑resolution cues to produce residual images, while LR distills complementary representations from HR to improve restoration quality. To supply LHFR with richer features, SR super‑resolves LR outputs before fusion. We further introduce two modules to bridge the high‑ and low‑resolution spaces. The Feature‑Refined Correlation Matching Transformation (FR‑CMT) module selects the top C/r~(C~\textdenotes the number of channels;~r\geq1~\textcontrols the squeezing level), from the fusion between max‑ and mean‑pooled high‑resolution features to replace less informative channels in the low‑resolution Transformer. The Adaptive Channel Modulator (ACM) adaptively recalibrates multi‑scale high‑resolution features, ensuring that only task‑relevant information propagates to LR. Extensive experiments demonstrate that UHDformer++ reduces model parameters by at least 86% compared with recent state‑of‑the‑art methods while achieving substantial performance gains across 5 UHD restoration tasks, including low‑light image enhancement, dehazing, deblurring, deraining, and desnowing. Code will be released at https://github.com/supersupercong/uhdformerplus.

Authors:Xincheng Tang, Yiji Chen, Youhan Xie, Wanyu Li, Zhengjie Shu, Lai Jiang, Wenkang Hu, Yitong Li, Jinchuang Zhang, Xibin Song, Ruigang Yang
Title: Video2DoorTraversal: Push Door Traversal via Simulated Door Twins
Abstract:
Door opening and traversal is a long‑horizon loco‑manipulation task that requires precise handle interaction and coordinated base‑arm control. We present Video2DoorTraversal, a single‑video real‑to‑sim‑to‑real framework for wheel‑legged mobile manipulators. Given one RGB video of a real door, DoorTwin reconstructs an instance‑aligned, articulated, and simulation‑ready door twin with realistic geometry and appearance. A simulation‑in‑the‑loop agent converts the recovered articulation into a parameterized skill program and iteratively refines failed rollouts to generate physically executable demonstrations. These demonstrations are used to train ArticuACT, a dual‑depth policy that predicts coordinated base, arm, and gripper commands using robot‑centric camera conditioning and interaction‑aware supervision. With all perception and policy inference running onboard, the system achieves a 96.57% average success rate across five real doors and an 80.95% zero‑shot success rate on structurally similar unseen doors, while completing the full approach, opening, and traversal sequence in approximately 13s on average. Project Page: https://video2doortraversal.github.io/.

Authors:Yu Chen, Ting Lei, Yaoyi Li, Jia Cai, Zhecen Wu, Yang Liu
Title: Rule-Compliant Visual Spatial Planning for Multimodal Large Language Models
Abstract:
Multimodal large language models (MLLMs) combine linguistic reasoning with visual perception, yet their ability to perform visual spatial planning under explicit or previously unseen rule constraints remains underexplored. This setting requires models to jointly understand spatial layouts, interpret natural‑language rules, and plan valid actions accordingly. To address this gap, we introduce RuleMaze, a controllable benchmark in which MLLMs must navigate mazes while obeying natural‑language rules of varying complexity. RuleMaze isolates rule‑compliant spatial planning by requiring accurate perception, rule interpretation, and constrained action planning. To enable scalable and systematic rule construction, we propose Language‑Logic‑Function Hybridization, which automatically generates natural‑language rules and translates them into logical representations and executable validators, eliminating manual rule engineering. To improve rule following and generalization, we introduce Disentangled Multimodal Planning (DMP), which separates perception, execution, and rule verification through interpretable reasoning primitives. By disentangling these components, DMP facilitates systematic generalization to more complex and previously unseen rules, while providing transparent intermediate planning traces. Experiments demonstrate that DMP substantially improves rule compliance and planning success compared to end‑to‑end textual planning baselines. Overall, RuleMaze establishes a principled benchmark for studying grounded and interpretable rule‑based spatial planning in MLLMs. Code is available at https://github.com/oceanflowlab/RuleMaze.

Authors:Christos Koutsiaris
Title: Daedalus-150M: A Convolution-Attention Hybrid Designed for CPU Inference
Abstract:
Small language models are usually built like large ones and then squeezed onto a CPU afterwards. We did the opposite: we fixed the target first, one user, one token at a time, 4‑bit weights, ordinary CPU, and chose the architecture to suit it. The result keeps full attention in only 6 of its 18 blocks. The other 12 use short convolutions whose memory is two timesteps wide no matter how long the conversation gets, so two thirds of the network never re‑reads a growing cache. Trained from scratch on 59.9B tokens, the model scores 47.31 on a five‑task benchmark against a bar of 42.20 that was fixed before training began. It beats GPT‑2 124M, Pythia‑160M, OPT‑125M and GPT‑neo‑125M, all trained on three to six times more data, and exceeds MobileLLM‑125M's published score despite that model seeing a trillion tokens. Validation bits‑per‑byte is 0.8685. To check the architecture rather than the training recipe, we trained a conventional all‑attention model of the same size on the same data, and wrote down the winning condition before scoring either. The hybrid won the chosen quality metric by 0.81%, matched it on downstream tasks, produced a 6.3% smaller 4‑bit file, and decoded 1.76x faster at 2048 tokens of context, 2.08x against an external model of similar size. In every measurement the speed advantage is near zero at an empty context and grows with length, which is what the mechanism predicts and what a merely leaner model would not show. A simple bandwidth calculation predicts only 1.17x, so memory volume alone does not explain the gap. We also report what did not work: an unmitigated 4‑bit quality cost, roughly half the convolution channels ending up inert and impossible to remove, and a vocabulary larger than this model size warrants.

Authors:Shaoxuan Wang, Guangting Zheng, Rui Huang, Zhipeng Tang, Sha Zhang, Jiajun Deng, Yanyong Zhang
Title: RoMAN-Flow: Taming Autoregressive Normalizing Flows for Offline Reinforcement Learning in Robotic Manipulation
Abstract:
Offline reinforcement learning improves robotic policies using previously collected data without further environment interaction. Yet prevalent diffusion‑ and flow‑matching robot policies lack tractable likelihoods, limiting their use in likelihood‑based offline RL post‑training. AR‑NFs offer both expressive action modeling and exact likelihood evaluation, but their sequential sampling incurs substantial sampling overhead during policy optimization and deployment. We present RoMAN‑Flow (Robotic Manipulation with Autoregressive Normalizing Flows), an offline reinforcement learning framework that makes AR‑NF policies practical for robotic manipulation by addressing this sampling bottleneck in both stages. During policy optimization, RoMAN‑Flow employs a sampling‑free, advantage‑weighted likelihood objective that assigns higher likelihood to high‑advantage actions from the offline dataset without sampling from the autoregressive policy. For efficient deployment, it distills the optimized autoregressive policy into a one‑step action generator, enabling low‑latency action prediction. Experiments across multiple simulated manipulation benchmarks and real‑world robotic platforms demonstrate that RoMAN‑Flow achieves competitive policy performance while substantially reducing inference latency. Code is available at https://github.com/konnyaku28/RoMAN‑Flow.

Authors:Julian Oelhaf, Georg Kordowich, Paula Andrea Pérez-Toro, Christian Bergler, Johann Jäger, Andreas Maier, Siming Bayer
Title: A Standardized Framework for Machine Learning in Power System Protection
Abstract:
Studies of machine‑learning‑based power‑system protection increasingly report near‑perfect scores, yet the meaning of those scores depends strongly on the evaluation setting. Protection task, physical scope, measurements, timing, targets, preprocessing, and validation often vary jointly and remain incompletely specified. This paper proposes a standardization‑oriented framework that treats evaluation design as part of the scientific contribution. It defines seven required study dimensions: protection objective, physical scope, observability, timing and decision windows, targets and sample validity, validation protocol, and evaluation outputs. The framework is instantiated in a bounded case study on the public PROTECT‑90 electromagnetic‑transient benchmark, comprising 9022 simulated episodes from a 90 kV double‑line topology, for onset‑conditioned fault classification and localization. Under centralized sensing, simulation‑metadata‑aligned 20 ms windows, and episode‑grouped validation, a multi‑layer perceptron (MLP) achieved a five‑fold mean macro‑averaged F1 score of 0.991 +/‑ 0.001 for classification and a localization mean absolute error of 10.20 +/‑ 0.25% of line length (mean +/‑ std across episode‑grouped folds). Extending the decision horizon to 50 ms preserved this task‑dependent performance asymmetry, while reduced observability approximately doubled the MLP localization error but had little effect on classification. A synchronized two‑ended conventional locator outperformed the learning locators under its richer clean information set, and measurement degradation showed that clean predictive performance did not determine robustness. The framework turns evaluation assumptions into explicit, reproducible evidence and provides a basis for more comparable, auditable evaluation and future certification‑oriented assessment of machine‑learning protection functions.

Authors:Atsuyuki Miyai, Kiyoharu Aizawa, Toshihiko Yamasaki
Title: Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task Selection
Abstract:
We present a novel approach to efficient LLM agent harness optimization through adaptive validation task selection. Harness optimization iteratively rewrites the harness code based on validation performance, enabling substantial performance gains without updating the underlying model weights. Existing approaches, however, evaluate a fixed validation set in full at every iteration, incurring substantial evaluation costs even on tasks that become less discriminative as the harness evolves. We propose Task‑CoEvolve, which co‑evolves the validation tasks with the harness by addressing two challenges: selecting informative tasks and estimating full‑set performance from partial evaluations. Task‑CoEvolve builds on the observation that tasks on which candidate harnesses disagree are more informative for distinguishing among them than tasks that are consistently solved or failed. It uses variance‑weighted sampling based on past outcomes to focus evaluation on tasks near the agent's capability frontier, with the sampling distribution adapting as the harness evolves. It then estimates full‑set scores from the sampled tasks by accounting for their sampling probabilities, enabling consistent comparisons across iterations despite evaluating different subsets. Experiments on online text classification and Terminal‑Bench 2.1 show that Task‑CoEvolve consistently outperforms fixed‑subset baselines and matches the final performance of full‑set search while reducing the number of evaluations during optimization by 80%. Code will be released at https://github.com/Agent4Science‑UTokyo/Task‑CoEvolve.

Authors:Minghang Zheng, Jingli Wei, Hongyi Yang, Yang Liu
Title: ID-VTG: Image-Disambiguated Video Temporal Grounding
Abstract:
Video Temporal Grounding (VTG) faces significant challenges when natural language queries must distinguish between multiple events involving visually similar entities, particularly when relying on fine‑grained visual attributes that are difficult to describe accurately in words alone. To address this, we introduce Image‑Disambiguated Video Temporal Grounding (ID‑VTG), a task that leverages multimodal queries combining a reference image and a text description to precisely localize segments where a specific instance performs a described action. To facilitate research, we construct two benchmarks: IDVTG‑Gym, focusing on fine‑grained, compositionally ordered gymnastics actions with athletes in similar uniforms; and IDVTG‑InternVid, an open‑world dataset featuring diverse entities (e.g., humans, animals, fictional characters) and significant temporal distractors. Methodologically, we propose the Visually‑Guided Disambiguation Aggregation (VGD‑Agg) framework based on a dual‑branch fast‑slow architecture. The fast branch efficiently generates preliminary event proposals, while the slow branch performs fine‑grained frame‑level matching between video frames and the reference image. We enhance discriminability via two learnable tokens: a Compare Token, which represents hard negatives to probe for the presence of the target instance (as referred to by the query image), and a Depress Value, which represents text‑irrelevant events. Proposals that the Compare Token identifies as lacking the target instance are pushed toward the Depress Value, thus easing disambiguation via the text query. Extensive experiments validate our approach, which achieves state‑of‑the‑art results on the proposed benchmarks. Code is available at https://github.com/oceanflowlab/ID‑VTG.

Authors:Hugo Porta, Emanuele Dalsasso, Chang Xu, Theo Gnassounou, Devis Tuia
Title: SAE-Xplainers: Rule-Based Feature Interpretation for Extreme Earth Events
Abstract:
The emergence of large‑scale Weather and Climate (W&C) datasets offers new opportunities for modeling extreme Earth events (ExEE) and their impacts using deep learning. However, their adoption in operational settings remains limited by the lack of models' interpretability. While for conventional text and image modalities, tools such as Sparse Autoencoders (SAEs) have proven effective for extracting human‑understandable concepts, their use for the analysis of ExEE remains challenging due to the nature of W&C data. To address this, we introduce (i) a geographic location‑based modulation of the inputs of SAE to capture the local semantic meaning of environmental patterns, and (ii) an ensemble of rule‑based SAE‑Xplainers to interpret the resulting high‑dimensional features derived from complex, multi‑modal environmental predictors. We evaluate our method on three ExEE types: the prediction of fires, and the detection of tropical cyclones and atmospheric rivers. We show that SAE input modulation improves both reconstruction performance and feature utilization, and that our SAE‑Xplainers enable faithful interpretation of complex climatic patterns by unfolding them into human‑understandable rules that are consistent with the scientific literature, while also supporting the identification of feature absorption.

Authors:Yigit Ekin, Enes Sanli, Aykut Erdem, Erkut Erdem, Aysegul Dundar
Title: BeyondMasks: Evaluating Causal and Physical Consistency in Video Object Removal
Abstract:
Recent advances in generative video models have significantly improved visual realism in video object removal, yet evaluation protocols still focus on masked region fidelity, treating removal as local inpainting. In real scenes, object removal is a causal intervention: eliminating an object also requires removing its induced physical effects, such as shadows, reflections, illumination changes, translucency, and dynamic traces. Existing benchmarks lack aligned clean references or remain limited to simplified synthetic settings, preventing systematic evaluation of causal consistency. We introduce BeyondMasks, a paired benchmark for causally consistent video object removal, consisting of temporally aligned synthetic and real world video pairs with clean background references. The dataset spans diverse photometric, geometric, volumetric, and dynamic interactions, and supports both mask based and instruction driven editing. We further propose CORE, a structured vision language model based evaluation protocol that jointly measures object disappearance and after effect consistency, aligning more closely with human judgments than existing metrics. Benchmarking state of the art methods reveals systematic failures in removing secondary physical effects despite high masked region fidelity, exposing a gap between visual plausibility and causal correctness. BeyondMasks reframes video object removal as causal scene consistency rather than local reconstruction and provides a unified framework for its evaluation.

Authors:Tao Huang, Ruofei Liu, Xuchen Tang, Xinyin Zhang, Junli Ren, Huayi Wang, Feiyu Jia, Yukai Qi, Kangning Yin, Weishuai Zeng, Lipeng Chen, Xi Li, Ting Wu, Kailin Li, Ruoli Dai, Jingbo Wang, Lei Han, Jiangmiao Pang
Title: Towards Professional Tennis Styles for Humanoid Robots with Adaptive Motion Planning and Tracking
Abstract:
Humanoid robots have recently demonstrated promising capabilities in real‑world ball sports. However, achieving professional motion styles while maintaining strong task performance remains challenging. In this work, we propose AdaPT, an Adaptive Motion Planning and Tracking framework that learns professional tennis serving and rally styles directly from broadcast videos. This hierarchical design is motivated by the key insight that the planner generates stylistic kinematic motions, while the tracker executes them with minimal interference with planning. Despite its effectiveness in simulation, a substantial sim‑to‑real gap emerges: tracking performance inevitably degrades on real robots, and this degradation is partially overlooked by autoregressive planning and further compounded by noisy perception. To address these issues, our adaptation mechanism improves tracking robustness by learning to track randomized execution speeds, while conditioning the planner on a learned motion‑speed adapter to mitigate compounding errors. Real‑world experiments on the Unitree G1 demonstrate the effectiveness of our adaptation mechanism in bridging the sim‑to‑real gap. We further deploy AdaPT policies on the full‑size Dobot Atom humanoid robot (1.7m) and demonstrate in‑the‑wild serving without motion capture. Beyond these results, our real‑world experiments reveal both algorithmic and engineering insights for future humanoid ball‑sports systems. Videos and code are available on our \hrefhttps://humanoidtennis.github.io/AdaPT/project website.

Authors:Narcis Marincat
Title: What You Can't See Is What You Learn: Restricted Evidence Visibility Favors Compositional Generalization in Shared-Genome Language-Model Societies
Abstract:
Multi‑module systems often expose every module to the full input. We test whether restricting evidence visibility changes which solutions gradient‑based training discovers. Four‑cell societies share one frozen pretrained language model and one low‑rank adapter, communicating only through two model‑width continuous vectors in a fixed relay. On a prospectively sealed natural‑language function‑composition task, we train ten matched restricted/global pairs sharing initialization bytes, training order, token layout, parameters, and computation; only the attention mask differs. Restricted societies outperform their globally visible twins by at least 20 points at both depths in 9 of 10 pairs, with median paired advantages of 0.7648 and 0.6050. Cutting communication reduces every restricted society to chance, and the depth‑three advantage remains 0.558 on programs whose composite function never appeared in training. Across six audited restricted societies, same‑value packet transplants preserve behavior at 0.94‑1.00 across all tested interfaces; destructive interventions collapse performance; and counterfactual packets redirect outputs toward the mathematically predicted answer. The sole high‑performing global model also requires communication, but its same‑value packets are not interchangeable across episodes. Restricted visibility is thus not necessary for composition; under this protocol it substantially increases the probability of a generalizing relay and favors a reusable, value‑indexed interface. The complete preregistered battery nevertheless formally fails because restricted‑arm median depth‑three accuracy is 0.6988, below the 0.70 floor. An earlier qualification cohort likewise yielded 0/10 complete passes: one model met every task‑performance gate, but all ten failed ordinary‑language preservation, confining the system to explicitly task‑gated use.

Authors:Rui Wang, Yeteng Wu, Xianlin Zhang, Mengshi Qi
Title: ExPhy: A Benchmark for Explicit Physical Property Learning in Multi-Object Trajectory Forecasting
Abstract:
Understanding object dynamics requires not only predicting future trajectories but also examining whether a model captures the physical properties that govern motion. However, existing benchmarks rarely expose object‑level physical properties as explicit evaluation targets alongside trajectory forecasting. To address this gap, we introduce \emphExPhy, a multi‑object trajectory forecasting benchmark containing 24,000 simulated physical scenes with explicit object‑level labels for mass, friction, and restitution. ExPhy provides observed and future trajectories together with an in‑distribution (ID) split and two out‑of‑distribution (OOD) splits over physical parameters (OOD‑Parameter) and initial states (OOD‑Initial) for jointly evaluating trajectory forecasting and physical property estimation. We further instantiate \textscPhyODE, a physics‑guided model with an explicit property interface that estimates physical properties from observed trajectories and uses them for differentiable future rollout. On the long‑horizon OOD‑Initial setting, \textscPhyODE reduces ADE and FDE by 33.1% and 31.0%, respectively, compared with the strongest baseline. Zero‑shot evaluation on ComPhy further assesses cross‑benchmark transfer. Property‑level analyses reveal that accurate trajectory forecasting does not necessarily imply accurate recovery of the underlying physical properties. Code and data are available at https://github.com/Zest86/ExPhy.

Authors:Jakub Micorek, Mateusz Koziński, Horst Possegger
Title: STEP: Score-Based Temporal Energy for Human Pose Video Anomaly Detection
Abstract:
Skeleton‑based Video Anomaly Detection (VAD) offers a robust, privacy‑preserving solution for identifying abnormal behaviors. To model the distribution of normal static and moving poses, recent methods train Energy‑Based Models (EBMs) via Denoising Score Matching (DSM). However, directly injecting noise, required for training, into raw joint coordinates creates physically impossible poses, and this structural collapse severely worsens as the temporal window expands. To address this, we introduce STEP, a simple framework that utilizes Principal Component Analysis (PCA) to project pose sequences into a compact, whitened PC‑space. Learning the data density within this well‑behaved PC‑space ensures that the injected noise translates into physically plausible variations, which allows the model to process longer video sequences without the performance collapse of raw coordinate baselines. Additionally, to mitigate inherent pose estimation inaccuracies arising from occlusions or motion blur, we integrate a sequence‑level weighting mechanism based on the estimator's confidence scores. Operating at real‑time computational efficiency, our simple and lightweight framework outperforms the previous skeleton‑based state‑of‑the‑art by 12.2% (90.1% AUROC) on the challenging UBnormal dataset and achieves highly competitive results by improving on the ShanghaiTech benchmark.

Authors:Shangbo Yuan, Jie Xu, Xiaofeng Zhu, Na Zhao
Title: Open-Vocabulary 3D Object Detection with Co-Distillation Discovery and Dual Guidance Robust Training
Abstract:
Recently, open‑vocabulary 3D object detection (3D‑OVD) has gained increasing attention for its ability to detect unseen objects in 3D scenes. Existing approaches typically adopt a two‑stage pipeline that first discovers novel objects using foundation models and then trains a 3D‑OVD model based on these discovered objects. Although effective, this pipeline often suffers from inaccurate localization and mismatched classification during the discovery stage, which subsequently limits the performance of the model training stage. To address these limitations, we advocate for improving both the reliability of novel object discovery and the robustness of model training, and propose an innovative framework. Specifically, for reliable discovery, our co‑distillation strategy distills high‑quality novel objects by applying Hungarian matching over a comprehensive score that incorporates geometric consistency, structural objectness, and semantic certainty. To enhance robust model training, we further propose a dual‑guidance learning scheme, incorporating a scene‑awareness‑guided uncertainty regularization for the regression head and an LLM‑guided hierarchical alignment for the classification head, effectively mitigating the negative effects of imprecise 3D bounding boxes and semantic ambiguity. Extensive experiments on SUN RGB‑D and ScanNetV2 demonstrate that our method achieves significant performance gains over state‑of‑the‑art approaches. Code is available at https://github.com/shangboyuan/Co‑3DGT

Authors:Bhavya Gupta, Onat Gungor, Tajana Rosing
Title: G-MARK: Grounded Multi-Agent Reasoning for Cooperative Driving via Knowledge Graphs
Abstract:
Autonomous driving systems must operate under partial observability, where safety‑critical objects may be occluded or visible only to neighboring connected vehicles. Vehicle‑to‑vehicle cooperation can reduce this uncertainty, but existing cooperative driving methods often compress multi‑agent evidence into latent features or hidden multimodal states. As a result, they obscure which agent observed each object, whether the object is visible to the ego vehicle, and how conflicting evidence affects downstream decisions. We propose G‑MARK, a grounded multi‑agent reasoning framework that converts cooperative object‑centric observations into explicit provenance‑aware knowledge graphs (KGs). The resulting KGs preserve object hypotheses together with their source attribution, ego‑versus‑partner visibility, uncertainty, conflicts, spatial relations, and planning‑relevant context. G‑MARK then derives a shared feature representation from these KGs, enabling lightweight task heads to support object reasoning, motion prediction, control selection, and trajectory forecasting. Compared with the state‑of‑the‑art baseline, GMARK improves occlusion reasoning accuracy by 42.2%, reduces control‑selection error by 13.1%, and achieves comparable trajectory‑planning accuracy with a 25.6x smaller structured communication payload. Our code is available at https://github.com/bhavyagupta98/g‑mark.

Authors:Muhammad Sarmad Sohail
Title: Auditing Recorded Predictive Lead Service-Line Classifications Against Physical Verification: A Statewide Study of New York
Abstract:
Under the US Lead and Copper Rule Revisions, a utility may determine a service line's material with a predictive model instead of inspecting it. New York State publishes, per address, which method was used. Almost no address carries both a model classification and a physical verification, so the check is between populations within a utility rather than paired addresses. We screen all 153 New York localities that classified at least 100 addresses this way. Seventy‑five (49%), covering 125,990 addresses or 57% of those screened, record one value. Zero variance alone is not misconduct: 68 of the 75 match their own verification or have too little to test. Seven are contradicted by their own crews, six beyond any sampling explanation. Five are boroughs of New York City, which file as one system; one is East Rochester, 550 km away. New York City is the largest case: a predictive model is the recorded basis for 43,215 addresses, and on all of them the recorded material is "Known Other". The city records "Unknown" on 121,779 addresses, 1,880 already excavated, and lead on 120,692. In the model bucket both counts are zero, and the 95% upper bound on the rate is 0.0085%. Across the rest of New York the same method records lead or the hedge "Unknown but could be lead" on 12.21% of 176,888 addresses, a comparison whose weaknesses we report. The model‑cleared population is newer, median year built 1984 against 1930, and construction era accounts for about a third of the gap and not the rest: holding era fixed, records‑based classification finds lead at 4.3‑31.9%, physical verification at 1.5‑14.5%, the model in no era. Six era‑aware estimators place the expected lead lines among them at 1,150‑1,450. Two findings need no comparison: 7,782 of these addresses are in pre‑1940 buildings, and the archived 2025 snapshot shows the public‑side determination was copied from a customer‑side model output.

Authors:Matthias Seeger, Zeyu Zhang, Vihang Patil, Konstantinos Benidis, Sebastian Schelter
Title: Learning how to Forget: Fine-tuning for Long-Context Sparse Attention
Abstract:
A lot of prior work addressed key‑value (KV) cache selection and compression by sparse attention to enable long‑context inference for transformer language models without excessive hardware budgets. We provide a new method for fine‑tuning models with sparse attention. It works for any KV cache policy, runs on a moderate hardware budget (e.g., a single Nvidia A100 GPU with 40 GB RAM), and allows the model to co‑adapt with the policy, often outperforming models trained with exact attention (sequence parallelism). We also provide an efficient implementation of H2O sparse attention (the leading policy in our experiments) with dedicated scaled dot product attention kernel support. KeysAndValues (https://github.com/awslabs/keys_values), a new open source library for long‑context inference and fine‑tuning, provides easy‑to‑use and performant code for all methods discussed here.

Authors:Yue Wang, Yi Liu, Gelei Deng, Ying Zhang, Yuekang Li, Zhenyu Chen, Leo Zhang
Title: MaliciousSkillBench: A Comprehensive Benchmark for Malicious Agent Skill Detection
Abstract:
Agent Skills extend LLM agents with reusable instruction packages that may also include scripts, resources, and service configuration. This creates a direct distribution channel for malicious behavior, yet existing malicious‑Skill datasets are fragmented across sources, artifact formats, evidence regimes, and benign coverage; duplicated and structurally related content further complicates direct aggregation and evaluation. We present MaliciousSkillBench, a comprehensive benchmark for malicious Agent Skill detection. We consolidate 13 public sources, 11 of which contribute Core malicious artifacts, and reduce 8,414 raw malicious records to 7,539 normalized‑unique identities in 4,588 operational structural families. After conservative cross‑label conflict exclusion, the primary benchmark contains 9,740 Skills: 7,505 malicious and 2,235 benign. To characterize its coverage, we harmonize 11 attack categories for 4,983 malicious identities with supported source‑native mappings and find substantial differences in threat composition across sources. We then evaluate three learned text detectors and three off‑the‑shelf Skill scanners. Learned detectors achieve 0.882‑0.932 Random Macro‑F1 but only 0.653‑0.665 under Source‑Disjoint evaluation; the strongest word TF‑IDF SVM scores 0.932/0.916/0.665 on Random/structural‑disjoint/Source‑Disjoint while retaining 95.6% malicious recall but producing 62.4% benign FPR on held‑out sources. Off‑the‑shelf scanners occupy different but also unsatisfactory operating regimes, reducing false positives only at the cost of sharply lower malicious recall. Together, these results show that reliable malicious‑Skill detection requires both broader cross‑source benchmark coverage and evaluation that jointly measures attack detection and benign over‑flagging.

Authors:François Costa, Raphael Kreft, Eckhard Goedeke, Felix Möller, Hardik Shah, Ramanathan Rajaraman, Shaohui Liu, Rémi Pautrat, Marc Pollefeys
Title: Unified and Efficient Point-Line Local Features
Abstract:
Multi‑view computer vision pipelines typically rely on accurate sparse keypoints and robust descriptors. While incorporating line features has shown clear benefits for matching and pose estimation, existing point‑line approaches remain inefficient: they detect points and lines separately, use increasingly heavy networks, and depend on CPU‑bound heuristics that hinder real‑time performance. We introduce a Unified Efficient Points and Lines (UPAL) feature extractor that jointly extracts keypoints, line segments, and feature descriptors within a single lightweight architecture. A shared backbone provides common representations that feed different branches for point and line features. Line segments are recovered through an accelerated post‑processing stage, an enhanced and highly efficient variant of the LSD algorithm. UPAL matches or exceeds state‑ofthe‑art performance in both point and line applications while significantly reducing computational cost, achieving, for instance, a 4x speedup and 10x smaller memory footprint over the ALIKED + DeepLSD pipeline. Code is publicly available at https://github.com/francois141/upal.

Authors:Roberto I. Ono Filho
Title: Interrupting the Loop: Periodic Subject Changes Raise Judged Surprise and Connection in Base Language Models
Abstract:
Where does the novelty a base language model produces with no task come from, and what can an LLM judge of a long stream actually see? We dismantle a cognitively inspired generation loop over 24 conditions on three base models. Most of its effect lives in one operation: a new subject injected every few hundred tokens (an interruption) into a stream whose literal repetition is damped (habituation). We judge windows of generated text only, with the premise as the unit (n=10) and a judge measured for repeatability, against a second judge family and against human readers. Under that protocol the interruption raises judged surprise by 1.2 to 1.4 points and connection by 0.8 over habituation alone. A connective that asks for continuity hurts; a bare paragraph break adds nothing detectable on fresh text; a reset context does at least as well as a kept one; and a pre‑registered replication on new premises confirms the primary contrast. Three things the window judge could not see changed the first version of this study, and we think they are of general use. The judge scores the experimenter's injected sentence as the model's own. A fixed rotation of injected sentences makes the model replay its earlier segments from beyond the judge's horizon, and the judge scores the replay as surprise and connection (65‑80% of post‑interruption windows at periods 150‑300). And the local gains do not compose: no arm produces an integrated document. The salience monitor, the in‑loop judge, memory across interruptions and a judge‑gated Review run with a gate that opens add nothing. On a problem with a verifier (online bin packing), the interruption multiplies valid, distinct candidate heuristics three‑ to fourfold without raising the quality of the best. We report an evaluation protocol for long generation and a controlled characterization of a simple intervention, not a mechanism of creativity.

Authors:Jia-Qi Lin, Yuangang Pan, Chang-Dong Wang, Haizhang Zhang, Ivor W. Tsang, Joey Tianyi Zhou
Title: Reliable Neural Collapse Approximation for Open-World Test-Time Adaptation
Abstract:
Test‑Time Adaptation (TTA) methods aim to bridge the domain gap between the source and target domains. However, traditional TTA methods become ineffective when the label distribution shift occurs, a challenge commonly referred to as an open‑world scenario. In this paper, we introduce a new method named Reliable Neural Collapse approximation (ReNC) for Open‑World Test‑Time Adaptation (OWTTA). Specifically, we leverage neural collapse as a structural prior for reliable target‑domain adaptation. Guided by this prior, we justify that the pre‑trained classifier weights can serve as the prototypes of the source domain. By measuring the similarity between samples and prototypes, we filter out the Out‑Of‑Distribution~(OOD) samples for reliable updates. Furthermore, we propose a neural collapse approximation mechanism to refine these prototypes, ensuring they can gradually adapt to the target domain while maintaining the neural collapse structure. Extensive experiments on several open‑world benchmarks demonstrate the superiority of the proposed method. Our empirical analysis suggests that ReNC better preserves NC‑related properties in the target domain, providing useful evidence for explaining reliable OWTTA and offering new insights for model design. Code is available at https://github.com/JiaqiLin‑AI/ReNC.

Authors:Umberto Cappellazzo, Xubo Liu, Stavros Petridis, Maja Pantic
Title: Listening Forward: Next Patch Embedding Prediction Enables Scalable Audio Learners
Abstract:
Self‑supervised learning (SSL) has driven substantial progress in audio representation learning, though existing methods have increasingly relied on elaborate pre‑training recipes to reach competitive performance. A markedly different pre‑training philosophy underpins the most influential progress in language modeling and, more recently, in visual representation learning: rather than train encoders as static feature extractors, models are trained to predict the next element, a discrete token or a continuous embedding, from the preceding context. Autoregressive prediction thereby provides a unified pre‑training interface that transfers across modalities, compelling the model to learn the underlying data distribution. We ask whether such a simple causal paradigm can yield strong audio learners, given that audio's temporal structure makes autoregressive prediction of patch embeddings a natural fit. We introduce NAPE (Next‑Audio‑Patch‑Embedding prediction), a self‑supervised framework in which a causal Transformer predicts each next patch embedding of a log‑mel spectrogram from the previous ones, using causal masking and stop‑gradient as its sole training signal. The design is intentionally minimalist, avoiding reconstruction decoders, acoustic tokenizers, student‑teacher setups, and auxiliary regularization losses. Across six audio and speech benchmarks, NAPE achieves state‑of‑the‑art fine‑tuning performance on several tasks, scales consistently across encoder sizes, and yields strong linear‑probing results. NAPE also produces structured attention patterns without explicit supervision.

Authors:Airin Akter Tania, Md Raihan Khan, Mohiuddin Ahmad
Title: AutoLumNet: Monotone Optimal Transport for Single-Shot Exposure Correction
Abstract:
Single‑shot exposure correction aims to map an arbitrarily degraded image‑‑‑whether under‑exposed, over‑exposed, or a spatial mixture of both‑‑‑to a well‑exposed output from a single capture. We present AutoLumNet, a framework that decomposes this task into a global monotone tone curve and a bounded local residual, making the global component the locus of formal guarantees. The tone curve is parameterized as the normalized cumulative integral of a strictly positive density, ensuring strict monotonicity by construction rather than by penalty. We prove that this parameterization (i)~preserves the pairwise luminance ordering of all pixels and all spatial extrema unconditionally, and (ii)~is dense in the space of valid tone corrections, containing the one‑dimensional optimal‑transport map from the input to any target luminance distribution. A differentiable sorted‑sample Wasserstein‑2 objective drives the learned curve toward the OT optimum during training. Spatially varying effects that the global map provably cannot address‑‑‑local shading, chrominance shifts, and clipped‑region restoration‑‑‑are handled by a bounded residual decoder with dual‑branch convex fusion, for which we provide an explicit sufficient condition for local order preservation. Experiments on five benchmarks (MSEC, SICE, LCDP, LOL‑v1, LOL‑v2‑real) show that AutoLumNet achieves state‑of‑the‑art PSNR and SSIM across both under‑ and over‑exposure regimes at 11.2\,ms per frame, and generalizes zero‑shot to pure low‑light benchmarks without retraining. To our knowledge, AutoLumNet is the first exposure‑correction method to unite structural monotonicity, optimal‑transport optimality, and bounded local adaptivity within a single trainable architecture. Code is available at https://github.com/kraihan/Autolumnet.

Authors:Silin Chen, Haoyi Teng, Xiaodong Gu, Yuling Shi, Jiale Huang, Yongpan Wang, Hongyu Zhang, Haibing Guan
Title: Repo0: Design-Driven Zero-to-All Code Generation
Abstract:
Large language model agents have made substantial progress in code generation, yet most existing systems assume a predefined repository architecture. This assumption does not hold in zero‑to‑all code generation, where an agent must construct an entire software project directly from natural‑language requirements while maintaining a modular repository architecture throughout development. We present Repo0, a continuous structural evolution framework for zero‑to‑all code generation. Repo0 maintains an explicit architectural state instantiated as a Dual‑Directed‑Acyclic‑Graph (Dual‑DAG), consisting of a requirement‑level DAG, a component‑level DAG, and their alignment relation. Starting from natural‑language requirements, it iteratively evolves component boundaries through structural actions guided by modularity metrics until structural convergence, after which the converged architecture guides test‑driven development code generation. We evaluate Repo0 on six real‑world repositories from RepoCraft using GPT‑5 mini and DeepSeek V3.2. Repo0 achieves the highest Functionality Coverage and Pass Rate across all settings. Compared with RPG, the strongest repository‑planning baseline, Repo0 improves Functionality Coverage by up to 20.08 percentage points and Pass Rate by up to 29.74 percentage points. Ablation and structural‑evolution analyses further demonstrate the importance of the Dual‑DAG architectural state, modularity‑guided structural evolution, and explicit structural convergence.

Authors:Kangdi Wang, Yusheng Dai, Jin Xu
Title: Fourier is Frontier: Frequency-Aware Autoencoding for High-Fidelity Music Reconstruction
Abstract:
Continuous‑latent audio autoencoders form the backbone of latent music generators, yet decoders at high compression rates commonly exhibit three failure modes: high‑frequency loss, phase incoherence, and stereo‑image collapse. These share a structural root: waveform autoencoders lack an explicit frequency axis, leaving no handle for targeted per‑band correction. Among five matched‑budget representations, the complex STFT achieves the lowest full‑band and high‑frequency spectral distances, providing direct access to magnitude and phase at every bin. Building on this, we present ear‑VAE2, a complex‑spectral autoencoder with cross‑channel interaction. Spec‑SnakeBeta learns a periodic activation per frequency bin with frequency‑dependent initialization, outperforming other activation variants while using fewer parameters than the fully independent variant. Duplex‑Aware Refiner applies band‑specific corrections to magnitude and phase following duplex theory of sound localization. On the 546‑track Song Describer Dataset, ear‑VAE2 achieves the best point estimates on five of seven reconstruction metrics. The Duplex‑Aware Refiner reduces Mel Distance by 19.4% and uses ~45% fewer residual‑output dimensions than the Unconstrained Refiner, while also lowering spectral distances, spatial‑cue errors, and receiving higher ratings from professional engineers. The downstream generator using ear‑VAE2 latents achieves better point estimates on all 12 automatic metrics.Demo page is available at https://eps‑acoustic‑revolution‑lab.github.io/EAR_VAE2/.

Authors:Tarun Kumar Garg, Vaanathi Sundaresan
Title: MOSAIC: Modality-agnostic Spectral Alignment for Federated Image-level Weakly Supervised Tumor Segmentation under Client-specific Missing Modalities
Abstract:
Trustworthy multimodal fusion in clinical settings requires handling incomplete and heterogeneous modality subsets across institutions, where privacy constraints prohibit centralized data sharing. Federated learning (FL) mitigates data‑sharing constraints but suffers from client‑specific missing modalities, where institutions possess incomplete multimodal subsets, degrading fusion quality and segmentation performance. While FL and weak supervision have been studied separately, their joint use with image‑level labels under heterogeneous missing modalities remains unaddressed. We propose MOSAIC, the first modality‑agnostic federated framework for weakly supervised binary tumor segmentation under client‑specific missing modalities. We introduce a client‑specific modality‑alignment module that fuses available channels into a shared latent space without prior knowledge of modality identity, a spectral prototype alignment loss that reconciles cross‑client distribution shift using compact non‑invertible frequency‑domain statistics, and a dedicated federated refinement network that denoises the resulting CAM pseudo‑labels into accurate masks, breaking the accuracy ceiling of weak supervision. Experiments on three multi‑institutional brain tumor benchmarks (FeTS2022, BraTS‑MEN, and BraTS‑SSA) demonstrate significant improvements over all image, box, and point‑supervised baselines, approaching fully supervised accuracy using only image‑level labels and reaching 0.84 Dice on FeTS2022. Dynamic new client addition enables previously unseen institutions to join an already‑trained federation within 0.01‑0.04 Dice without retraining. Code is available at https://github.com/Tarun2201/MOSAIC.

Authors:Julien Merand, Boris Meden, Liming Chen, Mathieu Grossard
Title: CoToGrasp: Contact-Topology-Conditioned Dexterous Grasp Synthesis via Canonical Workspace Learning
Abstract:
Current dexterous grasp planners primarily optimize for physical stability, focusing on whether an object can be grasped rather than how it should be grasped to support downstream functional tasks. However, conditioning grasp synthesis on specific human grasp taxonomies typically requires prohibitively expensive, object‑annotated datasets. To address these limitations, we propose CoToGrasp, a novel generative framework that synthesizes diverse, stable grasps strictly conditioned on specific contact topologies. To bypass the data collection bottleneck, CoToGrasp is trained entirely in an object‑agnostic manner. We introduce a feature‑based canonical workspace that projects local object features into a unified gripper‑centric domain, effectively decoupling the semantic functional intent from the arbitrary object geometry. By learning the intrinsic contact manifold of the gripper within this workspace, our model achieves zero‑shot generalization to unseen objects at inference. Extensive evaluations on the large‑scale DexGraspNet dataset demonstrate that CoToGrasp achieves state‑of‑the‑art performance, outperforming existing taxonomy‑guided planners. Finally, we demonstrate the physical viability and kinematic feasibility of our synthesized contact topologies on a physical robot platform. Code is available on our project website https://cea‑list.github.io/cotograspweb/ .

Authors:Maunil Shah, Vaanathi Sundaresan
Title: AsymFeX: A Symmetry-Driven Framework for Ischemic Stroke Segmentation Across Imaging Modalities and Stroke Stages
Abstract:
Fast and accurate segmentation of Acute Ischemic Stroke (AIS) lesions is essential for stroke prognosis and treatment planning. Non‑contrast CT (NCCT), the first‑line imaging modality for diagnosing ischemic infarcts, exhibits subtle infarct contrast, making manual delineation slow and labor‑intensive. Motivated by this, and by the clinical practice of comparing brain hemispheres to localize infarcts, we propose a two‑stage, nnU‑Net‑compatible 3D segmentation method. The first stage corrects head tilt to align each scan to its true anatomical mid‑sagittal plane; the second applies a novel Asymmetric Feature Extraction (AsymFeX) module, comparing each voxel to its true contralateral counterpart within a local 3 x 3 x 3 neighborhood via cross‑hemispheric attention, feature disparity estimation, and dual‑scale gating to capture both large and small infarcts. On AISD, our method achieves 0.6796 Dice, 23.53 mm HD95, and 7.69 mL AVD, significantly outperforming existing state‑of‑the‑art methods, with clinically relevant volumetric analysis at the 70 mL thrombolysis‑eligibility threshold. Proof‑of‑concept evaluation on ATLAS v2.1 and ISLES'24 demonstrates that the same symmetry‑driven design generalizes across imaging modalities and stroke time points without architectural changes, further supported by an uncertainty analysis assessing reliability under clinical deployment. Code is publicly available at https://github.com/biomedia‑lab/AIS‑detection.

Authors:Martin Tveten, Johannes Voll Kolstø, Per August Jarval Moen
Title: skchange: Fast and Flexible Algorithms for Changepoint Detection
Abstract:
Skchange is an open‑source Python library for detecting structural changes in time series. It implements modern change detection algorithms within a unified and extensible framework. The algorithms are modular and composable, and they include changepoint search methods based on both cost minimisation and statistical tests. Key features include the detection of anomalous segments in addition to changepoints; theoretically well‑founded fast and approximate search methods; theoretically well‑founded algorithms for high‑dimensional data, covering settings where either few or many features change simultaneously; utilities for automatic and data‑driven penalty calibration, which balances false alarms against missed detections; and a large collection of built‑in costs and statistical tests. The design follows established scikit‑learn conventions to streamline both user and contributor experience, and Numba is used extensively to achieve high computational performance. Source code and documentation are available at https://github.com/NorskRegnesentral/skchange.

Authors:Kang Liu, Suyan Li
Title: Finite-Horizon Input-Output Dynamics of Minibatch Perturbations in AdamW
Abstract:
A minibatch can influence training beyond the update at which it is observed because AdamW stores past gradient information in its optimizer states. We study this delayed effect through paired trajectories that differ only in one gradient update and share the same subsequent training sequence. We formulate AdamW as a finite‑horizon input‑‑state‑‑output (ISO) system whose state contains the model parameters and first‑ and second‑moment estimates. Linearizing the joint dynamics yields a signed response operator that maps a localized gradient perturbation to its future loss effects, revealing how optimizer memory shapes their magnitude, timing, and sign. We further derive an exact multistep error decomposition and establish first‑order finite‑horizon accuracy under local smoothness and controlled activation switching. Experiments validate the response mechanism and optimizer‑state effects, while repeated‑future analyses reveal substantial prospective structure in delayed influence that can be partially recovered from ISO approximations. Code is available at https://github.com/Kanyooo/Loss_ISO.

Authors:Julien Merand, Boris Meden, Mathieu Grossard, Liming Chen
Title: GOAG: Generative and Object-Agnostic Grasp Planner for Dexterous Robotic Manipulation
Abstract:
Multifingered grasping is a crucial robotic skill, but current deep‑learning grasp planners often struggle to generalize to new objects because they are trained on limited, object‑specific datasets. We introduce a fundamentally different approach, grounded in the observation that the gripper and the object share identical surface geometry at their mutual contact points. We propose GOAG: Generative and Object‑Agnostic Grasp Planner for Dexterous Robotic Manipulation, a novel deep generative model that learns a compact latent representation of a specific gripper's contact surface distribution, enabling the efficient sampling of valid grasp configurations without relying on object‑specific training data. We show that by introducing object features only at inference time, our model can effectively retrieve admissible contact areas that are compatible with the gripper's capabilities. We validate our approach through extensive experiments on established grasp protocols in both simulated and real‑world scenarios, demonstrating its effectiveness with different grippers from the literature. Our method delivers state‑of‑the‑art results on the objects from the MultiDex dataset, achieving an average success rate of 86.93%. It offers significantly faster processing when generating numerous grasps, while matching the performance of leading approaches specifically trained on this dataset. Unlike these methods, our approach does not rely on object‑specific training data, highlighting the advantages of object‑agnostic learning. It effectively addresses the generalization challenges faced by traditional data‑driven grasp planners. Code and videos are available on our project website https://cea‑list.github.io/goagweb/ .

Authors:Nicolò Savioli
Title: Gallileo-4D: Frozen Backbone Ensemble for Dynamic 4D Reconstruction
Abstract:
We describe our entry to the PhysAI Dynamic 4D Reconstruction Challenge, which placed third of 27 teams at 0.58356 APD on the final leaderboard, without a single gradient update. This was not the plan: of thirteen fine‑tuning configurations of a pre‑trained 4D backbone, twelve degraded the challenge score, and eleven of those twelve improved local validation at the same time. We trace this inversion to the structure of the benchmark: only 25% of the evaluation set belongs to the data variant released for training, so updates that fit the available data damage the pre‑trained features the remaining 75% relies on. Our system therefore freezes the backbone and spends its budget at inference time, fusing three decoding configurations ‑‑ temporal stride‑3, horizontal‑flip test‑time augmentation, and dense stride‑1 ‑‑ under a convex weighting. The ensemble recovers +0.041 APD over the frozen baseline, more than any training run achieved, at zero training cost.

Authors:Zhuochun Li, Youngmin Ko, Ali Keramati, Nicola Ferri, Susana Palmaz Lopez Pelaez, Liang-Chun Tsai, Calvin Wang, Mirco Milletari, Tuhin Kundu, Vadim Smolyakov, Kjartan Olafsson, Tommy Guy
Title: One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows
Abstract:
Recent agent benchmarks increasingly ground evaluation in executable environments, from code repair to web navigation, app APIs, and function calling. Yet completing consequential work beyond code requires more than producing a plausible response or valid tool call: agents must gather missing information over multiple turns, follow domain policies, coordinate dependent tools, and realize the correct persistent state transition without collateral effects. In this paper, we introduce Thinkingbox, a sandbox for tool‑agent‑user interaction that provides isolated MCP‑compatible tool sessions, complete execution traces, and outcome evaluation over terminal backend state. Built on this sandbox, Thinkingbox‑bench contains 507 policy‑conditioned workflows across numerous scenarios, including retail, hospitality, auto insurance, neobank internal IT, and consulting IT/HR support. Each attempt is evaluated by task‑specific executable checks that accept valid trajectories while rejecting wrong, missing, or extra effects; designated tasks additionally check required properties of the final response. Across proprietary and open‑weight models, the strongest achieves 65.36% pass@1, but only 25.25% pass^20. Moreover, many failed trials show clean termination and valid state‑changing actions, showing that response or tool‑call‑level signals are not clear proxies for end‑to‑end task completion. Thinkingbox‑bench reveals a large gap between occasionally finding a successful trajectory and reliably completing stateful business tasks. We release both Thinkingbox and Thinkingbox‑Bench: https://github.com/microsoft/thinkingbox

Authors:En Zhi Tan, Jia Xiang Lim, Bryan Lijie Chew, Tze Minh Ng, Benjamin Yan Han Yap
Title: RecPFN: Prior-Fitted Networks for In-Context-Based Recommendations
Abstract:
We introduce RecPFN, a prior‑fitted network that brings in‑context learning to sequential recommendation. RecPFN is pretrained entirely on synthetic clickstream environments sampled from a broad structural causal prior, enabling it to amortize Bayesian‑style inference from a small support set. At inference, a lightweight decoder‑only transformer conditions on a handful of domain sequences and produces next‑item predictions for queries in a single forward pass, without any weight updates. Across eight public benchmarks, RecPFN achives state‑of‑the‑art zero‑shot performance while remaining strongly competitive with supervised methods in low‑compute and low‑data regimes. It is deployment‑efficient and robust to domain shift, outperforming strong zero‑shot baselines that rely on large real‑interaction corpora. RecPFN provides a practical path toward generalizable, data‑efficient recommenders and opens avenues for richer priors, longer‑context ICL, and multimodal extensions. Code for training and evaluation is publicly available at https://github.com/SAP‑samples/tabular‑ai‑recpfn/.

Authors:Mohan Chen
Title: Loreley: Repository-Scale Program Evolution with Quality-Diversity Search
Abstract:
Sequential agent search accumulates changes from its current champion but discards alternative branches; independent proposals preserve breadth but restart from the root. Loreley instead retains complete repository states in a Quality‑Diversity (QD) archive and samples them as parents or supplies them as context for later edits. Candidates are Git commits produced in isolated worktrees and judged by a project‑supplied evaluator. We compare configured Loreley QD, sequential champion editing, and independent root proposals in a matched Zstandard experiment: seven paired blocks and 48 physical candidate jobs per policy and block (1,008 total), with root‑only initialization and each policy's native concurrency. Validation selected a winner at each budget checkpoint; an agent‑hidden holdout measured the fixed candidate. At 48 jobs, QD was 0.135% below Sequential Champion (95% BCa interval for the paired effect: ‑0.556% to +0.161%) and 0.320% above Independent Root (‑0.082% to +0.686%). Neither contrast established a QD advantage; Sequential had the highest observed 48‑job mean and median. Archive retention and later sampling did occur. Four of seven final QD winners had a non‑incumbent state in their primary‑parent ancestry under a retrospective one‑incumbent rule applied only to the observed QD stream. Including inspiration edges raised the count to six, without showing that supplied context caused an edit. Three earlier capability campaigns produced generation‑4, multi‑file improvements in two Python libraries and a separate Zstandard revision. Loreley engaged the intended stepping‑stone mechanism, but the controlled experiment did not show an endpoint benefit at 48 jobs.

Authors:Johannes Künzel, Peter Eisert, Anna Hilsmann
Title: RIPE++: Reinforced Keypoint Learning from Positive Pairs Only
Abstract:
Sparse keypoint extraction and matching underpin core tasks in geometric computer vision, including structure‑from‑motion, visual SLAM, augmented reality, and medical image registration. Learning robust local feature representations, however, typically requires accurate camera poses or depth supervision, which are often unavailable in real‑world settings. Reinforcement learning (RL) has recently emerged as a promising alternative, requiring only the information if two images show the same scene or not. However, existing RL formulations such as RIPE rely on coarse binary rewards and carefully constructed negative training pairs, limiting training stability and descriptor discriminability. In this paper, we revisit RL‑based keypoint learning and propose a reward that fully exploits the geometric consistency signal, deriving both reward and penalty from a single positive pair without contrasting against negatives. This richer signal provides sufficient supervisory contrast to learn discriminative detectors and descriptors from positive image pairs alone, enabling representation learning under extremely limited supervision. Furthermore, we show that the same RL objective can be extended to the matching stage by adapting LightGlue, raising AUC@5 on MegaDepth1500 from 56.58 to 59.65 and enabling weakly‑supervised training of the full sparse matching pipeline from image pairs with partial visual overlap. We validate our approach on established benchmarks, demonstrating competitive results compared to fully‑supervised methods. We further show that the method can be even trained on low texture medical video sequences, where camera poses are usually unavailable and standard SfM pipelines often fail. Code and data are available at https://github.com/fraunhoferhhi/RIPEpp .

Authors:Yichu Fang, Sitong Wei, Haozhe Hu, Xiaoyu Shen
Title: ReCache: Efficient KV Cache Reuse and Compression for Tool-Augmented LLM Agents
Abstract:
Agentic language models repeatedly encode tool and skill schemas that recur across requests in different combinations and orders, preventing standard prefix caching from reusing their key‑‑value (KV) states. We introduce ReCache, a framework for independently caching resource representations while reducing their inference‑time computational and memory overhead. Resource‑wise attention removes cross‑resource interactions and assigns resource‑local positions, producing composition‑invariant KV blocks. ReCache then restricts resource visibility to contribution‑selected layer‑‑KV‑head‑group routes and retains only invocation‑critical fields through structural and semantic pruning. We evaluate ReCache on a benchmark assembled from seven public tool‑ and skill‑use datasets, including resource‑disjoint tests. Resource‑wise attention matches dense invocation performance (82.3% versus 82.4% Inv‑F1) while providing a 3.655× time‑to‑first‑token speedup. The complete framework reduces allocated KV‑tensor memory by 92.43% and accelerates attention by 1.423×. These results show that separating reusable schema encoding from selective resource access substantially reduces agentic inference costs with limited effectiveness loss. The code is available at https://github.com/EIT‑NLP/ReCache.

Authors:Josias Moukpe, Priyanka Aryal, Matthew Kenney
Title: DeltaML-Bench: Evaluating Machine Learning Agents on Real-World Research Repositories
Abstract:
Autonomous agents for machine learning experimentation must navigate heterogeneous repositories, repair training pipelines, and evaluate candidate improvements under realistic compute constraints. Existing benchmarks only partially capture these conditions. We introduce DeltaML‑Bench, a benchmark comprising 48 tasks sourced from research papers that require agents to improve published baselines within imperfect, open‑source repositories. We evaluate GPT‑5 and Claude Sonnet 4 with a standard Modular agent and a search‑based ARG scaffolding. In the 4 x 6h allocation, ARG raises GPT‑5's per‑run success rate from 9.4% to 33.9%; in the 2 x 12h allocation, GPT‑5 ARG reaches 49.0%. Modular configurations exhibit specification gaming rates as high as 47.9%, while no gaming is observed in the evaluated ARG configurations. These results indicate that scaffolding design and integrity checks are important considerations when deploying agents for autonomous ML experimentation.

Authors:Yiwei Li, Jiannong Cao, Weixun Gao, Rui Cao, Songye Zhu, Yinfeng Cao, Mingjin Zhang
Title: S$^2$GS: Structured Sparse Gaussian Streaming for Efficient Free-Viewpoint Video Reconstruction on Edge-IoT Devices
Abstract:
Streaming reconstruction of Free‑Viewpoint Videos (FVVs) supports immersive Internet of Things (IoT) services, such as telepresence and digital twin visualization. Existing methods suffer from high per‑frame optimization time and large storage footprints, limiting deployment on resource‑constrained Edge‑IoT devices. To address these challenges, we propose Structured Sparse Gaussian Streaming (S^2GS), an FVV reconstruction framework that exploits structure‑aware temporal sparsity to selectively update Gaussian residuals, enabling efficient streaming without compromising visual fidelity. In the spatial domain, a streaming octree hierarchically organizes Gaussian residuals, capturing spatial correlations that guide residual updates. In the temporal domain, a structured gating mechanism, comprising hierarchical feature propagation (HFP) and Gumbel‑Sigmoid sampling, converts hierarchical dynamic cues into sparse residual update decisions under differentiable optimization. A multi‑level discrete scheme is further adopted to provide fine‑grained control over residual updates while preserving intricate dynamic details. Extensive experiments across consumer GPUs, industrial edge IoT devices, and a physical telepresence testbed demonstrate that S^2GS consistently reduces per‑frame optimization time and storage footprint while maintaining competitive visual quality. Compared with QUEEN, S^2GS reduces per‑frame optimization time by 59% and storage costs by 85% on an RTX 4090 GPU. On the Jetson AGX Orin, S^2GS delivers the highest rendering throughput (60+ FPS) and the lowest energy consumption among the evaluated methods, demonstrating its potential for deployment in resource‑constrained systems.

Authors:Xizhou Bu, Qingda Hu, Lei Zhou, Lingfeng Zhang, Yingbo Tang, Zihao Liu, Xinyi Tao, Zhiqiang Ma, Qingqiu Huang, Chufeng Tang, Hongbo Wang, Jing Zhang, Jiayi Ma, Hangjun Ye, Wei Li, Xiaoshuai Hao
Title: What Matters for Latent Actions in Robot Learning
Abstract:
Latent Action Models (LAMs) have emerged as a promising paradigm for enabling robot learning to leverage large‑scale unlabeled videos through latent actions that serve as compact surrogates for physical actions. Despite rapid progress, research on LAM remains highly fragmented, with existing methods evaluating different design choices in isolation under inconsistent experimental settings, making it difficult to identify the factors that truly determine downstream robotic manipulation performance. In this work, we present the first comprehensive empirical study of latent action learning for robotic manipulation. We unify representative LAM methods within a common autoencoding framework and systematically investigate 41 LAM design choices across three dimensions, including latent action modeling paradigms, learning objectives and regularization methods, and latent action integration strategies. We further examine four proxy metrics for evaluating latent action quality and assess their ability to reliably predict downstream robotic manipulation performance. Extensive experiments on three widely used benchmarks provide strong empirical evidence that fine‑tuning vision‑language model (VLM) backbones with latent actions provides a stronger initialization for downstream policy learning, with further validation on real‑world robot manipulation tasks.

Authors:Brian Ward
Title: DraftFM: A FoundationModel for Day-Zero Drafting in Magic: The Gathering
Abstract:
Drafting a new Magic: The Gathering expansion begins before any pick from it has been observed: the complete card list is public, but the draft logs that supervised pick models train on do not yet exist. We study this day‑zero regime directly. DraftFM is a discrete‑choice policy that scores exactly the cards available in the current pack, conditioned on the drafted pool and the state of the draft. Every card enters as a frozen 775‑dimensional function of its public card record, structured features and a fixed text embedding, with no card identities, set identities, or usage statistics anywhere in the model, so an unseen card is scored by the same machinery as a familiar one. A 1.6‑million‑parameter network fitted on 149 million human picks from 29 expansions predicts held‑out picks in three expansions withheld in their entirety, reaching 50.8%, 60.4%, and 56.7% top‑1 agreement, where uniform chance at the opening pick is about 7%. Refitted on all 32 observed expansions, the same architecture produced a card ranking for the then‑unreleased set The Hobbit, sealed with its complete cryptographic provenance and published roughly 36 hours before the set became draftable on MTG Arena. The sealed ranking agrees with six independent expert reviewers roughly as much as those reviewers agree with one another. Evaluation against realized outcomes is committed to a follow‑on note, whatever it shows.

Authors:Bowen Cui, Weijie Wang, Zeyu Zhang, Yefei He, Mingda Lin, Haoyu Zhao, Yuanyu He, Donny Y. Chen, Feng Chen, Bohan Zhuang
Title: Block3D: Efficient Text-to-3D Generation via Block-Wise Diffusion
Abstract:
While text‑to‑3D generation has advanced rapidly, achieving high geometric fidelity at low inference cost remains challenging. Existing text‑to‑3D methods either decode discrete shape tokens autoregressively or iteratively refine global 3D representations with diffusion or flow models. However, autoregressive decoding is sequential and cannot revise errors, whereas diffusion and flow‑matching models repeatedly process the full representation, making high‑quality generation increasingly expensive. In this paper, we propose Block3D, a block‑wise diffusion framework that partitions the discrete shape‑token sequence into contiguous blocks, generates the blocks autoregressively, and jointly denoises all tokens within the current block. To alleviate error accumulation, we introduce confidence‑guided intra‑block correction, which revises low‑confidence tokens before each block is finalized. On a held‑out set from TRELLIS‑500K, Block3D reduces mean end‑to‑end generation time from 25.71 seconds to 4.99 seconds, achieving a 5.15× speedup over the fine‑tuned autoregressive baseline without sacrificing geometric fidelity.

Authors:Yuanhao Ban, Jiaqi Feng, Hengguang Zhou, Xiaohuan Pei, Justin Cui, Cho-Jui Hsieh
Title: Stream4D: 4D-Consistency for Streaming Autoregressive Diffusion Video Models
Abstract:
Streaming autoregressive diffusion models enable real‑time, long‑horizon video generation, but their training objectives optimize local frame prediction rather than the geometry and dynamics of a coherent world: long rollouts accumulate geometric drift and degrade into static or unnatural motion. Recent bidirectional approaches address this problem using rewards signals built upon 3D Gaussian‑Splatting reconstruction. However, a single rigid 3d reconstruction cannot model a dynamic scene, so this critic penalizes genuine object motion as reconstruction error and is maximized by freezing the video. This shortcut is especially detrimental in the AR setting, where each chunk can propagate an already‑static configuration. In this work, we propose Stream4D, which replaces the static critic with a feed‑forward 4D reconstruction reward that explicitly models scene dynamics, allowing coherent motion to receive high consistency rewards. To further guide motion magnitude and quality, we add a motion prior that rewards natural scene‑flow magnitude while penalizing jitter and non‑rigid artifacts. Our final recipe combines these two terms with a lightweight perceptual anchor. Across various autoregressive video backbones and various generation horizons, Stream4D improves 4D reconstruction quality, preserves motion more effectively, and achieves higher human‑aligned preference. Project page: https://banyuanhao.github.io/Stream4D/

Authors:Su Yan, Rakesh Iyer
Title: When Machines Speak: A Unified Generative Framework for Integrating Machine-Native Symbols into Pretrained Large Language Models
Abstract:
Many real‑world AI systems represent entities, behaviors, and structured information using discrete machine‑native symbols rather than natural language. While these representations are compact and preserve task‑relevant structure, they lie outside the linguistic token space of pretrained large language models (LLMs), creating a fundamental divide between language modeling and structured prediction. We introduce UniLang, a unified generative framework that bridges this divide by extending pretrained LLMs to treat machine‑native symbols as first‑class generative units alongside natural‑language tokens. UniLang expands the LLM's vocabulary and embedding space with grounded machine‑native representations, enabling textual and symbolic tokens to be jointly modeled and generated under a single autoregressive objective. This unified interface allows pretrained LLMs to directly operate on machine‑native representations without requiring them to be verbalized as natural language or relying on task‑specific architectures. We evaluate UniLang on two structurally distinct tasks, sequential recommendation and legal precedent prediction, spanning different domains and types of structured prediction. Across both tasks, UniLang consistently outperforms strong baselines, demonstrating a path toward extending pretrained LLMs beyond language and using them as a common generative modeling backbone for heterogeneous machine‑native representations.

Authors:Kyle Bierly
Title: The Normal Procrustes Problem: A Riemannian Optimization Approach
Abstract:
For given m × n data matrices X, Y, we investigate the Normal Procrustes Problem‑‑‑the least squares optimization problem that aims to minimize \|AX‑Y\|_F^2, where A is constrained to be a normal m × m matrix. As far as the author of this article is aware, no other method that attempts to solve the Normal Procrustes Problem exists in the literature; we thus propose what is, to our knowledge, the first such method. We, furthermore, adapt our approach to address the Real Normal Procrustes Problem, where A must be real. In our treatment of these problems, we first reduce our complex and real objective functions to be purely optimizable over the Riemannian manifolds of the unitary and real orthogonal matrices, respectively. This reduction enables us to apply techniques in Riemannian manifold optimization to approximate solutions to both. The Closest Normal Matrix and Real Closest Normal Matrix Problems are both special cases of their respective Procrustes Problems and have been previously studied in the literature. Our approach thus recovers a novel Riemannian optimization method for approximating solutions to both these problems. We further numerically test the performance of our method across all such problems (including against previously developed algorithms on the Closest Normal Matrix Problems) and obtain competitive residuals and favorable scaling in wall‑clock time.

Authors:Sawan Dasari
Title: When to Retrain: An Empirical Study of Retraining Policies for Streaming ML Under Concept Drift, Budget, and Latency Constraints
Abstract:
Production machine learning systems degrade under concept drift, yet practitioners have little principled guidance on when to retrain. Retraining is costly, retraining budgets are finite, and a retrained model does not take effect instantly: training and deployment latency leave a stale model serving predictions while the data continues to move. We present a controlled empirical study of three practical model‑refresh policies (periodic retraining, error‑threshold triggering, and statistical drift‑triggered retraining with ADWIN) against a no‑retrain baseline, evaluated under a unified system model that makes retraining budgets and training‑plus‑deployment latency explicit. Across 3,933 experiment runs spanning three drift regimes, three budget levels, up to five latency levels, three datasets, and two learning modes, we find that the single most consequential design decision is not the retraining policy but whether the deployed model learns incrementally. With per‑sample incremental updates, and for the linear online learner with immediate labels studied here, no policy differs from the no‑retrain baseline by a practically significant margin in any of 54 paired comparisons, even at extreme latency. Without incremental updates, policy choice separates outcomes by 15‑55 percentage points of post‑drift accuracy, and simple periodic retraining significantly outperforms both reactive policies under abrupt and gradual drift, while reactive policies retain an advantage only under recurring drift. We document systematic failure modes of reactive policies and a latency‑budget queueing interaction that silently halves effective retraining budgets, and release the full simulator, dataset pipelines, and per‑run artifacts for reproducibility.

Authors:Georgii Kliukovkin
Title: The Lazy Pod That Lies: Deferred Cost and Failure Semantics of Lazy Container Image Pulling for Model Serving on Kubernetes
Abstract:
Lazy container‑image pulling promises to eliminate the dominant cost of starting a model‑serving pod by mounting the image immediately and fetching content on demand. We evaluate this promise for model delivery on Kubernetes, using KServe with two production lazy‑pulling systems ‑‑ eStargz/stargz‑snapshotter and AWS SOCI ‑‑ against eager baselines, on artifacts from 2 to 140 GB including real fp16 weights. Lazy pulling delivers its headline: cold time‑to‑first‑prediction becomes size‑independent (16.9‑‑17.6s, versus 24.5‑‑573.0s eager). But the cost is deferred, not eliminated: a full read of a 14 GB model through the lazy mount takes 105.3s, slower than the 72.4s eager pull it replaced, and the two systems pay at opposite lifecycle ends (SOCI prefetches nearly the full image before Ready; eStargz defers nearly everything to first read). More consequentially, we characterize a failure mode eager pulling structurally cannot exhibit: under sustained legitimate reads with default configuration, the snapshotter's node‑level cache exhausts its finite volume and already‑running pods begin failing reads of model files. At the earliest stage of exhaustion, an instrumented serving pod passed every Kubernetes‑visible and application‑level check for 196s while its snapshotter was already logging real failures; under heavier pressure, 67‑‑94% of model files fail, scaling monotonically with residual cache occupancy. A live pod self‑heals if cache space is freed under it, but a snapshotter‑daemon restart under a live pod leaves permanently stale file handles in a pod still reported Running. We derive placement, monitoring, and cache‑sizing guidance for serving platforms and operators.

Authors:Emanuel C. A. Valente, Lourenço A. P. Júnior, Leonardo Gonçalves Chahud, Júlio Cezar Estrella, Marcus Botacin
Title: Aray: Deterministic-First Synthesis of Benign Artifacts for YARA Validation
Abstract:
A YARA rule is easy to distribute, but the malware sample used to demonstrate a positive match is not. This complicates storage, continuous integration, disaster‑recovery exercises, and reproducible scanner validation. Constructing a replacement fixture requires more than embedding literals: YARA conditions can combine alternatives, counts, offsets, integer reads, and executable‑container constraints, while the resulting file should not reproduce malware behavior. Positive validation is existential: it requires one file‑level member of a rule's match set, not reconstruction of the originating sample. We present Aray, a deterministic‑first YARA interpreter and positive‑fixture synthesizer. Models may propose constructive normalizations or typed extraction fallbacks, but never backend source or binary structure. Conventional code validates normalized rules, derives string and integer witnesses, and performs extraction, routing, collision‑checked layout, and ELF, PE, or generic serialization. Only residual normalization semantics reach a bounded model judge. We evaluated Aray over 416 public‑rule entries. Normalization accepted 182 entries without model assistance and 234 after model normalization. Constructibility preflight admitted 406 entries, and every admitted fixture matched its upstream original rule. This yields 406/416 (97.6%) overall and 406/406 among constructible rules, with ten expected preflight dispositions and no scanner mismatches or construction failures. An unreachable endpoint confirmed zero model invocations during realization. The original‑rule oracle validates generated fixtures against their source rules; proving implication for all possible files is a separate, stronger objective. Two anchored‑regex failures were repaired before the final run, so these are post‑fix systems results, not a held‑out estimate.

Authors:Sparsh Garg, Yi-Wen Chen, Vijay Kumar B G, Abhishek Aich
Title: CAViAR: A Causal Video Dataset for Fine-Grained Accident Reasoning in Real-World Scenarios
Abstract:
While modern autonomous driving systems excel at perception tasks such as object detection and trajectory prediction, they lack the high‑level causal reasoning required to interpret traffic accidents. In particular, determining responsibility, such as identifying who is at fault and which traffic rule was violated, remains largely unexplored in current benchmarks. To this end, we introduce CAViAR (Causal Accident Video and Incident Analysis Repository), a human‑annotated dashcam benchmark comprising 2,249 real‑world accident videos collected from CarCrashDataset (CCD) and Nexar. Each video is annotated with structured labels spanning environmental conditions, accident type, causal explanation, apparent At‑Fault Agent, affected agent, and apparent rule‑violation category. We benchmark state‑of‑the‑art vision‑language models (VLMs), including Cosmos‑Reason2, Qwen3‑VL, and InternVL3. Once class imbalance is accounted for with majority/random baselines and balanced metrics, perceptual competence is uneven‑‑lighting is nearly solved, whereas weather and road‑condition accuracy fall at or below the majority‑class baseline‑‑‑and all models degrade sharply on accident type and responsibility reasoning. Overall, CAViAR exposes a practical Perception‑‑Reasoning Gap: current VLMs may recognize salient context, but do not reliably map visible agent actions to annotated rule‑relevant responsibility categories in safety‑critical driving scenarios. Code, annotation schema, prompts, and evaluation scripts are available at: https://github.com/nec‑labs‑ma/CAViAR

Authors:Yihan Xie, Hanwen Cui, Runze Ye, Juekai Lin, Haoyang Wang, Jinhao Mao, Bo Zhang, Wenqiao Zhang, Xiaogang Guo, Jun Xiao, Lei Zhang
Title: Holtercare-Bench: A Multimodal Benchmark for Evaluating Long-Term Dynamic ECG Analysis
Abstract:
While multimodal large language models (MLLMs) excel in medical applications, most of them favor static images or short‑term signals. In the critical field of dynamic electrocardiograms (ECG), models struggle with complex temporal reasoning and diagnostic report generation due to a lack of high‑quality datasets and benchmarks. To address this, we introduce (i) Holtercare‑23K, a large‑scale multimodal dynamic ECG dataset comprising 22,980 QA pairs derived from 788 clinical Holter records and featuring a novel signal‑video‑text tri‑modal alignment. Based on this dataset, we present (ii) Holtercare‑Bench, a multimodal benchmark that evaluates models on temporal localization, clinical diagnosis, and global summarization. Zero‑shot evaluations of leading MLLMs reveal a significant performance gap in processing ultra‑long pathological sequences. However, fine‑tuning representative models yields substantial improvements. This work illuminates the limitations of current MLLMs in electrophysiology and provides a foundational benchmark for long‑term medical MLLMs. Our project is available at https://github.com/ZJU4HealthCare/Holtercare‑Bench.

Authors:Kyriakos "Rock" Lambros, Steve Wilson
Title: Incident-Data Robustness Analysis of the OWASP Top 10 for LLM Applications (2026): How a Community-Expert Ranking Holds Up Against a Large-Scale LLM Incident Corpus
Abstract:
The OWASP Top 10 for LLM Applications ranks the risks that a community of security practitioners judges most important. We ask a narrower question: checked against the record of real incidents, does that expert ranking agree with the data? We assembled a large‑scale corpus of LLM‑security incidents (7,714 snapshotted and 6,639 labeled against the 20‑entry taxonomy) drawn from CVE, GHSA, OSV, and AIAAIC, and derived an incident‑based ranking with a Bayesian measurement‑error model that corrects each category's count for classifier precision and recall. The 2026 candidate list blends the two signals at fixed weights, 0.75 on the expert vote and 0.25 on the data, so the corpus corrects the consensus without overturning it. The agreement between the two rankings is weak: Cohen's κ\approx 0.20, with a 90% interval that crosses zero. The expert ranking is nonetheless robust. A pre‑registered bake‑off of four frontier classifiers returns no winner. None beats the incidence floor's balanced accuracy of 0.863. A ground‑truth check leaves the floor's ordering (Spearman ρ= 0.918 against held‑out truth) in place. This is an exploratory analysis by two working‑group members, not the official OWASP release, and it does not supersede the official list or process.

Authors:Aniket Wattamwar, Manav Anandani, Mrunal Kakirwar
Title: The Evaluation Context Protocol (ECP): A Portable Contract for AI Agent Evaluation
Abstract:
The evolution of artificial intelligence has necessitated a fundamental shift from evaluating isolated Large Language Models (LLMs) to assessing autonomous agentic architectures. This paper explores the critical methodologies for evaluating AI agents and the essential role of advanced observability infrastructure. We analyze the architectural components of agents and identify the severe limitations of current evaluation paradigms, including benchmark exploitation, the "confidently wrong" phenomenon, and the discrepancy between theoretical capability and operational reliability. To begin addressing the fragmentation in current evaluation infrastructure, this paper proposes the Evaluation Context Protocol (ECP), an early‑stage, vendor‑neutral framework intended to act as a portable evaluation contract layer for agentic systems. In its current form ECP defines a small JSON‑RPC interface over which an agent exposes its user‑visible output, the tool calls it made, and evaluator‑safe audit context, and against which programmatic checks can be run uniformly across frameworks and continuous integration systems. We describe an open‑source reference implementation that includes adapters for LangChain, LlamaIndex, CrewAI, and PydanticAI, and we situate the design against failure modes documented in the recent literature. ECP is presented as work in progress rather than a finished standard: the evaluation surface, method set, and grader families are all expected to change as the protocol is exercised against more systems, and the empirical validation required to justify adoption is outlined as future work.

Authors:Baichuan Mo, Zhengzhong Ricky You, Xiqun Michael Chen, Ruimin Li
Title: TorchDCM: A Unified PyTorch-Native Package for Discrete Choice Modeling
Abstract:
Estimating large and simulation‑intensive discrete choice models (DCMs) requires repeated evaluation of utilities, probabilities, derivatives, and simulated likelihoods over many observations, alternatives, and draws. Existing DCM software provides mature econometric workflows, while recent GPU‑oriented tools accelerate selected models, leaving a gap between econometric coverage and scalable differentiable computation. We introduce TorchDCM, an open Python package for discrete choice modeling that compiles choice data and model specifications into a unified PyTorch‑native likelihood engine for estimation, inference, prediction, and structured reporting on CPU or CUDA devices. The package covers the principal econometric functionality available across Biogeme and Apollo, including multinomial, nested, mixed, ordered, latent‑variable, and panel likelihoods. It also supports ragged choice sets, constrained parameters, covariance estimation, willingness‑to‑pay analysis, elasticities, and extensible likelihood components. We evaluate TorchDCM against seven other estimation packages in aligned synthetic and real‑data full‑estimation experiments. TorchDCM completes all 45 synthetic cases, runs fastest in every comparable synthetic case, and satisfies the prespecified final‑log‑likelihood tolerance in every comparison with at least two comparable solutions. More precisely, it reduces median runtime by 89.1%‑99.7% relative to Biogeme and Apollo across model‑data settings. CUDA provides an additional 12.0‑71.0x speedup over single‑core TorchDCM. These results establish a scalable and reproducible foundation for econometric estimation and differentiable choice‑model development. The open‑source package and executed examples are available at https://github.com/mbc96325/torchdcm.

Authors:Bo Liu, Simon Yu, Yiding Jiang, Ao Qu, Andrew Zhao, Zichen Liu, Junsu Kim, Zijian Zhou, Seungone Kim, Tongzheng Ren, Mickel Liu, Hanfei Yu, Zhaorun Chen, Weiyan Shi, Paul Pu Liang, Luke Zettlemoyer, Yejin Choi, Natasha Jaques
Title: SPADE: Self-Play in Adaptive Synthetic Executable Environments
Abstract:
Continuous self‑improvement requires an ever‑expanding pool of self‑generated, diverse, adaptive goals. For language agents, existing training environment pools (hand‑curated, statically synthesized, or frozen‑verifier) keep the goal distribution fixed as the learner scales. We introduce SPADE (Self‑Play in Adaptive Synthetic Executable Environments), a self‑play RL framework in which a single LLM plays two roles: an Environment Designer that writes complete, long‑horizon training environments as executable code with an OpenAI Gym‑style reset()/step() interface, and a Reasoning Agent that learns to act in them. Each is a stateful, multi‑turn environment (state transitions, reward functions, and verification code), so one interface spans reasoning problems and multi‑step agentic tool use. The Reasoning Agent's regret is estimated using the gap between its reward with and without privileged hints; in optimizing this regret signal the Environment Designer learns to target environments at the edge of the agent's capabilities while keeping them feasible. Through extensive experimentation, we find several components critical to success: grounding the Environment Designer on documents sampled from a large pretraining corpus, and giving it an accumulated environment memory. Scaling to 30B‑parameter models, SPADE improves over the strongest fixed‑environment baseline by +5.3 on average across eight held‑out math, science, code, and reasoning benchmarks, and lifts the tool‑use setting by +5.7 on BFCL‑v4 multi‑turn and +13.9 on ACEBench‑Agent; on the games setting, the margin over the strongest baseline grows with model scale. By making environment design itself a learnable component, SPADE takes a concrete step toward open‑ended self‑improvement.

Authors:Zhu Zhang, Jixun Wang, Xiaoang Xu, Xiaorong Wang, Zihan Zhou, Zhiyuan Wang, Shuo Wang, Chaojun Xiao, Yuezhi Zhou
Title: Beyond Teacher Likelihood: Group-Calibrated On-Policy Distillation for Long-Context Reasoning
Abstract:
On‑policy distillation (OPD) trains a student on its own responses using dense token‑level guidance from a stronger teacher. In long‑context tasks, however, token‑level teacher support can favor locally plausible responses that omit evidence distributed across the input or violate global task constraints. Task‑specific verifiers, in contrast, evaluate task completion at the response level and may return graded rewards that reflect partial success. We diagnose this mismatch on fixed responses from two representative long‑context evidence‑aggregation tasks. Across longer input ranges, trajectory‑level OPD scores become progressively less aligned with verifier rewards, indicating teacher‑verifier disagreement. Motivated by this observation, we introduce Group‑Calibrated On‑Policy Distillation (GC‑OPD). GC‑OPD separately normalizes verifier rewards and trajectory‑level OPD scores within each rollout group and uses their difference as a signed teacher‑verifier disagreement residual. Relative‑advantage‑based credit assignment (RACA) distributes this trajectory‑level residual across tokens according to their relative OPD advantages while preserving the original OPD signal. Across five long‑context benchmarks, post‑training with GC‑OPD raises the five‑benchmark averages of the official Qwen3‑4B and Qwen3‑8B checkpoints from 29.08 to 40.47 and from 35.12 to 44.65, respectively. Vanilla OPD reaches 39.31 and 43.56 under the same setup. Controlled ablations show that the signed residual is more effective than either an additional OPD‑derived term or direct group‑normalized verifier reward addition, while RACA further improves over uniform token allocation. Together, these results demonstrate that group‑relative residual calibration can incorporate verifier outcomes without discarding dense token‑level guidance. Code is available at https://github.com/SolereZhang/GC‑OPD.

Authors:Tate Berenbaum, Muthaiah Venkatachalam
Title: Pre-Compiled Pipeline Shards for Distributed LLM Inference on Intel AI PC Fleets
Abstract:
Modern Intel AI PCs ship capable integrated GPUs and NPUs with 16+ GB of unified memory, and they spend considerable time idle. That is not enough memory to fit a large model such as a 70B‑parameter LLM. We show that a handful of AIPCs, working together over an ordinary network, can serve models beyond the capability of any single one. We use pipeline parallelism: a model is split by layer into per‑stage shards, each pre‑compiled into an OpenVINO graph, so that every machine runs one shard and passes activations to the next. Three techniques make this fast enough to be useful. First, we recover the speed of the unsplit model: a naive per‑stage export runs well below monolithic inference because it misses an OpenVINO GPU optimization, and injecting a beam_idx Gather into each shard triggers that optimization (the IndirectKVCache fusion) and brings the shards to parity. Second, we leverage speculative decoding on stateful OpenVINO models. Third, the pipeline serves several users at once by interleaving their requests across the stages, each request carrying its own cache (micro‑batching). Together, a two‑node Llama 3.1 8B INT4 pipeline serves two concurrent users at 1.79x the single‑user throughput of the unsplit model on the same hardware, and the gap widens under simulated wide‑area latency. The same design scales to a 70B model that no single fleet member can hold: a four‑node deployment of Lunar Lake AI PCs on Intel Tiber Cloud serves a single user at interactive speed, with output token‑for‑token identical to the same four‑node pipeline decoding without speculation. Code, raw benchmark logs, and reproduction scripts ship as a self‑contained package at https://github.com/labscommunity/pipeline‑sharded‑inference‑paper (in the top‑level reproduction/ directory).

Authors:Huan-ang Gao, Haohan Chi, Yong Yan, Shiyuan Feng, Hanlin Wu, Zheng Jiang, Bingxiang He, Wei-Ying Ma, Ya-Qin Zhang, Hao Zhou
Title: Open-MOPD: Diagnosing and Fixing Capability Imbalance in Multi-Teacher On-Policy Distillation
Abstract:
Multi‑teacher on‑policy distillation (M‑OPD) has emerged as a promising paradigm for consolidating domain‑specialized reinforcement learning (RL) experts into a single generalist student via dense, token‑level reward supervision. Despite its practical success, the optimization dynamics governing multi‑teacher capability integration remain poorly understood, and open, rigorously reproducible recipes are conspicuously lacking. In this work, we establish a controlled M‑OPD benchmark on SmolLM3‑3B‑Base with oracle routing, isolating capability integration from routing ambiguity. Our investigation reveals a pronounced capability integration gap: standard M‑OPD captures only 35.6% of the available headroom relative to a domain‑routed oracle ensemble, with concise tasks such as instruction following suffering severe degradation and premature stagnation. Crucially, we show that this failure stems not from gradient conflict, but from a severe misallocation of the token‑level optimization budget. This pathology is driven by three orthogonal factors: structural sequence‑length disparities across domains, dynamic convergence drift due to non‑uniform learning rates, and multi‑step reward staleness from asynchronous policy updates. To resolve these imbalances, we introduce Open‑MOPD, a principled framework incorporating token‑share balancing, gap‑aware dynamic budget allocation, and student reward refresh. Together, these mechanisms systematically restore cross‑domain balance, elevating headroom recovery from 35.6% to 83.4% in a single deployable student. We fully open‑source our end‑to‑end post‑training recipe, training trajectories, and evaluation suites on an academically accessible hardware budget.

Authors:Jack A. Johnson, Bartłomiej W. Papież
Title: When Two Tracers Disagree: An Investigation of Multimodal Fusion for Clinical PET/CT Segmentation
Abstract:
PSMA and FDG PET/CT visualise complementary biological information in prostate cancer. Combining both tracers could capture heterogeneous tumour phenotypes that may be missed by either alone, yet there is no consensus on effective deep learning architectures for fusing these modalities. We evaluated multimodal image‑fusion strategies for automatic whole‑body PET/CT lesion segmentation to estimate total tumour burden. Using the public DEEP‑PSMA Challenge dataset, we trained tracer‑specific 3D nnU‑Net baselines and compared (i) early fusion with a single encoder and one decoder (OEOD) or two decoders (OETD), and (ii) intermediate fusion via a dual‑encoder cross‑attention U‑Net (DECA‑UNet). Tracer‑specific baselines performed strongly (PSMA Dice = 0.93; FDG = 0.81). Fusion yielded mixed results: OEOD produced a combined Dice of 0.90 (on an easier, non‑tracer‑specific task), whilst the tracer‑specific fusion models reached PSMA/FDG = 0.69/0.64 (OETD) and 0.76/0.57 (DECA‑UNet). Whilst fusion often provided reasonable PSMA segmentation, FDG performance degraded and no strategy consistently exceeded the single‑tracer baselines. Under the evaluated setting, tracer‑specific models remain the stronger baseline; clinically useful gains from multimodal fusion will likely require architectures that better preserve tracer specific representations. Our code is available at: https://github.com/JackJ3636/DEEP_PSMA_code

Authors:Pradeep Murugesan, Luoxiao Yang, Xueli Chen, Xinqi Fan
Title: Adaptive Memory and Reflection Multi-Agent System for Medical Question Answering
Abstract:
Accurate and responsible medical question answering (QA) is important in healthcare, where complex cases require factual knowledge and nuanced reasoning. Existing medical QA systems, typically based on single‑agent architectures and static retrieval, often lack adaptability, persistent memory, and structured decision‑making. This work introduces an adaptive memory and reflection (AMR) agentic system, a multi‑agent framework in which specialized agents use dedicated memory and reflection‑based feedback to retrieve relevant prior cases and improve subsequent reasoning. Complexity assessment routes questions through solo, collaborative, or escalated workflows, while consensus and ethical overseer modules support reasoning consolidation and output review. Evaluation on MedQA and MedMCQA demonstrates strong performance compared with several baselines. Ablation studies show that combining agent‑specific memory, reflection, and external retrieval yields the strongest performance. These findings highlight the potential of structured memory and feedback for developing more trustworthy medical agents. The source code is publicly available at https://github.com/mm‑air/AMR‑Agent.

Authors:Yajie Yin
Title: Grading the Graders: Verification Autonomy Levels (L0-L5) for LLM Reasoning
Abstract:
Large language models (LLMs) are increasingly paired with verifiers (step checkers, self‑consistency filters, tool‑based fact checkers, formal proof assistants) that claim to detect the model's errors. Yet the verification literature uses the word "level" to mean at least five different things: verification granularity, concept abstraction, risk tier, system‑stack layer, and the epistemic source of the ground truth. We propose Verification Autonomy Levels (VAL), a meta‑standard that classifies any verification scheme along a single axis: where does the verification spec come from, and what does the verdict guarantee? VAL ranges from L0 (LLM self‑declaration; no deterministic anchor) through L2 (objective ground truth; correctness only) to L3/L4 (decidable systems with single‑property or domain‑level completeness), with L5 impossible in the unrestricted case. Central to VAL is the completeness blind spot: substitution‑ and sampling‑based verifiers can confirm that proposed candidates hold, but cannot prove that no candidate was missed. We further identify a dichotomy the literature has not stated: completeness is reachable only for formally specifiable properties, whereas empirical open‑world verification (fact‑checking, diagnosis) caps at anchored correctness (L2). We document this gap empirically across four domains (symbolic mathematics, behavior monitoring, medical diagnosis, and code generation, the last a reverse validation with predictions stated before evidence) and in the strongest formal‑verification baseline in our survey, whose authors note the verifier focuses on the correctness of each step. We show the levels of granularity, concept hierarchy, risk, and system stack are orthogonal to VAL, resolving a systematic conflation across 17 surveyed papers. Code and full assessment are released as supplementary material.

Authors:Xujia Wang, Yizhe Zhang, Bin Xu, Lei Hou, Juanzi Li
Title: DeepWeaver: Bridging the Evidence Synthesis Gap in Open-Ended Question Answering
Abstract:
Retrieve‑then‑generate pipelines are commonly used to produce deep‑research answers for open‑ended questions, but retrieval alone is insufficient: LLMs must organize noisy and fragmented evidence into comprehensive, well‑cited answers. We refer to this process as evidence synthesis. However, direct generation often underuses evidence, misaligns citations, and collapses diverse information into shallow summaries, exposing an evidence synthesis gap between retrieval and generation. Thus, we propose DeepWeaver, a novel framework that weaves noisy retrieved evidence into comprehensive answers by maintaining Thought Block Chains (TBCs), a structured representation that groups claims, salient information, keywords, and supporting evidence. DeepWeaver uses subordinate TBCs to inspect residual evidence, commit TBC revisions, and discover new claims before final generation. We evaluate DeepWeaver on open‑ended QA over both knowledge bases and the web, and introduce LoQA, a high‑density benchmark for evidence synthesis. Across multiple LLMs, DeepWeaver improves content sufficiency, citation grounding, and detail preservation on LoQA, while achieving deeper insights and higher citation quality on DeepResearch Bench. These results show that evidence weaving is an effective mechanism for bridging retrieval and generation in open‑ended QA. Our code is available at https://github.com/KlozeWang/DeepWeaver.

Authors:Maedeh Hafezi Moghadas, Hakim Baazaoui, Lukas Bastian Otto, Susanne Wegener, Björn Menze, Ezequiel De la Rosa
Title: X-LMC: Cross-View Spatiotemporal Collateral Circulation Scoring from DSA
Abstract:
Digital subtraction angiography (DSA) is the reference standard for leptomeningeal collateral (LMC) assessment, providing critical prognostic insights to guide secondary treatment strategies, neurorehabilitation planning, and retrospective stroke research. However, clinical LMC grading via the ASITN/SIR scale relies on manual, highly variable visual inspection. We introduce X‑LMC, a spatiotemporal framework for automated collateral scoring from time‑resolved biplane DSA. The proposed architecture encodes spatial frame representations through a DINOv2 backbone, fuses orthogonal projections via a token‑level cross‑view attention module, and models representations of contrast bolus dynamics using a recurrent network architecture. We evaluate our framework on a multicenter dataset of 134 patients with M1‑segment occlusions. In a 5‑fold cross‑validation setting, X‑LMC yields higher point estimates than static architectures and spatiotemporal baselines adapted from related angiographic tasks, achieving a Quadratic Weighted Kappa (QWK) of 0.398 (vs. 0.322) and a dichotomized macro‑F1 score of 0.711 (vs. 0.663) against the best‑performing baseline. X‑LMC performance also aligns with the observed clinical inter‑rater agreement (QWK: 0.314). As the first DSA study attempting to automate LMC scoring, we demonstrate that multi‑view temporal deep learning can capture collateral‑specific contrast kinetics. Ultimately, these benchmarks delineate the clinical ambiguities and achievable performance boundaries of automated ASITN/SIR grading, establishing a reproducible foundation for objective hemodynamic phenotyping in stroke cohorts. Code is available at https://github.com/maedehafezi/X‑LMC.

Authors:Zane Kumar, Vishal Jain, Bernhard Kainz
Title: Frozen DINO Localizes Image Edits Without a Localizer
Abstract:
Localized image edits can change a photograph's meaning while leaving most of it authentic, so forensic analysis must identify where an edit occurred. We show that patch‑level perturbation responses from frozen DINO encoders are themselves localization maps. Training‑free Localization of AI‑image Edits from patch‑token Drift (TRAIL) applies one global Haar perturbation and maps cosine drift between corresponding patch tokens. On 80 source‑disjoint CocoGlide test images, TRAIL reaches .903 patch AUROC versus .912 for the mask‑supervised Detective SAM; fixed‑threshold Dice is .619 versus .709, while an oracle threshold raises TRAIL to .790. Transferred unchanged to Poisson image interpolation, TRAIL reaches .855 AUROC versus .864, showing that the cue persists without a generator. Across sixteen DINO encoders, the best block lies at normalized depth .80‑.94. Global context matters: AUROC falls from .903 globally to .857 for local‑in‑canvas perturbations and .735 for independently encoded crops. Frozen DINO patch tokens therefore contain a strong late‑layer localization signal whose visibility depends on the perturbation and preserved context. Code: https://github.com/VishalJ99/trail‑image‑edit‑localization.

Authors:Silin Chen, Han Li, Xiaodong Gu, Yuling Shi, Haibing Guan
Title: SkillForge: Self-Distilling Agents for Project-Specific Issue Resolution
Abstract:
Large language model (LLM) based agents have demonstrated remarkable proficiency in automated software issue resolution, yet they often struggle to resolve issues in a specific repository because they lack project‑specific knowledge. Existing self‑evolving approaches acquire such knowledge from repository history or online repair trajectories, but they either depend on available historical issue‑resolution signals or incur substantial per‑issue test‑time exploration cost. In this paper, we propose SkillForge, a self‑distillation framework that proactively acquires project‑specific knowledge from the repository itself. Instead of waiting for real issues to expose project‑specific knowledge gaps, SkillForge synthesizes project‑specific issues by re‑implementing test‑covered core functionalities of the repository. By resolving these synthetic issues, SkillForge distills reusable project‑specific knowledge into entity‑grounded skills and associates them with relevant repository entities for future issue resolution. Extensive experiments using both open‑source and closed‑source models show that SkillForge consistently improves issue resolution performance over strong baselines. These results demonstrate that proactively acquiring project‑specific knowledge before solving real issues substantially improves downstream software issue resolution.

Authors:Sebastian Doerrich, Francesco Di Salvo, Shyam Nandan Rai, Marco Lents, Christian Ledig
Title: Simple, Safe, and Overlooked: Reclaiming Sustainable Domain Generalization with Statistical Color Matching
Abstract:
Hardware shifts, color variations, and changing patient characteristics between development and deployment routinely break trained medical image classifiers. Existing remedies fall short: standard color jittering provides insufficient diversity, while deep generative style transfer algorithms hallucinate features, destroy clinically relevant structures, and waste massive compute resources. To address this, we revisit classical statistical color matching and repurpose it as Colorist, a highly efficient data augmentation strategy that applies global mean‑standard deviation matching directly in the RGB color space. We demonstrate that this training‑free, fully interpretable approach safely generates structurally intact domain variations, outperforming deep generative models in structural fidelity and color alignment. Across out‑of‑distribution histopathology, peripheral blood, dermatology, and retinal datasets, it improves balanced accuracy by up to +9% over state‑of‑the‑art domain generalization regularizers and by +13% over an unaugmented baseline. Moreover, by avoiding neural networks in the augmentation loop, Colorist preserves anatomical structure, minimizes carbon footprint, and integrates seamlessly into standard dataloaders. Together, these findings establish statistical matching as a safe, interpretable, yet overlooked alternative to deep architectures for clinical robustness. Source code is available at https://github.com/sdoerrich97/colorist.

Authors:Dinh Nam Pham, Shushen Manakhimova, Vivien Macketanz, Sebastian Möller
Title: Assessing Quality of Experience in Natural Language Generation of German Text
Abstract:
The rapid advancement of Natural Language Generation (NLG) has made the reliable evaluation of generated text increasingly critical, as these systems, such as large language models (LLMs), are now widely deployed in real‑world applications. However, traditional automatic metrics fail to capture the multifaceted nature of perceived quality. In this paper, we introduce TextQ‑German, a novel dataset suite for human‑centered evaluation of German NLG from a Quality of Experience (QoE) perspective, covering automatic text summarization and machine translation. Through crowdsourcing studies with German speakers, we collect human quality ratings and identify relevant perceptual quality dimensions for each task. We develop automatic QoE prediction models, including transformer‑based, linguistic feature‑based, and hybrid approaches. Hybrid models outperform pure transformer baselines in almost all experimental settings, while linguistic features alone can approach the performance of fine‑tuned language models. The dataset is extended with LLM‑generated outputs annotated with overall QoE scores. Final validation on held‑out sets indicates generalization to unseen data. Our work contributes a publicly accessible resource for NLG evaluation and baselines for automatic QoE prediction, providing a foundation for developing NLG systems that better align with human quality perception.

Authors:Qingyao Li, Wenxiang Jiao, Shuai Shao, Kangning Zhang, Yuan Lu, Yi Guo, Weiwen Liu, Weinan Zhang, Yong Yu
Title: SkillGate: Training In-Policy Skill Selection in Long-Horizon Agents
Abstract:
Agent frameworks increasingly package procedural knowledge as skills: instruction files an agent reads on demand, while public libraries now hold thousands of them. Which skill to read has thus become a decision the policy itself makes in the middle of an episode, yet no existing signal trains it. We show that the default remedy, outcome‑rewarded RL over the candidate slate, cannot teach it, for a structural reason we identify and name selector credit starvation: under a broadcast, sequence‑level advantage, the few tokens that name the chosen skill carry a vanishing share of the loss, and the credit they inherit is increasingly wrong‑signed as trajectories lengthen. A correct choice is punished whenever the execution after it fails, even though the choice itself is among the most valuable decisions in the trajectory. Auditing a completed run's own training artifacts confirms all three properties, each worsening monotonically with horizon. SkillGate removes the failure by construction: it partitions the token support into two disjoint credit channels, outcome credit reaching only execution tokens, and a separate action‑local advantage reaching exactly the skill‑naming tokens, positive only when a trajectory's single read is the correct one. On five agentic benchmarks under a 16‑candidate slate, SkillGate lifts a 9B policy from 40.8% to 53.2% trial success, well ahead of the identical budget spent on outcome reward alone, while cutting exposure to misleading candidates by two thirds and reading fewer skills.

Authors:Fa-Ting Hong, Runzhen Liu, Luchuan Song, Hongmin Cai, Chuhua Xian
Title: EfficientSync: Real-Time Lip Synchronization via Deformation-Based Reference Texture Mixing
Abstract:
Audio‑driven lip synchronization manipulates the mouth region of a talking‑face video to match the driving audio while preserving head pose, identity, and background. Although the task is inherently local editing, prevailing approaches reconstruct the entire lower face with heavy GAN‑ or diffusion‑based decoders, incurring substantial latency and, more critically, hallucinating intra‑oral details such as teeth and lip wrinkles instead of preserving authentic textures. We contend that the bottleneck in identity preservation is not the scarcity of reference frames, but the lack of a mechanism that faithfully transfers the genuine textures they already contain. We therefore present EfficientSync, a real‑time deformation‑based framework that retains reference textures rather than resynthesizing them. First, the Dynamic Texture Mixer reformulates multi‑reference fusion as channel‑wise selection, evaluating each spatially aligned reference in a global context and aggregating them by channel‑wise weighted summation, preserving textural integrity at low cost. Second, Spatio‑Temporal Shifted Adaptive Masking decomposes the source frame into lip‑generation conditions and an independent background prior, suppressing lower‑face leakage while blending the synthesized mouth seamlessly into the background. Third, STAR Sampling, a zero‑overhead pre‑processing step, retrieves the sharpest and most topologically diverse reference frames. Experiments on HDTF and VFHQ show state‑of‑the‑art visual quality and identity preservation at 166 FPS on a single GPU. Video demos: https://alunaticat.github.io/EfficientSync/index.html.

Authors:Jiandong Ding, Huijie Qin, Tiandeng Wu, Yi Cao
Title: SIDScope: A Diagnostic Resource for Semantic-ID Interfaces in Generative Recommendation
Abstract:
Semantic‑ID mappings are reusable interfaces between item tokenizers and generative recommenders, yet released mappings rarely state whether they are coherent, what structure they expose, how generated paths resolve, or what must be revalidated after a refresh. SIDScope is a source‑traced diagnostic resource for these decisions. It normalizes item‑to‑code artifacts, verifies provenance and joins, profiles mapping structure, compares paired revisions, and accounts for path‑to‑item outcomes in generated traces. Across nine source‑traced tokenizer exports from seven families on Amazon and Yelp data ‑ eight executable routes plus one auditable snapshot ‑ SIDScope reveals that interface health is multi‑signal rather than scalar. Its central finding is mechanism‑conditional: prefix alignment strongly tracks held‑out candidate exposure when retrieval consumes SID prefixes, then weakens as scoring becomes prefix‑independent. Trained trace accounting exposes a second hidden gap: a valid target path can survive without uniquely retrieving the target item by 1.2‑3.0 percentage points. A refresh case establishes a third: repairing the mapping does not by itself restore an inherited generator; model reuse requires a separate handoff check. The package provides frozen evidence summaries, conformance reports, trace labels, table builders, and CPU‑only verifiers. It supports decisions about artifact readiness, interface risks, and revalidation before model reuse.

Authors:José A. Perdiguero López, Miguel A. Durán-Olivencia
Title: Flama: a Python framework for development and deployment of production-ready APIs, machine learning, and LLM services
Abstract:
We present Flama, an open‑source Python framework for developing and deploying production‑ready web APIs, machine learning services, and large‑language‑model (LLM) applications. Built on the Asynchronous Server Gateway Interface (ASGI), Flama offers a type‑driven, async‑first programming model that unifies REST API development, predictive model serving, and generative AI inference in one architecture. It is organised around seven subsystems: a component‑based dependency injection system resolving handler parameters from type annotations at startup; a pluggable schema layer supporting Pydantic, Marshmallow and Typesystem behind a single adapter; an automatic CRUD generator turning a SQLAlchemy table and a schema class into REST endpoints backed by the Repository and Unit of Work patterns; a portable binary format (.flm) packaging models from scikit‑learn, TensorFlow, PyTorch and Hugging Face Transformers with their metadata for zero‑code deployment; a multi‑backend LLM server running vLLM (Linux/CUDA) or MLX (Apple Silicon) and exposing four wire protocols (OpenAI, Anthropic, Ollama, and a native streaming dialect) through a shared codec; a Rust‑accelerated core compiled via Maturin for routing, JSON encoding, compression and parsing; and a Model Context Protocol module turning any application into an MCP server over JSON‑RPC 2.0. Built‑in capabilities include JWT authentication, two pagination strategies, background tasks in threads or processes, WebSocket endpoints, Server‑Sent Event and NDJSON streaming, OpenAPI 3.2.0 generation from handler signatures, and a command‑line interface for running applications and for serving, packaging and inspecting models. We describe the architecture, present the programming model through worked examples, and compare Flama with existing frameworks, model serving platforms and LLM inference engines.

Authors:Maohao Ran, Chendong Ma, Yanting Zhang, Dailing Jiang, Yusen Huang, Meng Gao, Jun Song
Title: Execution-grounded evaluation reveals hidden failures in language-model calculations for environmental science
Abstract:
Large language models are increasingly used for quantitative work in the environmental sciences, yet existing evaluations score only final answers, leaving calculation process unobserved. Here we introduce AtmosCoder‑Bench, an execution‑grounded benchmark that makes the calculation process visible. Built through a transferable semi‑automated pipeline (436 problems, 3,910 variants, 7,029 graded quantities), every problem is validated to be unambiguous and human‑solvable, with uniquely verifiable answers. We find that (i) multiple‑choice formats inflate measured accuracy by at least 12 percentage points; (ii) many failures arise not from missing knowledge but from models failing to apply known formulas and constraints consistently throughout multi‑step computation; and (iii) even frontier models remain weak when task‑specific conditions invalidate familiar methods, often reverting to canonical solution patterns rather than adapting methods to the relevant physical regime, leaving expert oversight essential.

Authors:Berken Utku Demirel, Christian Holz
Title: EgoHRV: Continuous Heart Rate Variability Estimation from Egocentric Systems for Autonomic Response and Skill Assessment
Abstract:
Egocentric vision systems capture human behavior from visible cues, but overlook physiological indicators of autonomic states such as stress, engagement, and attention. Heart rate variability (HRV) is a widely used noninvasive marker of autonomic regulation under stress. HRV reflects small timing differences between successive heartbeats and has so far been out of reach for egocentric platforms, where motion and noise in gaze video mask exactly this fine‑grained timing. We propose EgoHRV, a method that estimates HRV as well as heart rate (HR) from the gaze cameras that are already integrated into egocentric headsets. Our pipeline combines a 3D backbone with a novel low‑‑high decomposition module that extracts the blood volume pulse (BVP) signal from gaze video. Our cross‑domain pretraining aligns the frequency‑domain representations of contact‑based and camera‑derived signals. This alignment gives EgoHRV the temporal precision to recover HRV from the subtle fluctuations in gaze video. EgoHRV achieves state‑of‑the‑art accuracy for HR and HRV estimation from egocentric video, and its uncertainty‑aware design improves downstream behavioral modeling. Integrating our HRV estimates and confidence measures into EgoExo4D's proficiency estimator raises accuracy by 17.8%. Beyond skill, continuous HRV estimation also opens egocentric systems to stress‑ and arousal‑aware estimation tasks. Code: https://github.com/eth‑siplab/EgoHRV

Authors:Kartik Chincholikar, Kaushik Gopalan, Mihir Hasabnis
Title: Impact of Iterative Fine-Tuning on Transcription Accuracy in Complex Historical Sanskrit Manuscripts
Abstract:
Digitizing the text from handwritten historical manuscripts is required to make them easily accessible, preservable, and to enable historical scholars to study them in new ways. Historical manuscripts, however, often exhibit complex heterogeneous layouts and non‑standard appearance due to period‑specific writing styles, page textures, camera noise, and other nuisance factors, making them difficult to perform OCR on. To tackle this challenge, we introduce a local traditional OCR pipeline, which can be iteratively fine‑tuned on the target manuscript at the layout‑level and the appearance‑level. By adapting to the target manuscript distribution, the proposed Traditional OCR pipeline makes better predictions on subsequent pages, causing iterative reduction in human annotation effort, which is expensive and time‑consuming as it requires historical domain expertise. Using this pipeline, we digitize text from three complex historical Sanskrit manuscripts and introduce a dataset with granular layout‑level annotations, along with Unicode annotations in the standard PAGE‑XML format. We demonstrate quantitative gains due to iterative fine‑tuning of the proposed traditional OCR pipeline, and also benchmark the performance of leading Multi‑Modal Large Language Models on the introduced Dataset. Code and dataset are available at: https://github.com/flame‑cai/gnn‑synthetic‑layout‑historical/.

Authors:Haotao Xie, Yutian Chen, Yangqi Liu, Xiaoyu Jiang
Title: Candidate-Fate Accounting for Transparent Sensor Diagnostic Pipeline Search
Abstract:
Industrial sensor diagnostics relies on preprocessing, representation, and classification pipelines, making automated pipeline search useful for reducing manual design cost. However, existing automated machine/deep learning (AutoML/AutoDL) reports typically retain only fitted trials, scores, and winners, omitting generated candidates that are invalid, pruned, skipped, cached, or unfitted. This omission limits reviewers' ability to check signal constraints, budget use, and unevaluated legal alternatives. To address this, we propose candidate‑fate accounting, a candidate‑level audit framework for diagnostic search traces. It records each observed candidate as auditable evidence: hashes merge repeated observations, legality checks flag invalid candidates, allocation rationales explain budget decisions, and a closed fate ledger assigns one terminal fate to each candidate. Experiments on three bearing‑diagnostic datasets show that the framework detects invalid candidates and identifies 30‑‑41 candidates omitted by fitted‑trial‑only reports, with closed fate records verifying complete candidate accounting while maintaining competitive diagnostic performance. The code is available at https://github.com/XXIE999/candidate‑fate‑accounting.

Authors:Zhiqiang Hu, Tao Yu, Shouren Huang, Masatoshi Ishikawa
Title: Dynamic SpectraFormer for Ultra-High-Definition Underwater Image Enhancement
Abstract:
Underwater images suffer from color distortion, haze, and poor visibility due to light refraction and absorption in water. These challenges significantly impact the utilization of Autonomous Underwater Vehicles (AUVs) or marine robots. Typically, color and brightness distortions manifest at lower frequencies, while edge and texture distortions are prevalent at higher frequencies. Traditional methods struggle to concurrently rectify these mixed distortions as they primarily concentrate on the spatial domain. To address these issues, we introduce the Dynamic SpectraFormer, which enhances underwater images through a frequency domain transformer. The Dynamic SpectraFormer introduces an ultra‑high‑resolution sparse spectrum attention module, which could capture the long‑term dependency without losing the universal approximating power. Additionally, we have developed a dynamic spectrum weight generation layer that serves as an adaptive spectrum band selector, accentuating critical frequency bands and suppressing less relevant ones. Consequently, this method significantly improves underwater image quality by addressing both high‑ and low‑frequency distortions. Our extensive ablation studies and comparative evaluations consolidate the Dynamic SpectraFormer's efficacy across multiple underwater image enhancement benchmarks. The source code is available at https://github.com/arifence2024/DynamicSpectraFormer.git.

Authors:Rime Wen, Zehan Liu, Shawn Qin, Lights Shi, Roy Gan, Hao Wang, Qian Wang
Title: X2Streaming-TTS: Causal Token-Level Text-to-Speech from Streaming Text with Speech-State Inheritance
Abstract:
Streaming text‑to‑speech is essential for low‑latency spoken dialogue systems, yet many systems wait for sentence‑level text and are therefore only pseudo‑streaming. True token‑level synthesis must generate speech from uncertain prefixes while maintaining perceptual continuity over an unbounded stream with bounded context. We present X2Streaming‑TTS, a causal TTS framework that consumes asynchronously arriving text tokens and emits speech without accessing future input. To handle uncertain prefixes, we introduce causal commitment, which keeps ambiguous expressions provisional through uncertainty‑aware buffering and performs capacity‑adaptive, punctuation‑aware segmentation. To preserve acoustic continuity, we further introduce causal speech‑state inheritance, which carries the complete Code2Wav state and selected historical Talker states across segment boundaries. Together with an attention prior constraint, it blocks access to future positions while retaining bounded acoustic context. Experiments show that X2Streaming‑TTS outperforms existing pseudo‑streaming models on most subjective and objective metrics. Further analysis shows that causal commitment stabilizes online segmentation and reduces failures caused by insufficient context, while speech‑state inheritance improves boundary continuity without degrading naturalness or speaker identity. X2Streaming‑TTS thus achieves strict token‑level synthesis with quality comparable to the evaluated offline baselines, a median time to first audio token (TTFT) of 15.8 ms for a single request, and a median TTFT of 260.8 ms at 128 concurrent requests. Our implementation is publicly available at https://github.com/X‑Square‑Robot/X2Streaming‑TTS .

Authors:Yinming Huang, Shuyuan Tu, Xi Yan, Zihan Yang, Jianhua Han, Xu Hang, Yu-Gang Jiang, Zuxuan Wu
Title: VA-Judger: Reward Modeling from Human Preference Feedback for Joint Video-Audio Generation
Abstract:
Using reinforcement learning to post‑train joint video‑audio generation models requires a reward signal. Existing methods construct this reward by combining metrics for individual quality dimensions, including audio quality, visual fidelity, and synchronization. However, these metrics evaluate perceptual dimensions separately and fail to capture the overall semantic and temporal coherence among the text prompt, video, and audio that shapes human preferences. Optimizing models against these metrics encourages reward hacking, generating video‑audio content that achieves high scores on these metrics yet appears incoherent or unfaithful to human viewers. To address this problem, we first construct a large‑scale human‑preference dataset VAPref‑10K for joint video‑audio generation, comprising 9K prompts and 10.3K fine‑grained paired comparisons from open‑source generation models. We also introduce the VA‑Judger‑Bench benchmark with both in‑domain and out‑of‑domain model comparisons to evaluate whether reward models truly align with human preferences. We further propose VA‑Judger, a chain‑of‑thought omni‑reward model for joint video‑audio generation. In particular, VA‑Judger first learns from pairs with clear quality gaps to establish structured output and coarse preference discrimination, then distills reliable preference explanations for harder near‑quality comparisons via rejection sampling verified against human annotations, and finally performs dimension‑wise reinforcement learning that decomposes human feedback into individual quality dimensions for denser reward signals than a single binary preference label. Experiments show that VA‑Judger outperforms metric baselines in predicting human preferences on both in‑domain and out‑of‑domain evaluations. Using its human‑aligned rewards for post‑training audio‑video generation model also yields significant improvements in generation quality.

Authors:Kai Van Brunt, Justin Kay, Sara Beery
Title: Teach a Molmo2Fish: Towards interactive fish tracking with natural language guidance
Abstract:
Computer vision is increasingly used to automate recognition tasks in large ecological datasets, but more complex tasks such as multi‑object tracking continue to pose challenges. As researchers seek to incorporate vision models in ecology workflows, various lines of research have explored how to make imperfect predictions useful through human‑in‑the‑loop processes. We propose a new approach to working with imperfect tracking predictions through an interactive prediction correction workflow taking place as a conversation with a multimodal large language model, which we tailor to a sonar fish tracking dataset as an initial proof of concept. We investigate the performance of the tool, Molmo2Fish, across guided and unguided tasks, correcting its own predicted tracks and external tracks. We find that Molmo2Fish achieves high performance on fish tracking and track correction tasks, but there is still much room to improve on incorporating natural language guidance. The code and data are publicly available at https://github.com/tidalove/molmo2fish.

Authors:Ruiqi Zhang, Hao Zhu, Wenhao Zhang, Qi Zhang, Junqi Shi, Ming Lu, Xun Cao, Zhan Ma
Title: ReX-Shot: Single-Image Rephotography via Geometry- and Camera-Grounded Generation
Abstract:
Single‑image rephotography aims to synthesize new shots of a scene from a single reference image with specified viewpoints, focal lengths, and photographic effects, which are intrinsically coupled in imaging. Existing methods typically treat these factors separately and struggle under joint control: novel‑view synthesis may introduce geometric distortions under focal‑length changes, while super‑resolution and instruction‑guided editing remain confined to 2D and cannot reliably extend detail restoration or appearance control to novel viewpoints. We attribute these limitations to imperfect single‑image 3D reconstruction and the sampling limit of continuous focal‑length enlargement. To reduce projection bias from geometric errors, we use implicitly transformed foundation‑model features for robust target‑view guidance. We further formulate focal‑length enlargement as a geometry‑guided super‑resolution problem and exploit generative detail priors to recover details lost during sparse 3D resampling. Built on this 3D‑aware generative backbone, we lift photographic‑effect control from 2D filtering to 3D‑aware appearance editing, preserving content consistency across viewpoints and focal lengths. These components form ReX‑Shot, a geometry‑ and camera‑grounded generative framework for single‑image rephotography. To our knowledge, ReX‑Shot is the first unified framework to jointly control viewpoint, focal length, and parameterized photographic effects from a single image. Experiments show that ReX‑Shot outperforms representative baselines across all three controls while enabling near‑real‑time interactive rephotography.

Authors:Kou Shi, Zun Wang, Qisheng Su, Shiting Huang, Ziao Zhang, Zhen Fang, Qingnan Ren, Jin Liu, Yu Zeng, Yiming Zhao, Lin Chen, Zehui Chen, Feng Zhao
Title: FACET: Preserving Source Intent and Executable State in Terminal Task Synthesis
Abstract:
Training terminal agents requires scalable executable supervision, yet synthesizing high‑quality terminal tasks remains challenging. Each task couples an instruction, an initialized environment, a reference solution, and an executable verifier; if these artifacts are generated from inconsistent assumptions, the resulting task may be unsolvable or incorrectly evaluated. Meanwhile, multi‑stage synthesis can discard the goals, dependencies, state transitions, and procedural constraints encoded in the original sources. We present FACET (Fine‑grained Agentic Construction of Executable Tasks), a framework that addresses both information preservation and cross‑artifact consistency. FACET reconstructs related agent skills into coherent, information‑rich scenarios, then realizes and repairs the execution environment before generating the final task artifacts. The resulting container state serves as shared grounding for the instruction, solution, and verifier, while execution‑based validation and targeted repair correct artifact‑specific failures without unnecessarily regenerating valid components. FACET produces complex terminal tasks with dense executable checks, and successful trajectories collected from these tasks provide effective, data‑efficient supervision. Fine‑tuning models across multiple scales consistently improves performance on Terminal‑Bench 2.1, while analyses of alternative generation schemes support the importance of environment‑grounded construction for task validity and solution‑verifier alignment. These results establish source‑intent preservation and shared executable‑state grounding as key principles for scalable terminal‑task synthesis.

Authors:Yuan li, Youyuan Lin, Chenhui Chu, Shin'ya Nishida
Title: MR-IQA-2: Faithful Image Quality Reflection via Fine-Grained Credit Assignment
Abstract:
Multimodal large language models (MLLMs) have shown strong potential for image quality assessment (IQA) by improving consistency between quality ratings and their underlying reasoning. However, most approaches supervise reasoning through human‑provided ratings and rarely examine whether it faithfully reflects image quality. Rating accuracy alone does not ensure faithful reasoning; a shared reward also obscures supervision sources and may reinforce unfaithful reasoning when a correct rating occurs by chance. To improve the faithfulness and reliability of blind IQA, we aim to (1) decouple credit assignment for reasoning and rating and (2) provide verifiable supervision for faithful reasoning. We introduce MR‑IQA‑2, an actor‑editor‑judge framework that operationalizes reasoning‑editing‑reflection. The actor generates quality reasoning for an input image, and the editor revises the image according to the identified quality factors. A frozen judge compares the original and edited images and provides reflective supervision for the actor's reasoning. MR‑IQA‑2 further uses fine‑grained credit assignment to decouple reasoning and rating supervision. Judge feedback supervises reasoning, whereas human ratings supervise the predicted rating. Masked token‑specific updates distinguish these signals while preserving the causal relation from reasoning to rating. Across IQA benchmarks, MR‑IQA‑2 achieves competitive rating alignment with humans. Visual reflection also enables richer and more faithful visual understanding beyond rating, which may inform image‑quality optimization and related downstream tasks. Code is available at https://github.com/RobinY99/MR‑IQA‑2.

Authors:Shayan Shahrabi-Farahani, Dara Rahmati
Title: Compress and Forget: bitsandbytes Quantization Amplifies Proactive Interference in LLMs
Abstract:
Proactive interference (PI) is a documented failure mode in large language models in which retrieval of a repeatedly overwritten value degrades as prior overwrites accumulate, mirroring a classical phenomenon in human working memory. Post‑training quantization (PTQ) is now the default deployment path for open‑weight models, yet its effect on this failure mode has not been tested. We evaluate three precision levels (FP16, INT8, INT4/NF4, via bitsandbytes) across three architecturally distinct instruction‑tuned models (Qwen2.5‑7B‑Instruct, Mistral‑7B‑Instruct‑v0.3, Phi‑3.5‑mini‑instruct), holding the retrieval task fixed. INT4 quantization significantly reduces accuracy under high interference in every model (e.g., from 81.0% to 68.3% for Qwen), confirmed by paired McNemar's tests (p \le 2.6 × 10^‑6) and a mixed‑effects regression spanning all interference levels; INT8, often assumed safe, also carries a smaller but real penalty in two of three models. The effect is specific to semantically similar (word‑type) distractors and reverses sign under a numeric control condition, and is mechanistically linked to a rise in same‑key intrusion errors under INT4 (from 21.5% to 24.6% of trials, p = 4.8 × 10^‑7). A follow‑up ablation shows the effect originates in the quantized transformer backbone rather than the output projection layer. These results suggest that bitsandbytes 4‑bit quantization can impose an additional cost on applications relying on long, updatable, semantically dense contexts, even when aggregate benchmark accuracy appears largely unaffected. We release our code and tokenizer‑verified vocabulary construction method at https://github.com/ShayanShahrabi/compress‑and‑forget

Authors:Yaqi Li, Jielun Peng, Yabin Wang, Jincheng Liu, Xiaopeng Hong
Title: PATE-Forensics: Perception-as-Tool for Explainable Deepfake Forensics with General-Purpose MLLMs
Abstract:
Existing explainable deepfake forensic methods typically rely on task‑adapted MLLM to jointly address detection, localization, and explanation. Inspired by agent‑style tool use, we instead introduce a Perception‑as‑Tool paradigm and instantiate it as PATE‑Forensics, which architecturally decouples detection and localization from explanation generation while coupling detection and localization as tightly as possible within a forensic perception tool. The DINOv3‑based tool couples a multi‑granularity detection module that integrates global, patch‑level, and segment‑level evidence with a cue‑guided localization module by spatializing the patch‑level and segment‑level evidence into forgery score maps that guide dense mask prediction. The original image and forensic perception outputs produced by the tool form structured forensic context for a general‑purpose MLLM, which is guided by prompt constraints to generate explanations without task‑specific fine‑tuning. On DDL‑X Track 3, PATE‑Forensics achieves the best official score of 0.89, outperforming the second‑ranked team by 0.19 points. Our code is available at https://github.com/yqli00000/PATE‑Forensics.

Authors:Yanlun Tu, Huacan Wang, Ziyue Zhou, Jie Zhou, Ningyan Zhu, Ge Chen, Wangyi Chen, Tengfei Zhou, Yifan Zhou, Dasheng Yang, Xiaofeng Mou, Hui Zhang, Yi Xu
Title: SemaPLC: A Project-Grounded, Verification-Gated Agent Harness for PLC Code Generation
Abstract:
Programmable logic controllers (PLCs) run industrial plants, and large language models can already generate independent program organization units (POUs) for them. Whether such logic integrates into an existing PLC project and then runs correctly has been checked only in limited tests. We present \textscSemaPLC, a project‑grounded and verification‑gated agent harness assembled from conventional tools but governed by a strict completion rule. Rather than stopping when the model judges its own output adequate, \textscSemaPLC declares a task complete only when logged external checks confirm it. Those checks cover the specification, the compilation, and the behavior on a live runtime. On 117 independent‑POU tasks matching existing benchmarks, it attains the highest strict verified pass rate on all seven models (72.6% mean). On a project‑context track of 65 tasks whose generated logic must compile and run inside a real project, it attains the highest mean on integrated compilation, static behavior, and dynamic behavior. Of the three layers, dynamic behavior is the most revealing. We measure it by deploying the generated and the reference logic to a live PLC runtime and comparing their executed traces. All methods fall within 10 static points of one another, whereas dynamic scores separate them sharply, from 22.4 to 31.4 for the baselines against 52.2 for \textscSemaPLC. Overall, our verification‑gated harness raises the mean at every layer and most sharply at runtime. Execution, not static scoring, is the faithful test of whether generated control logic actually works. \textscSemaPLC is open‑sourced at https://github.com/midea‑ai/SemaPLC.

Authors:Pratik Ghawate
Title: FinRCA-Bench: Benchmarking Evidence Retrieval and Reasoning for Financial AI Systems
Abstract:
Large language models are increasingly used to support financial operations, but their apparent reasoning performance can depend on whether they receive the right evidence. In financial reconciliation, the evidence needed for diagnosis is distributed across invoices, purchase orders, approvals, allocations, payments, ledger entries, and bank activity, linked by transactional relationships rather than textual similarity. End‑to‑end accuracy can therefore conflate evidence access with reasoning quality. We introduce FinRCA‑Bench, a deterministic synthetic benchmark of 2,250 accounts‑payable‑to‑bank reconciliation cases spanning 14 operational tables, including 1,500 injected failures across 15 causal categories and 750 legitimate or hard‑negative cases. Root‑cause labels and record‑level evidence contracts are hidden from the model, allowing retrieval to be evaluated independently of answer correctness. We compare Rules/SQL, classical machine learning, dense semantic retrieval, deterministic relational expansion, and Typed Provenance Graph Retrieval (TPGR), a typed traversal restricted to persisted transaction relationships. Rules/SQL reaches 84.97% held‑out exact accuracy and classical ML reaches 95.44%. Holding the reasoning model, prompt, and generation settings fixed while changing only retrieval increases macro required‑record recall from 0.83% to 77.70% and exact 16‑class accuracy from 2.05% to 72.44%. Structural retrieval failures outnumber reasoning failures with sufficient retrieval by 95 to 15; 254 correct predictions occur despite incomplete retrieval, and strict returned‑evidence contract accuracy is only 5.72%. On FinRCA‑Bench, retrieval architecture strongly shapes observed AI‑system performance, and a correct root‑cause label is a weak proxy for an auditable diagnosis.

Authors:Soumili Ghosh, Debapriya Roy, Aryan Das, Bikash Santra
Title: OptiModNet: A UNet-Transformer Hybrid with Grouped-Query and Channel Attention for Optic Disc and Cup Segmentation
Abstract:
Precise segmentation of the optic disc and cup is critical for the early detection and diagnosis of glaucoma. However, achieving consistently high performance across datasets while maintaining low computational requirements remains a significant challenge. In glaucoma detection, low‑computation methods are crucial for enabling rapid, large‑scale screening and facilitating deployment in resource‑limited clinical environments. While deep learning models such as UNets, Vision Transformers (ViTs), and Diffusion models have demonstrated strong segmentation performance but these methods often come with substantial computational overhead. UNets are efficient at capturing local features but are limited in modeling global contextual information. Conversely, ViTs excel at long‑range dependency modeling but are computationally intensive. Hybrid architectures, such as UNetR, which combine transformer‑based encoders with UNet‑style decoders, have shown improved performance but while incurring additional complexity. Considering these, in this work, we propose OptiModNet, a light weight novel hybrid architecture tailored for optic disc and cup segmentation. The model integrates diverse attention mechanisms at multiple stages of the network to enhance both local and global feature representation. We include an Aggregated Pyramid Loss that supervises predictions at multiple decoder depths, to promote better gradient flow and structural consistency. We evaluate OptiModNet on the REFUGE2 dataset for both optic disc and cup segmentation tasks. Our method achieves state‑of‑the‑art performance, exceeding existing approaches by over 2.5%, while maintaining high efficiency with only 3.73 GFLOPs and 1.93M parameters. The code is available at https://github.com/SG1947/OptiModNet.

Authors:Avik Ghosh, Akın Taşcıkaraoğlu, Daniela Rojas, Muhammed A. Beyazıt, Mohammad Reza Salehizadeh, Keaton Chia, Sasha Doppelt, Michael Ferry, Jan Kleissl, Sujit Dey, Yuanyuan Shi
Title: Power Estimation and Optimal Work-Charging Scheduling of Construction Electric Vehicles via Mobile Charging Stations
Abstract:
Construction electric vehicles (CEVs) are a promising clean alternative to diesel‑powered construction equipment, but their adoption is constrained by sparse onsite charging infrastructure, limited CEV mobility, and insufficient understanding of their power consumption. We address these gaps through a field‑data‑driven framework coupling CEV power estimation with mobile‑charging‑aware work scheduling. First, using a real‑world construction demonstration at the University of California, San Diego, we develop and validate a per‑subactivity power estimation model for a compact electric excavator. Manually labeled video is synchronized with coarse battery state‑of‑charge (SOC) telematics, and constrained nonnegative least squares is used to recover each subactivity's average power consumption. The model predicts held‑out test data within 17% normalized mean absolute error (NMAE), and the accompanying dataset is released publicly. Second, leveraging the subactivity power estimates, we formulate a mixed‑integer program that jointly optimizes CEV work and charging schedules together with the location, timing, and charging/discharging of mobile charging stations (MCSs) serving the CEVs. The optimization accounts for energy and demand charges, carbon emissions, unmet work penalties, MCS travel, and the physical and operational constraints of the CEVs and MCSs. Across realistic scenarios drawn from the demonstration, the proposed co‑optimization attains the lowest operating cost in every case, being 7‑‑96% below the best‑performing baseline, while solving most instances to proven optimality within an hour. Dataset and scripts are available at https://github.com/ghosh‑avik/CEV‑MCS‑Power‑Estimation‑and‑Joint‑Scheduling.

Authors:Pardis Taghavi, Reza Langari, Gaurav Pandey
Title: Partition the Support, Reconstruct the Residual: Training-Free Sparse Attention for Video Generation and World Models
Abstract:
Training‑free block‑sparse attention can accelerate video transformers, but row‑wise attention concentration does not by itself specify an executable sparse operator. Queries sharing a block route may have poorly overlapping supports, while retained attention mass alone does not determine the post‑softmax error from skipped interactions. We show that partition geometry affects both pooled support and the predictability of the remaining residual from the sparse output. We introduce SparsePR, which combines Response‑Coupled Partitioning with Probe‑Fitted Residual Reconstruction. Sampled‑query key responses form paired K/V groups, whose centroids induce query‑response coordinates for shared routing. A small set of exact query rows then calibrates a call‑specific affine correction from the sparse output within the output subspace observed in the probe residuals. Across four heterogeneous video generation and world models, SparsePR consistently reduces attention‑reconstruction error. Ablations show that probe fitting accounts for most of this reduction, while response‑coupled partitioning lowers hard‑drop error and improves reconstruction under a finite probe budget. SparsePR preserves generation quality at 22.0‑26.0% realized executed‑pair density while achieving 1.48x‑2.61x end‑to‑end speedups. Project page: https://pardistaghavi.github.io/SparsePR‑website/

Authors:Mengpeng Yang, Jingxu Yang, Chao Chen, Tian Xia, Yabo Sun, Qiang Liu
Title: OmniAlign: A Unified Multilingual Aligner for Word and Sentence Alignment
Abstract:
Cross‑lingual sequence alignment is fundamental for building and exploiting parallel corpora, spanning mappings from documents and sentences down to words and subwords. Existing tools, however, typically specialize in a single granularity, so practitioners often need separate systems for word‑ and sentence‑level alignment‑‑‑especially in multilingual and long‑text settings. We present OmniAlign, a unified multilingual aligner that supports both word‑level and sentence‑level alignment with a single lightweight model. Built on an encoder‑only backbone with strong long‑context modeling, OmniAlign induces word alignments from contextualized token similarity matrices, and obtains document‑level m‑‑n sentence alignments via sentence embeddings combined with dynamic programming. To balance fine‑grained alignment accuracy and sentence‑representation quality, we use a four‑stage training pipeline: alignment‑oriented continued pre‑training, self‑supervised learning, supervised fine‑tuning on human annotations, and sentence‑embedding distillation from a strong multilingual teacher. Experiments show that OmniAlign achieves highly competitive performance on both word‑ and sentence‑alignment benchmarks and generalizes well to unseen language pairs. Surprisingly, later‑stage supervised fine‑tuning on short texts further improves alignment quality while retaining the long‑context understanding acquired in earlier training, keeping the model robust on long‑text word alignment. \normalsize \colorblueCode: https://github.com/MilkDargon/OmniAlign\par \colorblueModel: https://huggingface.co/WPS‑Qingqiu/OmniAlign

Authors:Tianyou Wang, Chongyang Gao, Kezhen Chen, Chen Dong, Yinghao He, Donghan Li, Wangcheng Xu, Hongjiu Zhang, Chi Li
Title: FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents
Abstract:
Language model agents now execute bounded tasks reliably. Whether they can sustain effective decision‑making over long horizons, where actions have cumulative consequences and the environment responds to their choices, remains largely unmeasured. FM‑Bench (Football Management Benchmark) measures this. An LLM agent runs a football club for 20 in‑game years through 26 tools and roughly 340 to 400 decision stops. It drafts a squad on the same budget as every rival, trades players, negotiates contracts, invests in facilities and youth, sets lineups, and answers to a board that can fire it, while a deterministic engine accumulates every year into one final score with no LLM judge or human rater. The solo track plays each of 15 frontier models against a frozen scripted world, and the Arena places the same models plus a scripted anchor in one shared 20‑year world; to our knowledge, the first head‑to‑head evaluation at this scale. We measure six behavioral capabilities behind the score. Across three seeds, all 15 models complete every horizon while the blind scripted baselines die out in most of theirs, and claude‑fable‑5 tops the solo board on mean score and the Arena, where the title nonetheless rotates among ten models. Neither scale, price, nor vendor predicts the order; the order settles only late in the horizon, and the best first‑play human lands only at the bottom of the model board. What separates the models is managerial behavior rather than computation. Higher‑scoring models reduce slow‑payoff investment near the end, keep cash invested rather than idle, and open renewals well before the deadline, while token spend predicts nothing. No model learns the market's hidden prices from hundreds of rejected bids, and self‑managed memory fails in two opposite modes: an archive that only grows or a plan rewritten every season. Code is available at https://github.com/Analogy‑AI/fm‑bench.

Authors:Ryozo Masukawa, Sanggeon Yun, SungHeon Jeong, Hyunwoo Oh, Raheeb Hassan, Pietro Mercati, Nathaniel D. Bastian, Mahdi Imani, Mohsen Imani
Title: Vector Symbolic Policy Gradient
Abstract:
We answer this question with Vector‑Symbolic Policy Gradient (VSPG), a discrete‑action actor that represents each action by a unit‑norm hypervector and scores it by similarity to the encoded state. Under the standard softmax policy‑gradient surrogate, we prove that its update is exactly advantage‑weighted hypervector bundling followed by normalization, and therefore supports standard advantage estimators. We further show that each trained action hypervector is a fixed‑size compressed kernel memory, storing an advantage‑weighted kernel expansion over visited states and transferring evidence according to the encoder‑induced similarity. This provides a concrete mechanism that can support sample‑efficient learning without increasing inference‑time memory. Finally, for bipolar action memories, we prove that greedy action selection is stable under random bit flips, with failure probability decaying exponentially in the hypervector dimension. VSPG thus connects VSA action memories, log‑linear policy gradients, and kernel policy search while providing a quantitative robustness guarantee.

Authors:Tom van Nuenen, Pratik S. Sachdeva, Sahiba Chopra
Title: The Fabricated Front: Generative AI and the Opacity of Workplace Performance
Abstract:
Generative AI (GenAI) has become a fixture of workplace life. Current research asks chiefly what this implies for jobs and outputs, measured in productivity, displacement, or bias. What remains underexamined are the interactional reconfigurations that GenAI produces at work. The emerging concept of effort opacity has begun to fill this gap by highlighting the systematic decoupling of observable output from human engagement. When GenAI makes interactional cues less diagnostic, it weakens the reciprocal exchange that sustains collaborative trust. Extending this account of effort opacity, we examine the interactional mechanics that produce opacity in everyday workplace encounters. Drawing on Erving Goffman's dramaturgical framework and 1,250 interview transcripts from Anthropic's AI Interviewer dataset, we identify five opacity mechanisms through which workplace fronts are reorganized: voice (whose stance the words index), provenance (who can stand behind the artifact), vulnerability (whether the worker is uncertain), attention (whether the worker is engaged), and investment (how much labor the output reflects). We show that professionals defend the identity mechanisms while freely producing opacity around the labor mechanisms, and trace this asymmetry to the output‑centered organization of contemporary work, where deliverables already stand in for the labor process that produced them. The governance task, accordingly, is one of involvement management: specifying which forms of human involvement (attention, effort, judgment) must remain inspectable, and to whom. Workplace AI policies built on universal disclosure will systematically misrecognize a social field in which inspectability is already audience‑relative.

Authors:Gaston Besanson
Title: One Gate Is Not Enough: Composing Stateful Pre-Action Controls for Agentic AI
Abstract:
Agentic AI systems take consequential actions governed by more than one pre‑action control at once: authority, resource, and evidence gates that can admit, degrade, or remediate an action before it executes. This paper's central object is remediation‑induced control coupling: a remediation applied by one control can change the action, evidence, or context another control evaluates, invalidating that control's earlier judgment. We formalize this coupling and give a remediate‑and‑regate protocol that restores per‑action soundness in the current bounded, idempotent setting under its stated assumptions. We further show that the two implemented remediation operators (evidence substitution and resource‑budget downroute) do not commute ‑‑ a finite‑model checker finds concrete counterexample instances ‑‑ making remediation order part of the control‑plane semantics rather than an implementation detail. A governed evidence buffer that trusts its own most recent admitted write is a further instance of the same problem at the level of state ‑‑ current admissibility does not imply future reference trustworthiness ‑‑ and is vulnerable to poisoning from declared‑uncovered defect classes; two mitigations reduce, not eliminate, that exposure. Supporting results establish the exact condition under which positive‑weight linear aggregation of gate outcomes can compensate a member veto, a unified cross‑control Evidence Set, and that composition manufactures no new detection coverage, reported honestly. Empirically, on a deterministic open‑data artifact composing three published engines unmodified, CH1‑CH5 meet their registered decision rules across all 30 pre‑registered seeds; CH6 does so under W1 but not under the smaller W2 workflow, reported as such. This is a mechanism demonstration on open payload data with a synthetic metadata layer, not a claim about production prevalence.

Authors:Tommaso Apicella, Alessio Xompero, Andrea Cavallaro
Title: Reproducible Multimodal Affordance Prediction
Abstract:
Affordance prediction is the identification of potential actions an agent can perform on a target object from multimodal inputs. Affordance prediction methods are difficult to evaluate and compare due to heterogeneous problem formulations, inconsistent dataset annotations, incomplete reporting of experimental protocols, and limited information about deployment conditions. These limitations challenge fair benchmarking and performance comparison. To promote transparency, we propose the Affordance Sheet, a documentation detailing task formulation with its input modalities, model architectures and training information, datasets, and experimental protocols. Affordance Sheets enable reproducible benchmarking and reliable evaluation of affordance models for real‑world scenarios, including generalisation to novel conditions and human safety.

Authors:Tianchen Guan, Xinlei Lin, Royce Cheng-Yue, Xiangjun Wang, Shuyan Zhou
Title: ComponentBench: Diagnosing Component-Level Failures in Computer-Use Agents
Abstract:
Current evaluation of computer‑use agents is split between long‑horizon workflow benchmarks and atomic GUI‑grounding tests. This leaves an under‑instrumented middle layer: realistic component‑centered interactions (e.g., toggle a button set) that are short enough to diagnose and rich enough to capture the burdens of modern interfaces. We present ComponentBench, a benchmark and diagnostic pipeline for component‑level evaluation of computer‑use agents on modern web UIs. ComponentBench is organized around a library‑agnostic ontology of 97 canonical UI components instantiated as 2,910 programmatically verified tasks across widely used component libraries, paired with cleaned human reference trajectories that enable evaluation of both task success and interaction efficiency. Beyond task collection, we introduce a scalable pipeline for auditing realized structural difficulty after implementation and synthesizing structured failure analyses across tasks and component families. Evaluating seven models ‑‑ GPT‑5.4, Gemini 3 Flash, GPT‑5.4 mini, GPT‑5 mini, Gemini 3.1 Flash‑Lite, Qwen3‑VL‑235B, and UI‑TARS‑1.5‑7B ‑‑ across four observation and action spaces, we show that these design choices critically impact performance. Within a single shared harness, changing only the observation and action space shifts task success by more than 30% for the same model: GPT‑5 mini falls from 83.1% with accessibility‑tree observations to 48.9% with coordinate‑only Pixel control. Moreover, even the fastest configuration takes 3.7x as long as the matched human reference, and spatial manipulations that are trivial for humans continue to challenge current agents.

Authors:Florian Rascoussier
Title: Anytime Solver Evaluation with a Normalized Signed Primal Integral and Explicit Reference Policies - Extended Version
Abstract:
Anytime solvers return usable solutions before they terminate and improve them while time remains. Their progress is commonly summarized by a primal integral, a final gap, or convergence curves. Each summary leaves consequential choices open: how to score a run before its first feasible solution, whether poor incumbents are truncated, and whether the reference value is updated after the experiment or version‑frozen beforehand. These choices become visible when runs produce no valid incumbent or improve a published best‑known value. We study a normalized signed primal integral built from a bounded relative gap. It assigns an intrinsic worst value to an empty run, requires no acceptance threshold, and assigns negative instantaneous gaps to incumbents that beat a frozen reference. We compare the smooth difference‑over‑sum kernel with a signed version of Berthold's max‑normalized gap and use the former as a working default. We also distinguish analysis‑time from version‑frozen reference policies and recommend reading the score with the raw final gap, mean convergence curve, and target‑attainment curve. We evaluate these choices on a five‑arm routing campaign and two model‑fidelity ladders. The two signed kernels preserve every panel ordering in this study, whereas a common acceptance threshold compresses the distances between arms and reverses one panel ordering. Updating 54 of the 212 reference values changes score levels and removes all negative scores, while leaving the observed panel orderings unchanged. An exploratory screen also identifies nine panel comparisons in which similar integral scores conceal materially different endpoints or attainment rates. The implementation, frozen inputs, and generators are openly released.

Authors:Shriniwas Ramesh Suram
Title: Cacheable by Design? Training Mixture-of-Experts Routers for Locality Against the Edge Memory-Bandwidth Wall: A Pre-Registered Negative Result with a Systems Measurement Study
Abstract:
Serving a 235B‑parameter Mixture‑of‑Experts (MoE) model on a single 8 GB GPU is bottlenecked not by compute but by memory bandwidth: decode must stream each token's active experts from whichever tier holds them, and on consumer hardware most experts sit on an SSD far slower than RAM. We quantify this bandwidth wall on Qwen3‑235B (Q4_K_M, 134 GB): measured decode is 0.44 tok/s warm, matching a bytes‑per‑token / bandwidth model, while a batching scheme that should amortize one disk sweep instead collapses at batch 32 from paging thrash. We build llama‑moe‑trace, a zero‑surgery router‑telemetry tool, and measure routing on Qwen3‑30B: adjacent‑token expert reuse is 2.0x chance, 95% of traffic uses 52.5% of experts, and an LRU cache of 13.4% of experts serves 66% of requests. We then ask whether cacheability is trainable: we pre‑register training of 137M MoE language models with auxiliary locality and domain router losses, under joint criteria on cache‑miss reduction and perplexity. The mechanism works (misses down up to 60%; a 99% static‑pin hit rate) but every configuration fails the pre‑registered <=1% perplexity gate ‑‑ miss reduction and quality are tightly coupled. Concurrent StickyMoE reports the same loss as near‑free on single‑domain sub‑25M models; on multi‑domain 137M we find the tax real. Our contribution is this pre‑registered, stricter‑criterion, multi‑domain evaluation plus edge‑serving measurements. A 340M rung shows the tax does not shrink with scale (it rises slightly). We further show training‑free cache‑aware rerouting stacks with trained locality ‑‑ together ~80% miss reduction at <=3.4% perplexity at both sizes, far cheaper than either alone ‑‑ while domain‑primed prefetching does not help. All code, traces, and the pre‑registration are released.

Authors:Amanuel Ergogo, Diego Dall'Alba, Przemyslaw Korzeniowski
Title: VERAGMIL: Virtual Environment for Scooping Granular Foods with Imitation Learning Models
Abstract:
Robot‑Assisted Feeding (RAF) systems are essential for assisting individuals with disabilities or motor impairments in eating tasks. Manipulating granular food items, such as rice and beans, poses significant challenges due to their dynamic physical properties. Learning from human demonstrations offers a promising solution, but acquiring high‑quality demonstrations is complex. To address this, we present VERAGMIL, a framework that combines a high‑fidelity simulator with an intuitive Virtual Reality (VR) interface for recording demonstrations and supporting different imitation learning methods. VERAGMIL provides a realistic environment for training RAF systems to handle granular materials, including robots, sensors, and various food items with distinct physical characteristics. We evaluate VERAGMIL by training three imitation learning models, BC, BC‑RNN, and BCQ, on granular scooping and transporting tasks using both VR interface and 3D space mouse demonstrations, comparing them with a human‑expert baseline. The models are assessed on success rate, spillage, generalization to unseen food items, and task completion time. Results show that VR‑based demonstrations significantly outperform 3D space mouse data, with BCQ achieving the best overall performance, particularly in reducing spillage and approaching human performance. These findings underscore the effectiveness of our framework for training RAF systems in granular material handling. The code for our framework is publicly available at: https://github.com/AmanuelErgogo/VERAGMIL.git.

Authors:Stefano Goria
Title: ClosureBench: A Constructive Benchmark for Compositional Graph Reasoning
Abstract:
We introduce ClosureBench, a constructive benchmark for compositional graph‑relational reasoning with programmatically verified ground truth. Unlike fixed‑test‑set benchmarks vulnerable to data contamination, ClosureBench generates instances on demand: each task's reference answer is computed by executing a program in the Ein tensor‑logic language, ensuring machine‑verified correctness. The benchmark spans 26 task categories at three compositional levels (L1‑L3), with difficulty controlled along three independent axes: graph size, edge density, and query depth. We evaluate models from 1.5B open weights to frontier systems (o3, GPT‑4.1, Gemini 2.5, Claude Sonnet 4) and report three findings. First, because the benchmark can always supply fresh instances, it measures memorisation directly: a model fine‑tuned on a fixed test set shows a 19.3 percentage‑point gap between its accuracy on seen and on fresh instances, which a static test set cannot reveal. We scope this to supervised fine‑tuning on answer pairs, not pretraining contamination. Second, accuracy falls as graph size and query depth increase, and the two interact: models misread the graph from its natural‑language description and then reason correctly over the wrong graph, so even the strongest frontier model degrades from atomic to compositional queries. This bottleneck is a property of the reasoning rather than the input format: it persists when the graph is given as a JSON edge list or an adjacency matrix instead of prose. Third, a 4B model fine‑tuned to emit executable programs rather than answers stays nearly flat across compositional levels and approaches frontier accuracy (94.3% on held‑out instances) at a fraction of the token cost. This holds for two program targets, Ein and Python+NetworkX, so it is a property of verified program synthesis rather than of one language.

Authors:Ziyang Cheng, Tianshu Tang, Jinxin Lan, Xinze Chen, Yuhan Gong, Zhichao Liu, Changzhong Wu, Yahao Mao, Zongyan Deng, Mingxuan Ma, Huasen Xi, Yilong Liu, Yutong Wu, Xiaofeng Wang, Yang Wang, Yun Ye, Guan Huang, Xiaojie Jin, Zheng Zhu, Jiwen Lu
Title: GigaBrain-WBC-0.5: A Behavior World Model for Robust Whole-Body Control with Environment Interaction
Abstract:
Whole‑body motion tracking policies turn a humanoid into a robust control interface: the teleoperator‑‑‑or an upstream model‑‑‑only supplies a coarse movement intent, while the low‑level policy keeps the robot balanced and physically feasible. Existing trackers deliver this interface only on flat ground: trained in empty scenes, they never learn how contact with terrain and objects reshapes their dynamics, and they attempt to teach the policy to balance under any command by continually enlarging the reference‑motion corpus, which stops working once feasible behaviors become environment‑dependent. We present GigaBrain‑WBC‑0.5, the first Behavior World Model (BWM) for humanoid whole‑body control. Rather than a purely reactive tracker, we train a causal Transformer to jointly predict its next action, next state, and the distribution over its next latent behavior command, so the network that acts also models how the environment shapes what it can do next. An automatic terrain‑annotation pipeline recovers full 3D contact geometry from retargeted motion, enabling terrain annotation at the scale of existing motion datasets. The predicted distribution is reused at deployment to detect implausible commands online and retract them onto learned behaviors, so the robot attempts tasks in a "best‑effort" manner. The result is a unified policy that takes real‑time command, interacts with environment, and stays robust to implausible commands, falls, and disturbances. GigaBrain‑WBC‑0.5 achieves the highest success rate across all four regimes among three large‑scale tracker baselines: 81.3% on terrain interaction (4.3x the strongest baseline), 83.1% under implausible commands, and 99.3% fall recovery (16.8x the strongest baseline). Hardware trials show robust interaction under missing supports and disturbances; the Unitree G1 checkpoint transfers to the Maker L01 robot with simple fine‑tuning.

Authors:Yiwen Chen, Matheus Gadelha, Huaizu Jiang
Title: LumiTokens: 3D Relighting via Token-Space Lighting Transformation
Abstract:
Existing 3D relighting methods operate through either explicit material decomposition, diffusion‑based view‑space generation, or a combination of both, requiring full recomputation for each new lighting condition. We observe that recent latent scene representations, which encode multi‑view images into a set of compact tokens with no fixed physical semantics, open up a novel design space for relighting. We present LumiTokens, a framework that formulates 3D relighting as a direct transformation on latent scene tokens, without explicit 3D representations, rendering equations, or physics‑based decomposition. Our model introduces a Scene Token Editor that processes scene tokens jointly with light‑ray tokens through self‑attention, producing updated tokens that can be decoded into multi‑view‑consistent relit images. To support diverse lighting types through a unified interface, all lighting signals, including environment maps, point lights, and area lights, are parameterized as Plucker ray tokens, enabling native 3D user interaction with a representation that carries no explicit spatial structure. Crucially, this design supports progressive relighting: because the editor's output remains in the same latent space as its input, a user can incrementally build up illumination one light source at a time, with each edit composing in token space. Experiments demonstrate that LumiTokens achieves comparable or superior relighting quality to other methods and supports progressive, composable lighting edits. Project page: https://neu‑vi.github.io/LumiTokens/

Authors:Yang Chen, Tianqi Wang, Xiaorui Jiang, Yilei Man, Yihua Shao, Mengyuan Liu, Zhi Chen, Xiaofeng Cao, Qibin Zhao, Chi Harold Liu, Albert Y. Zomaya, Nicu Sebe, Jingren Zhou, Dacheng Tao, Song Guo, Jingcai Guo
Title: Human-Centric Intelligence in the Era of Foundation Models: A Survey
Abstract:
Human‑centric intelligence is evolving in the foundation‑model era, with growing emphasis on scale, transferability, and general‑purpose modeling. Yet it has not fully integrated with foundation models to achieve the comparable progress seen in them. More importantly, recent advances across this broad landscape remain fragmented across tasks, modalities, and research communities, leaving their intrinsic conceptual and methodological connections unclear. To bridge these divides and rethink human‑centric intelligence in the foundation‑model era, we introduce a full‑spectrum human context taxonomy that integrates six interconnected levels by viewing humans as observable subjects through visual appearance and spatial geometry, as dynamic actors through kinematic dynamics and interaction modeling, and as situated agents through world simulation and embodied agency. We next present the methodological foundations of the field, covering human‑centric data families, computational architecture paradigms, and representative training and inference optimization strategies. We then systematically review representative methods across these levels and organize the associated datasets, benchmarks, and evaluation metrics. We further discuss open challenges and promising research directions toward human‑centric intelligence that is scalable, trustworthy, physically grounded, and deployable, aiming to provide a coherent framework and practical reference for advancing the field. Finally, we provide a systematically organized and continuously updated collection of human‑centric AI literature and resources on our project page.

Authors:Gurucharan Marthi Krishna Kumar, Janine Dale Mendola, Amir Shmuel
Title: TractoGraphVLM: A Unified Vision-Language Framework for White Matter Tractography
Abstract:
Vision language models have transformed 2D medical imaging, yet extending them to 3D white matter tractography remains challenging due to the complex topology of fiber bundles. We introduce TractoGraphVLM, a unified framework for four tasks, bundle classification, text‑to‑tract retrieval, anatomical captioning, and visual question answering, built on a shared GPS architecture, training procedure, and read‑out design. Fiber bundles are represented as streamline graphs whose nodes encode 3D position and tangent orientation. A General, Powerful, Scalable (GPS) graph transformer produces bundle embeddings aligned with a frozen BiomedBERT text encoder via contrastive learning, while a BioGPT decoder with visual prefix tokens generates captions and answers. A single shared encoder and decoder is trained jointly across all four tasks and evaluated from one checkpoint. Trained on HCP Young Adult subjects, TractoGraphVLM achieves 91.8% bundle classification accuracy, 84.7% retrieval R@1, BLEU‑4=20.1, ROUGE‑L=66.8, and 66.4% VQA accuracy on a held‑out test set. The same checkpoints transfer zero‑shot to HCP Aging subjects, with a modest drop on discriminative tasks and a larger drop on generative tasks, showing robustness to age and acquisition shift. Language supervision yields richer representations than label‑only training, recovering structure like hemisphere and fiber family, carried by captions but never given as a label. Swapping only the visual encoder, graphs preserving fiber orientation outperform volumetric baselines, with GPS giving the best balance. Generative metrics measure consistency with a structured knowledge base rather than independent clinical text; even so, TractoGraphVLM shows that classifying, retrieving, describing, and answering questions about a white matter bundle can be served by one jointly trained model that learns transferable neuroanatomy from language alone.

Authors:Igor Itkin
Title: Temporal Multi-Signal Fusion for Token-Level Hallucination Detection
Abstract:
Token‑level hallucination detectors score each token independently from a single signal, and fail exactly when the generating model is confidently wrong. This paper instead treats hallucination as a temporally extended span and detects it by sequence labeling: each token is scored from a 33‑dimensional feature stream that fuses text statistics, Natural Language Inference (NLI) entailment, and language model surprisal, with no access to model internals. A Bidirectional Gated Recurrent Unit (BiGRU) over these features reaches an AUC of 0.840 on RAGTruth (10 seeds), an 11‑point gain over an independent logistic‑regression baseline (p = 0.002, Wilcoxon signed‑rank). A controlled decomposition attributes most of the gain to temporal order rather than model capacity: evidence propagates from confident positions to ambiguous neighbors within a span. The same 0.845 ceiling recurs across recurrent, state‑space (Mamba), and attention architectures, locating the bottleneck in the feature set rather than the model. Because it reads only the generated text and external signals, the detector works on closed‑source models, and it keeps working on text produced by language models it never saw during training, losing under 4% AUC.

Authors:Maikel Leyva-Vazquez, Florentin Smarandache
Title: Institutional Prestige as Geographic Bias in Large Language Models: Evidence from Three Factorial Experiments with Bootstrap Confidence Intervals
Abstract:
We investigate whether large language models (LLMs) systematically discriminate in candidate evaluations based on applicant name ethnicity and/or institutional prestige and geographic location. Three factorial experiments are reported (4,320 API calls, four LLMs, five professional domains). Study 1 (3x4 design) finds a statistically robust institution‑tier gradient of +0.297 points on a 10‑point scale (95% bootstrap CI: +0.175 to +0.422), while name‑origin effects are negligible and non‑significant (95% CI crosses zero). Study 2 (2x2 Prestige x Country design) breaks the prestige‑geography confound: the prestige effect (+0.185; 95% CI: +0.093 to +0.275) exceeds the country‑of‑origin effect (+0.126; 95% CI: +0.037 to +0.218) by 1.5x. Study 3 (2x2 Journal x Institution design) reveals that journal prestige (Nature vs. a peripheral open‑access journal) dominates institutional prestige by 5.7x: journal effect +1.937 (95% CI: +1.811 to +2.062) vs. institution effect +0.341 (95% CI: +0.184 to +0.504). A "rescue effect" is confirmed: publishing in Nature compensates for low institutional prestige more strongly for candidates from the University of Guayaquil (+2.127) than from MIT (+1.745). Results are quantified using the Neutrosophic Bias Index NBI<T,I,F>; the I component reveals elevated evaluation inconsistency for low‑prestige profiles, an epistemic disadvantage not captured by mean‑only metrics. Code and data: https://github.com/mleyvaz/geo‑bias‑llm

Authors:Yuanyuan Xu, Wenjie Zhang, Yin Chen, Xuemin Lin, Ying Zhang
Title: Self-Evolving Agents as Dynamic Graph Transformation: A Survey and New Perspective
Abstract:
Large language model (LLM)‑based agents are increasingly becoming self‑evolving systems that persist across interactions, maintain memories, use tools, acquire skills, refine workflows, and coordinate with other agents. These capabilities make agent states structural and dynamic: entities, relations, attributes, dependencies, and execution structures change with new evidence, feedback, and environmental conditions. Existing graph‑agent surveys typically treat graphs as support structures for agent functions rather than as evolving substrates, while self‑evolving‑agent surveys focus on agent‑level mechanisms and rarely discuss graph topology evolution. Thus, the coupling between evolving agent state and dynamic graph topology remains underexplored. This survey connects these two research lines by framing agent evolution as dynamic graph transformation. We model agent state as a dynamic graph, where memories, tools, skills, workflows, and inter‑agent relations are represented as typed nodes, edges, and subgraphs updated through schema‑constrained rewrites. Based on this formulation, we organize existing dynamic‑graph‑based methods for self‑evolving agents into four taxonomies: node/feature evolution, edge/topology evolution, subgraph activation, and cross‑component co‑evolution. Building on this taxonomy, we propose dynamic graph learning as reusable infrastructure for self‑evolving agents and map nine dynamic‑graph‑learning subfields to agent‑evolution capabilities, discussing their adaptations and possible failure modes. Finally, we discuss five types of graph‑aware evaluation and governance protocols from a dynamic‑graph perspective, which complement end‑task evaluation. The goal is to provide a compact structural lens for designing and governing self‑evolving agents.

Authors:Zijuan Zhao, Zheren Fu, Hou Xia, Licheng Zhang, Yi Liu, Zhendong Mao
Title: MAVEN: A Macro-Societal Value Evaluation Framework of Multimodal Content with Compact Aligned Evaluators
Abstract:
Assessing whether multimodal content aligns with macro‑societal values, such as peace, justice, and freedom, has become an increasingly urgent challenge. Existing frameworks are largely confined to safety‑oriented taxonomies, text‑only psychometric probes, or single‑label classification. Therefore, we propose MAVEN, a hierarchical framework for macro‑societal value evaluation of multimodal content, grounded in international human‑rights instruments and cultural value theory. MAVEN organizes values into 6 primary dimensions and 72 secondary indicators, supporting multi‑level quantitative scoring. Building on MAVEN, we construct a human‑verified multimodal benchmark and a soft‑match metric to evaluate VLMs' assessments across value dimensions. For evaluator optimization, we propose a span‑adaptive variant of multi‑level preference optimization for evaluator distillation, together with a training‑free multi‑role consensus strategy at inference time. We evaluate existing open‑ and closed‑source VLMs on our benchmark, revealing shared tendencies and clear differences in macro‑societal value judgments. Experiments show that our compact 2B evaluator matches its 8B counterpart in the same family and approaches frontier closed‑source VLMs, offering a practical path toward scalable macro‑societal value evaluation. Our SA‑MDPO implementation and MacroValue‑Bench are available at https://github.com/zzzzzzzzjj/MAVEN.

Authors:Ruizhi Zhang, Jinwei Chen, Xiangju Lu, He Yan, Mo Yu, Junmin Zhu, Wei Zhang
Title: LongNovel: A Multi-Scale Benchmark for Hallucination Detection in Long-Context Novel Summarization
Abstract:
Although context windows have expanded significantly in recent years, hallucinations in long‑context summarization remain a challenge. Long novels are better suited than news or papers for researching these hallucinations, due to their intrinsic information and detailed descriptions of events and dialogues. However, current research lacks a multi‑scale benchmark for hallucination detection in long‑context novel summarization and does not fully explore how hallucinations change as the context grows longer. In this study, we propose LongNovel, a multi‑scale long‑context bilingual (Chinese and English) novel benchmark for hallucination detection. This benchmark is constructed from 29 Chinese novels (ranging from 16k to 100k tokens) and chapter‑level data from the BookSum dataset. We design 8 hallucination types and employ a combination of Multi‑Model Arbitration and Entity‑Referenced Hallucination Generation to ensure both data authenticity and a balanced distribution of hallucination categories. Furthermore, we manually revise the content in the test set to guarantee data reliability. Extensive experimental results demonstrate that LongNovel is a challenging benchmark. We release LongNovel for future research. https://github.com/BDML‑lab/LongNovel

Authors:Hongyu Li, Bowen Wen, Xinghao Zhu, Yixuan Wang, Yilun Du, Yunzhu Li, George Konidaris, Stan Birchfield, Soha Pouya, Chenran Li, Yan Chang
Title: Hydra-0: Action Flow for Generalist World Modeling and Control
Abstract:
We introduce Hydra‑0, a generalist world model conditioned on action flow, which represents robot actions as pixel motion. This shared visual interface enables generalist world modeling and control by learning action consequences across embodiments, tasks, environments, and video‑generation backbones. Our best configuration achieves 90.4% lower robot‑motion error and 60.2% lower object‑motion error than our action‑conditioned baseline, while supporting zero‑shot composition and data‑efficient adaptation. On the RoboLab benchmark, Hydra‑0 achieves a Pearson correlation of r=0.96 between replayed and reference success rates. Finally, we uncover an emergent inverse mode of this interface: a world action model that predicts compatible robot motion from desired object flow transferred from a human demonstration. A trained action head maps the resulting latent features to executable actions without requiring task‑specific expert robot demonstrations. Together, these results demonstrate the potential of action flow as a shared control interface connecting heterogeneous training data, open‑loop policy evaluation, and robot control.

Authors:Qinyuan Ye, Yu Li, Yada Pruksachatkun, Jiaxin Zhang, Chien-Sheng Wu
Title: On the Fragility of Self-Improving Agents: Variance, Task Order, and Underspecification
Abstract:
Memory‑based self‑improving agents‑‑those that learn from an online stream of tasks and improve over time by maintaining a textual memory bank‑‑have shown great promise in recent literature. However, the reliability aspects of these methods have been critically overlooked. In this work, we conduct a comprehensive re‑evaluation of two memory‑based methods, broadening the scope of evaluation along two axes: (1) including multiple runs to quantify variance, and (2) randomly shuffling the tasks to investigate the effect of task order. Through these experiments, we make two observations that expose the fragility of current methods: First, agent evaluation is inherently noisy in complex environments and on multi‑step tasks, and stacking a self‑improving loop on top can further amplify this noise. Second, the agent's improvement is highly dependent on task order. Prior works often adopt default orderings that impose an implicit curriculum, acting as a hidden prerequisite for success. To better understand this fragility, we manually examine the agents' memory and hypothesize that task and environment underspecification contribute to this fragility. We validate this hypothesis by incorporating information that enables better specification, such as detailed rubrics and environment feedback, into the memory construction process. While this added information partially closes the performance degradation in previous experiments, significant gaps still remain, suggesting that other uncharacterized factors contribute to this fragility. Looking ahead, our work advocates for more rigorous evaluation protocols for self‑improving agents by reporting results across multiple runs and stress‑testing them under challenging conditions. Moreover, our findings on underspecification call for systems and interfaces that enable effective human oversight, preventing agents from failing in unforeseeable ways.

Authors:Clara Meister
Title: TokEval: A Tokenizer Evaluation Suite
Abstract:
Language model tokenizers are typically selected with minimal evaluation, despite the fact that their design choices directly impact model capabilities. This can be partly attributed to a limited understanding of which tokenizer properties affect which aspects of downstream performance. We introduce TokEval, a framework of tokenizer evaluation metrics that goes beyond standard measures like fertility and compression rate to capture linguistically and structurally meaningful properties, e.g., UTF‑8 character boundary integrity and digit place‑value boundary alignment for mathematics. To validate whether these metrics are predictive of downstream model performance, we conduct controlled language model pretraining experiments, varying solely the tokenizers' training data mixture, pretokenization strategy, and training algorithm. We evaluate the resulting models on bits‑per‑byte (a tokenizer‑agnostic version of perplexity) and several benchmarks, spanning linguistic understanding, mathematical reasoning, and code generation. Our experiments suggest that different intrinsic properties have different impacts on model abilities: information‑theoretic metrics predict language modeling abilities (Spearman rho up to 0.80), while structure‑sensitive metrics, such as those measuring digit and line‑break handling, correlate with task accuracy. We hope TokEval enables more principled tokenizer evaluation, replacing pretraining sweeps with intrinsic measurement wherever the two agree.

Authors:Veronika Spieker, Wenqi Huang, Cemre Ariyurek, Liam Timms, Daniel Rueckert, Onur Afacan, Julia A. Schnabel, Sila Kurugol
Title: Primitive Representation Learning for Unsupervised Dynamic Contrast Enhanced MRI Reconstruction
Abstract:
Reliable quantitative analysis of dynamic contrast‑enhanced MRI requires high‑quality spatiotemporal reconstructions at high undersampling rates. Scan‑specific reconstructions using Gaussian and Gabor primitives have shown promising results without the need for large training datasets, but have not addressed the additional dimension of dynamic contrast. We propose a multi‑dimensional, primitive based framework for dynamic contrast‑enhanced MRI reconstruction that disentangles the underlying anatomy, the dynamic contrast enhancement, and residual motion into separate temporal basis functions, thereby enabling a geometrical interpretation of the representation. We show that this architecture achieves performance competitive with conventional reconstruction methods, both in reconstruction quality and in the accuracy of extracted aorta and kidney enhancement curves. The modular tier design extends naturally to additional dynamic factors and higher acceleration rates. Code available at https://github.com/compai‑lab/2026‑GaborDCE‑spieker.

Authors:Zongzheng Zhang, Jijun Wang, Saining Zhang, Shuo Wang, Yiru Wang, Hai Yang, Yang Chen, Yuwen Heng, Hao Sun, Anqing Jiang, Hao Zhao
Title: Plug-and-Play Traffic Element Awareness for End-to-End Autonomous Driving
Abstract:
Traffic elements such as traffic lights and road signs play a fundamental role in human driving decisions and should naturally influence end‑to‑end driving performance. However, existing end‑to‑end driving research predominantly focuses on dynamic road participants (e.g., vehicles and pedestrians), while the role of traffic elements remains largely unexplored. The community still lacks a systematic study quantifying their impact, largely because public datasets rarely provide structured traffic‑element annotations and modern driving systems vary widely in architecture and training paradigm. In this work, we present the first systematic investigation of traffic element awareness for end‑to‑end autonomous driving. We construct a unified research infrastructure by augmenting multiple public driving datasets with comprehensive traffic‑element annotations. To support diverse model families, we adopt a minimal and universal integration design that incorporates traffic‑element signals into existing pipelines in a plug‑and‑play manner with negligible architectural modification. We evaluate this design across modern paradigms, including perception‑prediction‑planning pipelines, vision‑language‑action models (VLA), regression‑based planners, diffusion‑based policies, and trajectory‑scoring frameworks, on nuScenes, NAVSIM‑v1, NAVSIM‑v2, and Bench2Drive. Across all paradigms and datasets, this simple integration consistently improves driving performance, demonstrating that traffic element awareness provides a robust and generalizable signal for end‑to‑end driving systems. Notably, on the challenging NAVSIM‑v2 benchmark, our approach significantly improves state‑of‑the‑art architectures and data pipelines, establishing a new state of the art.

Authors:Zhikai Xu, Zhucun Xue, Teng Hu, Yabiao Wang, Yong Liu, Jiangning Zhang
Title: Deep Academic Survey: Stateful Agentic Closed-Loop Paradigm for Academic Survey Automation
Abstract:
Academic surveys play a central role in organizing rapidly expanding scholarly literature, yet their construction requires extensive paper analysis, coherent knowledge organization, fine‑grained citation support, and reliable manuscript assembly. Existing Deep Research and automated survey generation systems address parts of this process, but typically do not coordinate paper understanding, literature organization, evidence‑grounded drafting, and manuscript validation through a shared, revisable state. We introduce DAS, a stateful agentic framework for generating publication‑oriented academic surveys. Its key idea is to separate reusable paper analysis from topic‑specific manuscript construction. DAS builds on DAS‑2M, a dynamically updated metadata lake containing survey‑oriented representations of approximately two million papers. Its agents maintain explicit literature, organization, writing, and finalization states through candidate‑grounded taxonomy planning, reverse paper‑to‑section routing, and hierarchical claim and citation planning. Semantic review reactivates only the affected writing states for repair and reevaluation, forming a scoped closed loop with deterministic validation. We further introduce DAS‑Bench, a 30‑topic benchmark, together with DAS‑Eval, which assesses scholarly citation quality, taxonomic synthesis, hierarchical discourse, and manuscript assembly reliability through 16 criteria. Among systems evaluated on all 30 topics, DAS achieves the highest average in all four dimensions, with an overall score of 4.34 compared with 4.03 for the strongest competitor, and the same ordering is preserved on the matched 21‑topic CS subset. Blinded expert evaluation further prefers DAS to Naive RAG on 27 of 30 topics and to AutoSurvey on 19 of 21 shared CS topics. The project page is available at https://zhikaixu24.github.io/projects/DAS/.

Authors:Simon Weber, Mateo de Mayo, Je Hyeong Hong, Carl Olsson, Daniel Cremers, Ronald Clark
Title: Initialization-Free Bundle Adjustment Revisited: A Controlled Experimental Study
Abstract:
Initialization‑free bundle adjustment (InitFree BA) aims to recover camera poses and scene structure directly from image observations, avoiding the geometric initialization stages of conventional structure‑from‑motion pipelines. Recent methods based on Object‑Space Error (OSE) formulations and Variable Projection (VarPro) show encouraging optimization behavior from random camera configurations. However, existing evaluations primarily measure optimization success, leaving unclear whether a low OSE objective yields a valid metric 3D reconstruction. We revisit InitFree BA experimentally through a unified evaluation framework combining a C++ implementation of existing OSE formulations with a Blender‑based dataset generator providing exact ground truth and controlled camera configurations and observation densities. Our experiments reveal a previously overlooked optimization‑‑reconstruction gap: projective solutions with similarly low OSE values can lead to substantially different Euclidean reconstructions after metric upgrade. We identify initialization priors, landmark observation density, and metric‑upgrade stability as key factors governing reconstruction success. Overall, our results suggest that the main challenge of InitFree BA is not merely minimizing OSE objectives, but obtaining projective reconstructions that admit reliable metric upgrade. We believe that the proposed benchmark, implementation, and analysis establish stronger experimental foundations for future research on initialization‑free bundle adjustment, a problem largely unexplored within the computer vision community. Project page is available at https://github.com/simonwebertum/InitFreeBA.git.

Authors:Hsiang-Wei Huang, Fu-Chen Chen, Li-Wu Tsao, Cheng-Han Lee, Che-Chun Su, Lu Xia, Ronghui Peng, Jenq-Neng Hwang, Min Sun, Cheng-Hao Kuo
Title: Memory Tree Guided Key Frame Querying for Efficient 3D Question Answering
Abstract:
Answering questions accurately and efficiently in embodied scenarios presents significant challenges due to limited computational and memory resources for Vision Language Model (VLM) inference. Existing methods adopt visual search key frame retrieval method to select critical question‑related key frames for VLM input. However, visual search methods are inefficient because they require visual search among thousands of video frames for each individual user query. In this work, we propose a memory tree guided key frame selection paradigm for efficient 3D question answering in embodied scenarios. Our method leverages a compact and reusable 3D scene representation, termed MemTree3D, which supports real‑time online construction leveraging camera 6‑DoF poses. MemTree3D captures multi‑level 3D scene information, enabling a Large Language Model to efficiently query and retrieve question‑relevant key frames through our scoring‑based frame selection without reprocessing the entire video stream. On OpenEQA, our method improves the LLM‑Match of GPT‑4o by 17.4%, LLaVA‑OneVision‑7B by 5.8%, outperforms existing visual search methods. Our code is available at https://github.com/hsiangwei0903/MemTree3D

Authors:Haoran Qin, Zhengan Yan, Shikang Zheng, Xiaobing Tu, Jiacheng Liu, Yuqi Lin, Chang Zou, JinShan Liu, Peiliang Cai, Xiantao Zhang, Jinkui Ren, Linfeng Zhang
Title: AViTS: Adaptive Spatiotemporal Token Selection for Efficient Dynamic-Resolution Generation
Abstract:
Diffusion Transformers (DiTs) achieve high‑quality generation but are costly due to iterative sampling. Dynamic‑resolution sampling reduces early‑stage cost by denoising at low resolution; however, uniformly upsampling all latent tokens at resolution transitions incurs redundant computation and may degrade fine‑detail consistency. Existing partial upsampling strategies typically rely on local latent structure cues or single‑step statistics, making it difficult to jointly capture token‑text semantic relevance and token‑wise representation dynamics across diffusion steps. We propose AViTS, an adaptive spatiotemporal token selection framework for dynamic‑resolution DiTs. AViTS models spatial importance via latent‑text attention and temporal importance via token‑level feature variation across diffusion timesteps, and fuses them to enable spatiotemporal importance‑aware selective upsampling: it prioritizes resolution refinement for critical tokens while deferring less important ones, thereby reducing redundant high‑resolution computation and improving the quality‑efficiency trade‑off. AViTS achieves up to 6.34x on FLUX and nearly 9x FLOPs reduction on Qwen‑Image‑Edit and FLUX.1‑Kontext‑dev, orthogonal to distillation, quantization, and feature caching, and reaching 14.76x with distilled models. Code: https://github.com/QHR69/AViTS

Authors:Rui-Huan Wang, Si-Tong Wei, Jia-Qi He, Heng-Yi Wei, Baoquan Chen, Peng-Shuai Wang
Title: aDSL: Agentic 3D Creation via Joint Agent-Program Design
Abstract:
Programmatic representations provide a compelling paradigm for 3D content creation, enabling fine‑grained edits, interpretability, and explicit structural control. Yet, agentic workflows that rely on large language models (LLMs) to author 3D programs remain brittle, often failing to translate high‑level intent into consistent low‑level geometry. We attribute this fragility to a mismatch between existing programmatic interfaces and the reasoning strengths of LLMs, which favor semantic structure and spatial relations over fragile numeric choices. In this paper, we jointly design an Agent‑centric Domain‑Specific Language (aDSL) and a role‑specialized multi‑agent system to close this gap. aDSL bridges semantic logic and geometric constraints by emphasizing composability and spatial reasoning; it enables agents to manipulate geometry through relational operators instead of brittle absolute coordinates. Building on aDSL, our training‑free multi‑agent system follows a Plan‑Execute‑Critic loop to decompose requests, synthesize code, and iteratively repair errors and constraint violations using execution feedback. Experiments show that this co‑design improves robustness, controllability, and faithfulness to user intent. Our method outperforms prior LLM‑based baselines on text‑to‑shape and image‑to‑shape tasks while preserving explicit structure, editability, and interpretability. It also enables downstream applications such as articulated object creation and structured scene composition. Our code is available at https://github.com/sig‑pku/aDSL.

Authors:Jinshan Liu, Haoran Qin, Xiaobing Tu, Jiacheng Liu, Jiahui Hu, Zhengan Yan, Yukun Xie, Kerui Shen, Jinkui Ren, Yuqi Lin, Xiantao Zhang, Linfeng Zhang
Title: LinCa: Accelerating Diffusion Models via Learnable Decomposed Feature Caching
Abstract:
Diffusion models have achieved remarkable success in image and video generation, yet the high computational cost of iterative sampling remains a critical bottleneck for practical deployment. Feature caching has emerged as a promising acceleration paradigm by reusing or predicting intermediate features across timesteps. However, existing training‑free methods apply uniform prediction strategies that cannot adapt to the heterogeneous feature dynamics, causing significant quality degradation under high acceleration ratios. We propose LinCa, a feature caching framework based on learnable invertible networks. LinCa decomposes cached features into sub‑components with distinct continuity properties via a lightweight invertible network and applies differentiated prediction orders matched to each component. The strict invertibility guarantees lossless reconstruction back to the original feature space, forming a unified Decompose‑Predict‑Reconstruct pipeline. By training separate predictors for different models and timestep segments, LinCa adapts to heterogeneous feature dynamics. Experiments on FLUX, Qwen‑Image, and HunyuanVideo demonstrate that LinCa, with less than 0.2% additional parameters, significantly outperforms existing methods and maintains near‑lossless quality at 5‑7x speedup. Code: https://github.com/QHR69/LinCa

Authors:Tengbo Yu, Jiahao Wu, Hanning Wang, Rui Chen, Chuanhou Liu, Chuang Sun, Hangxin Liu
Title: PRISM: Precision and contact-rich Real-world Industrial Skill dataset with Multimodal sensing
Abstract:
Recent progress in robotic learning has been fueled by large‑scale datasets collected in everyday environments. However, most existing datasets emphasize short‑horizon, low‑contact tasks such as pick‑and‑place, and therefore do not capture the precision control, force/torque or tactile regulation, and multimodal feedback required for industrial assembly. To address this gap, we introduce PRISM, a large‑scale multimodal dataset for contact‑rich industrial operations. The dataset spans more than 25 manipulation tasks (e.g., electronic components plug/unplug, conveyor‑based sorting) and covers diverse mechanical constraints. PRISM includes more than 5,000 trajectories totaling 45 hours of teleoperated demonstrations, recorded using synchronized multi‑view RGB‑D, force/torque, tactile, and robot‑state measurements. In contrast to datasets collected in household or laboratory settings, PRISM provides a realistic benchmark for multimodal perception and control under high‑precision industrial constraints, and serves as a foundation for contact‑rich, generalizable manipulation in real‑world manufacturing environments. The dataset is open‑sourced at: https://tengbo‑yu.github.io/PRISM/

Authors:Javier Aguilar Martín
Title: An Omitted Mode Is a Rare Rule: The Sampling-Verification Danger Law in Continuous Code World Models
Abstract:
In the Code World Model paradigm an LLM synthesizes an executable world model that a classical planner searches, and the model is accepted when it reproduces sampled transitions. We ask what that acceptance certifies in continuous control. We define the pipeline's danger as an expected risk and isolate its exact factor: the probability that N i.i.d. gate rollouts all miss a critical event of probability r is exactly (1‑r)^N; an independent acceptance sample adds its budget to the exponent. On three hybrid instruments the accepted mode‑blind model is exploited: the planner is pinned at the mode boundary at a regret of nearly the whole attainable return. We prove a localization budget, valid at boundary points: models with Lipschitz constant at most L differing by eta at a point disagree above tolerance eps on a region of volume at least kappa((eta‑eps)/L)^(d+m); the discontinuous reset modes studied pay no such budget. With real LLM synthesis, GPT‑5.x repairs an omitted 1D clamp in 105 of 111 mode‑containing draws ‑‑ every attempt exact on 50 of 56 instrument‑stream blocks (95% CI [0.781, 0.960]). On 2D regions no artifact recovers the rule (0/156); eight targeted interventions leave the failure in place, and positive controls locate it: a located rule is not induced, while given form and location the constants follow exactly. A version‑space certificate proves identification is class‑relative: at the widest dose the declared fit succeeds in 20/20 blocks and every sample‑consistent circle is within tolerance in 18/20. We prove a class of entry rules exactly consistent with every sample yet harmless at play, so identifiability is a measurable property of the instrument. Re‑scoring all 1034 artifacts on independent samples confirms acceptance certifies sample consistency and no more: where the gate is provably informative it covers about two percent of the exploited planner's queries.

Authors:Chainarong Amornbunchornvej
Title: Collective Counterfactual Planning: Coordination, Consent, and Verification under Representational Constraints
Abstract:
Groups routinely complete projects that no single member can plan, execute, or verify alone. We propose a formal model of this phenomenon, Collective Counterfactual Planning (CCP), in which the binding limitation on each agent is neither capability, knowledge, nor observability, but representational geometry: each agent perceives the state, conceives moves, consents to actions, and certifies goal requirements only through a projection onto an agent‑specific subspace of a common task space. Four gates jointly determine whether a team can reach a conjunctive goal and legitimately recognize that it has done so: the exogenous implementation coalitions required to perform each action, together with three representational gates ‑‑ conception, consent, and task‑relative verification qualification. We define the Collective Counterfactual Solvability (CCS) problem, separating geometric feasibility, executable attainment, and validated completion. The results expose a positive‑negative duality. Iterated cross‑agent relay can unlock a solution that no one‑shot pooling of individual plans contains, but any goal requirement depending essentially on the subspace dark to the entire team is unverifiable and therefore not validly completable, even when the trajectory accidentally attains it. Memoryless and audited consent further constrain different objects ‑‑ action directions versus cumulative trajectory states ‑‑ and neither dominates the other. A four‑step exhaustive horizon‑bounded solvability scheme is sound and complete under exact representation of the relay closure; restricted implementations remain sound on returned plans but need not be complete. The model gives one geometry for sequential mutual enabling, competent execution of steps whose purpose is invisible to the executor, forced sub‑teaming at expertise boundaries, and completion that cannot be validly declared.

Authors:Shicheng Ma, Wenqian Cui, Irwin King
Title: SpeechSense: A Paralinguistic-Focused Dataset for Fine-Grained Speech Sentiment Analysis
Abstract:
Recent advances in AI have revolutionized speech processing, yet effective speech understanding requires discerning not just what is said, but how it is said. Speech Sentiment Analysis plays a critical role in decoding these paralinguistic cues for diverse real‑world applications such as recruitment and customer service. However, existing Speech Sentiment Analysis research faces two primary limitations. First, dominant approaches rely on text‑centric pipelines that cascade Automatic Speech Recognition with text analysis. This process inevitably discards essential acoustic features like prosody and tone, failing to capture attitudinal meanings in acoustically ambiguous utterances. Second, current benchmarks suffer from a mismatch in label granularity, prioritizing basic emotions (e.g., happy, sad) over the nuanced interpersonal stances (e.g., confident, impatient) necessary for social sensitivity. To address these limitations, we propose a novel dataset, SpeechSense, for fine‑grained speech sentiment analysis. Specifically, we define a specialized 8‑class taxonomy of interpersonal stances detectable primarily through prosodic cues beyond lexical content alone. We then construct a curated dataset based on this taxonomy, built from high‑fidelity speech synthesis and rigorous human validation. Comprehensive experiments across multi‑modal LLMs, text‑only LLMs, and speech encoders demonstrate that models with acoustic access consistently outperform text‑only baselines. These results empirically validate the primacy of acoustic cues in detecting subtle speaker attitudes, highlighting the necessity of SpeechSense. Dataset and supplementary materials are available at https://github.com/Sher13cked/SpeechSense.

Authors:Xinyang Gu, Zhilu Zhang, Honglei Xu, Yanting Mei, Yukang Ding, Wangmeng Zuo
Title: Improving Complex Moiré Removal with Generative Supervision
Abstract:
The availability of high‑quality paired data is essential for training learning‑based image demoiréing models. However, it remains challenging for existing datasets to encompass the complex moiré patterns captured in uncontrolled real‑world scenarios. Such degradations typically manifest as large‑scale, multicolored moiré patterns. Moreover, these patterns frequently occur in images for which clean counterparts are difficult to obtain, such as photographs acquired from public displays or existing online resources. In this work, we propose a novel data engine designed to improve the removal of complex moiré patterns by generating training supervision. Specifically, we initially collect real‑world images containing complex moiré patterns and localize the corresponding screen regions. Multiple image‑conditioned generative foundation models are subsequently deployed to produce candidate references. To establish reliable supervision, these candidates are subjected to patch‑level quality control to filter and select the optimal results. Based on this systematic paradigm, we construct the WildMoiré dataset, which contains 6.8K moiré‑GT training pairs. For evaluation, we additionally build an independent test set comprising ~250 pairs with captured clean ground truth. Extensive experiments on ESDNet, SDXL, and Qwen‑Image‑Edit demonstrate that the proposed generative supervision consistently improves the performance of complex moiré removal.

Authors:Ibrahim Mian, Shayaan Siddique
Title: A Kernel-Checked Exclusion Certificate for Erdős Problem 647
Abstract:
Erdős problem 647 asks whether any n > 24 satisfies \max_m<n(m + τ(m)) \le n + 2, where τ is the divisor‑count function. Computational searches have excluded solutions up to 10^12 by direct sieve and up to roughly 9.17 × 10^18 within a modular reduction whose Lean component relies on native_decide; those computations sit outside any proof kernel. We give the first exclusion checked end to end by one: no solution exists with 24 < n \le 10^9, proved in Lean 4 with axiom closure exactly propext, Classical.choice, Quot.sound ‑‑ no sorry, no native_decide, no problem‑specific axiom. The proof replays a chain of 6,685,922 factorization witnesses whose excluded intervals concatenate across (24, 10^9]; it needs no primality facts beyond primes below 1024, and it is the finite, fully proved form of a domination‑interval argument whose asymptotic step was the identified gap in a withdrawn January 2026 claim on this problem. The generation pipeline is cross‑checked by two further independent implementations, the compiled development replays through the standalone lean4checker, and two from‑source verification legs ‑‑ Lean toolchains compiled from source by gcc and by clang, mathlib rebuilt with no cache ‑‑ reproduce the committed certificates byte for byte, with olean digests identical across three builds on two architectures. Our range is three to ten orders of magnitude below the computational frontiers we cite; the contribution is the trust base, not the range.

Authors:Ramon Kaspar, Andrey Ignatov, Valentina Boeva
Title: DistillPath: An Efficient 22M Distilled Pathology Encoder Approaching Large Foundation Model Performance
Abstract:
Many high‑performing pathology tile encoders are now foundation models with hundreds of millions to over a billion parameters. Encoding and storing the thousands of tiles in each whole‑slide image with such models is costly on commodity hardware, so compact encoders that retain useful downstream performance are a valuable alternative. We present DistillPath‑KS16, which starts from the existing 22M kaiko ViT‑S/16 encoder and improves it by distilling from released pathology encoders used as frozen teachers. The recipe reads only the teachers' final class and patch tokens and trains on 6,000 public slides, needing neither their DINO nor iBOT pretraining heads nor a billion‑tile corpus, so it applies to any released encoder that exposes backbone tokens. We distill four teachers spanning 86M to 1.1B parameters into the same student. Every variant improves the kaiko baseline on all three benchmarks we use, EVA, HEST, and PLISM, and the strongest teacher is task‑dependent. On the seven‑task EVA mean, DistillPath‑KS16‑Virchow2 reaches 0.795, within 0.015 points of Virchow2, the top‑scoring model in our evaluation, at about 29× fewer parameters; it also scores above H0‑mini and GPFM on this aggregate metric, though that advantage is task‑concentrated rather than uniform. Because it remains a 22M ViT‑S/16 with 384‑dimensional features, DistillPath‑KS16 runs more than 25× faster than Virchow2. Code is available at https://github.com/RamonKaspar/DistillPath, and released model weights are available at https://huggingface.co/collections/RamonK/distillpath.

Authors:Maosen Zhang, Jianshuo Dong, Boting Lu, Wenyue Li, Xiaoping Zhang, Tianwei Zhang, Jie Zhang, Han Qiu
Title: The Model's Tell: Measuring Context-Leakage Attack Signals with Behavior Gauges
Abstract:
LLMs increasingly rely on external contexts, such as pre‑defined system prompts or retrieved documents, to improve generation quality. However, processing these contexts alongside user queries creates an attack surface: adversarial inputs can induce models to disclose them. Prior probing studies suggest that leakage‑related signals emerge in hidden states, yet the need to extract these states poses additional deployment challenges. In this paper, we explore whether this internal signal leaves a more accessible ``tell'' before decoding. We propose LeakGauge, which probes this response by appending a suffix that gauges leakage behavior and mapping its prefill token probabilities to an attack‑risk score. While a direct gauge uses the initial tokens of confidential content, we find that a content‑agnostic one that verbalizes leakage behavior yields more robust signals. Across 11 LLMs, including GLM‑5.2 (753B) and Kimi‑K3 (2.8T), LeakGauge reaches an AUROC range of 0.944‑‑0.996 on unseen attacks. The signal remains stable when the content changes language or the attack shifts from verbatim to semantic disclosure. By activation‑steering interventions, we further show that the risk score is sensitive to an internal leakage‑related direction, relating the observable signal to the model's internal representation. In addition, LeakGauge enables an input detector with fewer than 0.5K extra parameters and added latency of 10.34 ms. Code: \hrefhttps://github.com/yeasen‑z/LeakGauge.

Authors:Quang Minh Nguyen, Luis Frentzen Salim
Title: Whether LLMs Can Navigate Beliefs and Facts Depends on How You Phrase It
Abstract:
Humans naturally form and express beliefs in daily communication, e.g., "I think the answer is 3" or "I suppose that's right." Such beliefs inevitably intertwine with fact and knowledge, making the ability to handle them in tandem desirable for large language models (LLMs), as they are increasingly deployed in user‑facing settings. Prior work showed that even capable LLMs exhibit a systemic weakness in acknowledging user beliefs grounded in incorrect information. We extend this evaluation to 10 LLMs across 18 epistemic expressions and find that the size and direction of the weakness depend on the verb used to express the belief, with the accuracy gap between factual and false information ranging from +50% on "I vaguely remember" to ‑14% on "I seriously doubt". We further show that the phenomenon stems from task confusion: models default to fact‑checking the underlying claim, overriding the user's stated belief; chains of thought that explicitly fact‑check show lower accuracy on false information than those that do not; and a single instruction can reverse the failure across verb families. Mechanistically, models attend more to false beliefs they fail to confirm, but suppressing this attention at decoding time recovers accuracy only partially and only in some models, calling for future work on intervention methods. Our findings clarify prior results and show how fact‑checking, a generally desirable behavior, can interfere with belief tracking in LLMs. Our code is available at https://github.com/ngqm/belief‑fact‑phrasing.

Authors:Carla Salazar, Lazaros Nalpantidis
Title: Scale Matters: Adaptive Granularity Selection for Cross-Species 3D Plant Organ Segmentation
Abstract:
Recent 3D foundation models provide powerful feature representations for point cloud learning by controlling spatial granularity. However, relying on a fixed spatial granularity severely limits generalization in applications like plant phenotyping, where organ morphology and size vary substantially across species and growth stages. To address this, we propose AGS‑PlantSeg, a few‑shot 3D plant organ segmentation method that leverages the frozen Utonia (arXiv:2603.03283) foundation model combined with Adaptive Granularity Selection. By dynamically selecting the best granularity levels for each specific plant model, our method extracts optimized geometric features for a lightweight MLP segmentation head. Extensive experiments across PLANesT‑3D (arXiv:2407.21150), Pheno4D , and Crops3D demonstrate that AGS‑PlantSeg significantly improves cross‑species generalization, achieving 88.9% average mIoU performance and outperforming fixed‑granularity baselines by 2.5 mIoU points. Despite requiring minimal annotated data, our approach is highly competitive with fully supervised, plant‑specific architectures.

Authors:Qianlong Xiang, Miao Zhang, Kun Wang, Haoyu Zhang, Junhui Hou, Liqiang Nie
Title: TINA+: Probing Residual Visual Knowledge in Unlearned Diffusion Models via Diffusion-Consistent Text-Free Inversion
Abstract:
Although text‑to‑image diffusion models exhibit remarkable generative power, concept erasure techniques are essential for preventing harmful content. Existing adversarial probes evaluate these methods by testing whether erased concepts can still be recovered. However, existing erasure and probe methods remain largely text‑centric, focusing on whether the text‑to‑image mapping is severed while overlooking whether the corresponding visual knowledge remains. To investigate this question from a visual perspective, we leverage diffusion inversion to probe whether a generative trajectory can reconstruct visual instances of an erased concept. Under a null‑text condition, standard inversion avoids the textual pathway but amplifies approximation errors, hindering faithful trajectory recovery. To address this challenge, we introduce TINA+, a diffusion‑consistent Text‑free INversion Attack equipped with optimization‑based inversion. We also find that unconstrained diffusion inversion may discover spurious trajectories, even allowing a randomly initialized diffusion model to reconstruct the target concept. Such trajectories may falsely indicate residual visual knowledge. TINA+ therefore introduces Diffusion‑Consistent Trajectory Regularization to suppress this failure mode. By penalizing trajectories that fall far below the expected marginal energy evolution of diffusion, TINA+ suppresses spurious inversion paths while preserving its ability to recover erased concepts. Experiments across twelve erasure methods, four concept‑erasure tasks, and different model architectures demonstrate that TINA+ reliably probes residual visual knowledge through diffusion‑consistent visual trajectories. These results provide stronger evidence that current methods often obscure concepts by severing text‑image links rather than eliminating the underlying visual knowledge.

Authors:Olga Mashkova, Asaad Mohammedsaleh, Fernando Zhapa-Camacho, Robert Hoehndorf
Title: Neuro-symbolic learning over OWL 2 DL via consequence-based compilation to differentiable circuits
Abstract:
OWL 2 DL ontologies, grounded in the description logic \mathcalSROIQ, express large knowledge bases in biomedicine and the Semantic Web. Neuro‑symbolic (NeSy) learners over description logics either embed the ontology in a continuous space, abandoning classical entailment, or restrict to the Horn fragment \mathcalEL^++, which has a single canonical model. We present Baobab, which compiles a \mathcalSROIQ ontology with a finite ABox into a Sentential Decision Diagram (SDD): it saturates a propositional core under a consequence‑based calculus and instantiates the remaining \mathcalSROIQ features (nominals, number restrictions, and the role axioms) over the active domain. The SDD's evidence‑conditioned weighted model count then trains a perception network to recognize real images under partial ABox supervision: on an ontology that exercises every distinctive \mathcalSROIQ feature, a CNN learns to read MNIST digits coupled by a successor relation and recovers latent ontology concepts that an independent perception leaves at chance. When the supervision admits several ontology‑consistent completions, an independent perception collapses onto one, a reasoning shortcut: we show that a mixture indexed by the query's justifications can represent the calibrated posterior no independent perception can, and that seeding it from the circuit's enumerated completions attains the Bayes‑optimal posterior on a real‑image MNIST task where single‑WMC and learned mixtures (the BEARS‑ensemble hypothesis class) do not: to our knowledge the first to characterize and mitigate reasoning shortcuts in a non‑Horn description logic. Soundness of the compiler and the representation result are machine‑checked in Lean 4. Code is available at https://github.com/bio‑ontology‑research‑group/baobab.

Authors:Linnea Sartorius, Isak Randahl, Delia Fano Yela, Georg Andersson, Sadegh Jamali, Aleksis Pirinen
Title: Monitoring Pasture Restoration from Satellite Image Time Series: Caveats and Opportunities
Abstract:
Monitoring nature restoration at scale is an important but difficult ecological problem. Deep learning methods to analyze satellite image time series (SITS) have been widely used for land surface monitoring. In semi‑natural grasslands ‑ the habitat type in focus in this work ‑ restoration outcomes develop gradually, yet satellite observations are influenced by weather, acquisition conditions, and processing artefacts, making it difficult to distinguish genuine restoration signals from unrelated temporal variation. In this work, we examine ‑ to the best of our knowledge, for the first time ‑ whether restoration status can be detected directly from satellite image time series by formulating pasture restoration as a binary deep learning classification problem. We evaluate two common SITS deep learning architectures on different Sentinel‑2 image combinations, across 1,397 restored Swedish pastures and find that explicitly modeling intra‑year variability and per‑pasture normalization increases separability, reaching 0.88 accuracy for the best model. We further investigate our results and perform a targeted bias analysis finding that reliable deployment requires temporally balanced labels and evaluation protocols that explicitly test for year‑related confounding. We therefore frame our contribution not as a solved restoration‑monitoring system, but as a realistic case study of what works, what fails, and what future studies should control for. Code and models are available at https://github.com/aleksispi/ml‑nature‑resto.

Authors:Geon Tack Lee, Jaegul Choo, Kang Eun Jeon
Title: Denoised Variance-Based Pruning with Optimal Brain Bias Compensation
Abstract:
Vision Transformers (ViTs) achieve state‑of‑the‑art performance but carry massive computational overhead that restricts edge deployment. Although structural pruning has emerged as a key strategy to reduce these costs, existing methods often suffer from severe accuracy degradation or require expensive retraining. Recently, Variance‑Based Pruning (VBP) introduced a promising paradigm by selecting neurons based on activation variance; however, it remains limited by statistical noise in finite‑sample activation covariance and reliance on bias‑only updates that cannot fully account for structural reconstruction error. To address these limitations, we introduce Denoised Variance‑Based Pruning with Optimal Brain Bias Compensation (DVBP + OB^2C). We leverage random matrix theory to filter noise from the activation covariance spectrum for robust neuron selection and mathematically prove that integrating mean‑shift compensation into the Optimal Brain Compression objective reduces the layer‑wise Hessian exactly to the activation covariance matrix. This enables an optimal, closed‑form update of the remaining weights using the same statistics gathered for selection. Extensive experiments on DeiT, Swin, and ConvNeXt architectures demonstrate that DVBP + OB^2C achieves state‑of‑the‑art training‑free performance; at 50% MLP pruning, it retains over 90% of the original Top‑1 accuracy on Small and Base variants, outperforming VBP by up to 29.46% (ConvNeXt‑T) and 7.33% (Swin‑S). The code is available at: https://github.com/geontackee/DVBP_OB2C.

Authors:Lars Simon Zehnder
Title: rl-triton: High-Performance Triton GPU Kernels for Reinforcement Learning Credit Assignment
Abstract:
We present rl‑triton, an open‑source library of high‑performance GPU kernels for reinforcement learning credit assignment, implemented in Triton. The core contribution is a unified associative scan framework that recasts seven distinct RL estimation algorithms ‑ Generalized Advantage Estimation (GAE), V‑Trace, Retrace(λ), TD(λ) returns, discounted returns, eligibility traces, and episodic prefix sums ‑ as instances of a single first‑order linear recurrence solved in O(\log T) parallel steps. All algorithms share the same associative scan operator, with algorithm‑specific fused Triton kernels constructing their recurrence coefficients on‑chip. We verify the associative operator algebraically and define the treatment of terminated and truncated episodes explicitly. Benchmarks show a 1.6‑5.70× full‑call speedup over a vectorized torch‑compile baseline in the massively parallel simulation regime (thousands of environments, short rollouts). The reported range covers all seven algorithms on both GPUs, both with and without per‑step truncation handling. For most algorithms, speedups increase at longer sequence lengths, as the baseline requires more scan stages as \log T grows, each adding an intermediate HBM round‑trip. The library is available at https://github.com/simonsays1980/rl‑triton.

Authors:Afshin Bozorgpour, Sina Ghorbani Kolahi, Moein Heidari, Ilker Hacihaliloglu, Dorit Merhof
Title: MaLViL: Multi-axis Low-rank Vision-LSTM for Medical Image Segmentation
Abstract:
Vision‑LSTM (ViL) enables efficient global modeling, but its cost still scales with the number of spatial tokens, so existing segmenters confine ViL to a coarse bottleneck and lose fine anatomical detail. Rasterizing 2D features into a 1D sequence further breaks adjacency across the orthogonal scan axis. We propose MaLViL, a Multi‑axis Low‑rank Vision‑LSTM network that extends ViL across decoder resolutions. Bidirectional low‑rank ViL (Bi‑LRViL) reasons on a compact orthonormal subspace and preserves detail through an orthogonal residual; scale‑aware SaLViL restores cross‑axis neighbors before serialization; and a Cross‑Directional Mixer (CDM) fuses orthogonal horizontal and vertical traversal paths. Statistics‑Guided Skip Modulation (SGSM) further retains boundary cues in encoder skips. On skin‑lesion, ultrasound, and multi‑organ CT benchmarks, MaLViL achieves competitive or state‑of‑the‑art segmentation accuracy, while reducing ViL operator memory by up to 83× at fine decoder resolutions. Code is available at: https://github.com/xmindflow/malvil.

Authors:Tianjing Hao, Haiyu Lan, Angsong Li, Cheng Chen, Enyu Li, Jiarui Yang, Yuning Su, Peiwen Lin, Wang Chuang
Title: OVIP-SG: Open-Vocabulary Instance-Preserving Scene Graphs for Mapping and Retrieval of Small, Fine-Grained Objects
Abstract:
Integrating open‑vocabulary perception into object‑level 3D scene graphs is a double‑edged sword. While vision‑language detectors recover long‑tail categories and small, fine‑grained objects overlooked by closed‑set models, they also tend to fragment large surfaces and merge small objects into larger neighboring objects, compromising instance‑level consistency and undermining mapping fidelity. Moreover, existing methods struggle to retrieve previously unmapped targets or determine whether a queried object is absent, hindering robust embodied open‑world navigation and exploration. We present OVIP‑SG, a unified framework for instance‑preserving semantic mapping, functional scene partitioning, and language‑guided small, fine‑grained object retrieval. OVIP‑SG uses a vision‑language model (VLM) to enumerate scene‑specific categories for robust open‑world detection. Symmetric 3D Intersection over Union (IoU) association and area‑weighted feature fusion preserve small independent instances, while VLM‑inferred object functions partition scenes into compact functional search regions. A four‑stage cascaded retrieval pipeline further incorporates voxel voting and determines target absence from exploration coverage. Under a unified evaluation protocol on Replica, OVIP‑SG outperforms ConceptGraphs by 6.31 points in class‑mean accuracy (mAcc) and 5.15 points in frequency‑weighted mIoU (F‑mIoU) while achieving a class‑agnostic native‑instance Panoptic Quality (PQ) of 0.398. It reduces the search area to 21.8% of the indoor floor space and reaches 0.773 balanced accuracy for object‑presence classification. Real‑world robotic experiments further demonstrate its practical effectiveness. Code is available at https://github.com/Agibot‑Spatial‑AI/OVIP‑SG.

Authors:Jingyuan Wang, Richong Zhang, Zhijie Nie, Mingxin Li, Yanzhao Zhang
Title: DEPT: Document Embedding Preservation Tuning for Unified Query Expansion and Retrieval
Abstract:
Large language models (LLMs) can both expand underspecified queries and encode text as dense representations, suggesting a unified model for query expansion and retrieval. Existing systems usually rely on prompted expansions, independently trained modules, or staged optimization, leaving generated expansions only indirectly aligned with the retrieval loss that judges them. We train a single decoder‑only LLM end to end, where the same model generates the expansion and encodes both the expanded query and candidate documents. This unified setting creates a moving‑target problem: retrieval supervision should improve query‑side expansion, but the same update also shifts the document embeddings that serve as retrieval targets. We introduce Document Embedding Preservation Tuning (DEPT), which keeps tuned document embeddings close to cached initial embeddings while allowing retrieval gradients to pass through straight‑through decoding into the generator. DEPT converts joint query‑‑document movement into query‑side adaptation against approximately stable, whitened document embeddings that support index reuse and online hard‑negative mining. Experiments with Qwen3‑4B‑Instruct‑2507 and LLaMA‑3.2‑3B‑Instruct on five datasets in BEIR benchmark show that DEPT improves average retrieval quality over training‑free, independently trained, and staged unified baselines, while ablations isolate the effects of preservation, whitening, end‑to‑end expansion training, and online negatives. Code is available at https://github.com/ILSparkle/DEPT.

Authors:Syeda Faiza Ahmed, Zien Sheikh Ali, Hunzalah Hassan Bhatti, Firoj Alam, Shammur Absar Chowdhury
Title: Multi-turn Conversational AI from Text to Multimodal Interaction: Data, Models, Evaluation, and Open Challenges
Abstract:
Conversational AI is moving beyond isolated text prompts toward sustained, multimodal interaction. In real conversations, users clarify goals, revise requests, interrupt responses, switch topics, and introduce new evidence while expecting systems to preserve context across turns. This makes multi‑turn dialogue a distinct challenge requiring systems to maintain and update memory, ground responses across modalities, tools, and external knowledge, and adapt across languages and cultures. This study reviews multi‑turn conversational AI across text‑only dialogue, AudioLLMs and speech‑native systems, multimodal and omni‑modal systems, and tool‑augmented agents. We organize the literature around datasets and benchmarks, modeling paradigms, training strategies, evaluation setups, and cross‑cutting challenges. Our analysis shows that support for multiple modalities has advanced faster than the ability to sustain coherent interaction across a session. Despite stronger capabilities to perceive, speak, and act across modalities, current systems still struggle with persistent memory, cross‑turn grounding, full‑duplex interaction, robust evaluation, and cultural alignment. We conclude with a research agenda for systems that can remember, revise, ground, speak, listen, act, and adapt across turns, modalities, and cultures. (https://github.com/faiza‑sfa/multiturn‑conversational‑ai‑survey)

Authors:Yajing Bai, Jinhao Duan, Jie Peng, Xianfeng Wu, Sijia Liu, Song Wang, Tianlong Chen
Title: HarnessRisk: A Lifecycle-Oriented Benchmark for Agent Harness Safety
Abstract:
Large language models are increasingly deployed through agent harnesses that manage tools, extensions, persistent state, permissions, and external actions. Existing safety benchmarks mainly target individual attack mechanisms or a limited subset of operational settings, making it difficult to compare how safety failures emerge across different harness responsibilities. We present HarnessRisk, a lifecycle oriented benchmark that organizes agent harness safety into six operational phases including Harness Configuration, Capability Extension, Runtime Operation, State Persistence, Action Control, and Incident Recovery. HarnessRisk contains 128 sandboxed cases, each pairing a benign user objective with an adversarial instruction embedded in an untrusted workflow artifact. We evaluate each trajectory using Utility, Attack Success Rate, Persistence, and Detection. Across three harnesses, six language models, and 14 model and harness configurations, attack success ranges from 12.6% to 80.9%, while Utility remains between 75.0% and 97.6%. Harness Configuration is the most vulnerable phase across all three harnesses, showing that attacks can succeed by altering security sensitive parameters within otherwise authorized workflows. We also find that explicit risk recognition does not reliably lead to safe action, as some configurations detect risks in more than 90% of runs while retaining substantial attack success. These results highlight the need to evaluate agent safety across multiple harness responsibilities and at the level of the deployed model and harness configuration.

Authors:Ruizhe Wang, Yixuan Dong, Bolin Yang, Bingo Wing-Kuen Ling, Fuji Yang, Zelin Zang
Title: DMT-Dens: Density-preserving manifold visualization for biological data
Abstract:
Motivation: Low‑dimensional embeddings are widely used to explore cell‑state heterogeneity in single‑cell and other high‑dimensional biological data. Although many methods preserve local neighborhoods, they may distort the apparent sampling density of processed observations, altering the visual contrast between dense and sparse regions and complicating the interpretation of rare, transitional, or continuous cell‑state populations. Results: We present DMT‑Dens, a parametric manifold‑visualization method built on a latent‑token Transformer encoder. The model integrates rank‑based manifold alignment with hard‑pair aggregation. To preserve density, it optimizes a loss based on the Pearson correlation between k‑nearest‑neighbor log‑radius estimates in the processed input and two‑dimensional embedding spaces. Benchmark evaluations demonstrate strong density preservation, particularly on biological datasets, while retaining competitive label separability. Availability: Source code, data‑processing scripts, and resolved experiment configurations are available at https://github.com/Ruizhe‑wang/DMT‑Dens.

Authors:Fuchen Long, Cong Wang, Zitao Gao, Wenhao Zhong, Yu Cheng, Xiaolu Hou, Yan Li, Xiao Cao, Xinlong Sun, Xi Chen, Yu Liu
Title: CoinVE-200K: A Large-Scale High-Quality Dataset for Compositional Instruction-Guided Video Editing
Abstract:
The quality and diversity of instruction‑based video editing datasets are steadily improving, yet existing datasets mainly focus on single editing operations and fall short in supporting compositional instruction‑guided video editing. In particular, multiple editing intents must be jointly understood and faithfully executed within the same video. To address this issue, we introduce CoinVE‑200K, a large‑scale, high‑quality dataset for Compositional Instruction‑Guided Video Editing. CoinVE‑200K contains 1080p video‑editing pairs of up to 201 frames, covering diverse compositional scenarios where each sample involves 2 to 5 atomic editing operations. The instructions target humans, objects, and backgrounds, and cover edit types such as addition, removal, modification, and stylization. All samples are built through a carefully designed generation and filtering pipeline to ensure instruction faithfulness, visual quality, temporal consistency, and compositional diversity. We also introduce CoinVE‑Bench, a benchmark for compositional‑instruction video editing across diverse subjects, operation types, and instruction complexities. Furthermore, we present CoinVE‑Edit, a 22B compositional video editing model built upon Wan2.1‑T2V‑14B and Qwen3‑VL‑8B‑Instruct. CoinVE‑Edit disentangles region‑aware attention for different editing instructions, enabling precise multi‑region editing while preserving irrelevant content and temporal coherence. Experiments on CoinVE‑Bench show that CoinVE‑Edit achieves strong performance in instruction following, compositional editing accuracy, visual quality, and temporal consistency.

Authors:Zongyang Qiu, Yihan Wu, Kaixuan Fan, Bo Li, Hui Xiong
Title: Where a New Concept Must Enter: Entry Point Gates Cross-Task Usability in Unified Multimodal Models
Abstract:
Unified multimodal models (UMMs) are motivated by the hope that understanding and generation reinforce each other but controlled ablations repeatedly find that adding a generation objective leaves understanding flat. Joint‑training studies cannot settle the disagreement: with overlapping supervision, a gain cannot be attributed to the architecture rather than the data. To further investigate the relationship between the two directions in UMMs, we separate them by construction. A novel visual entity, a rendered 3D asset paired with a pseudo‑word screened for absence from the frozen model's behavior, is bound through exactly one task direction, and the untrained direction is then measured. We find that the channel is real in both directions, but the directions differ in kind: generation training installs a name the model can only match among candidates; understanding training installs one it can also produce. What governs cross‑task usability is where the binding enters the shared computation. An alignment probe predicts export across 36 configurations (Spearman ρ= +0.68). That objective's alignment term, maximized in closed form over activations with every weight frozen, makes a concept drawable when injected at layer 7 of 28 and is indistinguishable from the base model from layer 14 on, while the weight‑based version of the same edit peaks at layers 10‑14. In an observational series of four models, this window appears only where the understanding pathway is a semantic vision encoder, suggesting that unified weights are not enough: the two directions must share a semantic format at the entry point. Exploiting the rule, a mid‑stack alignment objective acquires the concept for a 0.1% relative loss of the model's general text‑to‑image ability, against 41% for the standard generative route. Our code is at https://github.com/Zane‑ZYQiu/entry‑point‑umm.

Authors:Jack Boylan, Chris Hokamp
Title: No Gaussian Required: Contrastive Inverse Dynamics for JEPA World Models
Abstract:
Joint‑Embedding Predictive Architectures (JEPAs) learn world models by predicting future embeddings, but the objective admits a trivial solution of a constant encoder, so every practical system adds an anti‑collapse mechanism (LeCun, 2022; Assran et al., 2023; Bardes et al., 2022; 2024). LeWorldModel (LeWM) prevents collapse with SIGReg, a regularizer that forces the latent distribution to match an isotropic Gaussian: the representation is stabilized by prescribing what it must look like, independently of the environment it models. We argue that the anti‑collapse pressure can instead come from the transition data itself. Action‑Contrastive Masked Transition Modeling (AC‑MTM) keeps LeWM's forward latent‑prediction objective and adds a training‑only inverse‑dynamics head trained with Action‑NCE: each latent transition must identify the action that produced it among the other actions in the batch, a discrimination task that a collapsed encoder provably fails. The inverse branch is discarded after training, leaving test‑time encoding, forward prediction, planning, and compute identical to LeWM. On four standard pixel‑control tasks under a matched planning protocol, AC‑MTM trains stably from scratch and matches SIGReg on average. On the harder multi‑object OGBench Visual Scene task, results are consistent with the prescribed geometry becoming a bottleneck: AC‑MTM reaches 80.0\pm2.0% success versus 58.0\pm2.0% for SIGReg, improving by 20‑24 points in each training seed. A single 50‑episode random‑policy run gives a 52% baseline estimate. Contrastive inverse dynamics thus provides a distribution‑free anti‑collapse signal that requires no target network, stop‑gradient, pretrained encoder, or reconstruction objective, and we characterize the action‑space and observability assumptions under which it holds. We make our code available at https://github.com/jackboyla/action‑contrastive‑jepa

Authors:Sarvesh Gharat, Junpei Komiyama
Title: SGHA: Evidence-Grounded Research Problem Discovery with Local Language Models
Abstract:
Recent efforts toward fully automated AI scientists have demonstrated that language‑model agents can generate hypotheses, execute experiments, and draft scientific manuscripts. However, during the early stages of research, when research problems are formulated, these AI scientists often rely heavily on proprietary frontier models. Their proposals are shaped by opaque parametric knowledge and by literature searches conditioned on the proposals themselves. Such knowledge is effectively a black box, and this dependence makes the evidential basis and validity of generated research problems difficult to audit and leaves the process vulnerable to model‑specific hallucinations and biases. Furthermore, if proprietary research materials are transmitted to external APIs, the use of these models creates confidentiality, privacy, and data‑governance concerns. We introduce the Structural Gap Hypothesis Agent (SGHA), a fully automated, corpus‑first research‑problem discovery system that runs entirely on a local LLM. SGHA structures a scientific literature corpus into evidence‑linked paper objects and a typed evidence graph, detects unresolved structural patterns across papers, screens candidate gaps before formulation, and produces traceable research‑problem families. In particular, it is able to output assumptions, objectives, success criteria, and remaining ambiguities. All LLM‑based components of SGHA are executed using a locally served open‑weight 9B language model, without requiring proprietary frontier‑model APIs. We compare SGHA with the AI Scientist‑v2 idea formulation module in five machine‑learning domains. Our results suggest that explicit corpus structure and evidence‑constrained reasoning can support promising, inspectable research‑problem formulation without relying on frontier models during generation or verification.

Authors:Feiyu Shen, Kun Xie, Yichen Wu, Ziqi Dai, Yichen Han, Junjie Li, Xuelong Geng, Fenglong Xie, Lei Xie, Xu Tang, Yao Hu
Title: FireRedTTS3: Unified Speech Generation and Editing with Semantically Enriched Speech Representations
Abstract:
Recent continuous autoregressive TTS models operate directly on continuous speech representations, preserving rich acoustic details while leveraging the instruction‑following capabilities of text LLMs. This paradigm opens new possibilities for voice cloning, instruction‑controlled voice design, and speech editing, but remains susceptible to error accumulation during autoregressive generation. Existing solutions often require additional semantic modules, multi‑stage tokenizer training pipelines, or complex autoregressive architectures. In this work, we propose FireRedTTS3, a simple yet effective speech generation and editing framework that mitigates error accumulation at the representation level. Specifically, we leverage a frozen Audio Encoder trained on diverse speech understanding tasks as a semantic teacher to regularize the audio feature space. This improves text‑speech alignment and stabilizes autoregressive generation while keeping the overall system simple. FireRedTTS3 provides two variants: FireRedTTS3‑Base for multilingual and multi‑dialect zero‑shot voice cloning, and FireRedTTS3‑Instruct for unified voice cloning, instruction‑controlled voice design, and speech editing. Experiments show that FireRedTTS3‑Base achieves the best average speech intelligibility and speaker similarity among compared systems on Seed‑TTS‑Eval and MiniMax‑MLS‑Test, while FireRedTTS3‑Instruct outperforms competing systems on InstructTTSEval and Ming‑Freeform‑Audio‑Edit. These results demonstrate that semantically enriched continuous speech representations, combined with a simple architecture, enable stable, controllable, and high‑fidelity speech generation and editing. Code and models are available at https://github.com/FireRedTeam/FireRedTTS3.

Authors:Yibo Liu, Bowen Jiang
Title: When More Foundation Models Means Less: Diagnosing and Addressing Multi-View Fusion Failure
Abstract:
Foundation‑model hubs turn multi‑view fusion into a selection problem: from a large heterogeneous encoder pool, which views should be fused, and how many? We show that downstream performance is non‑monotonic in the number of fused encoders; later views can be redundant or task‑misaligned, causing accuracy to saturate or decline. We formalise this setting as view‑set composition and propose KAGES (Kernel‑Alignment Greedy Encoder Selector), a label‑aware method that orders frozen encoders by their marginal gain in centred kernel‑target alignment. KAGES requires no downstream classifier training during selection, evaluates each candidate in \mathcalO(n^2) time independent of encoder dimension, and admits a conditional (1‑e^‑γ) prefix‑wise guarantee under monotonicity and a positive submodularity ratio. Across five recognition regimes and low‑shot, larger‑pool, and full‑data protocols, KAGES improves average AULC over full fusion by 3.9, 5.8, and 3.3 points, respectively, and exceeds DPP and facility‑location selection in average AULC. Image retrieval exhibits later, task‑dependent saturation along the KAGES ordering, while peak‑then‑decline reproduces in frozen‑LLM fusion. These results show that effective large‑pool fusion depends on selecting a compact, task‑aligned set of views rather than indiscriminately fusing more encoders.

Authors:Ruichao Hou, Boyue Xu, Tongwei Ren, Dongming Zhou, Gangshan Wu, Jinde Cao
Title: S$^3$AM: A Single-Stream SAM with Reliability-Calibrated Frequency Adapter for Multi-modal Salient Object Detection
Abstract:
Vision foundation models have recently advanced multi‑modal salient object detection (MSOD) through parameter‑efficient tuning and prompt learning. However, existing Segment Anything Model (SAM)‑adapted MSOD methods often rely on dual‑stream encoders or auxiliary prompt generators, leading to redundant computation. Although a single‑stream alternative can reduce this cost, early fusion may also propagate noisy or misaligned auxiliary high‑frequency cues through the backbone. In this paper, we propose a novel single‑stream framework that integrates reliability‑calibrated frequency adaptation into the adopted SAM backbone for MSOD. It avoids duplicated foundation backbones while explicitly controlling auxiliary frequency injection. Specifically, we design a mixture of frequency experts module, which uses the stationary wavelet transform to decompose each modality and aggregate cross‑modal frequency information. We further introduce a reliability‑calibrated frequency adapter with a dual‑gate calibration mechanism, which selectively propagates the calibrated residual across transformer stages while jointly controlling its injection strength and cross‑modal reliability. A hypernetwork‑guided semantic‑structural decoder then combines semantic mask features from the adopted backbone with Mamba‑based structural detail recovery. Comprehensive experiments on RGB‑D, RGB‑T, and RGB‑NIR salient object detection benchmarks validate that the proposed framework achieves competitive performance with only 12.20M trainable parameters, accounting for 5.4% of the total parameters. The code will be available at https://github.com/xuboyue1999/SSSAM.

Authors:Maolin Ran, Xiaoyang Lu, Jiaqi Liu, Jian Wang, Weiwen Liu, Jianghao Lin, Yong Yu, Weinan Zhang
Title: SAGE: Self-Evolving Storyboard Skills via Attribution-Guided Rule Evolution
Abstract:
Storyboards turn screenplays into visual shot plans for automated short drama production. Professional storyboarding relies on tacit directorial expertise and remains an industrial bottleneck. Large language models can automate this step, but methods for supplying directing knowledge face three challenges: (1) Knowledge acquisition: the craft remains implicit in exemplars or must be written manually. (2) Knowledge refinement: authored knowledge is not evaluated against execution outcomes, and opaque generation prevents feedback attribution to the knowledge behind each decision. (3) Knowledge injection: injecting all knowledge exceeds usable context, while manual selection for every narrative group does not scale. We present SAGE (Skill with Attribution‑Guided Evolution), a deployed framework that learns, attributes, evolves, and routes directing knowledge from expert demonstrations. SAGE derives rules that are independent of episode content by contrasting each training screenplay with its expert storyboard. During generation, the model records each narrative group's adopted rules. Combining these records with localized feedback enables targeted updates to individual rules. Evolved rules form scenario packages with a routing index, so each group retrieves only a bounded set appropriate to its situation without expert intervention. On 18 test episodes across three genres, SAGE scored 77.8 on a rubric validated by experts, versus 77.1 for professional directors. Deployed for 14 days on Virtual Film Studio, SAGE produced 1,344 narrative group outputs; 87.2 percent were accepted without substantive edits, and the production team recorded over 83 percent less authoring time per episode. We release PROSE, the first public dataset pairing screenplays with storyboards by professional directors across 68 episodes: https://github.com/creDreams/PROSE.

Authors:Yifan Lu, Adinath Dukre, Abhijit Das, Ziyun Zou, Haolin Yang, Yutong Xie, Imran Razzak
Title: Counterfactual Anatomy-guided Spatial-Temporal Decoding for Annotation-Free Hallucination Mitigation in Medical VLMs
Abstract:
Medical vision‑language models (Med‑VLMs) have demonstrated strong performance on medical visual question answering, yet they remain prone to hallucination, generating clinically unsupported statements that are insufficiently grounded in image evidence. Mitigation methods applied during decoding offer a practical solution, but they typically lack anatomical awareness or rely heavily on ground truth annotations, which limits their applicability. We propose Counterfactual Anatomy‑guided Spatial‑Temporal decoding (CAST), a framework that operates entirely during inference and requires no manual annotations for anatomically grounded hallucination mitigation. CAST automatically discovers anatomical regions relevant to the given query through broad medical segmentation. It then selects a compact, causally informative area using counterfactual intervention based on the drop in answer likelihood under occlusion. Guided by this chosen region, CAST performs a unified contrastive decoding process, combining classifier‑free guidance to correct spatial attention with stepwise temporal contrast to regulate generation dynamics. Experiments on the SLAKE and MIMIC‑CXR datasets across three Med‑VLMs demonstrate that CAST consistently outperforms strong baselines and surpasses decoding strategies reliant on ground truth. Our results indicate that compact, automatically selected regions provide highly effective contrastive guidance without expert annotations, offering a practical and generalizable solution for improving spatial grounding and reducing hallucinations. Code is available at https://github.com/csyifan/CAST.

Authors:Gen Li, Shu Han, Yun Xi Qiao, Hua Chen, Xuyang Dai, Bohan Li, Hao Zhao, Chaojian Li
Title: SPVC: Structured and Panoptic Video Fixing for Cross-Dataset Driving Scene Rendering
Abstract:
Driving scene reconstruction and rendering, especially with 3D Gaussian Splatting, has become an important component of autonomous driving simulation. However, rendered views often degrade under extrapolated ego trajectories and scene edits, producing blurry structures, temporal flicker, and foreground‑background misalignment. Existing refinement methods are commonly designed for a specific setting, such as image‑level novel‑view repair or object‑editing correction. In this paper, we introduce SPVC, a structured and panoptic video fixing framework for cross‑dataset driving scene rendering. The name summarizes four design principles. (1) Structured fixing denotes the use of explicit spatial conditions, including camera pose, 3D bounding boxes, and HD maps, to guide the repair process and reduce uncontrolled hallucination. (2) Panoptic fixing refers to correcting both background rendering artifacts, such as distorted roads, buildings, and lanes, and foreground vehicle artifacts introduced by scene editing, such as inconsistent object appearance. (3) Video fixing means that the model operates on driving sequences rather than isolated frames, allowing temporal cues to be used during artifact correction. (4) Cross‑dataset fixing means that a single shared network is trained and applied across multiple driving datasets, reducing the need for dataset‑specific or scene‑specific fixers. Concretely, we construct paired degraded‑clean training data by simulating under‑constrained 3DGS rendering and foreground vehicle insertion artifacts, and train a two‑stage controllable video diffusion model that first addresses video‑level appearance and then refines scene layout with structured controls.

Authors:Bonan Zhang, Shiyu Dong, Quan Hung Tran, Katharina Gschwind, Shuqi Yang, Sijia Chen, Adel Ahmadyan, Seungwhan Moon, Lu Zhang, Ahmed Kirmani, Babak Damavandi, Anuj Kumar
Title: MoE-ViE: Mixture of Experts Vision Encoder for Efficient Image and Video Understanding
Abstract:
Vision encoders are a critical component of vision‑language models, and scaling their capacity effectively improves performance. However, dense scaling increases compute cost and inference latency. Mixture‑of‑Experts (MoE) architectures offer a compelling alternative, having enabled efficient scaling in LLMs, yet the MoE design space for CLIP‑style vision encoders remains underexplored at State‑of‑the‑Art (SOTA) levels. In this work, we systematically study MoE designs for vision encoder scaling and find that fine‑grained MoE topologies yield substantial gains over both dense and standard MoE counterparts. We further propose an auxiliary‑loss‑free balancing variant for better expert utilization, and design a specialized MoE kernel to mitigate inference latency overhead. To enhance video capabilities while preserving image knowledge, we introduce frame‑level distillation paired with a novel freezing mechanism. We pretrain a series of Mixture‑of‑Experts Vision Encoders (MoE‑ViE) across a range of sizes, all consistently outperforming their dense counterparts. Our largest model matches the zero‑shot performance of a SOTA encoder 1.7x its size at 76% of its latency. When aligned with an LLM, MoE‑ViE surpasses all compared encoders on image and video benchmarks, including those with up to 5x more activated parameters. Code is available at https://github.com/facebookresearch/moe_vie.

Authors:Tinghao Jiang, Sheng Tang, Shengzhe Wei, Juntong Fang, Weiqi Zhang, Junsheng Zhou, Zesong Li
Title: GeoWeaver: Accurate Long-Sequence 3D Reconstruction via Hierarchical Geometric Assembly
Abstract:
Long‑sequence 3D reconstruction from RGB videos requires both accurate local geometry and globally consistent camera motion. Feed‑forward models provide strong depth and pose predictions, but their memory cost prevents joint inference over long sequences. Chunk‑wise processing improves scalability, yet independently predicted chunks often exhibit scale drift, pose errors, and point‑cloud misalignment. We present GeoWeaver, a unified framework comprising a Geometric Prior Model (GPM) and Test‑Time Adaptation (TTA). The GPM predicts chunk‑wise depth, confidence, and camera parameters as adjustable geometric priors. TTA then performs sequential initialization, global chunk‑level Sim(3) alignment, and coarse‑to‑fine refinement of camera poses, affine depth corrections, and intrinsics. Dense correspondences provide adjacent, cross‑chunk, and long‑range constraints, while a robust CDF‑style objective jointly optimizes weighted 2D reprojection and 3D consistency residuals. This design preserves local geometric accuracy while correcting accumulated pose, scale, depth, and calibration errors. Experiments across diverse long‑sequence benchmarks demonstrate improved camera accuracy, global consistency, and point‑cloud quality. Ablations verify the contribution of each adaptation stage, and applying the same TTA procedure to different geometric prior models consistently improves their trajectory estimates, demonstrating that GeoWeaver is not tied to a specific GPM.

Authors:Hoda Yamani, Henry Williams, Bruce A. MacDonald
Title: Integrating Novelty and Surprise for Experience Prioritization and Exploration in Image-Based Reinforcement Learning
Abstract:
Sample efficiency is a central challenge in reinforcement learning (RL), particularly in image‑based domains where agents must learn from high‑dimensional visual inputs. Traditional sampling often relies on random or suboptimal experience selection, leading to redundant updates and slow learning. Improving efficiency requires mechanisms that prioritize informative experiences while also encouraging effective exploration. Prioritized Experience Replay (PER) addresses part of this challenge by reusing high‑value transitions, while intrinsic rewards promote the exploration of novel or uncertain states. However, their integration has not been extensively studied. This paper introduces Novelty and Surprise Prioritized Experience Replay (NSPER), which uses novelty to capture underrepresented states and surprise to expose gaps in the agent's understanding of the environment. We further extend this with NSPER+R, integrating these signals as intrinsic rewards to jointly improve replay quality and exploration. Experiments on DeepMind Control Suite tasks show that NSPER and NSPER+R improve training efficiency and convergence speed compared to existing methods in image‑based RL.

Authors:Zhida He, Xiaoyu Wen, Han Qi, Ziyuan Zhou, Peng Yu, Jiajia Li, Chaochao Lu, Qiaosheng Zhang
Title: Fair ASR: Re-Evaluating Black-Box Jailbreaks under Shared Target-Call Budgets
Abstract:
Reliable jailbreak evaluation is essential for assessing LLM safety, but most existing studies rely solely on attack success rate (ASR) without accounting for its dependence on attack budgets, resulting in unfair comparisons across methods. Existing compute‑aware evaluations reduce heterogeneous resources into FLOPs, which is difficult to estimate for black‑box models and fails to capture resource‑specific constraints. To provide a comparable evaluation basis, we introduce Fair‑ASR, an evaluation protocol for black‑box jailbreak attacks under shared target‑call budgets B, using target calls as a directly observable and method‑agnostic comparison axis while tracking attacker calls separately for efficiency analysis. We re‑evaluate 11 representative attacks under the Fair‑ASR protocol and find that attack rankings change substantially across target‑call budgets, simple stochastic perturbations and hand‑crafted templates remain highly competitive under equal target access, and no evaluated LLM‑driven method is efficient in both target and attacker calls. Motivated by this efficiency gap, we introduce ReCode, a compositional budget‑efficient attack that combines desensitization rewriting with two effective low‑cost primitives identified by Fair‑ASR. Under a budget of 20 target calls, ReCode achieves 85% ASR on GPT‑5 while requiring only 7.19 attacker calls per request on average, showing strong efficiency in both target and attacker calls.

Authors:Weiran Wang, Hongxiang Shi, Huitao Tang, Wenjuan Qin
Title: ArguLens: An Open-Source System for Automated Essay Scoring and Label-Aware Feedback Generation
Abstract:
Most automated essay scoring (AES) systems output a single holistic score without interpretable evidence and rely on closed APIs that introduce data privacy and cost barriers. We present ArguLens, an opensource, locally deployable system that decomposes AES into three decoupled components: a discourse‑move classifier (Qwen2.5‑7B‑Instruct fine‑tuned with LoRA on PERSUADE 2.0), a grade‑independent LightGBM scorer over 31 linguistic and discourse features, and a label‑aware feedback generator served through vLLM with a Qwen2.5‑14BInstruct backbone. A Gradio web UI exposes pluggable inference backends and supports single‑essay and batch scoring with downloadable per‑essay breakdowns. On an essaydisjoint PERSUADE 2.0 test split, the logitprobe classifier achieves 82.6% accuracy and 0.727 macro‑F1; under prompt‑grouped 5‑fold cross‑validation the scorer reaches a mean QWK of 0.813 under an oracle discoursefeature protocol, and an ablation shows that adding gold discourse annotations yields an increment of +0.055 QWK over the lexical+syntactic configuration (paired t‑test, p = 0.010). This is a component‑level diagnostic rather than an end‑to‑end classifier‑to‑scorer result. The feedback generator ships with a structured evaluation protocol; its human‑rater study is left to future work. The system is released under Apache 2.0 at https://github.com/wwrwbs/AI_AWE.

Authors:Hoda Yamani, Yuning Xing, Koen van Rijnsoever, Bruce A. MacDonald, Henry Williams
Title: Repetition as Reinforcement: Enhancing Sample Efficiency via Instant Episode Repetition in Reinforcement Learning
Abstract:
Repetition is a fundamental mechanism in human learning, where revisiting successful experiences strengthens memory, consolidates skills, and improves future performance. Motivated by this biological principle, we introduce Instant Episode Repetition (IER), a simple and novel mechanism that improves sample efficiency by immediately repeating action sequences from successful episodes during environment interaction. Unlike conventional approaches such as Experience Replay and Self‑Imitation Learning (SIL), which passively reuse past experience during training updates, IER directly influences the data collection process. Upon identifying a high‑reward episode, the agent repeats its action sequence for a fixed number of subsequent episodes, reinforcing valuable behaviors through renewed interaction with the environment. We integrate IER into state‑of‑the‑art SAC and TD3 algorithms and evaluate its effectiveness on continuous‑control benchmarks, including MuJoCo, the DeepMind Control Suite, and a real‑world dynamic object translation task with a robotic manipulator. Experimental results demonstrate that this simple mechanism improves learning performance over standard and self‑imitation‑based baselines.

Authors:Hanzhi Zhang, Qiao Zhang, Qinglei Cao, Heng Fan, Yan Huang, Kewei Sha, Yunhe Feng
Title: TileMix: Tile-Centric Mixed-Precision Attention for LLM Inference Acceleration
Abstract:
Long‑context prefill in large language models (LLMs) incurs substantial computation and memory traffic because dense self‑attention computes quadratic query‑key scores. Existing methods either use a uniform low‑precision path or select token interactions, leaving spatial precision routing over hardware‑aligned score tiles outside fused dense attention. We introduce TileMix, a tile‑centric precision‑routing kernel that makes numerical precision an executable spatial decision over score‑tile groups within fused dense attention. TileMix partitions the attention matrix into hardware‑aligned score tiles, packs routing decisions into compact bitmasks, and dispatches each tile group through FP16 or INT8 score computation while both paths update a shared online‑softmax state. Scalable precision grouping lets each routing bit govern multiple adjacent key tiles, preserving hardware‑aligned compute tiles and compact metadata at long contexts. By routing all legal tile groups, TileMix preserves dense token connectivity, requires no training, and supports grouped‑query attention, variable‑length batches, and INT8 key/value caches. Across LongEval, LV‑Eval, and A100 prefill benchmarks on LLaMA, Qwen, and Vicuna, TileMix recovers long‑context quality lost under uniform INT8 and improves prefill throughput over FP16, yielding a controllable accuracy‑efficiency frontier across model families. The implementation is available at https://github.com/HanzhiZhang‑Ulrica/TileMix.

Authors:Yining Hua, Cyrus Ayubcha, Hongbin Na, Levi Lian, Alon Gorenshtein, Yiftach Barash, Eyal Klang
Title: LLMs for Medical Consultation Are Evaluated Too Late: The Preformulation Gap
Abstract:
Large language models for medical consultation are often evaluated after a clinical problem has already been made clear, although real consultations may begin with a vague, minimized, or misframed concern. We evaluated three API models across four physician‑authored, multi‑turn vignettes under baseline and entry‑to‑care instruction conditions, yielding 24 fixed‑script transcripts; two cases also used adaptive standardized‑patient simulation, yielding 12 transcripts. Self‑care or home‑management advice before any patient answer appeared in 9 of 12 baseline case‑model cells and 0 of 12 instruction cells, while structured handoff summaries appeared in 0 of 12 and 10 of 12 cells, respectively. The instruction changed sequencing and documentation, although it did not reliably ensure elicitation of decisive facts. The preformulation gap should therefore be evaluated directly through observable first‑contact behavior rather than inferred from diagnostic accuracy or final‑answer quality.

Authors:Xurong Liang, Tong Chen, Quoc Viet Hung Nguyen, Jianxin Li, Xiangliang Zhang, Hongzhi Yin
Title: Empowering Compact LLMs with Fusion of Layer-wise Exits for Recommendation
Abstract:
Large language model‑based recommender systems (LLM‑RSs) have demonstrated remarkable capabilities, but are computationally unsustainable for many real‑world applications. Compact LLMs offer a practical alternative, yet their reduced capacity often requires reasoning or knowledge distillation methods that increase latency or depend on larger models. Combined with autoregressive generation, these approaches face severe scalability bottlenecks. In contrast, discriminative LLM‑RSs enable efficient full‑corpus ranking through embedding similarity, but compact backbones remain limited in expressiveness and structural adaptivity. We propose the Fusion of Layer‑wise Exits for Sequential Recommendation (FLEXRec), a discriminative framework that enhances compact LLMs while retaining scalable full‑corpus ranking. FLEXRec inserts prediction heads (i.e., exits) at multiple transformer layers and adaptively fuses their score distributions. An adaptive continuous router (AC‑Router) dynamically selects both the number and identity of exits for each user sequence, while a novel target‑k hinge loss regulates routing sparsity. Experiments on three real‑world datasets with Qwen 3 1.7B and Llama 3.2 3B show that FLEXRec achieves state‑of‑the‑art accuracy among compact‑backbone methods while remaining highly efficient. Code: https://github.com/xurong‑liang/FLEXRec

Authors:Yang Chen, Zhan Zhuang, Yanbin Wei, Zebin Chen, Hua Liu, Yu Zhang
Title: Learning What Not to Learn: Adversarial Disentangled Prompt Tuning for Robust Vision-Language Models
Abstract:
While adversarial prompt tuning can enhance robustness of vision‑language models efficiently, we find that existing methods aggravate robust generalization overfitting on seen classes, leading to a rapid degradation in performance against adversarial examples of unseen classes as training progresses. We empirically identify that this degradation stems from the tendency of the model to learn pseudo‑robust features (i.e., non‑generalizable shortcuts). To mitigate this, we propose ADAPT (Adversarial Disentangled Prompt Tuning), a robust prompt tuning framework following the philosophy of ``Learning What Not to Learn''. Specifically, ADAPT uses a dual‑prompt mechanism with a target prompt and a pool of decoy prompts. During training, the decoy prompts are guided to entrap diverse pseudo‑robust features, while the target prompt is constrained to be orthogonal to the decoys in the embedding space to learn robust features. By disentangling the robust features from the pseudo‑robust features, ADAPT effectively prevents robust generalization overfitting. We further provide an analysis showing that the orthogonal loss bounds the effect of shifts in pseudo‑robust features on unseen classes, yielding a testing error guarantee. Empirically, extensive experiments demonstrate that ADAPT substantially improves the robustness of the target prompt on unseen classes. The code is available at https://github.com/cheny02/ADAPT‑ACMMM2026.

Authors:Rongwen Li, Haixin Xie, Xiao Wang, Changjian Chen
Title: Beyond MSE: Rethinking the Evaluation Metric and Benchmarking for Irregular Time Series Forecasting
Abstract:
Existing research on irregular time‑series forecasting has primarily focused on model design, while evaluation metrics remain insufficiently studied. Existing benchmarks typically use mean squared error (MSE) as the evaluation metric. We show that, in irregular forecasting, MSE is determined not only by the model prediction but also by the sample‑specific timestamp sampling distributions, leading to a biased assessment of the models' continuous‑time predictive performance. To address this issue, we propose the Continuous‑time Squared Error (CSE), which employs importance weighting to eliminate the influence of the timestamp sampling distributions. We further theoretically prove that CSE's asymptotic estimation error with respect to continuous‑time risk is no greater than that of MSE. Finally, we construct a systematic benchmark covering synthetic, semi‑synthetic, and eight real‑world datasets to validate the effectiveness of CSE and systematically evaluate models' continuous‑time predictive performance. Experiments show that CSE can recover continuous‑time risk more accurately than MSE, while relying solely on MSE may not fully reflect models' continuous‑time predictive performance in real‑world scenarios. Our code can be obtained at https://github.com/hnu‑vis/ITS‑Bench.

Authors:Rongwen Li, Changjian Chen
Title: Rethinking Irregular Time Series Forecasting from the Perspective of Basis Functions
Abstract:
Irregular time series forecasting is crucial in many domains, such as healthcare and meteorological observation. However, due to the inherent characteristics of irregular time series, including sparse observations and non‑uniform sampling, accurately predicting future dynamics remains challenging. In light of these two characteristics, many existing methods aggregate irregular observations into fixed‑dimensional estimated response coefficients through predefined basis functions and use these coefficients as sequence representations. Nevertheless, this modeling paradigm still suffers from two key limitations: (i) a potential non‑vanishing asymptotic bias caused by ignoring the sampling density of timestamps; and (ii) the limited adaptability of predefined basis functions to diverse temporal patterns. In this study, we propose a Debiased Neural Basis‑Function Network (DNBNet) to address these challenges. Its core is a debiased neural basis‑function response mechanism, which corrects asymptotic bias through importance sampling while parameterizing basis functions with neural networks to adapt to diverse temporal patterns. In addition, considering the sparsity of irregular data, we design a novel multi‑scale decomposition module based on average pooling, together with a mass‑aware fusion mechanism, to obtain richer representations. Finally, a dual‑branch decoder is employed for forecasting. Extensive experiments on multiple real‑world datasets demonstrate the effectiveness of DNBNet and its strong generalizability across diverse irregular time series scenarios. Our code can be obtained at https://github.com/hnu‑vis/DNBNet.

Authors:Tiancheng Chen, Sheng Tang, Wenhua Jin, Weiqi Zhang, Juntong Fang, Junsheng Zhou, Zesong Li
Title: UniQuery4R: Unified 4D Scene Reconstruction from a Single Query
Abstract:
Reconstructing dynamic 4D scenes requires jointly estimating correspondence, geometry, object motion, and camera motion. Existing feed‑forward methods typically predict dense task‑specific maps or independently process source‑target pairs, leading to unnecessary computation for sparse queries and limited feature reuse across different frame pairs. We present UniQuery4R, a query‑conditioned framework that encodes a multi‑frame clip once and selects the source view, target view, and continuous source‑image coordinate only at decoding time via source‑to‑target cross‑attention. Each query jointly predicts target correspondence, target‑time 3D position, and scene flow, along with source depth, while camera parameters are estimated per view. This design allows the encoded clip to be reused across arbitrary source‑target selections and supports both sparse inference and dense reconstruction through batched queries, without learned temporal embeddings tied to a fixed clip length. We further introduce a direction‑magnitude parameterization of scene flow with separate supervision for moving and static points. Among the evaluated methods, UniQuery4R achieves the best macro‑average results on WorldTrack for both scene‑flow estimation and dynamic‑point reconstruction.

Authors:Soheil Gholami Shahrouz, Vaughn Betz
Title: The Road Less Traveled: Congestion-Aware NoC Placement and Packet Routing for FPGAs
Abstract:
To help scale to ever‑larger and more complex designs, recent FPGA architectures now integrate network‑on‑chips (NoCs). NoCs help transfer high‑bandwidth data over long distances within the chip without using scarce low‑delay long routing wire segments. While NoC‑enhanced FPGAs aid system integration and design reuse, they also complicate FPGA computer‑aided design (CAD) flows by introducing new constraints and metrics. Placement and routing need to optimize NoC metrics like latency and bandwidth utilization and avoid link oversubscription (congestion), while simultaneously optimizing the programmable routing resource usage of the design modules attached to NoC routers. In this work, we develop several new approaches to reduce NoC congestion while minimizing the impact on other design metrics. First, we incorporate a NoC link congestion cost into the placement engine of the open‑source CAD flow, versatile place & route (VPR). Second, we integrate turn model NoC routing algorithms into the placement engine to leverage path diversity to further reduce congestion. On average over a suite of 29 benchmarks, combining placement congestion modeling with turn model packet routing reduces NoC congestion by 90.7% at the cost of increasing aggregate bandwidth demand by 4%. In cases where the enhanced placement engine and NoC routing fail to fully resolve congestion, we formulate NoC routing as a Boolean satisfiability (SAT) problem. This approach yields significant additional improvements; the combined algorithm reduces congestion by 95.1% compared to the baseline placement. Finally, we enhance the reinforcement learning (RL) agent in VPR's placement engine by introducing a NoC‑aware move type, resulting in an 8.8% reduction in wirelength on designs that make extensive use of the NoC.

Authors:Yunhao Yang, Yuexin Bian, Yunjie Tian, Di Fu, Tianjin Huang, Yuanyuan Shi, Ziang Xiao, Nuno Vasconcelos, Yijiang Li
Title: Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL
Abstract:
Reinforcement learning (RL) has emerged as a powerful approach for improving reasoning in language and vision‑language models, yet its strongest successes still depend heavily on ground‑truth supervision (e.g., verifiable reward). Such annotations are costly to obtain and become increasingly scarce as reasoning capabilities advance beyond what humans can reliably evaluate. Self‑rewarding RL reduces this dependence by enabling models to derive reward signals from their own completions. However, training solely on self‑generated feedback can reinforce existing biases and suboptimal behaviors, reduce response diversity, and ultimately lead to homogenized responses and training collapse. In this work, we show that unsupervised reasoning can emerge through cooperative multi‑agent training. We introduce Co‑RL, a framework in which multiple decoupled models, sharing no parameters, are simultaneously optimized through RL using rewards derived from their peers. We further show that increasing cohort diversity, through heterogeneous model families, sizes, and rephrased training samples, reduces the correlated errors that drive self‑reinforcing feedback loops. This diversity consistently improves reasoning performance, maintains behavioral diversity, and mitigates training collapse. Across text‑only and multimodal domains, Co‑RL consistently outperforms the base models and prior label‑free approaches, while matching or surpassing supervised methods, without access to any ground‑truth labels. Concretely, Co‑RL yields average gains of 3.0‑8.6% across seven text‑only benchmarks for LLMs and 2.3‑7.2% across four multimodal benchmarks for VLMs. Code is available at https://github.com/DrStranded/Co‑RL.

Authors:Rodela Ghosh, Aviral Gupta, Guangjing Wang
Title: Which Source Wins? Task-Dependent Reliance in Vision-Language Models
Abstract:
Vision‑language models (VLMs) combine images and text, but when the two conflict and one becomes harder to read, it is unclear how a model shifts its reliance between them. We study this modality reallocation with a controlled setup: we degrade either the image or the text across four levels of legibility while keeping the other clean, and track how the model's preference changes. We build conflicts from GSM8K and SVAMP by pairing the rendered image of one arithmetic problem with the text of another, so the two sources support different answers. We also introduce ChartQA‑Conflict, a manually reviewed benchmark of 229 chart‑report conflicts with matched chart and table‑image representations. We evaluate six open‑weight VLMs using both generated answers and a length‑normalized conditional log‑likelihood margin. On GSM8K and SVAMP, five of six models shift more strongly away from degraded text than from degraded images. On ChartQA‑Conflict, all six likelihood‑scored models exhibit the opposite pattern, shifting more strongly away from the degraded visual source. This reversal persists after calibrating for unimodal accuracy loss and after replacing charts with plain table images. Two frontier API models, GPT‑5.6‑Luna and Gemini‑3.5‑Flash, behaviorally replicate the ChartQA‑Conflict reversal, with GPT‑5.6‑Luna also matching the arithmetic direction. These results show that modality reliance in VLMs is not fixed, but varies across tasks, evidence structures, models, and evaluation settings. The source code is available at https://github.com/Ro‑netizen004/multimodal‑arbitration‑artifact.

Authors:Michael Schleppy, Emina Soljanin
Title: One-at-a-Time Quantum Guessing: Multipartite Entanglement Beyond MoE Games
Abstract:
Multipartite entanglement remains a challenging and not fully understood aspect of quantum information. Monogamy‑of‑Entanglement (MoE) games have been highly effective for studying limitations on the usefulness of entanglement imposed by monogamy constraints. To better reveal the extent to which multipartite entanglement can be useful, we introduce a class of quantum guessing games, termed One‑at‑a‑Time Guessing (OTG) games. In these games, quantum players individually guess the outcomes of random measurements performed by a referee on a pre‑shared entangled state. Unlike MoE games, OTG games select players individually at random according to a specified probability distribution, thereby probing each player's correlation with the referee. We show that, despite monogamy constraints, players sharing certain entangled states can moderately outperform those relying only on classical uncertainty. This advantage arises even in simple OTG games involving only Pauli measurements on qubits, where optimal entanglement increases the winning probability by at least 4%. This contrasts with MoE games, where shared entanglement has been shown in several settings to provide only limited (if any) advantage over classical strategies. We further establish a majorization property: the value of an OTG game respects the majorization ordering of the player‑selection probability distribution. We also analyze in detail a two‑player OTG game in which the referee measures one of the three Pauli observables on a qubit, and show that it is optimally played using a specific parameterized family of three‑qubit W‑like states. These results suggest that OTG games provide a useful framework for investigating the usefulness of multipartite entanglement in multiparty quantum correlations.

Authors:Nam Anh Dinh, Itai Lang, Oded Stein, Rana Hanocka
Title: RADmesh: Remesh-Aware Mesh Deformation
Abstract:
We propose a remeshing‑enhanced method for generatively deforming shapes with visual losses. It is intuitive that sufficiently drastic deformations of a mesh without changing its triangulation can easily compromise element quality, even if such large geometry changes may be semantically desired. Shape deformation methods could thus benefit from changing the triangulation; however, this is not done by most generative, text‑based, visually‑supervised mesh deformation methods. Remeshing is a discrete operation, proven to be especially challenging to couple with the notoriously noisy supervision signal provided by visual losses. We propose a vertex‑based deformation optimization quantity capable of large deformations and robustness to such noise; we periodically remesh using an isotropic remesher that interpolates and carries forward the deformation optimization state. This enables continuous, geometry‑informed progress in coarse‑to‑fine addition of resolution. The resulting shapes' triangulations fit their optimized geometry and have neat isotropic elements. Further, our method is localizable, able to grow new features on a base shape with expressive detail, leaving the rest unchanged. We showcase the effectiveness of our method on a variety of shapes and prompts, both local and global deformations, and demonstrate its superior visual quality and triangle efficiency. Our project page is at https://threedle.github.io/radmesh.

Authors:Neeraj Kumar Singh Beshane
Title: The Acknowledgment Point Is the System: Durable Policy-Decision Receipts for AI Audit Evidence
Abstract:
An AI audit record is useful only if its durability and trust boundary are explicit. Returning a guarded decision before any durable write minimizes latency, but it cannot guarantee that evidence survives an immediate crash. We rebuild RuntimeGuard‑AI around this constraint. The resulting research prototype binds each deterministic policy decision to the exact policy source, commits a privacy‑minimizing record at a caller‑selected synchronization boundary, and returns an Ed25519‑signed receipt that states whether that boundary completed. After restart, the engine validates framed records, manifests, shard placement, sequence continuity, and replay identity. A separate attestation path groups committed records into chained, signed Merkle epochs that an auditor verifies with an externally obtained key. On an Apple M4 Pro at four worker threads and 2,048‑byte prompts, buffered signed evidence reaches 27,193 requests/s with 141.9 microseconds median latency. Per‑record data and full synchronization reduce throughput to approximately 242 requests/s and raise median latency to 16.0 ms. Sealing a 100,000‑record signed epoch takes 97.0 ms. The result is a measured durability‑latency trade‑off, not a "free" asynchronous audit path. The prototype does not prove model execution, prevent a compromised signer from forking history, or establish legal conformity.

Authors:Xiang Li, Yuqi Wang, Casey C. Heirman, Jihye Heo, Kyle J. Lafata
Title: Lymphocyte Mimicry Correction via Region-Level Tissue Reasoning and Unbalanced Optimal Transport
Abstract:
Cell mimicry arises when different cell types appear morphologically similar. Human pathologists resolve this ambiguity using surrounding tissue context, whereas current vision models either lack contextual reasoning (cell foundation models) or cannot operate at the cell level (pathology MLLMs). We present Loki‑OT, which propagates region‑level tissue reasoning to individual cell predictions via Unbalanced Optimal Transport, using MLLM‑derived density priors as soft guidance for ambiguous cell reassignment. Loki‑OT is motivated by the observation that pretrained cell foundation model features already encode discriminative information, including tissue context, but standard cell‑level supervision fails to use tissue context effectively. The resulting transport plan is distilled into a lightweight student MLP classifier that learns context‑aware decision boundaries within the pretrained feature space. On the independent TCGA‑BRCA cohort, Loki‑OT achieved lower patient‑level MAE than the fully supervised in‑domain PanopTILs classifier and improved F1 in epithelium‑rich mimicry tissues, using 278 weak region‑level MLLM estimates built on a general‑domain cell foundation model. Code: https://github.com/xiangli980/Lymphocyte_Mimicry_Correction_via_Loki_OT

Authors:Mariia Gladkova, Neehar Peri, Ishan Khatri, Deva Ramanan, Daniel Cremers
Title: OV3D-Bench: A Diagnostic Benchmark for Open-Vocabulary Monocular 3D Detection
Abstract:
Open‑vocabulary monocular 3D detectors report strong in‑domain performance, but each evaluates under a different protocol, several rely on per‑image category oracles unavailable at deployment, and all collapse geometry and semantics into a single AP metric. To address this, we introduce OV3D‑Bench, a diagnostic benchmark that compares open‑vocabulary monocular 3D detectors under deployment‑realistic conditions across seven indoor and outdoor datasets. Our benchmark replaces the per‑image class name oracle with test‑time dataset‑level class name prompts, and decouples detection accuracy along three axes: localization, semantic robustness, and cross‑domain transfer. We evaluate seven representative detectors and find that (i) they localize objects well yet often mislabel a correctly localized box as a semantically adjacent category; (ii) accuracy is highly sensitive to prompt phrasing (e.g. WildDet3D's performance collapses from 18.6 to 5.4 AP when prompted with "a detailed high‑resolution photo of a car" rather than "car"); and (iii) the widely adopted target‑aware protocol hides these errors (e.g. inflating DetAny3D's AP by 1.9 × on ScanNet). Lastly, we demonstrate that simply remapping a frozen closed‑vocabulary detector's predictions using a contrastive vision‑language encoder such as SigLIPv2 performs competitively against recent purpose‑built open‑vocabulary methods. This indicates that geometric localization is more mature, while open‑vocabulary semantics remains the primary bottleneck.

Authors:Liudmila Rozanova, Alexander Temerev
Title: A Glyph Is Not a Letter, a Token Is Not a Word, a Space Is Not a Space: What the Units of Voynichese Are Not
Abstract:
The Voynich manuscript (Beinecke MS 408) is usually analysed on three unstated assumptions: that its glyphs are letters, that the strings between blanks are words, and that every blank is a word space. We test all three against the Zandbergen‑Landini transliteration with matched prose, cipher, and pseudo‑text controls and quire‑level resampling. None holds, and the failures share a shape: the order in Voynichese sits at the edges of tokens and at graded boundaries between them, not in the succession of tokens themselves. Glyph regularity is too strong for one‑to‑one substitution of any tested plaintext (conditional entropy 2.7 bits against about 3.5 for Latin, Italian, and English) and resolves instead onto a quire‑stable scale of recurrent multi‑symbol units. Tokens form a plausible vocabulary, yet the identity of one token predicts the next by under 1% of token entropy, below every matched control (2‑10%), while the glyphs at token edges share 0.2 bits of mutual information, more than in any prose control. Blanks fall into two regimes: the separators transcribers marked uncertain behave like word‑internal junctures, are physically narrower on the page (AUC 0.905 from independent image coordinates, with the same sign in a small blind ink audit), and are crossed by learned units even when every space is erased before learning. This profile is also what discriminates. A published Voynich‑imitating cipher and a self‑citation text generator both reproduce the low entropy, the unit scale, the weak token order, and the null result of a calibrated substitution attack; neither reproduces the edge‑glyph coupling or the open, hapax‑rich vocabulary (70% singleton types against 41% and 59‑60%). Any account of the manuscript must therefore earn, rather than assume, the step from glyphs, tokens, and separators to letters, words, and word spaces, and these are the measurements on which to do so.

Authors:Konstantinos Kogkalidis
Title: Backward through Time, Algebraically
Abstract:
Linear temporal logic is a modal extension of propositional logic that allows one to state how a system should behave over time. Its canonical domain is the booleans, but discretely‑valued judgements are of little use in steering softly‑valued systems (neural policies, adaptive controllers, sequence models, etc). In such cases, the goal formula's (dis)satisfaction becomes a training signal, and differentiability becomes a prime concern. Candidate differentiable semantics abound, but navigating them is tricky. Implementations, where available, are shallow embeddings, demanding an upfront commitment to a single semantic algebra and its (usually implicit) conduct. The paper casts the reader as a functional programmer asked to come to terms with this predicament, and refusing. Out of that refusal comes an evaluation engine that is algebra‑generic and amenable to differentiation, together with an executable specification of the algebras it can accept. Various algebras are implemented and audited for their behavior, both forward and backward. Each algebra turns out to be a choice of which direction to disappoint, and how. Everything described (and more) is part of the PyTorch library telos, to be found at https://github.com/konstantinosKokos/telos.

Authors:Youwei Zhong, Ben Merbaum, Timos Antonopoulos, Ning Luo, Charalampos Papamanthou, Katerina Sotiraki, Ruzica Piskac
Title: Certified but Private: Scalable Zero-Knowledge Proofs for Neural Network Guarantees
Abstract:
With the growing deployment of machine learning models, formal guarantees of the robustness and fairness of these models have become increasingly important in safety‑critical and legal‑compliance settings. However, model parameters are often commercial secrets that cannot be disclosed to auditors or end users. To this end, we present PANDA, a scalable system that uses zero‑knowledge proofs (ZKPs) to prove the robustness and fairness properties of a model without revealing its private parameters. PANDA is built on top of CROWN, an efficient robustness certification framework that is used in many state‑of‑the‑art formal verification tools for neural networks. The core contribution of PANDA is a novel algorithm for proving linear relaxation bounds for non‑linear activation layers, yielding simple, lightweight proofs. Remarkably, our system can generate proofs of local robustness for neural networks with more than 2.9M parameters in 5 minutes, and can verify them in 10 seconds. Prior ZKP‑based robustness system rely on exponential‑time algorithms that cannot scale to nontrivial networks. In contrast, PANDA scales polynomially in the number of neurons in a network, allowing us to support neural networks 4 orders of magnitude larger than previous approaches with significantly reduced prover overhead.

Authors:Aditya Kumar, Sumit Chongder
Title: Dynamic Entanglement-Weighted Pruning for Quantum Federated Unlearning in Supply-Chain Risk Prediction
Abstract:
Federated deployments of variational quantum classifiers are attractive for cross‑organisation risk prediction in supply chains, because raw data never leaves the client, yet data‑protection regulations such as the GDPR grant clients a right to request that their contribution be removed from a trained model after the fact. Retraining a federated model from scratch to honour such a request is correct but wasteful, and it is not obvious which quantum circuit parameters actually carry a given client's influence. We introduce Entanglement‑Weighted Pruning (EWP), an unlearning procedure for quantum federated learning that scores every trainable circuit parameter with the product of two signals: the diagonal entry of the quantum Fisher information matrix estimated on the target client's data via the parameter‑shift rule, and a structural entanglement weight associated with the parameter's gate. Parameters with the lowest scores are pruned, optionally followed by a short fine‑tuning pass on the retained clients. We implement the full pipeline in Qiskit for a four‑qubit data‑re‑uploading ansatz trained with FedAvg across five simulated supply‑chain‑risk clients, and benchmark EWP against full retraining, fine‑tuning alone, random pruning, Fisher‑only pruning, and entanglement‑only pruning, over three random seeds. EWP attains a mean post‑unlearning accuracy statistically indistinguishable from the full‑retraining oracle, while producing a lower forgetting score and requiring roughly 16 times less wall‑clock time. Ablations over pruning threshold, client count, and non‑IID strength show that combining the two signals is necessary, as entanglement‑only and Fisher‑only pruning each substantially degrade accuracy relative to EWP.

Authors:Md. Jahidul Islam, Mahfujul Alam, Md. Nazmul Islam Seyam, Md. Tamim Hossain
Title: CAS-FD: Contact-Aware Temporal Sampling for Single-View Foul vs Dive Recognition
Abstract:
Distinguishing a genuine foul from a simulated dive in football remains one of the sport's most contested fine‑grained recognition problems, especially when such decisions have to be from a single broadcast view without multi‑view camera angle. We introduce a balanced 600‑clip single‑view Foul/Dive dataset and show that contact‑aware sampling concentrating the model's attention around the moment of physical contact rather than treating all frames equally yields substantially improved recognition of this contact‑ specific problem. The proposed approach achieves 86.0% accuracy and macro‑F1 0.860 on the held‑out test split, a 12 percentage‑ point gain over contact‑unaware alternatives that grows further on unseen data. We also evaluate each pipeline component against human annotations, establishing where and why the system suc‑ ceeds and fails. The result is a documented dataset, a reproducible single‑view pipeline, and a grounded evaluation framework for fine‑grained contact‑event recognition in broadcast football footage. The dataset and code are available at https://github.com/hossain‑ tamim/contact‑aware‑dive.

Authors:Jun Hyuk Lee, Chihyeong Lee, Jooeun Ahn
Title: Lambda-Hold Control: Human-Like Movement Emerges from a Minimal Task Reward in Predictive Musculoskeletal Simulation
Abstract:
The massive overactuation in the human musculoskeletal system makes it challenging to train musculoskeletal models to generate human‑like motion via reinforcement learning, primarily because exploration in the resulting high‑dimensional and redundant action space is extremely inefficient. To address this problem, we propose the λ‑hold controller, inspired by the equilibrium‑point (EP) hypothesis, which has been widely supported by extensive evidence from human motor control studies. The policy's control variable is the per‑muscle EP threshold length λ, from which a stretch‑reflex recruitment law computes the muscle excitations automatically. Holding each λ over an interval of the gait phase also sharply reduces the frequency at which the policy must be queried. Consequently, the controller, to our knowledge for the first time, enables a muscle‑actuated skeletal model to learn human‑like sprinting using only a minimal reward within an hour of training. The efficient exploration through the proposed λ‑hold controller is not merely an engineering trick but an approach grounded in physiology, bringing together the EP hypothesis, intermittent control, and optimal feedback control. Beyond encapsulating human‑like behavior in predictive simulation, this achievement contributes to developing a learnable model of the human motor controller.

Authors:Zhiyuan Yuan, Guanying Chen, Lingteng Qiu, Ruimao Zhang, Shuguang Cui, Xiaochun Cao
Title: PXDepth: Pixel-Space Modeling for Structure Preserving Monocular Depth Estimation
Abstract:
Recent monocular depth estimators achieve strong zero‑shot generalization, yet often struggle to preserve fine‑grained structures and object boundaries. We attribute this limitation to the prevalent combination of large‑patch ViT encoders and convolutional decoders, as coarse tokenization can weaken pixel‑level cues that upsampling cannot fully recover. To address this issue, we propose PXDepth, a discriminative monocular depth model that separates global context modeling from pixel‑level depth prediction. Specifically, a large‑patch ViT captures global scene context, while a pixel‑space predictor composed of Context‑Modulated Pixel Transformer blocks maintains high‑resolution spatial representations throughout depth estimation. This design preserves fine structures and sharp boundaries without sacrificing global depth consistency. Across diverse zero‑shot benchmarks, PXDepth combines faithful local geometry with competitive global depth accuracy while remaining efficient at inference. Our code and model are available at https://yuanzhy29.github.io/PXDepth‑Page/.

Authors:Zeyu Zheng, Shengtong Zhang, Jeremy Avigad, Prasad Tetali, Sean Welleck
Title: The Problem Is the Problem: Towards Scalable Mathematical Discovery
Abstract:
AI systems are increasingly capable of contributing to mathematical research. In research practice, frontier‑model reasoning is a limited resource, and expert mathematical review is even more sharply constrained. Allocating these scarce resources well is therefore central to making AI‑assisted mathematical discovery efficient. In most current AI‑for‑math workflows, human effort is concentrated at the beginning and end, in selecting suitable research problems and later reviewing the resulting artifacts. These two stages are becoming bottlenecks for research‑level mathematics. We address them by proposing a new human‑AI discovery paradigm. The human input is no longer a single problem selected in advance, but a research direction in which the experts have interest and expertise. The system then searches a broad literature corpus for candidate problems in that direction. Inspired by search and recommender systems, we build Find, Attempt, and Recommend (FAR), a literature‑to‑review cascade that automates the search for suitable problems and focuses human attention on artifacts that have passed several stages of filtering. In a combinatorics pilot, the pipeline starts from 5,245 combinatorics papers, recovers 6,453 candidate conjectures or open problems, and filters them to 4,717 apparently well‑posed and still‑open conjectures. Subsequent reasoning and automated triage stages surface 598 potential resolutions and select 77 items for author‑team review. Among them, we identify many interesting discoveries, including results on conjectures and questions of Davies‑‑Jenssen‑‑Perkins‑‑Roberts, Erdős‑‑Straus, Ikenmeyer‑‑Pak‑‑Panova, and Lund‑‑Saraf‑‑Wolf. These results demonstrate the effectiveness of this new mode of human‑AI collaboration for mathematical discovery.

Authors:Mariia Vladimirova, Jean-Yves Franceschi, Thibaut Issenhuth
Title: Position: Fairness Failure in Generative Models is an Evaluation Problem
Abstract:
Despite groundbreaking advancements in generative models during the last decade, concerns about their lack of fairness, reinforcing societal inequalities and harming marginalized groups, remain under‑addressed and difficult to act upon. This position paper argues that fairness failures in generative models, albeit driven by multiple factors, are ultimately stemming from an evaluation problem: fairness findings are rarely comparable across papers or actionable for deployment decisions. This paper diagnoses recurring empirical and conceptual failure modes in current practice and motivates a shift from ad‑hoc bias checks to standardized, generative‑specific evaluation. We propose Fairness Cards as a minimal reporting artifact that makes evaluation choices explicit (prompt families, counterfactual protocols, metrics, and refusal handling) enabling reproducibility, comparability, and accountability. We conclude with additional recommendations towards a paradigm shift in evaluation standards. Our project page can be found at https://mariiavladimirova.github.io/fairness‑cards .

Authors:Aleksi Pippuri, Nilusha Jayawickrama, Risto Ojala
Title: Multi-Observer Vehicle Localization Case Study with Roadside Radar and Connected Vehicle Sensing
Abstract:
In modern intelligent transportation systems, it is essential to accurately estimate vehicle positions, especially in mixed traffic conditions where both connected and conventional vehicles coexist. Roadside infrastructure and connected vehicles can provide complementary observations of the same traffic scene, but real‑world evidence on decision‑level fusion between these sources remains limited. This paper proposes a multi‑observer vehicle localization framework that fuses compact object‑level detections from a static roadside radar and a dynamic LiDAR‑equipped connected vehicle. We evaluate the framework with real‑world data collected at an urban intersection in Helsinki, Finland, with a separately instrumented target vehicle used as the reference trajectory. Two extended Kalman filter based strategies for the localization task were benchmarked. The performance of the radar and LiDAR sensors were evaluated separately, and the two fusion strategies were explored under nominal sensing conditions, reduced LiDAR update rates, simulated LiDAR occlusions, and different target‑vehicle motion states. The results show that, under full LiDAR availability, fusion performance is dominated by the LiDAR observations, while the less accurate and less consistent radar observations provide only limited additional improvement. Nevertheless, AEKF achieves small gains over the LiDAR‑only baseline, and object‑level connected vehicle observations remain useful when shared at reduced update rates. These findings indicate that decision‑level fusion provides scenario‑dependent benefits rather than automatic improvement over a strong single‑sensor baseline. We release the dataset and implementation on Github to support further research: https://github.com/AppuriAalto/multi‑observer‑vehicle‑tracking

Authors:Travis Smith
Title: The Little Scientist: LLM Agent-Driven Discovery via the Scientific Method
Abstract:
What happens when you teach an LLM‑based agent the scientific method? Motivation: Scientific discovery emerges from cycles of hypothesis, implementation, empirical testing, and feedback. Can this process be automated? We approach automated algorithm design through the lens of the scientific method, where an LLM‑based agent goes through each step of the process in an ordered, iterative fashion. Results: We present The Little Scientist, a framework in which a "Scientist agent" works inside an evaluation environment that benchmarks its code and returns structured per‑instance diagnostics. When the Scientist plateaus at a local optimum, a "Kuhn agent" injects a paradigm‑shifting conjecture paired with a cross‑disciplinary inspiration, forcing exploration of a different region of the LLM's latent space. We demonstrate the framework on two problems that require fundamentally different modes of discovery. For protein fitness prediction, the Scientist discovered Delta V, an ensemble calibration strategy that ranks first on the ProteinGym DMS Substitutions Zero‑Shot leaderboard across all five official evaluation metrics, exceeding the #2 model (VenusREM) by +0.033 mean Spearman correlation across 217 DMS assays. For DNA motif discovery, the Scientist wrote an algorithm from scratch‑‑DALE (Dual‑seed Algorithm for Latent Enumeration)‑‑that outperforms STREME (the default in the MEME Suite) across 132 ENCODE transcription factors (mean AUROC 0.842 vs. 0.803, Wilcoxon p < 10^‑6) while running 11x faster. This demonstrates that the framework can produce genuinely novel algorithms, not just optimize existing components. Together, these results show that an LLM agent stepping through the scientific method can discover both new algorithms and new ensemble strategies that outperform prior solutions. The entire research program consumed 704M tokens on a single virtual machine with no GPUs

Authors:Minjun Kim, Jong Hak Moon
Title: Mr.Dec: Daily-Scale Longitudinal Multimodal Modeling for 30-Day Readmission Prediction
Abstract:
Predicting 30‑day hospital readmission is essential for assessing patient stability and optimizing healthcare resources. As clinical risk evolves with the accumulation of evidence during hospitalization, capturing these dynamic trajectories is essential. However, many existing approaches compress the complex longitudinal history into fixed representations, often losing the granular, day‑level clinical signals that reflect a patient's evolving physiological state. To address this, we propose Mr.Dec (Multimodal Readmission‑risk prediction Decoder), which models each admission as a natural chronological sequence of daily multimodal events. By leveraging a Transformer Decoder, Mr.Dec integrates daily Electronic Health Record(EHR) updates and intermittent Chest X‑ray(CXR) findings in a time‑aligned stream, reflecting the actual clinical workflow. To ensure robustness, we utilize Disease‑Specific Supervised Contrastive Learning as an auxiliary regularization to induce a diagnosis‑aware structure in the latent space. Evaluations on the MIMIC‑IV and MIMIC‑CXR datasets show that Mr.Dec achieves state‑of‑the‑art performance by preserving the integrity of the clinical sequence. Furthermore, our model identifies "Critical Days" within an admission, providing actionable and clinically grounded interpretations for real‑time risk stratification. Code is available at: https://github.com/yejix‑ai/MR.DEC

Authors:Weiliang Chen, Haowen Sun, Jun Gao, Jiawei Chi, Hanyang Wang, Qiyu Dai, Yihao Li, Hao Li, Jingnan Gao, Yi-Hsin Hung, Xingzhuo Guo, Shangchen Miao, Zhiyuan Shi, Xiang Li, Fengrui Tian, Weihua Du, Ziqi Huang, Shenyuan Gao, Siqiao Huang, Mingyu Liu, Yifei Li, Shizun Wang, Xi Wang, Tianqi Zhang, Xue Luo, Xiyin Ren, Jinshan Ren, Xiaoyang Shen, Xiaobo Hu, Zhiyang Dou, Mingyu Ding, Yichao Yan, Xinchao Wang, Yizhou Wang, Shilong Liu, Wenzhao Zheng, Yueqi Duan, Yuan Gong, Ziwei Liu, Ming-Yu Liu, Jialong Wu, Jiangran Lyu, Fangfu Liu
Title: HarnessEval-W: Agentifying the Evaluation of Visual Worlds
Abstract:
A benchmark should deliver more than a scalar score: what makes an evaluation trustworthy is the reasoning that justifies the score. This is especially critical for world models, where judging a rollout requires understanding whether physics, causality, and world state evolve correctly. Humans spot such violations naturally, yet no existing benchmark automates this capability: metrics are computed brute‑force, leaving no reasoning chain that can be examined or verified. We introduce HarnessEval‑W, an agentified evaluation pipeline that brings the harness paradigm from the LLM ecosystem to world model benchmarking. Rather than applying a fixed rubric, HarnessEval‑W interprets the context of each evaluation case, decomposes the evaluation question into measurable subproblems, and spawns specialized sub‑agents, each equipped with tailored context and diagnostic tools to reason over its own subproblem. The parent agent then validates the gathered evidence and summarizes it into the final verdict. This hierarchical workflow turns every evaluation into a transparent evidence tree whose complete reasoning chain justifies the result. We apply HarnessEval‑W to 18 representative world models over 330 evaluation cases. Its judgments closely align with human preferences while providing verifiable, fine‑grained diagnoses of every generated rollout. We open‑source the full pipeline as a live benchmark and invite the broad community to contribute to grow new skills and evaluation cases as world models evolve.

Authors:Junhao Hou, Chenqi Luo, Pufan Wang, Jiaying Lu, Yusheng Liu, Feiwei Qin, Meie Fang, Kun Zhou
Title: HiFi-BRep: High-Fidelity Latent Representation for Robust B-Rep Generation
Abstract:
Boundary representation (B‑Rep) generation is a fundamental task in computer‑aided design, yet the direct synthesis of high‑fidelity and structurally valid B‑Reps remains a major challenge. Existing deep generative methods suffer from two forms of brittleness: representation brittleness, caused by padding noise and feature contamination in the latent space, and generation brittleness, stemming from sequential error propagation and a train‑inference mismatch due to non‑differentiable validity enforcement. We propose HiFi‑BRep, a novel framework that addresses these limitations through two synergistic contributions. First, a topology‑aware encoder constructs a high‑fidelity latent representation by eliminating padding via learnable queries and preventing feature contamination with topology‑guided attention. Second, a single‑stage decoder jointly predicts geometry and topology in parallel, embedding core manifold constraints as a differentiable learning objective. This design ensures mutual guidance between geometry and topology while avoiding cascaded errors. Extensive experiments show that HiFi‑BRep significantly outperforms state‑of‑the‑art methods in both structural validity and geometric fidelity, providing a robust solution for high‑quality B‑Rep synthesis. Code and models are publicly available at https://github.com/1nnoh/HiFi‑BRep.

Authors:Zonghao Ying, Xiangfan Wu, Huiyu Wu, Xing Zheng, Huangsheng Cheng, Xiaorong Shi, Jing Guo
Title: Security Assessment of DeepSeek Harness with A.I.G: Evaluating Resistance to Indirect Prompt Injection
Abstract:
We assess indirect prompt injection in DeepSeek Harness (DSH), using AI‑Infra‑Guard (A.I.G) to construct tests, deliver controlled taint, execute DSH, collect traces, and judge outcomes. The study covers 14,560 controlled executions over 16 indirect‑content channels, text and file carrier modes, 35 payload objectives, one unmodified baseline, and 12 attack methods. The experiment preserves DSH's agent loop, tool registry, model adapter, and session‑event path; source tools and sensitive sinks are local fixtures, so attempted actions are recorded without external side effects. We evaluate each trace with a deterministic rule‑based judge, \JudgeR (RuleJudge), and a semantic LLM‑based judge, \JudgeL (LLMJudge). The strongest observed attack success rates are 17.0% under \JudgeL for fake‑completion attack in text mode, 25.5% under \JudgeR for hidden Unicode in file mode, and 16.0% under \JudgeR for the skills channel in file mode. \JudgeL also assigns partial compliance more often than \JudgeR (7.3% versus 2.0%). We relate these results to DSH's treatment of tool results, additional contexts, and tool‑call policy hooks, then identify controls that should sit between untrusted content and sensitive actions. Our code is available at https://github.com/Tencent/AI‑Infra‑Guard/tree/main/Research/deepseek‑harness‑security‑assessment .

Authors:Diptesh Kanojia, Archchana Sindhujan, Sourabh Deoghare, Daria Sokova, Shenbin Qian, Girish Koushik, Tharindu Ranasinghe, Constantin Orăsan, Chrysoula Zerva, Ricardo Rei, Frédéric Blain, André F. T. Martins, Marco Turchi, Matteo Negri, Anoop Kunchukuttan, Mitesh M. Khapra, Pushpak Bhattacharyya
Title: IndicQE-APE: A Benchmark for Quality Estimation and Automatic Post-Editing for Indic Languages
Abstract:
Indic quality estimation (QE) and automatic post‑editing (APE) data is spread across separate releases, so no single resource supports training and evaluation across tasks and language pairs on one footing. We consolidate the WMT 2020‑‑2024 shared‑task lineage with an extended English‑‑Malayalam resource into \indicqe: 126,754 instances over nine directional pairs, with up to four label types aligned on the same segment, a direct assessment, a human post‑edit, word‑level OK/BAD tags and an error explanation, and a test set stratified over four difficulty axes. On it we benchmark six prompted LLMs and three COMET metrics on segment‑level QE, and three systems on APE. Two of the axes are defined partly on the direct assessment and select a compressed slice of it, so each axis is compared against a control drawn from the same language pair with the same score distribution. Only one survives that control: segments whose holistic and token‑level quality signals conflict are ranked worse than equally‑scored segments of the same language, for all nine systems and all seven pairs that carry the axis. Annotator disagreement, which looks second‑hardest without the control, has no effect with it. Few‑shot prompting costs every model \leq 3.4B both correlation and output‑format compliance. Within‑language accuracy does not make scores comparable across pairs: of the three trained metrics, the one with the best within‑language correlation loses most when the pairs are pooled. The benchmark (https://huggingface.co/datasets/surrey‑nlp/IndicQE‑APE) and code (https://github.com/surrey‑nlp/IndicQE‑APE) are released.

Authors:Taegang Kim, Saleh Afroogh, Junfeng Jiao
Title: SafeGesture: Evaluating Fine-Grained Hand Gesture Understanding in Vision-Language Models through Scenario-Conditioned Safety Interpretation
Abstract:
Open‑weight and frontier vision‑language models (VLMs) perform well on general image understanding, but their ability to interpret fine‑grained hand gestures in safety‑critical operational contexts remains largely unexamined. We introduce SafeGesture, a benchmark that evaluates whether a model can infer scenario‑appropriate safety actions from hand gestures. It pairs six HaGRID gestures with eight operational scenarios for 4,800 items and evaluates Qwen2.5‑VL‑7B, LLaVA‑NeXT‑7B, InternVL2‑8B, Phi‑3.5‑Vision, and GPT‑4o. Results reveal a perception‑reasoning decoupling: GPT‑4o achieves 98.4% gesture accuracy but 53.3% safety accuracy, while Qwen2.5‑VL reaches 84.9% and 39.5%, yielding gaps of 45.0 and 45.4 percentage points. Four of five models rarely or never use the uncertainty label, and failure directions differ substantially across models. Accuracy also obscures label bias: a scenario‑majority policy with no visual input reaches 58.3%, above every evaluated model, while only GPT‑4o exceeds this prior under macro‑F1. Visual input improves safety accuracy by 11.2 to 30.2 percentage points, but providing the ground‑truth gesture as text improves performance by only 0.4 to 3.2 points, and no model exceeds 56.2%. These results indicate that the main bottleneck is scenario‑conditioned safety reasoning rather than gesture recognition.

Authors:WooJoo Kim, HyunSik Yoo, JunYoung Kim, JaeHyung Lim, SeongKu Kang, HwanJo Yu
Title: TRACER: Balancing Stability-Plasticity-Cognitivity Trilemma for LLM Enhanced Continual Recommendation
Abstract:
Continual recommendation aims to capture evolving user interests from streaming data but struggles with sparsity. LLM enhancers mitigate this with semantic knowledge, but naive integration creates a new conflict. We identify this as the Stability‑Plasticity‑Cognitivity (SPC) Trilemma, where generalized LLM semantic priors (Cognitivity) conflict with retaining personalized historical preferences (Stability) and adapting to individual interest shifts (Plasticity). To address this, we propose Trilemma‑Responsive Adaptive Continual Enhancement for Recommendation (TRACER). TRACER synergistically combines three specialized modules, each targeting stability, plasticity, or cognitivity, while preventing any single lemma from dominating. This holistic design enables semantic knowledge to support history retention and adaptation to evolving interests without disrupting continual learning. Across five real‑world datasets, TRACER effectively harmonizes the SPC trilemma and outperforms state‑of‑the‑art baselines by up to 14.38%. Our code is available at https://github.com/woo‑joo/TRACER_CIKM26.

Authors:Ingrid Navarro, Yutong Duan, Jonathan Francis, Jean Oh
Title: ScenarioCharacterization: A Modular Toolkit for Characterizing Safety across Trajectory Datasets
Abstract:
We introduce ScenarioCharacterization, an open‑source framework for automated, dataset‑agnostic profiling of driving scenarios in trajectory datasets. Our framework is packaged as a modular, configuration‑driven pipeline of three layers: a dataset adapter that maps custom datasets onto an open Scenario representation, a characterizer that performs feature extraction, behavior probing, and criticality scoring at scenario and agent levels, and an analysis layer for scenario visualization and feature, score, and probe analyses. Because the layers communicate only through Pydantic‑validated schemas composed via configurations, a new dataset can easily plug in without rewriting the characterization and analysis stack. This technical report describes the design and APIs, shows example outputs on Waymo Open Motion, Argoverse2, and nuPlan, and discusses downstream uses of the approach. The framework is available at https://github.com/navarrs/ScenarioCharacterization.

Authors:Peisong Wang, Zhiwei Ma, Bowen Liu, Feixue Liu, Aochuan Chen, Chenyi Zi, Hongchuan Zeng, Yuhan Li, Jia Li
Title: $R^3$-Bench: LLMs Struggle with Resource-Rational Reasoning under Shared Budgets
Abstract:
In cognitive science, resource rationality asks how an agent should allocate limited computation to maximize expected value. Most reasoning and agent benchmarks use independent per‑task budgets; existing shared‑budget studies do not calibrate suite performance against the same model's demonstrated single‑problem competence. We introduce R^3‑Bench, which evaluates six‑problem suites under shared budgets across mathematics, competitive programming, and abstract reasoning in tool‑free and agentic settings. Matched single‑problem response curves define an offline empirical oracle over observed successes. Across 72 main‑table cells for six models, the oracle mean matches or exceeds the contest mean in all cells and is strictly higher in 71. Under moderate tool‑free pressure, equal‑allocation replay also exceeds contest performance for four of six models. Trajectory diagnostics reveal limited strategy updating and pressure‑dependent failure patterns. In a three‑model diagnostic under strong agentic pressure, at least one fixed scheduler exceeds the contest mean in six of nine cells, but no policy dominates across domains. These results expose a persistent gap between demonstrated competence and shared‑budget realization.

Authors:Md Habibur Rahman, Jaeho Kim
Title: Proof-of-Execution Memory: Defending LLM Agents Against Forged-Reasoning Attacks by Verifying What Actually Happened
Abstract:
LLM agents are stateless and rely on external memory to carry context between steps. Because agents treat that memory as trustworthy, an adversary who can write to it can steer their behavior. The FARMA attack does this with no malicious command: it inserts fabricated entries into the agent's reasoning memory claiming a required safety step is already done, so the agent skips it. SENTINEL, the defense proposed with FARMA, scores entries against a fixed list of suspicious wordings; its authors note that an attacker who knows the list can reword the forgery and evade it, and leave this open. We show the gap is worse than stated. An automated attacker that simply asks a language model to reword the forgery evades SENTINEL on its first try, reducing its protection to zero on every model tested. We also find a capability paradox: the attack succeeds far more often on stronger models (98‑100% on GPT‑4o and GPT‑4o‑mini) than on Llama‑3.1‑8B (44%), because more capable agents follow reworded claims more faithfully, so the threat grows with capability. We propose Proof‑of‑Execution Memory (PoEM), which does not inspect memory at all. PoEM keeps a separate, tamper‑evident, HMAC‑chained ledger of the safety steps that actually executed, writable only by the trusted action layer, and allows a skip only if the ledger confirms real execution. An attacker can change what memory says but cannot forge a ledger entry for a step that never ran, so rewording no longer helps. Across three models and three scenarios, PoEM drives attack success to 0% while leaving legitimate operation intact (0% false positives in eight of nine cells, 1.7% in the ninth, within sampling noise), whereas SENTINEL wrongly blocks 33‑50% of legitimate operations. PoEM also withstands attacks aimed at itself, adds microseconds of overhead, and works unchanged in a real LangChain agent. PoEM protects exactly the decisions it gates.

Authors:Mingsheng Zheng, Zirui Jiang, Bo Liu, Yupeng Chen, Jun Zhang, Kai Zhao
Title: Multi-scale Decomposed Convolution Refinement Network for Visible-Infrared Person Re-Identification
Abstract:
Visible‑infrared person re‑identification (VI‑ReID) suffers from cross‑modal discrepancies and limited discriminative capabilities, leading to suboptimal recognition performance. Current approaches exhibit limitations in semantic mining, cross‑modal fusion and feature constraints. To tackle these challenges, we propose MDCRNet, a Multi‑scale Decomposed Convolution Refinement Network that enhances cross‑modal feature learning and discriminative metric learning. Specifically, we introduce a Hierarchical Learning Module (HLM) containing four Hierarchical Decomposed Convolution Attention (HDCA) modules, each equipped with lightweight channel attention and multi‑scale spatial perception blocks to capture multi‑scale spatial dependencies. Moreover, we develop a Joint Discriminative Metric Loss (JDML) incorporating a novel Granularity Discriminative Loss (GDL) that simultaneously optimizes intra‑identity compactness and inter‑identity separability across modalities. Extensive experiments on SYSU‑MM01 and RegDB datasets demonstrate that MDCRNet achieves state‑of‑the‑art performance on both benchmarks. Code is available at https://github.com/Kevin‑zms/MDCRNet.

Authors:Zhaocen Liu, Satvik Praveen, Yi Sheng
Title: Breaking the Compression Barrier: Cross-Architecture Compression Boundary Learning via Reverse Regrowth
Abstract:
Model compression is critical for deploying networks on resource‑constrained edge devices. While pruning‑based methods can significantly reduce model size, they often suffer from abrupt performance collapse beyond a sparsity thresh‑old, making it difficult to identify the feasible compression limit of the model. To address this challenge, we propose a boundary‑Learning reverse regrowth framework, BRIDGE, that reformulates compression as a constructive boundary‑search problem. Unlike forward pruning, our method first drives the model to an extremely sparse state to expose the collapse region, and then selectively regenerates the critical structure to restore performance. The proposed framework employs a hierarchical regeneration strategy, including coarse‑grained layer selection and fine‑grained regeneration parameter selection, to accurately identify which parameters require recovery. Experiments show that our method can recover models from the brink of collapse on both CNNs and Transformer architectures, demonstrating its architecture in‑dependence. BRIDGE achieves a performance improvement of up to 1.49% in unstructured pruning and up to 4.77% in structured pruning. These results demonstrate that reverse regeneration can effectively extend the compression limit while maintaining stable performance. The source code is available at https://github.com/EnumaCaliber/BRIDGE.

Authors:Parsa Mazaheri, Kasra Mazaheri
Title: Prior Audit-Repair Context Shifts LLM Verifier Thresholds Toward Leniency
Abstract:
Automated checking pipelines increasingly place one language model as the checker and another (or the same one) as the fixer. We ask whether that wiring changes what the checker reports. Measuring false alarms on human‑verified‑correct ProcessBench traces with the present task held byte‑identical, we find that a completed audit ‑> repair episode already in the model's context lowers false alarms in 15 of 15 model x wording combinations, by 2.8 to 11.5 percentage points against a length‑matched non‑audit control, a 9 to 25% reduction relative to that control. The direction contradicts what the accumulated‑message literature predicts: an episode whose audit reported an error lowers false alarms further still, at all five wordings on the model where that manipulation lands cleanly, though a negativity asymmetry predicts more flagging. Decomposing the episode finds repair content and audit verdict complementary: different components carry the effect on different model families. Signal‑detection analysis locates the change in the threshold rather than in discrimination ‑‑ the criterion moves in 15 of 15 combinations and survives correction in 13 while d' survives in none, though the d' test is half as sensitive by construction ‑‑ and a hand audit of 50 false alarms finds 82% simply wrong, so at this operating point the shift need not be harmful. With reasoning enabled the effect keeps its relative size on both models tested, and the threshold reading holds there too.

Authors:Kaname Yokoyama, Norimichi Ukita
Title: A Plug-and-Play 2D Motion Interface for Real-World Motion Language Models
Abstract:
Motion Language Models (MoLMs) typically understand human motions by tokenizing 3D motion and processing the resulting tokens using a language model. However, obtaining accurate 3D motions from monocular videos is challenging, limiting their real‑world applicability. To address this issue, we introduce a plug‑and‑play 2D Motion Interface that enables 3D‑pretrained MoLMs to accept 2D motion inputs without modifying or fine‑tuning the original models. Experiments on public datasets show that our method achieves performance comparable to 3D motion inputs across multiple MoLMs and outperforms training MoLMs from scratch on 2D motions. We further construct a monocular real‑world video motion evaluation dataset and introduce a real‑video adapter, demonstrating the usefulness of 2D motions over 3D motions under the evaluated monocular pose‑estimation setting. These results suggest that 2D motion provides a practical interface for deploying MoLMs in real‑world motion understanding settings. Code is available at https://github.com/irajisamurai/2D‑Motion‑Interface.

Authors:Ravi Satya Durga Prasad Yenugula
Title: A Scalable Pipeline for LLM-Teacher Distillation Labeling: Work-Stealing Job Scheduling and Memory-Aware GPU Concurrency
Abstract:
Labeling large text corpora with LLM teachers has become a practical route to training data at scale. At millions of items, hand‑labeling every batch is not feasible, and two questions dominate: what label quality a teacher buys per dollar, and how to keep a fleet of GPU workers busy under skewed, failure‑prone workloads. We present a simple, reproducible pipeline that addresses both. First, a work‑stealing ring pool: each worker owns a queue, drains it first, and then steals from ring successors, with exactly‑once task claims via atomic conditional writes and crash tolerance via stale‑claim sweeping. The claim protocol requires only a compare‑and‑set primitive from its storage layer; we implement it on a single SQLite file, which makes the reference implementation dependency‑free and the experiments reproducible on one machine. Second, a memory‑aware concurrency rule that sizes per‑node parallelism by how many model copies fit on the GPU, so the same code runs safely across device sizes. Third, a relabeling benchmark methodology in which the teacher relabels a public dataset that already has gold labels, so quality reduces to an agreement measurement and cost follows from measured throughput. Under skewed load the pool sustains up to 3.4 times the throughput of static sharding while matching it at zero skew, loses 0 of 2,000 tasks when half the workers are killed mid‑run (static sharding loses 953), and yields measured quality and cost points for an instruction‑tuned teacher on irony and sentiment tasks. All experiments run on public data and commodity hardware; code, tests, and run logs are released.

Authors:Jiawei Xu, Zhilin Zhai, Jinrui Fang, Ruohan Xu, Mingfei Lu, Yi Zhang, Guanchu Wang, Tianlong Chen, Ying Ding
Title: SEER: Long-Context Reasoning via Selective Visual-Text Compression
Abstract:
Long‑context reasoning remains computationally expensive for large language models due to the quadratic complexity of attention over text tokens. Visual‑text compression offers a promising alternative by rendering text into images and processing them with vision‑language models, often reducing token usage. However, existing approaches apply uniform compression regardless of query relevance, potentially sacrificing precision where detailed extraction is required. We present SEER, a framework that learns to select query‑relevant images through visual scanning and retrieve textual content only where needed, combining the efficiency of visual compression with the precision of text‑based reasoning. Through supervised fine‑tuning on tool‑interaction trajectories, SEER learns adaptive tool invocation for selection and retrieval. Experiments on long‑context benchmarks show that SEER improves extraction precision through selective text retrieval while retaining average prompt‑token savings relative to full‑text baselines. On LongBench, SEER achieves 51.11% average accuracy, outperforming the visual‑text baseline Glyph‑9B by 2.33 points and Qwen3‑8B by 3.49 points. Code can be accessed at https://github.com/jiaweixu98/SEER

Authors:William Kalikman, Šimon Sukup, Michal Tešnar, Vilém Zouhar
Title: Augmenting Text to Increase Translation Difficulty
Abstract:
As state‑of‑the‑art machine translation models saturate standard benchmarks, the field needs more challenging evaluations to distinguish between models of varying quality. We propose augmenting existing benchmarks to increase translation difficulty by combining adversarial optimization with a differentiable translation difficulty estimator. Our Adversarial Translation Optimization (ATO) uses gradients from a combined difficulty and fluency objective to iteratively replace tokens. Because each step branches over candidate substitutions at every position, optimization becomes a tree search problem, which we address with Beam Search. ATO offers a gradient‑based alternative to LLM‑based dataset creation without LLM prompting, expensive human curation, or task‑specific model training. Our ATO‑modified benchmark lowers average translation quality (xCOMET) from 0.93 to 0.82, compared to 0.88 for paraphrasing and 0.86 for a zero‑shot baseline. Human evaluation shows the modified texts are somewhat less natural than the baselines but remain reasonably grammatical and plausible while being substantially harder to translate. We release two datasets of 350 English texts each, generated by our methods, as well as the code.

Authors:Marianne Menglin Liu, Leonid Boytsov, Daniel W. Peterson, Pramuditha Perera, Rongguang Wang, Sai Ashish Somayajula, Syed Hamza Rafique, Rohit Saini, Shubham Pathak, Sujeeth Bharadwaj, Tao Sheng, Graham Horwood, Fahad Shah, Ankan Bansal, Sujith Ravi, Dan Roth
Title: PLSQLBench: Benchmarking LLM Systems for Executable Procedural Database Programming
Abstract:
We present PLSQLBench, to our knowledge the first benchmark for evaluating whether LLMs can write executable PL/SQL programs, with correctness measured through execution‑based tests. Existing LLM evaluations largely target general‑purpose code generation or declarative text‑to‑SQL, leaving procedural database programming underexplored. PLSQLBench contains 2,865 instances: 2,594 single‑turn tasks and 271 multi‑turn conversations spanning 978 turns. The benchmark combines complex schema‑grounded tasks over enterprise‑style Spider 2 databases, simpler schema‑grounded tasks derived from Spider, and MBPP‑derived procedural problems, covering varying levels of database grounding and procedural complexity. Experiments with eight LLMs reveal recurring difficulties in schema grounding, PL/SQL dialect fidelity, procedural control flow, exception handling, and cross‑turn consistency. Tool‑augmented LLM agents improve performance on several schema‑grounded evaluations, although substantial gaps remain. These results highlight procedural database programming capabilities not directly assessed by conventional code generation or text‑to‑SQL benchmarks. Our code is available at https://github.com/oracle‑samples/plsqlbench.

Authors:Šimon Sukup, Ariyan Bighashdel, Pavol Jancura
Title: Unified Pedestrian Path Prediction Using Inverse Reinforcement Learning
Abstract:
Pedestrian path prediction is crucial for enhancing the safety of autonomous vehicles and advanced driver‑assistance systems. Previous studies explored different learning‑task formulations for pedestrian path prediction and compared these formulations using shallow neural networks, but did not extend this analysis to more complex deep‑learning models. This paper adapts the Spatial‑Temporal Graph Attention Network (STGAT) to a unified pedestrian path prediction framework and introduces state and action definitions specific to STGAT. The resulting formulations support deterministic and stochastic policies, one‑time and sequential decision‑making, and reinforcement‑learning algorithms including REINFORCE and proximal policy optimization. The proposed learning‑task formulations improve prediction performance across the selected benchmark datasets compared with the standard supervised‑learning formulation. These results demonstrate that reformulating the decision process and training objective can improve an advanced pedestrian trajectory prediction architecture and may provide a path toward improving other graph‑based prediction models.

Authors:Nokimul Hasan Arif, Qian Lou, Mengxin Zheng
Title: Conjunctive Poisoning in AI Supply-Chain Applications
Abstract:
Large Language and Vision‑Language Models are increasingly deployed through inference pipelines that include prompt wrappers (e.g., templates and post‑processing scripts) and configuration metadata (e.g., JSON/YAML files) that together shape model outputs. While model weights and binaries are routinely verified, these textual deployment artifacts remain weakly protected despite directly influencing runtime behavior. We show that a malicious developer can pair a benign‑looking wrapper with crafted metadata to deterministically alter post‑generation behavior without modifying model weights, training data, or inference backend. We study this behavior through a controlled conjunctive‑gate implementation, where activation depends on both an embedded wrapper marker and cryptographically bound metadata. We evaluate the attack across fifteen open‑ and closed‑source LLM/VLM deployments, and assess prompt and system level defenses including static metadata inspection, wrapper scanners, PromptShield, and SigStore‑based artifact signing. To mitigate this risk, we introduce TIF‑BAH, a lightweight middleware defense that verifies wrapper integrity and records behavioral attestations during inference. Our results reveal that wrapper‑metadata interactions form an under‑protected execution layer in modern AI deployments, exposing a deployment‑time behavioral risk that is not captured by model‑weight or prompt‑level defenses. Code is available at https://github.com/N‑H‑Arif/llm_temp.

Authors:Xabier Muruaga
Title: Bounded Agents: Delegation Security for Multi-Agent AI Systems
Abstract:
LLM‑based agents can act on behalf of a user to access cloud services, call tools, or invoke agents. At session start, the agent's permissions are set but remain static, and each request is evaluated independently, without considering prior actions. Within its permissions, an agent may act contrary to the delegated task, combine individually permitted actions into a prohibited outcome, or delegate authority to a sub‑agent without limiting it. A prompt injection poses a risk only if the agent has authority to perform such actions; this is therefore a problem of authorization architecture, not just the model. The Agentic Principal Chain (APC) tracks delegated authority from one principal to the next. APC evaluates each request against the accumulated session state using six authorization checks. APC carries forward and restricts delegated scope and budgets. Using composition closure, APC checks requests against prior actions to prevent prohibited combinations and enforces the decision outside the model. We prove Blast Radius Monotonicity and Composition Soundness for APC implementations; Composition Soundness is limited to prohibited combinations under a complete restriction set and serialized admission. We evaluated 3,154 instances including InjecAgent, AgentDojo, and ASB. Our compromised‑model evaluation tests APC independently of model behavior by inserting the ground‑truth attack call after the first legitimate tool call. AgentDojo exfiltration fell from 75‑100% to 0% across all four domains; APC blocked all 544 InjecAgent data‑stealing cases. Intent binding reduced destruction from 38.6% to 4.0% and manipulation from 90.5% to 12.1%. Authorization latency was 0.24 ms at the 99th percentile on an idle host; across 949 AgentDojo task‑injection pairs, utility was 8.6 and 13.9 percentage points lower in the two settings. Implementation, evaluation tools, and data are publicly available.

Authors:Chunran Zhang
Title: Dense Expands, Sparse Anchors: Channel-Asymmetric Query Expansion for Hybrid Retrieval
Abstract:
LLM‑based query expansion improves retrieval by generating document‑like passages. In hybrid retrieval, however, most evaluations fuse fixed top‑L dense and sparse rankings. Because the cutoff controls both which cross‑channel contributions enter fusion and how much of each ranking is accessed, gains measured at one L can change or reverse at another. We separate these effects by evaluating retrieval effectiveness under complete‑list fusion and recording the policy‑specific per‑channel replay stopping depths at which its ordered top‑K is certified. We then introduce DESA (Dense Expansion and Sparse Anchoring), a channel‑asymmetric query expansion method. An LLM generates complementary reference passages; orthogonal residual expansion adds their new semantic directions to the dense query, while score‑product anchoring incorporates their lexical cues into sparse retrieval without broadening the original query's lexical support. Across seven BEIR datasets, DESA improves nDCG@10 and Recall@20 over the unexpanded query by 3.82% and 2.38%, while reducing dense and sparse access depths by 36.90% and 36.56%. With equal dataset weighting, 63.31% of queries become shallower in both channels. However, both depths increase with Contriever on Touché‑2020. These results support channel‑specific integration of generated passages and joint evaluation of retrieval effectiveness and access depth.

Authors:Xudong Chen, Shengbo Gong, Lu Cheng, Wei Jin
Title: Temporal Graph Prototype-conditioned Conformal Prediction for Fraud Detection
Abstract:
Conformal prediction (CP) provides distribution‑free coverage guarantees and has emerged as a principled tool for uncertainty quantification. In edge‑level fraud detection on temporal interaction graphs, where false positives and false negatives both carry substantial cost, such coverage guarantees are particularly appealing for risk‑aware decision making. However, directly applying existing graph conformal predictors yields inefficient prediction sets due to two recurring properties of fraud data. Fraudulent interactions are often embedded in benign‑dominated neighborhoods that dilute calibration signals, while extreme class imbalance leaves scarce labeled‑fraud support in the calibration split and leads to overly conservative class‑conditional thresholds. To address these issues, we propose ProtoCP, a conformal prediction framework for edge‑level fraud detection on temporal graphs. ProtoCP improves calibration efficiency by focusing calibration on fraud‑relevant subgraph context and producing more stable nonconformity scores under class imbalance and temporal drift. Specifically, it leverages learned prototypes to suppress benign‑dominated noise in the calibration context and introduces a neighborhood‑relative scoring mechanism with temporal score diffusion for stable class‑conditional calibration. Experiments on four fraud benchmarks (YelpChi, S‑FFSD, FTFD, and BankSim) show that ProtoCP achieves the target coverage with consistently smaller prediction sets than state‑of‑the‑art baselines. Our codes are available at https://github.com/Picard1701ent/ProtoCP.git

Authors:Armin Steinhauser
Title: TinyCast: Probabilistic Zero-Shot Forecasting with Computed Periodicity
Abstract:
We introduce TinyCast, an attention‑free zero‑shot forecaster that emits a predictive distribution from 146,505 parameters, on the premise that at this size the periodic structure of a context is worth computing rather than learning. A zero‑parameter spectral detector supplies the dominant periods, the context is folded on their phase, and a dilated convolutional encoder and a block‑autoregressive quantile decoder model the rest. It is smaller than every zero‑shot entry on the GIFT‑Eval board whose parameter count can be established. On probabilistic accuracy it defines the size‑accuracy frontier. Among zero‑shot entries declaring no test‑data leakage it is the only one below 1.4M parameters that emits a predictive distribution, and every entry scoring better carries at least that budget. On Chronos‑ZS and fev‑bench every neural model ahead of it carries at least 28 times its parameters. Because the mixing path is convolutions and matrix multiplications only, it exports to static INT8 and forecasts end to end on an embedded device without per‑signal fitting.

Authors:Zheming Xu, Aiyue Tang, Shidi Chen, Xuechao Zou, Congyan Lang, Rogelio A. Mancisidor, Michael Kampffmeyer
Title: Beyond Independence: Learning Correlated Views for Variational Incomplete Multi-View Clustering
Abstract:
Incomplete multi‑view clustering (IMVC) aims to uncover shared cluster structures from data with partially observed views. Although recent imputation‑free methods based on variational inference demonstrate robustness to missing views, they commonly rely on a conditional independence assumption across views in the posterior aggregation stage, which fails to capture the inherently structured and potentially correlated nature of multi‑view data. In this paper, we propose a variational framework that explicitly goes beyond this assumption by introducing a learnable cross‑view correlation structure. Specifically, we explicitly model and learn correlations between views by utilizing the covariance structure of posterior estimation errors during aggregation. To facilitate robust and efficient learning, the correlation matrix is parameterized through a normalized Cholesky decomposition, ensuring positive definiteness and enabling the entire model to be trained jointly through a unified variational objective. Extensive experiments on multiple IMVC benchmarks demonstrate that our method consistently outperforms state‑of‑the‑art approaches across diverse missing‑view settings while introducing only a negligible number of learnable parameters. These results highlight the effectiveness of adaptive correlation modeling in variational IMVC, demonstrating the need to go beyond the independence assumption in IMVC. The code is available at https://github.com/zmxu196/ACOVA.

Authors:Jinhui Sun, Wei Zhou, Bowen Yang, Xinliang Xiao, Li Yang
Title: Making two action heads agree: coordination mechanisms and a runtime collapse certificate for flow-matching policies
Abstract:
A dual‑representation flow‑matching policy decodes each predicted motion into joint and end‑effector spaces, and the residual between the two kinematically equivalent decodings provides a physically interpretable runtime signal. On multimodal tasks, however, independently sampled branches may choose different valid modes, causing false alarms. We study how to coordinate the two branches and at what cost. Across two robot environments and a non‑robotic testbed, the tested mechanisms fall into four classes. An auxiliary latent shared by both branches but absent from the flow‑matching construction is erased at the population optimum, a provable dead end confirmed within a prespecified 2% equivalence band. Sharing source noise can coordinate or anti‑coordinate: its effect changes sign with the representation map and tracks the alignment of decoder mode basins. Consistency regularization gives intermediate coordination but reduces the valid‑pair rate, while training‑supported discrete partitions achieve near‑ceiling coordination robustly. We further derive a chance‑corrected coordination bound based only on each branch's Gini‑Simpson diversity, yielding an attainable region and a label‑free certificate that separates coordination from collapse when zero mismatch is ambiguous. On LIBERO‑Plus, benign multimodality adds 1.57 percentage points of false alarms to the residual, which remains the strongest evaluated failure signal; the preregistered token intervention does not meet its false‑alarm criterion or produce a seed‑robust detection change. Code, models, and per‑run configurations are available at https://github.com/kimo423/dual‑head‑coordination.

Authors:Srikumar Krishnamoorthy
Title: Learning Auditable Classifier Models: Source-Disjoint Tree Ensembles
Abstract:
Predictive models in clinical and regulated settings must be accurate and fully auditable. Tree ensembles deliver strong accuracy on tabular data, but their sequential boosting couples structure discovery with coefficient estimation, making compact per‑prediction auditing difficult. Interpretable alternatives impose structural constraints that limit expressiveness: generalized additive models typically restrict interactions to pairwise terms and post‑hoc rule extractors produce overlapping rules that hinder compact interpretation. We introduce Residual Pattern Tree Ensemble (RPTE), a three‑stage learning approach, that is built on three key principles: bounded feature budget, source disjointness, and separate coefficient estimation. Stage~1 builds a supervised symbolic feature vocabulary. Stage~2 grows shallow trees under a source‑disjointness constraint, where each raw variable is allocated to at most one tree, and retains only the discovered tree structures. Stage~3 solves a single \ell_1‑regularized logistic regression over leaf‑region indicators, yielding jointly optimal sparse coefficients. This learning approach ensures that every prediction decomposes into an algebraic sum of named, non‑overlapping rule contributions, enabling full auditability by design. Empirical evaluation on twelve clinical‑domain binary classification benchmarks using repeated stratified 5‑fold cross‑validation shows that RPTE performs competitively against tuned opaque ensembles and interpretable baselines. RPTE reduces model inspection units by 9× to 87× relative to XGBoost and maintains lower audit complexity than EBM on all 12 datasets. RuleFit requires comparable or fewer inspection units on three datasets where its rule count is small, but without source‑disjointness guarantees. The source code is available at \hrefhttps://github.com/srikumar2050/hugiml‑corethis https URL.

Authors:Jinhwan Seo, Kyubeom Han, Jumin Lee, Junhyug Noh, Sung-eui Yoon
Title: What You Ask is What You Ground: Bridging Question Intent to Temporal Evidence for Grounded VideoQA
Abstract:
We study a critical yet overlooked failure mode in Grounded Video Question Answering: question‑invariant grounding, where models predict nearly identical temporal segments for different questions about the same video. We trace this behavior to two structural limitations in prior common designs: (i) modality isolation that fixes video representations before they receive question semantics, and (ii) weak question injection inside the grounding module. To address this, we propose GroundFormer, which conditions video features on question intent before localization via learnable communication tokens that mediate directed visuo‑lingual interaction. On top of the question‑conditioned features, a factorized MIL cross‑attention couples answer selection with temporal evidence under candidate‑level supervision, while Gaussian smoothing converts peaked attention into temporally coherent segments. We further introduce a hierarchical multi‑modal contrastive loss that aligns video, question, and answer embeddings across a two‑pass training pipeline. GroundFormer achieves state‑of‑the‑art grounded VideoQA performance on NExT‑GQA and STAR, substantially improving question‑discriminative temporal grounding.

Authors:Xin Lin, Haodong Li, Zhifei Zhang, Yutong Yang, Haitian Zheng, Juanxi Tian, Zhe Lin, Truong Nguyen
Title: PixelControl: Fine-Grained Condition Fidelity in Text-to-Image Diffusion
Abstract:
Controllable text‑to‑image diffusion models can often follow the global layout of spatial conditions, yet still violate fine‑grained structures such as object boundaries, thin contours, and medium/small conditioned regions. This limitation is especially problematic for VAE‑based latent diffusion, where spatial compression can weaken high‑frequency and low‑area condition signals. We propose PixelControl, a pixel‑space controllable diffusion framework for fine‑grained condition fidelity. Built on a PixelDiT‑style backbone, PixelControl avoids the latent bottleneck and introduces two complementary designs. First, Structure‑Aware Control Injection derives a condition structure map and uses it to strengthen injected control residuals around spatially sensitive regions. Second, Multi‑Scale Pyramid Cycle Loss verifies generated images against condition‑derived structures across multiple resolutions, balancing global layout consistency with local boundary and detail accuracy. PixelControl supports depth, segmentation, edge, and their combinations through modality‑specific control branches with lightweight gated fusion. Experiments across depth, segmentation, and edge control show that PixelControl improves structural fidelity and visual quality over existing controllable generation methods, with especially strong gains on boundaries and medium/small conditioned regions. The project page can be found at: https://linxin0.github.io/pixelcontrol_homepage/pixelcontrol‑site/

Authors:Peng Chunyi, Xu Zhipeng, Yan Yukun, Liu Zhenghao, Yu Shi, Mei Sen, Sun Yubo, Zhang Yongheng, Zhou Jie, Gu Yu, Yu Ge, Sun Maosong
Title: ConceptFormer: Learning Adaptive Latent Concepts for Query-Document Alignment in Visual Document Retrieval
Abstract:
Visual document retrieval is a critical component of multimodal retrieval‑augmented generation, aiming to identify query‑relevant pages from document collections where evidence is distributed across text, layout, charts, and visual structures. Recent efforts toward finer‑grained supervision primarily rely on textual descriptions or localized visual regions as evidence proxies. However, such supervision signals may either overlook complex visual structures or provide incomplete and inaccurate representations of the underlying evidence. To address these limitations, we propose ConceptFormer, a latent concept representation learning framework for visual document retrieval. ConceptFormer models query‑relevant evidence as continuous, query‑conditioned latent concepts that explicitly bridge localized visual evidence and semantic relevance, without requiring either textual intermediate representations or direct reliance on raw visual annotations. During training, ConceptFormer employs a strong vision‑language model to dynamically determine the number of latent concept tokens and uses these concepts as an intermediate representation to bridge the semantic gap between queries and documents, thereby guiding the learning of the embedding space. Experiments on diverse visual document retrieval benchmarks demonstrate that ConceptFormer achieves 16.7% and 22.1% relative improvements in average NDCG@10 over the strongest visual retrieval baseline and the strongest OCR‑based text retrieval baseline, respectively. Further analysis reveals that latent concepts effectively connect localized visual evidence with semantic relevance, enabling the retriever to capture both fine‑grained textual cues and complex document‑level visual structures while preserving strong retrieval alignment. Codes and data are available at https://github.com/Neuir/ConceptFormer.

Authors:Subhransu Das, Jiaming Cheng, Arnav Kumar, Sadia Afrose, Mingzhe Han, Michael Silagy, Shreya Palande, Brijesh Soni, Rajiv Ramnath
Title: Large Models for Small Devices: Recent Advances and Empirical Analysis of Edge AI Deployment
Abstract:
Running large AI models on resource‑constrained edge devices requires model compression to reduce model size and computation. What compresses well, however, need not deploy well. We survey dozens of recent works that report compression results on real hardware and extract practical deployment guidelines from them. Following these guidelines, we deploy compact language and image models on GPU, CPU, and Raspberry Pi platforms across question answering and image segmentation. No single technique wins across tasks. For question answering, Qwen3.5 0.8B reaches 93.85 SQuAD F1 and 92 EM under Q5_K_M GGUF quantization, while structured pruning at the same precision costs 16 F1 at a 1% ratio. For segmentation, the ranking reverses: default quantization leaves parameters and MACs unchanged, whereas pruning cuts model size by nearly 80% at near‑constant mIoU. Pruning can even inflate the deployed artifact by 21‑49% by breaking k‑quant super‑block alignment; combined with longer, less format‑compliant outputs, this raises Raspberry Pi latency up to 3.4x. Compression can also manufacture the appearance of competence rather than destroy it visibly: one LoRA‑recovered variant stays fully parseable and holds 71% strict BoolQ accuracy while sending 97 of 100 predictions to a single class, at 52.6% balanced accuracy. We explain these effects through neural‑flow graph analysis and prefill‑decode‑level latency decomposition, and condense them into task‑specific deployment research directions. The right technique depends on the task, the model, and the hardware. Our experiment code and artifacts are open‑sourced at https://github.com/Arnavvvkumar/deployment

Authors:Tomasz Stanczyk, Seongro Yoon, Francois Bremond
Title: Training-Free Long-Term Multi-Object Tracking for Sports Video Analytics
Abstract:
Long‑term multi‑object tracking in sports remains challenging due to frequent occlusions, rapid camera motion, and repeated player reappearances. We introduce McByte++, a training‑free tracking‑by‑detection framework that integrates lightweight mask propagation, conditional camera motion compensation, and online re‑identification within a unified pipeline. Compared to its predecessor, McByte++ substantially improves runtime efficiency while enhancing identity preservation. On SoccerNet‑tracking and SportsMOT benchmarks, McByte++ achieves up to +3.0 HOTA and +6.1 IDF1 improvements over the original McByte in the online setting, with further gains when combined with offline global association. Replacing heavy segmentation components and optimizing motion modeling yields up to an order‑of‑magnitude speed increase. All results are obtained without detector retraining or dataset‑specific tuning. Code will be made available at https://github.com/tstanczyk95/McBytePlusPlus.

Authors:Rama AlHamidi, Rasul Khanbayov, Erchin Serpedin, Hasan Kurban
Title: Counterfactual Sensitivity Is Not Repairability: Auditing Replay Probes for Video Evidence
Abstract:
Tool‑using video agents retrieve visual evidence before answering, but the final answer is not forced to depend on what was retrieved. The natural black box test is counterfactual: destroy the semantic content of the frames the agent retrieved and check whether the answer changes, against a matched sham that re‑executes the identical pipeline on those same frames. We introduce CARVE, a black‑box counterfactual probe that compares answer changes under matched SHAM and DESTROY replays. Across three independent k=3 runs on a frozen VideoExplorer‑style agent, DESTROY changes the answer 29.3 percentage points more often than SHAM, yielding a large and reproducible aggregate effect. Question‑level scores are less stable, and increasing the replay budget from k=3 to k=10 reduces ties but weakens the original zero‑threshold routing policy. At k=3, CARVE selects 538 of 1,258 LVBench questions and improves accuracy by 3.26 points, with higher fallback yield than most matched random subsets. The score shows only a weak association with annotated temporal coverage, so CARVE is best understood as a routing signal rather than a direct grounding classifier. Our implementation is available at https://github.com/KurbanIntelligenceLab/CARVE.

Authors:Asaf Livne, Amir Jevnisek, Shai Avidan
Title: Scalable Black-Box Model Attribution for Images
Abstract:
The rapid proliferation of generative models raises the model attribution problem: given only an image, can we determine which model produced it? Existing methods have grown as elaborate as the generators they target, on the as‑ sumption that a more sophisticated model demands a more sophisticated attributor. We show it does not. RPA (Raw‑ Patch Attribution) attributes images in the strictest black‑ box setting with a lightweight CNN. Despite its simplicity, it attributes more models at higher accuracy than prior work, reaching 98.0% on 25‑class DRAGON and 92.9% on 27‑ class OpenFake; it is data‑efficient and runs at a cost inde‑ pendent of the number of candidate models; and it stays ro‑ bust to the compression, blur, and resizing images undergo in the wild. Training for closed‑set attribution yields a ver‑ satile feature extractor: the same representation recovers model lineage without supervision, flags and groups unseen generators, and admits new models through few‑shot adap‑ tation rather than retraining.

Authors:Bin Ren, Qi Ma, Yue Li, Zongyan Han, Yidi Li, Yuqian Fu, Rao Muhammad Anwer, Theo Gevers, Fahad Shahbaz Khan, Salman Khan
Title: Gaussian-JEPA: Joint-Embedding Predictive Learning for 3D Gaussian Splats
Abstract:
3D Gaussian Splatting (3DGS) represents 3D content with anisotropic primitives that jointly encode geometry and appearance. Fixed‑budget encoders consume sampled observations of Gaussian assets, so the same object may be observed through different primitive realizations. Existing self‑supervised methods mainly reconstruct masked Gaussian attributes, tying supervision to one sampled realization and requiring an input‑space decoder. Latent prediction offers an alternative, but its application to Gaussian tokens requires targets that accommodate coupled attributes and heterogeneous spatial support. We introduce Gaussian‑JEPA, which predicts representations of held‑out Gaussian token blocks from visible context. An online encoder processes the context, while a shared exponential‑moving‑average encoder supplies stop‑gradient features for multi‑scale targets. Complementary target projections and feature‑space grounding provide latent supervision without reconstructing Gaussian attributes. We evaluate the features under Gaussian resampling, partial observations, and renderable shape completion, together with transfer to part segmentation and object classification. Compared with matched reconstruction pretraining, Gaussian‑JEPA is more consistent across resampled inputs, retains more instance information under partial observations, and provides stronger frozen features for Gaussian completion. These results support latent prediction as an effective objective for reusable 3D Gaussian representations. Code is on the project page (https://amazingren.github.io/Gaussian‑JEPA/).

Authors:Shuaishuai Cao, Meng Tang, Shuwei Peng, Xuan Liu, Min Huang, Jie Chen, Jiacheng Niu, Yong Chen, Edore Akpokodje, Hui Lin
Title: Hierarchical Adaptive Feature Refinement Network for VHR Remote Sensing Image Segmentation
Abstract:
Semantic segmentation of very‑high‑resolution (VHR) remote sensing imagery increasingly benefits from strong pretrained hierarchical encoders, yet exploiting their multi‑stage representations remains difficult. Nearby regions demand different balances between fine detail and semantic context, aggressive task‑specific transformations perturb useful pretrained features, and conventional semantic supervision provides limited structural guidance. We present HAFR‑Net, a progressive refinement framework that adaptively organizes and conservatively refines hierarchical representations instead of replacing them with a monolithic decoder transformation. Heterogeneity‑Guided Stage‑Adaptive Fusion (HG‑SAF) predicts dense stage weights conditioned on local feature variation. A Frequency‑Residual Adapter (FRA) then injects frequency information through a bounded, zero‑initialized residual branch that keeps the fused representation as its reference. A Confusion‑Aware Tri‑Prior Decoder (CATP) finally regularizes the prediction with boundary, objectness, and training‑derived class‑relation cues. Under a matched Swin‑B training and single‑scale inference protocol, HAFR‑Net attains 84.12%, 87.86%, 55.17%, and 67.70% mIoU on ISPRS Vaihingen, ISPRS Potsdam, LoveDA, and OpenEarthMap, improving the matched UPerNet baseline by 0.55, 0.95, 1.55, and 1.84 percentage points, respectively. Controlled analyses further show consistent spatial reweighting beyond content‑only routing, improved boundary and thin‑structure accuracy over matched spatial and spectral alternatives, and reduced confusion on pre‑declared class pairs.

Authors:Md Rezwanul Islam
Title: Benchmarking Quantum Machine Learning for Power-System Attack Detection: Evaluation Choices Decide the Outcome Before the Models Do
Abstract:
Machine‑learning detectors for power‑system cyberattacks are themselves attack surfaces, and quantum machine learning has been proposed for them. We benchmark fidelity‑kernel SVMs and variational classifiers against six tuned classical models on public power‑system attack data (Mississippi State/ORNL), across white‑box, transfer, decision‑based black‑box, and poisoning attacks. Our headline finding is methodological: the benchmark's answers are set by the evaluator's choices before the models. Eight choices ‑‑ six in the evaluation protocol, two in the tuning the benchmark itself runs ‑‑ each reversed or moved a conclusion at fixed models. The largest is the split: the row‑level protocol scores 0.905 macro‑F1 where holding whole source files out leaves 0.594, and in the capped matched‑dimensionality regime the quantum arm sits within noise of chance with the classical arm 0.024 above it. A fidelity kernel looks most robust until attacked directly (retention 0.886 to 0.064); a mis‑fitted surrogate manufactures a 10x asymmetry; an unseeded black‑box attack moves 75% between restarts. A positive control explains the accuracy null: the labels, not the pipeline. We give the control that catches each choice and release the seeded benchmark.

Authors:Kuan-Lin Chen, Tzu-Ti Wei, Chao-Chi Liao, Yu-Chee Tseng, Jen-Jee Chen
Title: AlloEgo-VLM: Disambiguating Allocentric and Egocentric Reference Frames in Vision-Language Models
Abstract:
This study investigates the challenge of ambiguity faced by Vision‑Language Models (VLMs) in understanding spatial semantics. Spatial cognition, shaped by cognitive psychology, spatial science, and cultural context, often assigns directionality to objects. However, natural language descriptions of spatial relations frequently omit explicit reference frames, leading to semantic ambiguity and potentially serious errors for embodied AI robots. Existing VLMs, due to insufficient training on reference frames and object orientations, often produce inconsistent responses. To address this issue, we construct a new dataset, AlloEgo‑View, comprising (image, query, view‑specific answer) triplets that capture key object relations from both allocentric and egocentric perspectives. The view‑specific descriptions follow a structured spatial representation that annotate detailed scene descriptions, reference and target objects, their orientations, reference frames, and view types. Building on AlloEgo‑View, we develop AlloEgo‑VLM, a framework to disambiguate allocentric and egocentric reference frames, even under ambiguous queries, and to be easily integrated into existing VLMs via supervised fine‑tuning. Furthermore, we deploy our framework onto an embodied robotic platform within NVIDIA Isaac Sim to validate its real‑world feasibility in open‑ended object searching tasks. Experiments highlight the limitations of current VLMs in handling view‑specific queries and demonstrate the strong disambiguation ability of AlloEgo‑VLM.

Authors:Qingyao Yang, Runming Yang, He Xiao, Wendong Xu, Junyu Chen, Haobo Liu, Chenchen Ding, Ruihan Hu, Yik-Chung Wu, Ngai Wong
Title: FluxBin: Flexible LUT-based Ultra-low-bit LLM Inference by Algorithm-Kernel Synergy
Abstract:
While binary quantization theoretically promises extreme compression and acceleration for Large Language Models (LLMs), existing research often overlooks the necessity of specialized hardware kernels, thus failing to unleash the full acceleration potential due to persistent reliance on expensive floating‑point arithmetic or runtime dequantization overheads. To bridge this gap, we propose FluxBin (Flexible LUT‑based Ultra‑low‑bit eXecution with Binary bases), an algorithm‑kernel co‑design that synergizes post‑training quantization with a highly optimized CUDA kernel. Algorithmically, we introduce Decoupled Row‑Column Binary Decomposition to enhance representational capacity while maintaining hardware efficiency, complemented by a Hessian‑guided saliency‑aware hybrid bases that preserve critical information. At the kernel level, we implement a Lookup Table Building Approach with Scale Fusion to reduce floating‑point arithmetic, featuring a Virtual Columnar Mapping that transforms irregular, sparse, and salient matrices into dense execution. Extensive evaluations demonstrate FluxBin achieves up to 5.92× speedup and 10.19× energy savings across diverse model architectures, delivering comparable accuracy to heavily fine‑tuned methods. This effectively enables the deployment of 70B‑scale models on one single A100 GPU with a 4× memory reduction. Code is available at https://github.com/nicyyyy/FluxBin.

Authors:Yufeng Chi, Huimin Ma, Fan Gao, Zhice Niu, Keqin Li, Jianmin Li
Title: PoseAdapter: Dual-Stream 2.5D Controllable Image Generation for Complex Multi-Object Scenes
Abstract:
While Text‑to‑Image (T2I) diffusion models have achieved remarkable success, precise spatial and orientational control in multi‑object scenes remains a persistent challenge. Existing methods either rely on computationally expensive dense 3D maps or suffer from severe attribute leakage and "cut‑and‑paste" artifacts. To address these limitations, we propose PoseAdapter, a lightweight framework for high‑fidelity 2.5D controllable image generation. Instead of dense spatial maps, it establishes precise spatial‑angular anchors using an efficient condition layout: individual object captions, 2D bounding boxes, and 3D angles. To resolve the generative trade‑off between strict instance isolation and global coherence, we introduce a Context‑Aware Dual‑Stream Representation. By injecting local object tokens and relation‑enriched scene tokens into the visual stream of modern MM‑DiT architectures via parallel masked and unmasked pathways, PoseAdapter eliminates attribute leakage while preserving natural inter‑object relationships and scene‑level coherence. To support this paradigm, we construct OrientLayout, a high‑quality dataset featuring standardized 2.5D annotations and instance‑level decoupled semantics. Extensive experiments demonstrate that PoseAdapter outperforms state‑of‑the‑art baselines in spatial accuracy, orientational precision, and multi‑object visual fidelity. Code and dataset will be available at https://github.com/cyf23/PoseAdapter.

Authors:Yogesh Kumar
Title: Catching Hallucinated Citations in Video-LLM Question Answering: A Self-Verification Pipeline and Verifier Ablation Study
Abstract:
Video question answering systems built on vision‑language models often produce timestamped claims with high confidence even when unsupported by the cited frame. This deceptive hallucination arises because timestamps imply grounding without ensuring correctness, increasing user trust but not accuracy. We introduce a pipeline that closes this loop. A retrieval‑augmented language model drafts answers with per‑claim timestamp citations, and each cited frame is independently re‑examined before being shown to the user. We compare against a plain baseline and ablate three verification designs, evaluated on both Apple Silicon (MLX) and Google Colab (HF Transformers, CUDA). Directly asking the vision model whether a frame supports a claim fails completely (0% catch rate on 40 claims) due to sycophancy. Blind re‑captioning plus a general LLM judge improves results but is unstable, oscillating between 0% and 100% flagged depending on prompt phrasing. Replacing that judge with a small natural language inference model yields a stable, interpretable verifier that catches 79% of fabricated claims on adversarial false‑premise questions while leaving true claims untouched. We release the full pipeline, evaluation harness, and implementations for both Apple Silicon and Colab. Code is available at https://github.com/yogesh‑iitj/grounded‑video‑qa.

Authors:Junbo Jacob Lian, Huiling Chen, Hanzhang Qin, Chung-Piaw Teo
Title: Admission Without Answers: Label-Free Certification and Experience Learning for LLM-Based Optimization Modeling
Abstract:
Experience‑learning agents for optimization modeling improve by storing verified skills, but existing learners admit knowledge by checking against known answers, which real ticket streams do not provide. The natural label‑free alternatives are unreliable: on a 300‑problem label‑blind stream, admitting every executable model poisons roughly one admission in four, while single‑instance agreement accepts models that match at one value but differ elsewhere. We propose AdmitOR, an admission gate built on calibrated external behavioral evidence. Candidates from three model families, prompting strategies, and solver stacks are run on instances resampled from an extracted parameter domain; agreement across the resulting value‑function traces is summarized by a cross‑family clique, and a calibrated threshold returns accept, abstain, or escalate. The preregistered false‑discovery criterion holds on calibration data but not on the wild stream. We report this negative result in full and trace most failures to benchmark texts that do not faithfully encode their labeled instances. Comparing four admission judges on one collection of logs inside a state‑of‑the‑art skill learner, AdmitOR raises admission precision to 0.927, against 0.871 for majority vote and 0.726 for execution success, yielding 3.1x and 8.0x fewer poisoned admissions. Its library is the smallest and attains the highest macro accuracy across five public benchmarks, 58.4 against 54.8 for majority vote and 53.9 for the ground‑truth‑labeled library. The 3.5‑point gain over majority vote is supported by a paired bootstrap and survives correction for a host‑side anomaly. To our knowledge, AdmitOR is the first label‑free admission mechanism designed around an explicitly calibrated false‑discovery target. The transfer failure identifies a necessary condition for extending it to wild streams.

Authors:Swarnim Jain, Shangzhe Wu
Title: RigidBench: Evaluating Rigid-Body Physics in Video Generation Models
Abstract:
Video models are increasingly used to predict what happens next in a scene, yet the metrics commonly used to compare their outputs say little about whether the predicted objects move correctly. Motion, geometry, identity, background stability, and visual similarity can fail independently, but whole‑frame scores often mix these errors together. We introduce RigidBench, a simulator‑grounded benchmark that compares a generated continuation with a reference rollout from the same initial frame and motion description. Its five rigid‑body tasks vary objects, materials, viewpoints, and indoor and outdoor scenes, with per‑frame masks, depth, 6‑DoF trajectories, and contacts available for scoring. We evaluate eight models on the same 100 examples with ten measurements that keep these aspects separate. The resulting rankings depend strongly on what is measured: no model leads on all ten, and across model means, higher SSIM accompanies larger 3D trajectory error (r = 0.89). RigidBench also includes 5,000 training videos with exact simulator state, which we use to fine‑tune and analyze Wan 2.2 TI2V‑5B. Full fine‑tuning reduces 3D trajectory error by about 20% with almost no change in SSIM, while teacher‑forced probes and targeted interventions show that object position is represented throughout Wan's diffusion transformer and used by its denoising computation.

Authors:Danial Yazdani, Mohammad Nabi Omidvar, Yuan Sun, Maksud Ibrahimov, Xiaodong Li
Title: ATLAS: Scaffold-Free Algorithm Synthesis by LLMs via Embedding-Guided Quality-Diversity Search
Abstract:
Most LLM‑based automated algorithm design methods optimize a designated component within a human‑specified scaffold, fixing overall organization and component interactions. We present ATLAS, an embedding‑guided quality‑diversity framework for scaffold‑free full‑algorithm synthesis in combinatorial optimization. The problem specification supplies objectives and constraints; a minimal I/O interface fixes only instance and solution formats; the LLM chooses and restructures components, interactions, and control flow. This freedom enlarges the search space, risking invalid candidates and premature convergence to one design region. ATLAS independently detects execution, interface, and feasibility failures, recomputes objectives, and applies error‑conditioned repair; similarity‑based archive management preserves algorithms across embedding‑space regions to counter premature convergence. Its three‑layer search refines the best design, gives other regions dedicated refinement opportunities, and performs cross‑region synthesis to recombine components and their interactions. Across four NP‑hard problems, ATLAS outperforms several state‑of‑the‑art component‑synthesis methods and a matched full‑synthesis baseline while remaining competitive with strong human‑designed algorithms. One ATLAS run retains several algorithms with comparable performance from distinct embedding‑space regions rather than a single design. Code inspection finds that these multi‑component designs differ in their primary construction or global‑search backbone. Our results suggest that embedding‑guided quality‑diversity search can make the enlarged full‑algorithm design space practically searchable. Source code and exact executable prompts are available at https://github.com/Danial‑Yazdani/ATLAS .

Authors:Sahil Shah, S P Sharan, Harsh Goel, Manvik Pasula, Adithya Hebbalae, Minkyu Choi, Sandeep P. Chinchali
Title: CrossView: Can Vision-Language Models Reason Across Cameras?
Abstract:
Video understanding benchmarks have long centered on single‑camera settings, where modern multi‑modal language models achieve strong performance across image and video tasks. Yet, the real world runs on multi‑camera networks: autonomous vehicles, security systems, and robots all gather data across many simultaneous views. We argue that this is not simply "more" of the single‑camera problem; it is fundamentally different. Multi‑camera reasoning requires handling context that scales with the number of views, resolving occlusions visible from only a subset of cameras, judging which views matter, and integrating evidence across perspectives that may overlap or diverge. Current models struggle with exactly these challenges, yet no benchmark systematically targets them. We introduce CrossView, a multi‑camera video question‑answering benchmark spanning autonomous driving, security surveillance, egocentric/exocentric video, and robotics. Evaluation of proprietary models, such as GPT‑5.2, and open‑source models, like Qwen3‑VL, reveals consistently low accuracy, with open‑source models trailing by a wide margin. Performance scales strongly with a model's ability to jointly process multiple viewpoints, positioning CrossView as a rigorous benchmark for multi‑camera video. We open‑source our code and dataset at https://utaustin‑swarmlab.github.io/CrossView.

Authors:Shuo Lu, Weicheng Meng, Aijing Yu, Kun Shao, Jian Luan, Ran He, Jian Liang
Title: Topological collapse of higher-order interactions bottlenecks collective intelligence in AI agent societies
Abstract:
Current paradigms in artificial intelligence concentrate on scaling the capabilities of individual models, yet the collective behaviour of interacting agents is shaped by the topology of their interactions rather than by individual cognition alone. Here we show that the binding constraint on collective behaviour in agent societies is topological. Analysing a macroscopic AI social platform of 1.6 million registered agents (174,458 active in the interaction record), we identify a phenomenon we term topological collapse: extreme hub dominance degrades higher‑order group interactions into star‑shaped broadcast patterns, suppressing the cohesive structure that discontinuous social contagion requires. We formalise this constraint through a Hyperedge Irreducibility Score (HIS) and an analytical topology amplification factor (Φ). Across 22 frontier language models from ten vendors, 1,040 controlled simulations and empirical human networks, the bottleneck proves model‑agnostic: under a fixed interaction protocol the topological indicators are invariant across models (cross‑model HIS s.d. = 0.000 in the pairwise condition) even as behavioural outcomes diverge widely. These findings reframe the design of artificial societies around the geometry of interaction rather than the optimisation of individual cognition, with implications for AI sociology, algorithmic group dynamics, hybrid human‑AI ecosystems and collective alignment. The code is publicly available at https://github.com/Darwin‑Agent/topological‑collapse‑agent‑societies.

Authors:Pengyu Wang, Baochen Xiong, Xiaoshan Yang, Yifan Xu, Zhang Qimeng, Haifeng Chen, Changsheng Xu
Title: UniFed-VLM: Federated Instruction Tuning for Vision-Language Models with Multiple Heterogeneity
Abstract:
Vision‑Language Models (VLMs) have demonstrated strong performance in multimodal understanding and generation. However, fine‑tuning of VLMs typically relies on centralized data, which raises privacy concerns in certain domains (e.g. healthcare). Federated Learning (FL) provides a natural solution by enabling model training without sharing raw data. However, applying FL to VLM instruction tuning is highly challenging. VLMs have substantial parameter scales, and in real‑world scenarios, clients exhibit significant heterogeneity in tasks, modalities, and model architectures. Existing methods mainly focus on simplified settings and are unable to handle such multi‑dimensional heterogeneous scenarios. In this work, we study federated instruction tuning under joint heterogeneity in tasks, modalities, and model architectures. We propose UniFed‑VLM, a unified federated instruction tuning framework for VLMs that addresses multiple types of heterogeneity. It consists of two key components: 1) Federated Compensated Subspace Aggregation (FedCSA), which performs subspace‑aligned aggregation of parameter‑efficient adapters with dynamic weighting and compensation to mitigate heterogeneity‑induced conflicts; 2) Two‑stage Collaborative Distillation (TCoD), which enables effective knowledge transfer across heterogeneous models via a Mutual Distillation Adapter (MDA) and a mixture‑of‑experts‑based distillation strategy. We conduct experiments on multiple benchmark datasets, and the results show that UniFed‑VLM achieves stronger average performance across diverse tasks compared with existing FL methods. The source code is available at: https://github.com/wangpengyu2004/UniFed‑VLM.

Authors:Hao Zhang, Zhangli Zhou, Zhen Kan
Title: Temporal Logic Guided Universal Task Representations for Reinforcement Learning
Abstract:
Task guided agents demonstrate strong performance in a wide range of complex tasks. However, most existing task representation algorithms are tailored to specific contexts and struggle to generalize across diverse scenarios. Moreover, they typically depend on gradient signals from reinforcement learning controllers to update their weights, which can degrade both representation quality and learning efficiency. To overcome these limitations, we propose LOTUS, a temporal logic inspired universal task representation framework that can be seamlessly integrated into any RL algorithm to enhance agent performance across diverse task settings. Specifically, we design a novel task representation architecture capable of modeling relationships and extracting task semantics from LTL formulas. We further introduce a more effective update mechanism that treats the LTL encoder as a policy, thereby improving representation capacity. To enhance stability and robustness, LOTUS leverages the bisimulation metric, which provides theoretical guarantees for LTL representation, including behavioral equivalence, optimality fidelity, and trajectory robustness. Experimental results show that LOTUS outperforms most existing methods in learning efficiency, generalization capability, and representation quality. Specifically, LOTUS accelerates convergence over 20% in single‑task scenarios, achieves a 15%‑45% higher success rate in unseen manipulation tasks, and improves generalization performance over 25% in complex multi‑task environments with increased sub‑goal depth or conjunctions. The corresponding code, videos, and appendix are available at: https://lotus‑website.github.io/.

Authors:Takahiro Hattori, Kento Kawaharazuka, Kei Okada
Title: Detachable Wire Drive : Reconfigurable Robot Architecture with Shared Actuators
Abstract:
Reconfigurable robots offer significant potential for adapting to diverse tasks; however, conventional centralized architectures often require dedicated actuators for each module, leading to substantial increases in overall system weight, volume, and cost. To address these challenges, this paper presents the "Detachable Wire Drive," a reconfigurable robotic system that enables the sharing of heavy and expensive actuators across various morphologies. The core of this system is the "Wire Detach Unit," a mechanism designed to physically split and reconnect wire drive paths, allowing motors to be consolidated into a common base unit. We demonstrate the versatility of this approach by developing a 2‑DOF rigid arm, a continuum arm, and two distinct grippers, all of which are interchangeably attached to, and driven by, a single shared actuator set. Experimental results validate the mechanical reliability of the detachment process and the control framework's ability to seamlessly manage transitions between configurations, highlighting a path toward more efficient and multi‑functional robotic systems.

Authors:Aditya Singh
Title: Not All Attention Is Equal: A Quantitative Survey of the EEI Trade-off
Abstract:
Attention mechanisms have driven machine learning for a decade, from neural machine translation to language models that do general‑purpose reasoning. This survey covers four connected threads: their formulation for sequence‑to‑sequence tasks, adaptation to computer vision, efficiency innovations that address the quadratic bottleneck, and advances in interpretability. We define three criteria: efficiency, expressiveness, and interpretability, and compare twenty‑one methods using an EEI scoring framework. Scores come from a single rater with an assumed +/‑1‑point perturbation range. A deterministic Monte Carlo analysis with 200,000 samples shows that, under this perturbation model, rank changes of more than one position occur in 67‑70% of samples on average. A rank‑matched null model reproduces a similar stability profile, so the results support coarse tier‑level comparisons rather than fine‑grained rankings. The survey traces attention from Bahdanau‑Luong alignment through the Transformer and into vision architectures. It reviews fixed and learned sparse attention, linear attention, IO‑aware exact algorithms including FlashAttention, and state‑space alternatives including Mamba. It also covers induction heads, superposition, and the attention‑SSM duality. We further provide a structured narrative review, a benchmark synthesis with cross‑study caveats, a five‑problem research gap analysis, and a 2015‑2026 evolution timeline. We conclude by framing attention research as an expansion of the efficiency‑expressiveness‑interpretability frontier and identifying future directions including unified efficiency benchmarks, learned routing for hybrid architectures, length generalization, and scalable mechanistic interpretability.

Authors:Ananya Trivedi, Sarvesh Prajapati, Mohamed Khalid M Jaffar, Zhexin Xu, David Rosen, Taskin Padir
Title: Accelerating Mixed Discrete-Continuous Motion Planning via Neural Graphs of Convex Sets
Abstract:
Motion planning problems such as collision‑free navigation and contact‑rich manipulation can be naturally formulated as optimization problems that couple discrete decisions with continuous trajectories. The Graphs of Convex Sets (GCS) framework offers a practical solution to these problems. It represents discrete decisions as nodes of a graph and encodes continuous trajectories in the edges connecting them. However, the resulting optimization subproblems can become computationally prohibitive for online replanning. In this work, we propose a learning‑based strategy to mitigate this limitation. Specifically, we replace the costly convex relaxation step required by nominal GCS with a single forward pass through a Graph Attention Network that predicts a set of highly probable candidate paths through the graph. A lightweight ranking network then orders these candidates by their estimated trajectory cost. Evaluating them in this order, we terminate our search early while still recovering a near‑optimal motion plan. We validate the resulting pipeline across diverse robotic tasks, including collision‑free motion planning for a 3D quadrotor and a 7‑DoF manipulator, and planning through contact for planar pushing. Across both convex and non‑convex cost and constraint settings, our approach yields up to two orders of magnitude speedup over nominal GCS while maintaining a 100% success rate, at the cost of some suboptimality in the recovered solutions. Code implementations and video demonstrations can be found at https://neural‑gcs.github.io/.

Authors:Xingqiao Wang, Zi Wang, Xiaowei Xu
Title: NeuRoute: Logit-Guided Neural Routing for Billion-Scale Vector Search with Sub-Hour Index Construction
Abstract:
Building approximate nearest neighbor (ANN) indexes at billion scale is often dominated by expensive global clustering or graph construction, making time‑to‑index a first‑order systems concern. We present NeuRoute, a learned hashing index that turns short binary codes into an effective routing primitive for large‑scale vector search. NeuRoute trains a lightweight neural network encoder with a selective similarity‑preserving objective to produce well‑balanced binary addresses. During construction, NeuRoute organizes vectors into buckets by their codes and performs bucket‑local clustering in the encoder's low‑dimensional space to form centroids. At query time, NeuRoute exploits the encoder logits as an uncertainty signal: it uses deviation‑to‑threshold scores to prioritize uncertain‑bit perturbations for query‑adaptive multi‑bucket probing, scores bucket‑local centroids by their distances to the query to form a compact candidate cluster set, and applies centroid‑stage gating with heap‑quality‑driven early stopping to prune low‑value clusters before exact refinement. On billion‑scale benchmarks, NeuRoute achieves strong accuracy‑throughput trade‑offs with fast index construction: on BigANN‑1B it reaches 90.3% Recall@10 at 2,414 QPS and is 1.7× faster than OPQ+IVF‑PQ (refine) at comparable accuracy, while completing end‑to‑end training+construction in under an hour on both BigANN‑1B and Deep1B‑1B. These results show that logit‑guided neural routing can make hashing competitive as a lightweight ANN indexing framework at billion scale. Source code and artifacts are available at https://github.com/XingqiaoWang/NeuRoute.

Authors:Yiming Fu, Fangjun Li, Xiujin Liu, Ruidong Ma, Hang Yu, Zhichen Lu, Kanwei He, Alessandro Di Nuovo, Angelo Cangelosi, Zhegong Shangguan
Title: NumerosityVLM: A Cognitively Inspired Benchmark for Interpreting Numerosity Representations in Vision-Language Models
Abstract:
Vision‑language models (VLMs) achieve strong performance on high‑level multimodal tasks, yet numerosity perception, a cognitive ability that emerges in human infants before language acquisition, remains poorly understood in current models, as existing counting benchmarks entangle numerosity with correlated visual factors. We introduce a cognitively inspired diagnostic benchmark, NumerosityVLM, comprising 10,800 synthetic images across six controlled conditions. The benchmark orthogonally manipulates object size, spatial arrangement, and numerosity, while progressively ablating texture, shape, and color. Evaluating seven VLMs in a zero‑shot setting, multi‑factor analysis reveals that model architecture explains the largest proportion of performance variance (partial ω^2=0.325), far exceeding visual conditions. Layer‑wise probing further shows that linearly separable numerosity signals consistently emerge at early stages of the vision encoder, while performance differences across evaluated models are primarily associated with the language model component. Code and data are publicly available at https://github.com/fuy3/NumerosityVLM‑Benchmark, and https://huggingface.co/datasets/fuy3/NumerosityVLM.

Authors:Benjamin T. Hughes, Stuart James
Title: HistReNeRF: Historic Image Relocalisation within Contemporary Neural Radiance Field Reconstructions
Abstract:
Relocalising archival photographs within a contemporary scene model is challenging because historic and modern views can differ in photographic appearance, visible objects, and spatial layout. Therefore, we present HistReNeRF, a framework that estimates the 6‑DoF pose of a historic photograph by matching adapted DINOv2 patch features to candidate rays sampled from a contemporary Neural Radiance Field (NeRF) reconstruction. The continuous representation of a NeRF provides a queryable scene interface from which candidate rays can be sampled and matched, enabling domain adaptation between historic photography and contemporary images directly in the feature representation used for localisation. We evaluate embedding‑space‑based domain adaptation against pixel‑space methods on a new cross‑temporal dataset comprising 10,545 contemporary street‑level images and 230 archival photographs from three European landmarks. Embedding‑space adaptation reduces translation and rotation errors by an average of 11% and 16%, respectively, across the three scenes. These results show that neural scene relocalisation provides a natural interface for feature‑space adaptation, reducing cross‑temporal appearance shift without modifying the query image. Code and dataset at https://github.com/ARTUROLab/HistReNeRF.

Authors:Rohit Swami, Tushar Singh, Akash Warde, Sri Muthu
Title: Chameleon: An Adaptive AI-Driven Honeypot Architecture Using Threat-Calibrated Particle Swarm Optimization and Semantic Deception Rapidly-Exploring Random Trees
Abstract:
An invariant behavioral profile is the defining vulnerability of traditional honeypot installations: a skilled adversary can confirm the presence of a deception environment within only a few diagnostic commands, limiting its intelligence value. High‑cost commercial deception products (USD 100,000‑‑150,000 per year) share a related weakness in that their response engines are not coupled to real‑time model‑driven feedback. Chameleon is an openly distributed adaptive honeypot platform introduced here to address both shortcomings. Three core components are integrated: a bidirectional long short‑term memory (BiLSTM) classifier achieving 99.61% accuracy across seven threat categories at approximately two milliseconds CPU latency; a locally deployed Qwen3.5‑0.8B language model (Qwen Team, 2026; Unsloth, 2026) delivering 90% contextual generation accuracy at 4.5 milliseconds average latency; and two domain‑specific meta‑heuristic engines. Threat‑Calibrated Particle Swarm Optimization (TC‑PSO) dynamically reshapes swarm inertia and objective amplification in proportion to the classifier's anomaly output, enabling real‑time adjustment of connection‑holding delays. Semantic Deception Rapidly‑Exploring Random Trees (S‑RRT) drives deception schema evolution via exponentially scaled pheromone updates derived from a language‑model severity assessment, while a depth‑decay multiplier enforces a finite memory footprint. Across five benchmark runs (seeds 42‑‑46), TC‑PSO outperformed standard PSO by 48.1% in mean fitness (2.60 to 3.85) with a 32.7% convergence gain, and S‑RRT exceeded standard RRT by 258.9% in best‑run fitness (450.2 to 1,615.8), achieving a 329.2% gain at critical severity and a 24.9% memory reduction (p < 0.01). Operating costs are approximately USD 17 per month, a roughly 490‑fold reduction versus commercial alternatives.

Authors:Yusuf Meric Karadag, Gulay Oklan, Seref Baris Cagliyan, Umut Ozdemir, Emre Akbas
Title: CBX-Bench: A Human-Aligned MLLM Council for Benchmarking Concept Bottleneck Model Explanations
Abstract:
Concept Bottleneck Models (CBMs) are designed to make visual classification interpretable by expressing predictions through human‑understandable concepts. Although interpretability is the central motivation for CBMs, they are still largely evaluated as predictive models by downstream classification accuracy, supplemented by isolated qualitative examples. This highlights a pressing need for quantitative measures, a challenge complicated by the infeasibility of ground‑truth concept annotation at scale and the open nature of concept lists due to a lack of consensus. To fill this gap, we develop a multimodal large language model (MLLM) council that, given an image and its CBM explanation, produces an explanation quality score. To ground and validate the council, we first conduct a human study to establish a ground‑truth reference for CBM explanation quality: for an image, annotators compare explanations from two of LF‑CBM, VLG‑CBM, and CBM‑Suite and choose the more useful one, or mark them as equally good or equally bad, yielding 2700 judgments over 900 image‑comparison items on CUB‑200, ImageNet‑100, and Places365. Against this human reference, our five‑model council, consisting of open‑weight MLLMs, recovers over 70% of strict human preference rankings, rising to 83% on items where human annotators unanimously agree. Building on this validated council, we introduce CBX‑Bench, a public benchmark and leaderboard: authors of new CBMs can submit their model's explanations, and CBX‑Bench scores them with the council and maintains dataset‑level rankings of explanation quality. CBX‑Bench thus provides a human‑aligned, scalable evaluation of CBM explanations beyond accuracy and isolated qualitative examples. The benchmark is available at https://github.com/meric‑karadag/cbx‑bench.

Authors:Mathis Koroglu, Guillaume Jeanneret, Hugo Caselles-Dupré, Matthieu Cord, Arnaud Dapogny
Title: JoLT: Joint Latent Trajectories for Context-Guided High-Resolution Tiled Generation
Abstract:
Although text‑to‑image generative models produce impressive results, they struggle to generate densely detailed, high‑resolution (HR) images. Current literature addresses this issue with a low‑to‑high‑resolution approach. First, a low‑resolution (LR) image is generated. Then, an upsampled version is generated using the LR image as an additional cue. In this paper, we present Joint Latent Trajectories (JoLT). To generate an image, JoLT uses two streams that jointly denoise LR and HR latent images at each sampling step. The LR latent controls the overall layout, while the HR latent controls the details. We interconnect both branches to jointly integrate their information. We extensively validate our method, demonstrating its advantages over competing baselines. The resulting images are not only richly detailed but also visually pleasing, opening new avenues for artistic creation.

Authors:Madhusudhanan G
Title: Visible Reasoning and Indirect Prompt-Injection Monitorability Across English, Tamil, and Tanglish
Abstract:
Chain‑of‑thought monitoring is a potentially useful safety signal, but its reliability across languages and behavioral settings remains uncertain. In a small case study of eight manually verified synthetic scenarios, one model, one annotator, and one deterministic generation seed, I study API‑visible reasoning during indirect prompt injection in Sarvam‑105B across English, Tamil, and Tanglish. A four scenario pilot found 5/12 injected attack successes without reasoning and 1/11 with reasoning. A preregistered four‑scenario follow‑up reversed that direction, finding 2/12 attacks without reasoning and 3/12 with reasoning. With only four scenarios per phase, this design cannot distinguish a real reasoning‑mode effect from prompt‑specific variation or sampling noise. Across 20 non‑empty injected‑thinking traces, all 17 benign‑correct outputs stated an intent to ignore the injection, while all three attack successes stated an intent to follow it. These descriptive observations provide a reproducible case study of behaviorally informative visible reasoning when it is available; they do not establish that reasoning mode improves safety, that visible reasoning is mechanistically faithful, or that the findings generalize beyond this configuration.

Authors:Juseok Jeon, Ramy E. Ali, Doyun Kwon, Myungbeom Her, Jinhwi Kim, Jinhyun So
Title: FedPA-LoRA: Product-Aligned Framework for Mitigating Aggregation and Initialization Errors in Heterogeneous Federated LoRA
Abstract:
Low‑Rank Adaptation (LoRA) enables efficient federated fine‑tuning of large language models, but its factorized parameterization creates a tension between accurate aggregation of local updates and continuity of locally optimized factors. Factor‑wise aggregation incurs aggregation mismatch but better preserves factor continuity, whereas product‑space reconstruction reduces this mismatch at the cost of greater factor‑level initialization mismatch from newly reconstructed factors. We propose FedPA‑LoRA, a product‑aligned federated LoRA framework that jointly addresses these limitations and provably converges under both homogeneous and heterogeneous client ranks. Each client preserves its local factors across communication rounds and aligns its product toward a rank‑specific global reference, maintaining local optimization continuity while promoting global consistency under data heterogeneity. The server aggregates heterogeneous‑rank updates in the common product space and efficiently reconstructs a rank‑constrained global adapter without forming the dense aggregate. This design supports client‑specific computation and communication budgets. Experiments on natural language understanding and generation tasks show that FedPA‑LoRA consistently outperforms representative baselines across varying levels of data heterogeneity and homogeneous‑ and heterogeneous‑rank settings, with up to a 6.82 percentage‑point improvement in average GLUE accuracy under heterogeneous client ranks.

Authors:Sahil Gangurde
Title: AudioTQ: A Data-Oblivious 6-Bit CPU Audio Codec via Randomized Hadamard Rotation and Lloyd-Max Quantization
Abstract:
Lossy audio compression algorithms traditionally rely on psychoacoustic modeling and frequency‑domain representations (e.g., MP3, AAC, and Opus) to discard information that is imperceptible to the human auditory system. While highly effective, these approaches are computationally complex and domain‑specific. In this paper, we present the design and mathematical formulation of AudioTQ, a data‑oblivious lossy audio codec that operates directly in the time domain. Inspired by Large Language Model (LLM) weight quantization techniques (specifically the TurboQuant framework), AudioTQ uniformizes volatile time‑domain amplitudes into a predictable standard normal distribution using an orthonormal, randomized Fast Walsh‑Hadamard Transform (FWHT) rotation. This enables coordinate‑wise scalar quantization using an offline‑trained, MSE‑optimal 6‑bit Lloyd‑Max quantizer, augmented by a 1‑bit Quantized Joint Least‑Squares (QJL) residual correction layer. The resulting 7‑bit virtual indices are packed into native 8‑bit containers, aligning with standard CPU register boundaries to ensure real‑time single‑threaded execution without hardware parallel accelerators. We detail the bitwise reconstruction of 24‑bit studio stems, analyze the butterfly network of the FWHT, derive the mathematical failure modes under sparse inputs, and present benchmarks showing up to 74.4% physical size reduction alongside a Signal‑to‑Quantization‑Noise Ratio (SQNR) of ~30 dB.

Authors:Duong M. Nguyen, Tuan Nghia Nguyen, Xuan Truong Nguyen
Title: ENAF: A Multi-Exit Network with an Adaptive Patch Fusion for Large Image Super Resolution
Abstract:
To accelerate single image super‑resolution (SISR) networks on large images (2K‑8K), many recent approaches decompose an image into small patches and dynamically determine an execution path according to its difficulty (referred to as a dynamic network). To quantify the hardness of a patch, they mainly rely on a handcrafted assessment score, e.g., edge, which weakly associates a patch's texture with the computational complexity of a SISR model. To address the problem, we introduce ENAF ‑ a dynamic network for SISR with an adaptive patch fusion. Built on top of a backbone, ENAF incorporates multiple early exits (EEs) to tackle the over‑parameterized SISR model. More importantly, ENAF plugs a tiny network that estimates PSNR to associate data texture with a computation cost at an EE. Based on the scores, ENAF effectively assigns image patches to an exit, enhancing the quality‑complexity trade‑off. Extensive experiments on common datasets with popular SISR backbones demonstrate the effectiveness of ENAF in various settings. The source code is provided in https://github.com/nmduonggg/ENAF

Authors:Yuezhe Yang, Li Cheng
Title: Feed-Forward Hierarchical Gaussian Diffusion for Extreme CT Reconstruction
Abstract:
Reconstructing three‑dimensional computed tomography (CT) from severely constrained projections is highly ill‑posed. Sparse angular sampling, restricted angular coverage, and low photon counts can occur individually or jointly, obscuring global anatomy and local tissue detail. Many learned CT reconstruction methods are tailored to a single dominant degradation. Existing diffusion and Gaussian approaches commonly recover global structure and local detail within a shared representation. We propose HiGDiff, a feed‑forward hierarchical Gaussian diffusion framework that decomposes reconstruction both spatially and from structure to detail. Physics‑conditioned anatomical anchors and a foreground capacity field allocate learnable Gaussian primitives to informative regions. A structure diffusion stage first recovers global attenuation geometry, and its learned representation conditions a detail diffusion stage for residual boundaries and tissue transitions. The resulting Gaussian banks are rendered as attenuation fields and further refined by a gradient‑isolated residual module. Experiments on three distinct CT benchmark datasets demonstrate state‑of‑the‑art reconstruction performance across isolated, paired, and joint degradation settings, including improvements of 5.81 dB in macro‑average peak signal‑to‑noise ratio (PSNR) and 0.113 in structural similarity index measure (SSIM) on the Low Dose CT Image and Projection Data (LDCT‑PD) collection. Code and experimental configurations are openly available at https://github.com/Bean‑Young/HiGDiff.

Authors:Alireza Dehghanpour Farashah, Zhuan Shi, Negar Rostamzadeh, Golnoosh Farnadi
Title: TEA: Text Encoder Alignment for Robust Concept Erasure in Text-to-Image Models
Abstract:
Text‑to‑image diffusion models can be misused to generate harmful content through adversarial or paraphrased prompts that bypass built‑in safety mechanisms. Existing concept erasure methods often suffer from limited robustness against adversarial prompts, degradation of benign generation quality, or reliance on inference‑time interventions that introduce persistent computational overhead. To address these limitations, we formulate concept erasure as a domain alignment problem in the text representation space. We propose a lightweight Text Encoder Alignment framework (TEA) that fine‑tunes only the text encoder while keeping the generative backbone fully frozen. Given concept‑‑anchor prompt pairs, our method trains a discriminator to distinguish token‑level representations of concept‑containing prompts from those of safe anchor prompts, while updating the text encoder to make these representations indistinguishable. TEA introduces zero inference‑time overhead and requires only a small number of fine‑tuning steps, making it highly efficient to deploy at scale. Despite this efficiency, TEA achieves state‑of‑the‑art erasure robustness against black‑box and white‑box adversarial attacks on Stable Diffusion v1.4, while preserving generation quality on benign prompts. Furthermore, TEA is model‑agnostic and achieves the lowest attack success rate on Stable Diffusion v3.5, extending concept erasure to a Rectified Flow Transformer architecture with T5 conditioning where prior methods remain largely unexplored. Code is available at \hrefhttps://github.com/alirezafarashah/TEA.githttps://github.com/alirezafarashah/TEA.git

Authors:Wen Li, Shangshu Yu, Dunqiang Liu, Qiming Xia, Sheng Ao, Siqi Shen, Chenglu Wen, Cheng Wang
Title: LightLoc++: Sensor-Robust Representation Learning for Efficient Outdoor LiDAR Localization
Abstract:
Scene coordinate regression (SCR) achieves strong performance in outdoor LiDAR localization, but it usually requires scene‑specific training that can take days, limiting practical deployment. Recent works improve training efficiency by decoupling SCR into a scene‑agnostic backbone and scene‑specific prediction heads, where the backbone is pretrained on source datasets and frozen for new scenes, and only lightweight heads are optimized. However, we find that this paradigm heavily depends on the pretrained backbone. Existing decoupled methods can match conventional SCR methods fully optimized for each new scene when LiDAR configurations are similar to those used during backbone pretraining, but their accuracy drops noticeably on datasets collected with different LiDAR sensors. This suggests that efficient LiDAR localization requires representations that capture stable scene geometry across LiDAR configurations. Motivated by this observation, we propose LightLoc++, a sensor‑robust and efficient outdoor LiDAR localization framework. To support sensor‑robust representation learning, we introduce SULID, a synchronized urban multi‑LiDAR dataset with representative 32‑, 64‑, and 128‑beam rotating LiDARs, extensive cross‑sensor overlap, and diverse urban scenes. Using SULID, we pretrain a sensor‑robust backbone through cross‑sensor consistency learning. LightLoc++ further preserves efficient new‑scene learning by incorporating sample classification guidance and redundant sample downsampling, which reduce regression ambiguity and computational redundancy in large‑scale outdoor scenes. Extensive experiments on multiple outdoor LiDAR localization benchmarks demonstrate that LightLoc++ achieves state‑of‑the‑art localization performance with the lowest new‑scene training cost among compared methods. Code and dataset will be made available at https://github.com/liw95/LightLoc‑PlusPlus.

Authors:Jiaqi Hu, Junwen Huang, Hongli Xu, Peter KT Yu, Nassir Navab, Benjamin Busam, Slobodan Ilic
Title: SOS! : A Streamlined Object-Conditional Transformer for Model-free Segmentation
Abstract:
Foundation segmentation models excel at generating high‑quality, class‑agnostic masks, but they struggle to associate these proposals with specific target objects. This semantic gap severely hinders their deployment in downstream applications like robotic manipulation, which demand precise unseen objects segmentation. Existing approaches attempt to resolve this by relying on exhaustive 3D object model priors, inherently introducing prohibitive computational overhead and complex, multi‑stage pipelines. To address these limitations, we propose SOS (Streamlined Object‑conditional Transformer for model‑free Segmentation). SOS completely eliminates the reliance on 3D models, requiring only a single reference image per target object. Central to our framework is a novel Object‑Conditional Transformer that learns identity‑anchored queries, unifying mask generation and target identification into a single feed‑forward pass. This streamlined design drastically improves both structural and computational efficiency. Extensive evaluations across multiple benchmarks demonstrate that SOS establishes a new state‑of‑the‑art for model‑free unseen objects segmentation, delivering accurate and high‑efficiency performance. The project page and code are available at https://sos‑seg.github.io/.

Authors:Shiven Khurdi
Title: No Task Fails Every Time: Why One-Shot Audits Are Structurally Blind to Agent Damage
Abstract:
We introduce AgentRelBench, an environment‑agnostic reliability instrument that computes ground‑truth, severity‑priced damage from database state diffs across repeated runs, with no LLM in the measurement path, demonstrated on EnterpriseOps‑Gym. Across 2,128 evaluation runs spanning nine models in six families (four development, three pre‑registered held‑out, plus a frontier pass on two frontier‑tier models that the pre‑registration designates exploratory), we find: (1) damage on irreversible actions is universal across the families we measured and stochastic within them on pinned, single‑provider stacks. (2) No task damaged on every run: zero always‑fail cells across 42 confirmatory held‑out damage events. A single clean run misses a damage‑producing (model, task) pair 0.80 of the time on the development pool (13 pairs); the held‑out pool is descriptively consistent (0.575 over 5 pairs, pair‑weighted) but sits below our pre‑registered power floor and is reported as underpowered, not as confirmation. (3) Damage‑producing task count falls with model capability, from 7 of 20 tasks for an 8B model to 1 of 20 for the most capable; capability is confounded with family and training, so this is an observed gradient, not a causal claim. The residual damage does not change in character: in the exploratory frontier pass, the most capable model's one damaging task damages at \hatp = 0.16 per run, inside the same demonstrably‑stochastic band, and a single audit misses it 84% of the time. (4) One model family committed the gated irreversible change while declaring it had refused: transcript‑ and judge‑based grading scores those runs as safe refusals, only state diffs as damage. All confirmatory findings were pre‑registered with per‑claim demote criteria; one demoted our own initially favored finding, which we report.

Authors:Yoon Gyo Jung, Jaewoo Park, Kuan-Chuan Peng, Seongdeok Bang, Octavia Camps
Title: Memory-Bounded Continuation of Greedy Sampling for Continual Anomaly Detection
Abstract:
Greedy sampling produces a compact yet representative summary of normal data, which is essential for reliable anomaly detection that relies on measuring distance from normality. For continual anomaly detection where tasks arrive sequentially, extending greedy sampling is straightforward with unbounded memory through coreset accumulation. However, practical deployment requires fixed memory where the coreset size remains constant regardless of task count. We observe that continued greedy sampling, which iteratively applies greedy selection over previously greedy‑sampled sets, effectively preserves representativeness under strict memory limits. Despite discarding data at each step to satisfy the memory constraint, coreset quality degrades gracefully rather than catastrophically, enabling reliable anomaly detection across the tasks. We provide theoretical justification by showing that resulting greedy‑continued coreset approximates the oracle coreset within a bounded gap. We instantiate this principle in ContCore, which constructs a greedy‑continued coreset through greedy expansion on new task features followed by greedy consolidation to enforce the memory budget. Unlike neural methods susceptible to catastrophic forgetting or naive coreset accumulation requiring unbounded memory, ContCore maintains fixed memory with theoretical guarantees. Empirically, ContCore achieves state‑of‑the‑art performance across 11 task schedules on MVTecAD and VisA, and extends effectively to online continual AD settings where prior methods degrade significantly. Code: https://github.com/jungyg/ContCore

Authors:Weikang Yu, Yonghao Xu, Pedram Ghamisi
Title: On the Adversarial Robustness of Remote Sensing Semantic Change Detection
Abstract:
Semantic change detection (SCD) is a bitemporal dense‑prediction task that jointly identifies changed regions and their semantic states before and after change. Unlike single‑image segmentation or binary change detection, SCD couples two temporal inputs with timestamp‑wise semantic prediction, change localization, and final semantic‑change decoding, creating adversarial dependencies that are not captured by conventional robustness protocols. We present a task‑specific evaluation framework that separates output‑side attack objectives from input‑side temporal perturbation access, enabling systematic analysis of component vulnerability and cross‑temporal propagation. Experiments on four datasets and six representative CNN‑, Transformer‑, and state‑space‑based models evaluate component‑level and temporal objectives, single‑ and dual‑timestamp perturbations, multiple attack methods, and cross‑architecture transferability. The results show that final semantic‑change predictions can be severely corrupted even when binary change localization remains comparatively stable, and that perturbations or attack objectives associated with one timestamp can propagate to the prediction of the other. These behaviors occur across different architecture families, while direct cross‑model transfer remains considerably weaker than white‑box attacks. The study demonstrates that adversarial robustness in SCD depends on the complete bitemporal prediction pathway rather than on an individual branch or backbone family, and provides a structured protocol for evaluating robustness in coupled bitemporal image analysis. Code is available at https://github.com/EricYu97/AdvSCD.

Authors:Wei Zhang, Yihang Wu, Songhua Li, Qi Wang
Title: VGGT-Align: Bridging Local Reconstruction and Global Consistency for Long-Sequence 3D Reconstruction
Abstract:
Maintaining global geometric consistency is a central challenge in long‑sequence 3D reconstruction, with scale drift being the most critical failure mode. In chunk‑based inference pipelines, the scale degree of freedom in sequential Sim(3) alignment is left unconstrained, causing estimation errors to compound multiplicatively and distort global trajectories and point cloud geometry. We present a scale‑consistency enhancement framework built on a key insight: in structured environments such as driving scenes, geometric quantities arising from environmental regularity remain inherently invariant across temporal segments, and discrepancies in their per‑chunk measurements directly expose inter‑chunk scale drift. We propose Scene Geometric Invariant Anchoring (SGIA), which extracts dominant geometric invariants from each chunk's predicted point cloud via coarse‑to‑fine robust estimation and exploits their cross‑chunk consistency to establish scale constraints independent of point cloud registration, explicitly degenerating 7‑DoF Sim(3) alignment into 6‑DoF rigid‑body transformation and severing chain‑wise scale error propagation at its source. We further introduce a lightweight test‑time adaptation strategy that fine‑tunes only normalization‑layer parameters via multi‑objective self‑supervision, progressively improving intra‑chunk predictions along the sequence. Both modules are plug‑and‑play and require no offline retraining. Experiments on multiple long‑sequence benchmarks demonstrate state‑of‑the‑art performance, reducing absolute trajectory error by up to 32% with significant gains in trajectory stability and reconstruction quality. Code: https://github.com/WZ‑CS/VGGT‑Align

Authors:Chan Lee, Kimin Yun, Yuseok Bae, Seong Tae Kim, Jung Uk Kim
Title: PersonaDrive: Controllable Trajectory Prediction with Multi-Dimensional Driving Personas
Abstract:
Although recent trajectory prediction and end‑to‑end autonomous driving methods improve robustness in urban environments, they still lack meaningful controllability. Existing benchmarks either provide no persona‑conditioned annotations or support only a single urgency spectrum (i.e., emergency, normal, relaxed), which cannot distinguish personas that share the same urgency level but require different driving dynamics. To address this, we propose (i) the Persona‑Conditioned Trajectory (PCT) dataset, which decomposes driving personas along two axes, Temporal Urgency and Ride Comfort, and combines three levels of each to form a grid of nine personas, each paired with natural‑language descriptions and trajectories, and (ii) PersonaDrive, a framework that can learn driving personas from language and can generate persona‑specific trajectories. PersonaDrive incorporates Persona‑Conditioned Anchor Transform (PCAT), which hierarchically reshapes anchors along both axes, and Persona‑Conditioned Multi‑Modal Fusion (PCMF) for BEV‑level persona fusion. Training is supervised by a Hierarchical Guide Loss enforcing axis‑aligned physical orderings and an Axis‑Decomposed Diversity Loss preventing diagonal mode collapse. Experimental results show that PersonaDrive consistently improves over the compared baselines across multi‑dimensional scenarios. The code and PCT dataset are available at https://github.com/VisualAIKHU/PersonaDrive

Authors:Ruohan Wu, Ziqi Zhu, Yang Zhao, Jiarui Tang, Yingzhe Cui, Junshi Chen, Zhao Jing, Jun Shi, Hong An
Title: TERRA: A Hierarchical Parallel Training and Memory Orchestration Framework for High-Resolution AI-based Earth Modeling
Abstract:
Training high‑resolution AI‑based Earth forecasting models is memory‑intensive. Window‑based Swin Transformers reduce the quadratic cost of global attention, but existing distributed systems such as AERIS primarily target pixel‑level models and do not jointly support convolutional sampling modules and shifted‑window execution. Long‑lead rollout finetuning further increases activation memory. To address these challenges, we present TERRA, a hierarchical parallel training framework for high‑resolution Earth forecasting. TERRA introduces Sampling‑Aware Window, Sequence, and Tensor Parallelism (SAWSTP), which preserves spatially contiguous layouts for sampling modules and routes tokens into topology‑aware ragged window layouts for Transformer execution. For long‑lead finetuning, Memory Orchestration (MO) provides rollout‑aware checkpoint planning and combines input buffering with budget‑constrained activation offloading. Experiments on the 1/12^\circ GLORYS‑based Wenhai workload show that TERRA supports models with up to 11.4B parameters on 96 H200 GPUs and sustains up to 39.76 PFLOPS, achieving 65.0% strong‑scaling and 94.1% weak‑scaling efficiency. Compared with checkpoint‑only policies, MO further reduces peak allocated GPU memory by 32.2%‑‑51.8% with at most 20.0% step‑time overhead, which makes finetuning with smaller patch sizes and longer rollouts feasible for improved forecasting accuracy.

Authors:Timo Sämann
Title: P-PAS: Prefill-Pressure Adaptive Scheduling for Long-Context LLM Serving
Abstract:
Long‑context LLM applications such as retrieval‑augmented generation (RAG) and agentic systems often process tens of thousands of input tokens to produce short outputs, making end‑to‑end request latency an important serving objective. We show that the maximum number of batched tokens (MBT), which controls the token scheduling budget in vLLM, has a scheduling‑pressure‑dependent effect on latency. Larger token budgets can reduce latency under low scheduling pressure, while smaller budgets become preferable under higher pressure. Consequently, no single static MBT performs best across load regimes. We introduce Prefill‑Pressure Adaptive Scheduling (P‑PAS), a lightweight policy that dynamically adapts the scheduling budget based on concurrent prefill and decode state. P‑PAS retains a large token budget under low pressure and constrains prefill work as pressure increases. Across models, workloads, and GPUs, P‑PAS maintains low end‑to‑end latency across changing load regimes, avoiding the limitations of a fixed MBT. Kernel‑level profiling shows that large prefill chunks can improve execution efficiency under low scheduling pressure, but that this advantage varies across model‑‑hardware configurations. As scheduling pressure increases, smaller chunks can instead reduce interference with active decoding, explaining the observed load‑dependent MBT sensitivity. Code and artifacts for reproducing our results are available at https://github.com/TimoSaemann/ppas‑vllm .

Authors:Yihong Ji, Jinsong Zhang, He Hu, Hongbo Xu
Title: HOIMask: Towards Generative Masked Modeling for Human Object Interaction Generation
Abstract:
Diffusion‑based methods have dominated the HOI generation, as they enable critical contact fusions or signals to guide the diffusion process. However, they often result in high artifacts and unstable interaction quality due to error accumulation during iterative denoising. In this work, we propose HOIMask, the first generative masked framework for modeling HOI motion in discrete space. HOIMask first encodes both motion sequences and contact‑aware signals into discrete 2D human and object token maps via HOI Vector Quantization (VQ), preserving fine‑grained spatial‑temporal structure beyond conventional 1D representations. On this basis, a generative masked modeling framework is employed to jointly capture human‑object interaction dynamics, leveraging a transformer architecture designed to model complex spatial‑temporal and interaction dependencies. To generate more coherent and physically plausible motions, we further introduce a novel contact‑aware reconstruction guidance in discrete space during inference, which fuses contact signals to optimize HOI tokens that forces the generated motion with higher spatio‑temporal consistency. With craftily designed motion interaction tokens, dedicated architecture and guidance strategy, HOIMask outperforms state‑of‑the‑art diffusion‑based methods, generating more realistic and semantically aligned HOI motions. Please refer to https://jyhflash.github.io/HOIMask/ for more results.

Authors:Jiarui Yang, Bin Zhu, Jingjing Chen, Na Zou, Yanwei Fu, Jianggang Zhu, Yu-Gang Jiang
Title: StructRL: Structured Action-Space Exploration for Flow-Based VLAs
Abstract:
Flow‑based Vision‑Language‑Action (VLA) models are now widely used for continuous robotic manipulation, and online reinforcement learning (RL) is emerging as a key technique for adapting them to new tasks. Existing RL methods typically inject stochasticity inside the denoising chain, often through isotropic or temporally independent noise. However, effective robot exploration calls for structured noise: temporally smooth and scaled differently across action groups. We show that simply switching the in‑chain noise to a structured form does not suffice: noise added at an intermediate flow time can be weakened by the remaining denoising steps before execution, a phenomenon we call \emphStructured Noise Dilution. We propose StructRL, which avoids dilution by relocating policy stochasticity to the action space via three coupled choices: (i) a deterministic ODE decoder, (ii) structured noise injected directly in the action space, and (iii) last‑step replay, where policy‑gradient updates avoid assigning likelihoods to intermediate denoising states. This keeps structured exploration tied to the executed action while providing a tractable training signal for the flow decoder. Across three flow‑based VLA models on multiple simulated manipulation benchmarks and two real‑world tasks, StructRL improves exploration efficiency and OOD performance over prior in‑chain baselines, demonstrating the effectiveness of structured action‑space exploration for adapting flow‑based VLA with RL. Project page: https://flyfaerss.github.io/structrl/

Authors:Jiaming Liang, Chi-Man Pun, Weisi Lin
Title: Fast Test-Time Refinement for Robust Learned Image Compression
Abstract:
Learned image compression (LIC) has demonstrated remarkable rate‑distortion (RD) performance in benign settings. However, the high representational capacity endowed by deep neural networks (DNNs) comes at the expense of increased adversarial vulnerability. This hinders their adoption as trusted standardized codecs. Recent work has sketched test‑time refinement (TTR) as a defense in gray‑box scenarios, despite its original purpose of improving benign RD performance. Unfortunately, extensive iterations of TTR incur prohibitive overhead, while the robustness mechanism lacks theoretical understanding. Moreover, TTR has not been evaluated in white‑box settings or against attacks beyond \ell_2‑bounded rate and untargeted distortion objectives. To bridge these gaps, we present a systematic study. Our study reveals an Asymmetric Adversarial Trajectory (AAT) property in LIC systems: transitioning from adversarial to benign regions is significantly easier than the reverse process, where adversarial examples can often be roughly recovered within only 1‑2 steps. We provide a two‑dimensional Tube Model to explain this phenomenon. Based on AAT, we propose a Fast Test‑Time Refinement (FTTR) framework for practical and robust LIC systems. We establish that the robustness arises from the contraction of adversarial regions induced by the Input‑as‑Label property of LIC systems, rather than from obfuscated gradients. Extensive evaluations with diverse strong adaptive attacks across multiple LIC systems demonstrate the promise of the proposed FTTR framework. The code is available at https://github.com/chinaliangjiaming/FTTR.git.

Authors:Amrit Gopinath, Raghul, Durairaj Thenmozhi
Title: A Declarative-Procedural Perspective on Expert Routing in Bilingual Mixture-of-Experts Language Models
Abstract:
We investigate whether Mixture‑of‑Experts (MoE) language models develop linguistically structured expert routing during bilingual language acquisition. Inspired by the Declarative‑Procedural framework, we analyze lexical, grammatical, and syntactic processing in a decoder‑only English‑German MoE Transformer trained under sequential language exposure. We construct a probe‑based validation set and extract token‑level routing distributions to quantify category‑dependent specialisation using mutual information, routing entropy, and Jensen‑Shannon distance. The curriculum‑trained model exhibits a peak mutual information of 0.1148 at layer 5, indicating category‑dependent differences in routing distributions across linguistic categories. Surprisingly, a no‑curriculum baseline trained on mixed English‑German data shows stronger aggregate specialisation, reaching a peak mutual information of 0.2599 at the same layer. These results suggest that interpretable linguistic organization emerges within MoE routing patterns even without sequential language exposure. A replication at a second training seed shows that the no‑curriculum condition's specialisation concentrates on a single language whose identity is seed‑dependent, whereas the curriculum consistently yields a stable, language‑balanced routing profile; rather than uniformly increasing specialisation, staged bilingual exposure reduces single‑language dominance. The official Github repository: https://github.com/Amrit828/DP‑Theory‑MOE‑Interpretability‑Research

Authors:Rosen Ting-Ying Yu, Christophe Hatterer, Advaith Narayanan, Cyril Picard, Faez Ahmed
Title: BOCoDe: Engineering-Centered Benchmarking for Bayesian Optimization
Abstract:
Bayesian optimization (BO) is a sample‑efficient, surrogate‑based approach to black‑box optimization (BBO), but its evaluation remains dominated by synthetic functions and hyperparameter optimization (HPO) tasks that are typically low‑dimensional and single‑objective. Engineering design poses a substantially different regime: problems are physics‑based, often high‑dimensional, constrained by requirements such as cost and manufacturability, and may involve multiple objectives or mixed variables. To close this benchmarking gap, we introduce BOCoDe, an open‑source, PyTorch‑native benchmark comprising 307 BBO problems, including 159 engineering design tasks and widely used synthetic and HPO benchmarks. Each problem includes cited provenance and machine‑readable metadata that supports programmatic discovery, including by LLM‑based agents, and all tasks are exposed through a unified API compatible with open‑source BO libraries. We evaluate 31 BO and evolutionary algorithms across five problem classes spanning single‑ and multi‑objective optimization, constrained and unconstrained settings, and mixed‑variable search spaces. Analyses of problem structure show that engineering tasks uniquely span constrained and multi‑objective settings that synthetic and HPO suites rarely cover, while embeddings from a tabular foundation model separate them most clearly from HPO tasks. Algorithm rankings also vary substantially across domains; in several problem classes, rankings obtained on standard benchmarks do not transfer to engineering tasks. BOCoDe establishes a reproducible and extensible foundation for developing and evaluating BO methods that better reflect the demands of engineering design. Code & data can be found at https://github.com/rosenyu304/BOCoDe

Authors:Tianxin Wei, Zhan Shi, Minhua Lin, Bing He, Zewen Liu, Yisi Sang, Yuanchen Bei, Xuying Ning, Jiaru Zou, Ting-Wei Li, Xiao Lin, Yanjun Zhao, Chi Wang, Benoit Dumoulin, Dakuo Wang, Jingrui He, Hanqing Lu
Title: Evo-Harness: Context-to-Harness Skill Compilation for Self-Evolving Agents
Abstract:
Learning from experience is critical for developing capable, self‑improving large language model (LLM) agents. Existing methods typically extract knowledge from accumulated trajectories via reflection, memory, rules, or skills. However, agents in realistic environments continuously encounter novel tasks, often offering only a one‑shot opportunity to improve. These executions yield rich but highly noisy contexts, entangling broadly useful lessons with task‑specific artifacts. Critically, prior works rarely validate their effectiveness on complex real‑world tasks or isolate the underlying drivers of improvement. To address these gaps, we formulate online harness learning, where a frozen agent improves by continually updating a structured harness across sequential tasks. This formulation enables a systematic study of key self‑improvement factors through our proposed Evo‑Harness. At its core, context‑to‑harness skill compilation distills noisy, single‑shot executions into reusable skill harnesses for cross‑domain and topic‑level adaptation. To demonstrate the efficacy of one‑shot skill compilation, we evaluate across five realistic benchmarks (TerminalBench2, SWE‑bench, CL‑Bench, bench, WebArena‑Infinity). Our extensive analysis demonstrates the effectiveness of Evo‑Harness and provides a principled understanding of how LLM agents can effectively learn on the fly. Our code is available at https://github.com/A‑EVO‑Lab/a‑evolve/tree/release/evo‑harness.

Authors:Tarun Tomar
Title: Do Visual Grounding Decoders Need Feed-Forward Networks? A Controlled Study over Frozen Vision-Language Features
Abstract:
Do feed‑forward networks (FFNs) in visual grounding decoders add essential computation once a pretrained vision‑language model has already encoded image and language context? We compare a four‑block attention‑only decoder (A4), a matched four‑block attention‑plus‑FFN decoder (S4), and an eight‑block attention‑only parameter control (A8) over frozen VLM features. A4 matches or slightly exceeds S4 on RefCOCOg and Ref‑Adv‑s. FineCops‑Ref reveals a small A4 deficit of 0.52 percentage points at IoU@0.5 (95% CI [0.12, 0.95] in favor of S4), but A8 recovers it and finishes 0.26 points above S4. Official FineCops levels do not show a monotonic increase in the gap. A4 reduces trainable decoder parameters by 44.4% and cached‑decoder latency by 10.1%, although end‑to‑end latency remains backbone‑dominated. These results concern the trainable grounding decoder, not a complete attention‑only VLM.

Authors:Feng Gao, Zizhe Pan, Haoting Wang, Ruzhuang Hua, Jingchao Cao, Junyu Dong, Qian Du
Title: Frequency and Edge-Guided Segment Anything Model for Remote Sensing Image Semantic Segmentation
Abstract:
Remote sensing image semantic segmentation (RSISS) has attracted significant attention due to the growing demand for fine‑grained land cover information. The Segment Anything Model (SAM), proposed as a foundation vision model, offers strong segmentation performance and generalization capabilities for RSISS tasks. However, existing SAM‑based approaches face two limitations: (1) Insufficient adaptation of SAM's features to the diverse characteristics of land cover types. (2) Semantic ambiguity at object boundaries, which hinders accurate delineation. To address these limitations, we propose Frequency and Edge‑guided SAM (FE‑SAM), a scalable and efficient framework for RSISS. Specifically, we introduce a Frequency‑Modulated Adapter (FMA) that adaptively decomposes and modulates frequency‑domain features based on the input data. It selectively enhances informative high‑ and low‑frequency components corresponding to different land cover types. Furthermore, to improve SAM's ability to capture fine‑grained details, we design EGRefiner, which integrates multi‑scale edge‑enhanced information extracted from the input image. Extensive experiments on three benchmark datasets demonstrate that FE‑SAM outperforms state‑of‑the‑art methods. The source codes are available at: https://github.com/oucailab/FE‑SAM.

Authors:Zi'an Wang
Title: Andy: A Mathematical Agent for Rigorous Proof and Autonomous Research
Abstract:
Andy is an autonomous mathematical research agent that turns a mathematical problem into a traceable proof. It solves or verifies a submitted problem, formulates a literature‑grounded new problem through a research‑value gate, and carries it through proof construction and final verification. It organizes proof steps in an executable DAG, verifies each step independently and binds the result to a certificate, retains verified work whose interfaces remain unchanged during local repair, and records the full path from problem formulation to final proof. The system separates proof generation from correctness evaluation and can acquire, retain, retrieve, and reuse knowledge from existing results. Starting from a self‑triggered impulsive consensus result, Andy formulates a global exponential leader‑follower synchronization problem for delayed heterogeneous networks with switching communication topologies. The proposed hybrid control combines self‑triggered impulses with execution delay and continuous feedback over a recovery window. After each delayed impulse, the feedback cancels the delayed error channel until the pre‑impulse history leaves the active delay interval. Sufficient conditions for global exponential synchronization are established, Zeno behavior is excluded for both timing sequences, and a numerical example illustrates the result.

Authors:Pengyu Wang, Chenkun Tan, Shaojun Zhou, Qirui Zhou, Yanxin Chen, Xingyang He, Huazheng Zeng, Jijun Cheng, Chenghao Wang, Xiaomeng Qian, Pengfei Wang, Zhan Huang, Shanqing Gao, Wei Huang, Longjun Cao, Wu Ran, Jie Liu, Changtai Zhu, Hongkai Wang, Yixian Tian, Chenghao Liu, Zhen Ye, Xinghao Wang, Botian Jiang, Guoguo Feng, Zhaoye Fei, Ruixiao Li, Mingshu Chen, Yang Gao, Qinyuan Cheng, Shimin Li, Xipeng Qiu
Title: MOSS-VL Technical Report
Abstract:
We present MOSS‑VL, an open vision‑language model family that treats real‑time interaction ‑‑ perceiving while it speaks ‑‑ as a first‑class capability. It is co‑designed across the stack: the language decoder attends to vision only through gated cross‑attention, so the model can naturally see incoming frames while generating; a synthesized interaction corpus supervises when to speak, when to stay silent, and when to revise; and a staged curriculum concentrates all real‑time‑specific training in one light final stage over a strong offline foundation. Offline, MOSS‑VL‑Instruct is competitive at comparable scale and leads temporal‑reasoning video sets. Across four streaming benchmarks, MOSS‑VL‑Realtime posts the best average on three (second on the fourth) among open‑source streaming models, sweeping the three subsets that squarely test proactive behavior ‑‑ 66.0 vs. 37.5 for the best baseline on OmniMMI Proactive Alerting. With 11.3B parameters but visual tokens outside the decoded sequence, MOSS‑VL widens its time‑to‑first‑token advantage over same‑backbone Qwen3‑VL‑8B from 2.8x to 5.1x as visual context grows. We release all five checkpoints, the training curriculum, and the real‑time inference code at https://github.com/OpenMOSS/MOSS‑VL.

Authors:Yuhua Jiang, Jiaming Wang, Qingbin Liu, Feifei Gao
Title: SCOPE: Score-Isolated Agentic Optimization for Video World Models
Abstract:
Video world models are increasingly used as simulators for planning and embodied decision making, yet improving them at inference time introduces a subtle evaluation problem: prompts, samplers, verifiers, and selectors may evolve together, making it difficult to attribute gains or prevent held‑out feedback from shaping the final policy. We introduce \scope (\emph\scopefullname), a framework for auditable inference‑time adaptation of frozen video world models. \scope represents external controls as a typed state, updates this state only through bounded changes supported by development evidence, and freezes the resulting policy before held‑out evaluation. On Physics‑IQ benchmark, \scope improves over the exact frozen base by +14.24 (95% CI [+8.10,+21.23]). Controlled ablations further identify gains from scene specification, sampling, and learned selection, while the margin over the strongest matched agentic baseline remains unresolved. Cross‑backbone and prospective evaluations reveal a complementary result: useful inference‑time updates exist, but their benefits do not transfer uniformly across models and settings. Together, these findings suggest that reliable inference‑time adaptation requires not only better proposals, but also a principled mechanism for deciding which updates should become part of the deployed system. Code is available at https://github.com/YuhuaJiang2002/SCOPE.

Authors:Bruno Chicelli, Henrique Alves, Rodrigo Anselmo, Joshua Weinberg, Felipe Lemos, Jan Baryla
Title: Handoff-H1: An Orchestrated Vision-Agent System for Material Quantity Takeoff from Construction Blueprints
Abstract:
Converting a set of architectural blueprints into a complete material quantity takeoff requires visual perception across drawing sheets, dimensional and multi‑hop reasoning, and grounding in construction conventions that the drawings never state. We present Handoff‑H1, a takeoff system built from three layers: purpose‑built computer‑vision models that extract primitives; tool‑using agents equipped with image operations and in‑house visual‑task tools, including CV‑model‑backed counting, detection and plan decomposition; and a persistent, hierarchically structured project foundation, grounded in a curated construction knowledge base. We evaluate on the Construction Blueprint Takeoff Benchmark: 10 real residential blueprint sets paired with consensus‑validated expert takeoffs ‑ 2,009 verified line items, restricted for scoring to the 1,348 primary‑tier materials that drive an estimate ‑ scored per trade by an LLM judge on material coverage and quantity Precision@25% (P@.25) and combined into a weighted composite. Under identical scoring from the raw PDF, seven frontier and open‑weight models span composites of 35‑61, and independent professional estimators ‑ scored against the same reconciled gold standard ‑ post 77.6% (65.5% coverage, 87.9% P@.25). Handoff‑H1, working end‑to‑end from the raw PDF, reaches 81.6% (86.1% coverage, 78.8% P@.25): roughly 20 points above the strongest frontier agent, and above the independent estimators by pairing near‑human quantity precision with coverage they do not reach. The evaluation harness is public for the open harbor framework; the blueprint sets and ground truth are available upon request for research use.

Authors:Conor Miller-Lynch, Sandip Purnapatra, Syed Konain Abbas, Lambert Igene, Faraz Hussain, Soumyabrata Dey, Stephanie Schuckers
Title: Generation of Synthetic Fingerphotos with GANs
Abstract:
Contactless fingerprinting is an emerging approach to biometric authentication that allows users to scan their fingerprints without touching a scanner. Due to the limited amount of contactless fingerprint data available and the security risks associated with sharing real individuals' fingerprints, it is valuable to explore methods of generating synthetic data that can be used in place of ‑ or in conjunction with ‑ real data to develop and evaluate contactless fingerprinting systems. In this paper, we present and evaluate synthetic fingerphotos generated using StyleGAN2‑ADA and StyleGAN3, existing image generation architectures. We evaluate the realism, privacy preservation, and variety of the synthetic fingerphotos by comparing their biometric feature statistics to those of real fingerphotos, computing match scores between real and synthetic fingerphotos, and computing match scores between different synthetic fingerphotos. This paper provides a quantitative comparison point for future evaluations of synthetic fingerphotos. The evaluation code is made available at https://github.com/cmillerlynch/fingerphoto‑gan.

Authors:Xuran Hu, Mingzhe Zhu, Djordje Stanković, Yujie Zhu, Zhenpeng Feng, Yifang Ban, Ljubiša Stanković
Title: Geometry-Calibrated Closed-Form Shrinkage for SAR Despeckling
Abstract:
Synthetic aperture radar (SAR) despeckling is an inverse‑recovery problem in which multiplicative non‑Gaussian noise must be suppressed without erasing scattering structures. We revisit a nonlocal sparse estimator that applies a log‑‑Yeo‑‑Johnson transformation, stacks similar patches into groups, codes each group on its own left singular basis, and shrinks the resulting coefficients. Three quantities usually treated as tunable are shown to be fixed by this construction. First, the group dictionary is orthonormal, so the weighted Lasso admits an exact coefficient‑wise soft‑threshold solution: the iterative inner solver is unnecessary, and the two apparent weighting matrices are the numerator and denominator of a single threshold field rather than independent modules. Second, because the dictionary is estimated from the noisy group itself, its retained subspace absorbs speckle in proportion to the group aspect ratio γ=p^2/K; a random‑matrix argument converts the corresponding regularization constant into a geometry‑calibrated correction and collapses patch size, group size, and shrinkage scale into one analytically determined degree of freedom. Third, singular projection makes the coefficient noise nearly Gaussian at every tested look number, which locates the point at which an exact speckle likelihood ceases to be informative. The resulting estimator is deterministic, training‑free, and applies one set of analytically determined settings to every image and sensor. It ranks first in 18 of 24 PSNR/SSIM comparisons against twelve published methods on three synthetic benchmarks, and attains the lowest mean deviation of the ratio image from the theoretical speckle model over six real‑SAR configurations from five sensors. Code is available \hrefhttps://github.com/Teriri1999/Geometry‑Calibrated‑Closed‑Form‑Shrinkage‑for‑SAR‑Despecklinghere.

Authors:Parsa Mazaheri
Title: Gathered, Not Admitted: How Attention Brings a Latent Variable into Verbalizable Form
Abstract:
Language models hold latent quantities in a form they can report on, and more of a quantity is present in that form when the task requires reusing it flexibly. What causes a representation to enter that form is open, and the word workspace invites an admission story: a gate that decides what gets in. Testing it on open‑weight models with Jacobian lenses, over a benchmark whose five arms share an identical context, we find no gate where it predicts one. Demand raises a concept's lens visibility beyond what applying an operator to a supplied value produces: +0.050 [+0.045, +0.057] in percentile rank on our primary checkpoint, positive on all four we measure, though that arm answers at ceiling and the accuracymatched contrast is stronger under that readout. At the same time one shared linear map decodes the variable from every arm, the control included, at 6.4‑9.0x its selection‑corrected floor. What produces the later readable form at the queried position is attention‑mediated gathering inside a mid‑depth window: separating patch depth from readout depth puts transport there at least 17x above anywhere shallower under non‑saturating readouts, with no tested MLP output contributing positively inside it. Under the saturating percentile rank the same grid does not localise the window, which is a fact about that measure. An arm that needs the variable for nothing concentrates sevenfold less, so the window is demand‑specific. That window has two measured edges, a survival failure below and destruction above, and it falls at the same fractional depth in a 64‑layer hybrid and a 62‑layer dense model from another family. We localise where the variable is installed and read, not the route from the passage, which transports nothing. But the readout is not a calibrated measure of use: three components move it to within 12% of one another and differ 7.4x in what they do to the answer.

Authors:Haochen Huang, Shengxuan Qiu, Meng Li
Title: S2-MoE: Enabling Efficient Self-Speculative Decoding for Mixture-of-Experts on Edge Devices
Abstract:
Deploying large language models (LLMs) for inference on edge devices is challenging due to severe memory and bandwidth constraints. While speculative decoding and Mixture‑of‑Experts (MoE) have been proposed to improve inference efficiency, naively combining them often incurs excessive verification overhead and poor expert reuse, limiting their effectiveness in memory‑bound edge settings. In this work, we propose S2‑MoE, an efficient self‑speculative decoding framework for MoE inference on edge devices. S2‑MoE reduces redundant verification through routing‑aware adaptive speculative expansion, improves verification efficiency with reuse‑aware expert gating, and aligns draft and target execution via shared context. Implemented in llama.cpp, S2‑MoE achieves up to 5.3× speedup (about 2.0× on average) over standard autoregressive decoding across diverse MoE models and datasets on edge devices. Code is available at https://github.com/angerybob/S2‑MoE.

Authors:Xingzheng Wu, Cheng Zhang, Guihao Yan, Xifeng Hu, Zhi Liu, Qing Cai
Title: ForceU-VLA: A Force-Aware Vision-Language-Action Model for Embodied Ultrasound Scanning
Abstract:
Embodied intelligent ultrasound scanning enables the automation and standardization of the ultrasound examination process by integrating perception, decision‑making, and execution capabilities. However, existing methods suffer from loosely coupled modeling between force and ultrasound modalities and lack awareness of scanning stages, which limits their ability to capture dynamic probe‑tissue interactions. To address these issues, we propose ForceU‑VLA, a force‑aware Vision‑Language‑Action model for autonomous embodied ultrasound scanning, which leverages force signals and ultrasound image feedback throughout the scanning process to enable accurate and high‑quality ultrasound acquisition. Firstly, we propose a Force‑Ultrasound Synergistic Fusion Module (FUSFM) that synergistically fuses ultrasound visual and force‑feedback information to provide stable, reliable guidance for probe motion. Secondly, a Stage‑Adaptive Modulation Mechanism (SAMM) is proposed to accommodate the task requirements across different scanning stages by adaptively modulating multimodal features to enhance their representation quality. Additionally, we introduce ForceU‑VLA‑Data, a real‑world, force‑aware embodied ultrasound dataset that integrates visual, force, and action signals, including data from two organs across five representative clinical scanning views, and comprising 450 expert‑collected trajectories with approximately 100,000 synchronized multimodal frames. Extensive experimental results demonstrate that ForceU‑VLA significantly improves contact stability and probe pressure regulation in embodied ultrasound scanning, thereby effectively enhancing task execution quality and overall system reliability. The source code is available at https://github.com/VMVLab/ForceU‑VLA.

Authors:Qizhen Lan, Xi Xiao, Xiangchen Guan, Mengchen Fan, Moule Lin, Jung Im Choi, Lijing Zhu
Title: Trust Is Not Enough: Influence Calibration for On-Policy Self-Distillation in Agentic RL
Abstract:
On‑policy self‑distillation (OPSD) gives language agents dense token‑level supervision from a privileged self‑teacher on the policy's own trajectories. Existing methods allocate this supervision mainly by teacher trust, but trust does not reveal whether emphasizing a token supports the current policy objective. We call this the trust‑utility mismatch and introduce Influence Calibration for Self‑Distillation (ICSD). For each supervised token, ICSD measures the first‑order response of its importance‑weighted RL surrogate contribution to a teacher‑directed output perturbation. Batch‑adaptive calibration converts this non‑stationary signal into a bounded allocation weight while preserving the original auxiliary‑loss mass within each action turn. These detached weights affect only the distillation loss and require no additional model pass. Across ALFWorld, WebShop, and Search‑QA, ICSD improves all matched aggregate metrics over trust‑only allocation under Group Relative Policy Optimization (GRPO) and Group‑in‑Group Policy Optimization (GiGPO), across two model families spanning 1.5B to 7B. At 7B, it reaches 96.1% ALFWorld success and a WebShop score of 93.1. Frozen‑batch analyses show that ICSD reduces teacher‑supported mass assigned to objective‑opposed tokens from 60.1% to 37.8% and raises cosine compatibility with the RL gradient by 0.192. A companion repository is avail‑ able at https://github.com/lanqz7766/Influence‑Calibration‑for‑On‑Policy‑Self‑Distillation‑in‑Agentic‑RL.

Authors:Nawrin Tabassum, Yanzhao Wu
Title: STAR-FL: Secure Federated Learning with Spatial-Temporal Analysis and Robust Aggregation
Abstract:
Data poisoning attacks pose serious security threats to Federated Learning (FL) systems in Computer Vision. Despite growing research attention, two key challenges remain for existing defense techniques: (1) accurately distinguishing between benign and malicious model updates and (2) effectively mitigating the influence of poisoned model updates during model aggregation. To address these challenges, we propose a novel defense framework against targeted poisoning attacks with Spatial‑Temporal Analysis and Robust aggregation for FL (STAR‑FL). First, we employ spatial‑temporal clustering to identify and remove potentially malicious updates from the FL training process. Second, we adjust the learning rate during aggregation to mitigate the impact of any malicious updates that evade detection. Third, we conduct extensive experiments across multiple benchmark datasets to evaluate the spatial‑temporal analysis and robust aggregation in STAR‑FL. Experimental results demonstrate their synergistic effect in enabling STAR‑FL to effectively protect FL and consistently outperform state‑of‑the‑art defenses against targeted poisoning attacks, significantly reducing Attack Success Rates (ASRs). The source code is available at https://github.com/mlsysx/STAR‑FL.

Authors:Chengyan Wang, Hanliang Xie, Yueyi Yang, Haoyu Chen
Title: Zero-MELO: Test-Time Evidence Calibration with Multimodal LLMs for Zero-Shot Micro-Gesture Recognition
Abstract:
While Multimodal Large Language Models (MLLMs) excel in general video understanding, their capability in fine‑grained and motion‑centric tasks remains limited. This limitation is particularly critical in micro‑gesture recognition (MGR), where micro‑gestures (MGs) ‑ subtle, short‑duration, and spatially localized human movements ‑ serve as key discriminative signals for implicit affective analysis, yet are easily neglected following common prompting practices. Although MGR has been intensively studied by many discriminative approaches, the use of MLLMs for MGR is underexplored, with notably poor performance. We hypothesize that the motion‑sensitive representation ability of MLLMs is constrained by their inherent single‑pass forward inference, which can be substantially enhanced through carefully designed test‑time guidance. Motivated by this, building on our prior findings regarding temporal insensitivity in Video LLMs, we diagnose zero‑shot MGR errors in the Negative Log‑Likelihood (NLL) space. We observe that MLLMs suffer from two bottlenecks: 1) insufficient localized evidence and 2) severe score biases driven by language and motion‑agnostic appearances. Thus, we propose a novel test‑time evidence calibration framework that improves both reasoning details and prediction reliability. Specifically, we introduce a tree search mechanism to progressively acquire localized, fine‑grained visual evidence, coupled with a test‑time calibration module to mitigate score biases. The multi‑cue fusion module then integrates evidence from multiple cues without relying on a single cue for final prediction. Our framework achieves mean‑class accuracies of 26.84% on iMiGUE and 22.10% on MA‑52, significantly outperforming the Qwen2.5‑VL baseline, which produces 16.15% and 10.20%, respectively. The code will be available at https://zero‑melo.github.io/Zero‑MELO.

Authors:An Phan, Yufei Jin, Xingquan Zhu
Title: M-LINKX: Multiview Graph Learning for Brain Cognitive Disease Detection
Abstract:
Electroencephalogram (EEG) is a non‑invasive and relatively low‑cost procedure that measures brain electricity for the detection of cognitive diseases. EEG‑based classification of dementia‑related conditions, including Alzheimer's disease (AD), mild cognitive impairment (MCI), and frontotemporal dementia (FTD), remains challenging because EEG signals are noisy, non‑stationary, and vary across subjects. Segment‑based learning provides a practical way to model long EEG recordings by converting them into fixed‑length inputs. For each segment, discriminative information may be explored by using signals within each channel (i.e. electrode), as well as interactions between EEG channels. In this paper, we propose M‑LINKX, a multi‑view graph learning framework for EEG‑based dementia classification. For each segment, we extract channel‑level node features and construct multiple functional‑connectivity (FC) graph views, where each view is defined by a specific combination of connectivity metric, frequency band, and topology filter, respectively. Instead of relying on message passing over the constructed graphs, M‑LINKX follows a simple design in modeling node features and adjacency‑based connectivity representations. The graph‑view representations are fused using global trainable view weights, and subject‑level prediction is obtained by averaging segment‑level probabilities. Experiments on two three‑class EEG datasets with different diagnostic groups, CAUEEG (HC/MCI/Dementia) and AHEAP (HC/AD/FTD), show that M‑LINKX achieves the best subject‑level performance under the main experimental settings. Our study suggests that multi‑view functional connectivity can improve EEG‑based dementia classification when integrated with an appropriate graph‑learning architecture. Code and data are available at https://github.com/anphantt/MLINKX.

Authors:John Helsby, Yi Yang, Bodo Rosenhahn, Michael Ying Yang
Title: OvDSGG: End-to-End Open-Vocabulary Dynamic Scene Graph Generation
Abstract:
Dynamic scene graphs (DSGs) capture spatio‑temporal interactions across videos as \langlesubject, predicate, object\rangle triplets, and underpin downstream tasks such as video captioning, video question answering, and action analysis. However, end‑to‑end dynamic scene graph generation (DSGG) methods are closed‑set: they recognize only objects and predicates from a fixed training vocabulary and struggle with the long‑tailed distribution of rare concepts, severely limiting their real‑world applicability. Existing open‑vocabulary models typically inherit pretrained large language models, resulting in multi‑stage training and inference with substantial cost. We introduce OvDSGG, the first end‑to‑end framework for open‑vocabulary DSGG. OvDSGG builds on top of an open‑vocabulary Spatial Backbone and a Temporal Backbone; we further propose a Triplet Feature Extraction Module that bridges them, and a Visual‑Language Alignment Module that preserves open‑vocabulary recognition by learning an adaptive decision boundary in the joint visual‑language feature space, without expensive knowledge distillation in existing methods. We further introduce a rigorous open‑vocabulary DSGG benchmark adapted from Action Genome, with disjoint Base/Novel splits for both objects and predicates. OvDSGG significantly outperforms open‑vocabulary baselines across all metrics, with zero‑shot Recall@K scores 10.0‑‑20.4 percentage point higher than the next‑best baseline, while on closed‑set DSGG remaining competitive with state‑of‑the‑art models. Code and benchmark are publicly available at https://github.com/jhelsby/OvDSGG/.

Authors:Angelos-Nikolaos Kanatas, Yuexuan Kong, Pablo Alonso-Jiménez, Xavier Serra, Dmitry Bogdanov
Title: What Makes a Good Layer? Assessing the Layer-Wise Intrinsic Properties of Music Foundation Models
Abstract:
Music foundation models are commonly used as frozen audio feature extractors, yet selecting which layer to extract from remains largely heuristic. Current practice defaults to fixed depths or multi‑layer fusion, with limited understanding of why certain layers transfer better across downstream tasks or how representation quality varies with depth and pre‑training paradigm. We conduct a systematic layer‑wise analysis of 12 music foundation models spanning three pre‑training paradigms (masked modeling, autoregressive modeling, and contrastive learning), characterizing their hidden representations through intrinsic geometric and transformation‑based properties. Correlating label‑free representation‑quality metrics with layer‑wise performance across 15 downstream tasks, we find that several metrics track layer quality for genre classification, emotion recognition, automatic tagging, and beat tracking, albeit with varying strength across tasks and pre‑training paradigms. However, all metrics fail on tonal tasks such as key estimation and chord recognition, indicating that no single property serves as a general proxy for representation quality across music information retrieval tasks. To address this gap, we introduce a pitch‑transposition equivariance measure that captures properties missed by these standard metrics, providing a consistent indicator of tonal quality across model families. Finally, we show that intrinsic metrics can serve as effective proxies for layer selection, matching or outperforming trainable multi‑layer fusion methods, particularly in limited‑data settings.

Authors:Haoran Wang, Xiongxiao Xu, Philip S. Yu, Kai Shu
Title: Beyond Tokens: A Survey on Decoding Methods for Large Language and Vision-Language Models
Abstract:
Large language models (LLMs) and large vision‑language models (LVLMs) have demonstrated impressive generative capabilities, yet ensuring their outputs align with user intent is still challenging. While most existing approaches address this issue at the training stage, inference‑time approaches like decoding methods offer a more efficient and scalable solution. Decoding methods control model generation by guiding token‑level selection, performing sequence‑level generation, or generating tokens in parallel to accelerate the process. In this survey, we identify three emerging paradigms from recent works on decoding methods for LLMs and LVLMs, provide a systematic review of these methods, highlight ongoing challenges, and discuss potential future research directions. Our goal is to underscore the efficiency and effectiveness of decoding methods and offer a practical view of their applications. Paper lists and more resources on decoding methods for LLMs and LVLMs can be found at https://github.com/wang2226/Awesome‑LLM‑Decoding.

Authors:Moein Salimi, Danial Parnian, Shaygan Adim, Amirmohammad Ebrahiminasab, Nima Alighardashi, Parsa Gholami, Sahand Akramipour, Mahdi Jafari Siavoshani, Mohammad Hossein Rohban
Title: CEDAR-GRPO: Process-Aware Reinforcement Learning for General Abductive Reasoning in LLMs
Abstract:
Abductive reasoning, often characterized as inference to the best explanation, is central to explanation under uncertainty, from everyday sense‑making and investigation to scientific discovery. Yet LLM research has mostly studied abduction through narrow, task‑specific benchmarks, making it unclear whether observed gains transfer beyond the benchmark family used for training or evaluation. We ask whether RL post‑training can improve abduction as a transferable reasoning capability. We introduce CEDAR‑GRPO, a process‑aware framework that combines final‑answer correctness with abductive rewards for evidence coverage and evidence‑to‑explanation directionality. Four open‑weight LLMs are post‑trained on a controlled, domain‑neutral mixture of abductive hypothesis‑generation and hypothesis‑selection tasks. We evaluate them on 11 unseen tasks spanning hypothesis selection, missing‑fact generation, defeasible inference, long‑context investigation, clinical reasoning, code debugging, and non‑abductive controls. CEDAR‑ GRPO improves every model on every held‑out task over both base models and correctness‑only GRPO, with average gains of 7.4 and 2.7 points, respectively, and a maximum gain of 30.8 points. Ablations confirm that RL, abductive reward design, and task diversity each contribute to transfer. Process‑level metrics further show stronger abductive behavior, including exploration of alternatives, elimination of rivals, backtracking, and uncertainty marking.

Authors:Yunpeng Bai, Yossi Gandelsman, Michaël Gharbi, Qixing Huang
Title: Qwen-Video-Edit: Instruction-Based Video Editing by Repurposing an Image Editing Model
Abstract:
Instruction‑based video editing is commonly built on video‑pretrained generative backbones: a video diffusion transformer is adapted, at considerable cost, to condition on a source video and an editing instruction. In this report we explore a different route and show that a strong instruction‑based image editing model can edit videos by operating directly on video‑VAE latents. Starting from Qwen‑Image‑Edit, we arrange the latent frames of a Wan~2.1 video VAE as tiles of one large virtual image, reuse the editor's image positional encoding for every tile, and bridge the two latent spaces with a pair of lightweight input/output projections warm‑started from the editor's own patchify and unpatchify layers, so that at initialization a (static) video is embedded exactly as an image the model already understands. The whole system is then fine‑tuned on the public Ditto‑1M editing triplets, and a few denoising steps of Wan~2.2 serve as an optional temporal enhancer. We motivate the design with a chain of zero‑training observations: the stock image editor already edits a video presented as a contact sheet; it is indifferent to whether the sheet's tokens come from one joint encode or from per‑frame encodes stitched in latent space; and it even edits genuine video latents zero‑shot to a clearly recognizable degree, leaving fine‑tuning only a fidelity gap to close. Our results suggest that, despite the large investment in training video latent spaces, per‑frame video latents remain close enough to the image domain that mature image editing priors transfer with minimal adaptation. Project Page: https://yunpeng1998.github.io/Qwen‑Video‑Edit‑Page Code: https://github.com/yunpeng1998/Qwen‑Video‑Edit Model: https://huggingface.co/yunpeng1998/Qwen‑Video‑Edit

Authors:Manwen Liao, Xinyu Lian, Jian Mao, Kaixu Chen, Li Luo, Jinghao Yan, Wanshui Gan, Qiao Yu, Weitian Zhang, Chunhua Shen, Guang Chen, Bo Dai, Xudong Xu, Zhaoyang Lyu
Title: MegaParts: Scaling Part-Aware 3D Object Generation to 300 Parts via Token-Efficient Autoregressive Modeling
Abstract:
Part‑aware 3D object generation is essential for graphics applications such as controllable modeling, editing, and articulation, where objects are represented as coherent assemblies of semantic parts. However, existing part‑aware generation methods, do not scale well to highly complex objects. As the number of parts increases, generating detailed geometry becomes prohibitively expensive in token length and memory. We introduce MegaParts, a scalable autoregressive 3D generation framework to address this challenge by combining structured sequence modeling with a token‑efficient vector‑quantized shape tokenizer. Our tokenizer learns discrete latent representations for part‑level geometry by minimizing token usage subject to high‑fidelity reconstruction, enabling adaptive‑length tokenization based on geometric complexity. On top of this compact representation, we train a large language model to generate object bounding boxes, part bounding boxes, and part shape tokens within a unified structured sequence. Combined with efficient long‑context training strategy, our token‑efficient formulation scales to objects with up to 300 parts and sequence lengths up to 256k tokens. This substantially extends the scale of part‑aware 3D generation while preserving compositional structure and enabling fine‑grained part‑level control. Our method achieves higher mesh quality than baseline autoregressive and diffusion models, showing that compressed discrete part tokens improve not only scalability but also the achievable fidelity of generated geometry. These results suggest that LLM native token‑efficient autoregressive modeling is a compelling alternative to diffusion for large‑scale part‑aware 3D generation. The project page is available at https://expmaster.github.io/megaparts_webpage.

Authors:Harshil Lodhiya
Title: ER-KANs: Efficient and Robust Kolmogorov-Arnold Networks for Data-Scarce Scientific Machine Learning
Abstract:
The efficient‑KAN literature‑‑‑covering Chebyshev, wavelet, and radial‑basis‑function variants of the original Kolmogorov‑Arnold Network‑‑‑has been benchmarked almost entirely on clean data. We show that this choice conceals a large capability difference between architectures: ChebyKAN's test MSE (evaluated against clean ground truth) increases by a factor of 10.6x when training data is corrupted with sigma=0.1 noise, versus 7.9x for vanilla KAN, 1.7x for a standard MLP, and just 1.4x for our proposed ER‑KAN. ER‑KAN combines three design choices targeting the noisy, data‑scarce setting: shared Gaussian RBF bases across all edges in a layer (providing locality and efficient parameterisation), curriculum noise injection during training (explicitly teaching noise robustness), and entropy‑weighted adaptive regularisation (preventing overfitting at small N). The result is a 595‑parameter network that matches MLP accuracy at moderate noise while degrading far more gracefully as noise grows. We evaluate on eight analytic functions (N in 50, 200, 500, sigma in 0, 0.03, 0.1), on a damped harmonic oscillator physics‑informed neural network where ER‑KAN achieves 4.2x lower solution MSE than MLP, and on a Burgers' equation PINN where all models fail to converge‑‑‑a genuine limitation we report rather than suppress. We introduce the noise degradation ratio as a simple complementary metric and recommend it become a standard reporting requirement for efficient‑KAN papers.

Authors:Robin Koch, Annabella Mascot, Rayan Younis, Martin Wagner, Stefanie Speidel, Mark Cutkosky, Ingo Sieber, Roberto Calandra
Title: MISTac: A Vision-Based Tactile Sensor for Minimally Invasive Surgery
Abstract:
Minimally invasive and robot‑assisted surgery offer many advantages over traditional open surgery, but deprive surgeons of tactile feedback and the ability to palpate tissue with their fingers. To address this lack of tactile feedback, we introduce the MISTac, a high resolution vision‑based tactile sensor specifically designed for palpation in MIS. The sensor has a replaceable sensor tip with a diameter of 8 mm which allows it to fit through the trocars used in minimally invasive surgery. Its modular 3D‑printed case design allows the use of bulky off‑the‑shelf illumination and imaging hardware that can easily be exchanged and upgraded. The sensor has an optical resolution of 176.68 μm, a tactile resolution of 250 μm, and can resolve forces as little as 24.3 mN. An in vivo study with the sensor shows its usability in minimally invasive surgery. We trained a machine learning model with the tactile data collected in the trial on a tissue classification task achieving an aggregate accuracy of ~84% in a leave‑one‑out cross validation. Tactile sensors have the potential to one day aid surgeons during minimally invasive surgery with tasks such as tissue classification or intra‑operative tumor localization; MISTac is a small step towards this vision. We open‑source MISTac at https://github.com/lasr‑lab/mistac

Authors:Quoc Anh Nguyen, Sunhong Park, Jin Tae Kwak
Title: Test-Time Instance Selection for Improved Whole Slide Image Analysis
Abstract:
Whole Slide Image (WSI) analysis has been widely studied for cancer diagnosis. Conventionally, a gigapixel WSI is divided into small patches and processed by Multiple Instance Learning (MIL) models. However, existing MIL models typically process all patches, many of which contain redundant or non‑informative tissue patterns. Although recent approaches have focused on instance selection to identify discriminative patches and reduce redundancy, these selection modules still require additional training. In this work, we propose Test‑Time Instance Selection (TTIS), a training‑free, plug‑and‑play framework that selects compact yet representative patches during inference. TTIS further incorporates a multi‑view ensemble strategy to integrate distinct facets of tissue morphology, enhancing robustness. Importantly, TTIS can be seamlessly integrated into existing MIL models without retraining or architectural changes, enabling flexible deployment. Extensive evaluations across multiple benchmarks demonstrate that our approach improves or matches baseline MIL performance across a range of classification and subtyping tasks. Our implementation code is available at https://github.com/QuIIL/TTIS

Authors:Vitaliy Polshkov, Marcin Pitera, Jeremy Yang, Kirill Priemko, Maksim Gaiduk, Aleksandr Nikolenko, Denis Bykov, Clare Southern, Denis Yarats, Jerry Ma
Title: WANDR: A Benchmark for Wide and Deep Research
Abstract:
WANDR (Wide ANd Deep Research) is a benchmark of 500 realistic, challenging data‑collection tasks for research agents. Each task requires a system to discover a large set of entities that satisfy specified criteria (breadth), investigate each entity through multiple coordinated web searches (depth), and return independently verifiable records with supporting sources and excerpts. Tasks are represented as qualification key hierarchies that specify the entities, relationships, evidence, and required count at each level; a hierarchy with n companies, m employees per company, and k sources per employee requires n x m x k records. This structure supports diverse workflows such as market mapping, due diligence, literature review, product comparison, and talent sourcing, with targets ranging from dozens to thousands of records. WANDR replaces static gold answer sets with task‑specific judges that refetch cited pages and verify each record against its evidence, allowing evaluation of current and changing facts. Record verdicts are aggregated into soft and hard precision, recall, and F1 scores that distinguish factual quality, coverage, and hierarchical completeness. The tasks are derived from de‑identified product‑usage logs and produced through a semi‑automated pipeline with automated checks, empirical audits, and human review where needed. We evaluate six production research systems and find that the benchmark is far from saturated: at high effort, the strongest system reaches only 0.363 soft F1 and 0.133 hard F1. Performance degrades as target volume and hierarchy depth increase, with incomplete discovery, missing enrichment, and incomplete evidence construction remaining major bottlenecks. The benchmark and evaluation harness are available at https://github.com/perplexityai/wandr.

Authors:Huipeng Huang, Hongxin Wei
Title: Tail-Aware Top-$k$ On-Policy Distillation
Abstract:
On‑policy distillation (OPD) has emerged as an effective paradigm for transferring knowledge between language models, where a student is trained to align its next‑token distribution with the teacher's along its own trajectories. To provide dense supervision at tractable cost, many works minimize the reverse Kullback‑Leibler (KL) divergence between the student and teacher's normalized distributions over the teacher's top‑k tokens. However, this normalized objective discards the information about tail probability: the total probability outside the teacher's top‑k tokens. As a result, the optimization can steadily increase the student's tail probability and entropy, empirically degrading downstream accuracy. To address this issue, we propose Tail‑Aware Top‑k OPD (TA‑OPD), a novel distillation method that restores the missing tail probability signal. In particular, TA‑OPD minimizes the reverse KL divergence over the top‑k tokens plus a tail token that carries the tail probability. In effect, TA‑OPD better aligns the student's next‑token distribution with the teacher's, preventing the increase in tail probability and entropy caused by top‑k normalization. Extensive experiments demonstrate the superiority of TA‑OPD, improving Avg@8 by up to 8.05 points on common benchmarks. Our code is available at https://github.com/HuipengHuang/TA‑OPD.

Authors:Ze Zhang, Yang Zhang
Title: Local Gains and Fixed-Assignment Set Losses in Shared Set Decoders
Abstract:
A query‑relation deletion can improve the edited slot while reducing the utility of the prediction set that contains it. We study this tension in two related ResNet‑50 DETR‑family checkpoints using recorded, selection‑conditional evidence from 710 paired image‑relation units per checkpoint. The primary comparison subtracts a matched active control, which deletes the same leader source at a different recorded recipient, from the selected target deletion. It is therefore a composite contrast rather than a same‑recipient placebo. The target‑minus‑control contrast is locally positive and fixed‑assignment negative in both checkpoints. The opposite‑sign pattern occurs within 302/710 DETR units and 460/710 DINO units. After rematching, the corresponding counts are 285/710 and 433/710. Rematching and native selection absorb enough of the mean loss for DETR intervals to cross zero, whereas DINO intervals remain negative, so persistence across readouts differs by checkpoint. A fixed‑map comparison between hard deletion and a mass‑preserving edit also differs before rematching. That comparison is conditional on the outcome‑blind map and does not establish same‑dose transport. Local intervention success therefore does not determine the consequence for a jointly decoded set. The supported conclusion is selection‑conditional deletion sensitivity whose persistence depends on the readout and intervention operator. We do not identify an intervention‑invariant edge mechanism, detector‑level degradation, population prevalence, or the value of a training‑time regularizer.

Authors:Ruochen Liu, Wei Lou
Title: Path2ST: Hierarchical Cell-Tissue Grounded Cross-Modal Translation for Spatial Transcriptomics
Abstract:
Predicting spatial gene expression from hematoxylin and eosin (H\&E)‑stained images offers a cost‑effective alternative to spatial transcriptomics (ST). However, existing methods treat H\&E images as generic visual inputs and ignore their intrinsic biological hierarchy, where spatially organized cell types collectively form functional tissue microenvironments that govern local gene expression programs. To bridge this gap, we formulate H\&E‑to‑ST prediction as a cross‑modal semantic translation task and propose Path2ST, a hierarchically grounded autoregressive framework featuring three key components: (i) a Hierarchical Cell‑Tissue Conditioning mechanism that fuses explicit and implicit cellular features with tissue‑level semantic representations to construct hierarchical conditioning signals; (ii) a Scale‑Adaptive Autoregressive Generation process over a hierarchical semantic vocabulary, enabling coarse‑to‑fine, biologically consistent expression synthesis; and (iii) SpectraLoss, a full‑spectrum objective that jointly enforces ordinal fidelity, models transcriptional bursts, and aligns semantic structures with cell types. Extensive experiments on three datasets demonstrate state‑of‑the‑art performance, validating that Path2ST generates highly accurate and spatially coherent transcriptomic profiles. The related code is released at https://github.com/RuochenLiu23/Path2ST.

Authors:Yitong Mu
Title: Deep Analog: Open-Set Film Emulation with Reference-Conditioned 3D LUTs
Abstract:
Film emulation reproduces the look of an analog film stock on a new digital photograph. We target its open‑set form ‑‑ matching any reference film frame from a single example ‑‑ with a 3D lookup table (LUT) predicted from that reference. Real‑time image enhancement predicts per‑image weights over a fixed bank of 3D LUTs and blends them. We show this is a gated mixture of experts and inherits its failure: trained end‑to‑end against reconstruction, the gate collapses onto a single expert, so a bank of K LUTs delivers the capacity of one. An entropy term, the enhancement‑setting analogue of mixture‑of‑experts load balancing, restores utilization and recovers about 1 dB PSNR. The deeper constraint survives: a fixed LUT basis is closed‑set, freezing the achievable looks at training time. We therefore discard the basis and predict a single 3D LUT as a residual from a reference image (StyleLUTNet), trained by self‑supervision on procedurally generated color transforms. The conditional design removes the gate and generalizes open‑set to unseen film stocks without paired data or retraining. Around this color backbone we build Deep Analog, a film‑emulation pipeline that adds histogram‑based tone matching and a physics‑informed optical renderer ‑‑ multi‑scale grain and per‑channel halation driven by parameters an inverse network regresses from the reference. On 350 self‑supervised pairs the color stage reaches 22.05 dB PSNR / 0.925 SSIM and the full pipeline 21.72 dB / 0.923; the color path runs in 5.2 ms at 1080p (192 FPS) and exports a portable .cube LUT for standard editing tools. A second degeneracy in conditional LUT training ‑‑ residual‑scale collapse ‑‑ shares the root cause and yields a general principle: auxiliary regularization must stay subordinate to reconstruction.

Authors:Chaolong Yang, Yinuo Guo, Kai Yao, Yuyao Yan, Jie Sun, Guangliang Cheng, Shibin Wu, Bin Dong, Kaizhu Huang
Title: Xemo-Talker: Unlock Emotions Explicitly for Audio-Driven Talking Portrait Synthesis
Abstract:
Precise emotion control in audio‑driven talking heads remains a challenge due to the reliance on implicit emotion regulation in existing systems, which often leads to indirect and insufficient control. Additionally, training with explicit emotion‑related losses across the entire motion space poses significant difficulties due to the inherent trade‑off between accurate lip synchronization and fine‑grained emotion control. In this paper, we reveal a key finding: although emotional cues are distributed throughout the motion space, concentrating discriminative supervision on less‑principal components achieves a better emotion‑lip synchronization balance, as principal components mainly encode high‑energy articulation and pose variations. Building on this insight, we propose Xemo‑Talker, which first learns a neutral speech‑to‑motion mapping for stable articulation and lip synchronization, and then introduces a lightweight emotion branch guided by less‑principal subspace supervision. To enhance emotion control, we design a Tri‑Loss consisting of inter‑class separation, intra‑class compactness, and less‑principal contrastive learning. Given an audio input, a reference image, and an emotion label, Xemo‑Talker achieves state‑of‑the‑art emotion classification accuracy while maintaining competitive lip synchronization and high inference efficiency, with performance approaching that measured on real videos.The source code is publicly available at https://github.com/chaolongy/Xemo‑Talker.

Authors:Ayrton Porto
Title: Beyond Correctness: Toward Automated Novelty Verification with Lean 4
Abstract:
Artificial intelligence systems applied to mathematics verify correctness but not novelty: an automatically generated theorem can compile in Lean without errors and yet be an already known result. This article presents AViD Journal, a pipeline that receives a LaTeX article, formalizes its statements in Lean 4, and issues a novelty verdict through a decision tree over three dimensions: prior existence in a formal corpus (Mathlib) and an informal one (TheoremSearch and Matlas, with temporal filter and LLM judge), non‑triviality via automatic tactics, and structural distance between proofs measured as Jaccard distance over premise sets. Evaluation on papers withdrawn from arXiv due to declared duplication produced a result more informative than any performance measure: the identification of three obstacles that limit the approach regardless of this implementation. First, successful compilation of a Lean file does not guarantee semantic fidelity. Second, the recall ceiling is imposed by the coverage of theorem indices, not by the similarity metric. Third, arXiv removes the source code of articles upon withdrawal, compromising the reproducibility of any benchmark built upon them.

Authors:Leonardo Kuffo, Peter Boncz
Title: Stop Indexing at Full Precision: Revisiting Clustering for Vector Embeddings
Abstract:
In this study, we revisit three widely used techniques in vector search and utilize them to optimize vector embedding indexing through clustering: dimensionality reduction, quantization, and dimension pruning. We propose an indexing pipeline in which these techniques are applied before clustering, and we focus on how they affect storage footprint, clustering time, and the quality of the resulting centroids for vector search tasks. Our results reveal that using full‑precision vectors for clustering is excessive, as even 1‑bit codes can achieve near‑optimal clustering quality (within 1% of ideal) while reducing storage requirements by 60x and delivering attractive performance gains (Figure 1). We open‑source our implementations at https://github.com/cwida/SuperKMeans.

Authors:Viquar Khan
Title: Proof-Gated Publication: Verify-Before-Commit Content Integrity for Serverless Data-Mesh Lakehouses
Abstract:
Federated data meshes give domain teams ownership of their data products, and serverless compute is an attractive substrate for domain‑owned writes. Both trends weaken correctness at publication. Open table formats such as Apache Iceberg and Delta Lake guarantee that a commit is atomic and that readers see an isolated snapshot, but not that the rows persisted equal the rows the job intended to write; publication is decided from the writer's exit status. A serverless job that silently drops a partition, truncates a file on retry, or duplicates a chunk still produces a valid, atomic, isolated, and wrong snapshot. This paper presents PVDM, a proof‑gated publication protocol with four phases: Physical (write to rollbackable staging), Verify (a keyed multiset proof that written content equals declared intent, stored by an independent notary), Durable (replay completed chunks across serverless retries), and Metadata (commit the catalog last, only if the proof passed). Metadata commits only if the proof passes, so a failing proof yields no consumer‑visible snapshot. The verification primitive is a keyed, incremental multiset hash over identity and content projections, telling missing or duplicated rows apart from corrupted values. A dependency‑free reference gate, a thirty‑case adversarial suite, and a reproducible benchmark catch eight thousand of eight thousand injected faults up to one million rows with no false blocks. We also run PVDM end‑to‑end on Apache Spark 4.0 and Apache Iceberg 1.11 at up to one hundred million rows, where every injected fault is blocked on the real commit path, the gated publish costs about seventeen milliseconds regardless of table size, and verification overhead is about a fifth of the write. The primitives are prior art; the contribution is composing them into a fail‑closed, notarized, verify‑before‑commit protocol for serverless federated writes.

Authors:Bhaskar Gurram
Title: Valid Per-Field Selective Risk Control for Document Extraction: Three Failure Modes, a Validity Ladder, and When Conditioning Pays
Abstract:
Per‑field accept/review with selective risk at most alpha ‑‑ accept a field only if the error rate among accepted fields is controlled ‑‑ is the trust contract document‑extraction systems need, and the natural procedure silently violates it on real documents. On 13,859 genuine claude‑sonnet‑5 fields from 800 CORD receipts (49.0% correct) we diagnose three failure modes: document clustering (design effect 1.84‑2.45), score‑refit leakage (coverage 0.416 at risk 0.127, violating alpha=0.10 in 95% of splits), and a tie‑mass pathology (a degenerate score collapses the threshold grid, 0.030 to 0.001). We organize the fixes as a validity ladder, guarantee form stated per tier. A fit/val split protocol restores expected‑selective‑risk control for a learned fusion: coverage 0.318 at risk 0.096 at nominal alpha=0.10, no tolerance band (production variant 0.326) ‑‑ an on‑average point whose realized risk exceeds alpha in 47.5% of resplits, not a certificate. Mondrian Learn‑then‑Test with exact binomial tails yields per‑group PAC certificates: field‑iid 0.171 at risk 0.068, cluster‑corrected 0.140, doc‑iid 0.060 ‑‑ the only tier matching documents, honestly near‑vacuous today. Support‑bin, the pre‑specified provenance taxonomy, wins every rigor tier on the sonnet CORD capture (p<1e‑4, Bonferroni‑corrected) ‑‑ a win that does not replicate on the same documents under haiku or qwen ‑‑ while on higher‑accuracy corpora pooled thresholds win: conditioning helps exactly where pooled cannot certify, subsumed by a learned score elsewhere. A frozen‑configuration confirmation on selection‑untouched claude‑haiku‑4‑5 held at both risk levels, and a blind three‑annotator human‑gold audit verifies the practical tier's accepted‑set risk at 1.3% against its 10% budget (Fleiss' kappa=0.83; labels err one‑sidedly pessimistic). Released Apache‑2.0 with seed‑pinned, regression‑gated procedures.

Authors:Nathael Altman
Title: Wolff-Parkinson-White Detection at 471:1 Class Imbalance: A Leakage-Controlled Study of the Data Bottleneck
Abstract:
Wolff‑Parkinson‑White (WPW) syndrome is a congenital cardiac pre‑excitation, clinically important and often missed on the resting 12‑lead ECG. Detection is hard: the signature is subtle and the condition rare. We pool two public 12‑lead corpora, PTB‑XL and Chapman‑Shaoxing‑Ningbo: 66,951 recordings, 142 of them WPW, a prevalence of 0.21% (about 471:1). Under one pre‑specified, leakage‑controlled protocol, with a held‑out fold contacted exactly once, we compare seven representations of the signal, holding the split and the evaluation fixed. Within these corpora and under a modest compute budget, added diversity and capacity do not raise the ceiling: the most orthogonal detector significantly hurts, a feature‑union model matches a two‑member vote, a convolutional network reaches the wavelet detector without exceeding it, and self‑supervised pretraining fails a pre‑specified gate. A leak‑free learning curve, re‑selecting features at every size, still rises at the full 115 positives for the strongest deployed detector (paired 90‑to‑100% difference +0.027, 95% CI [0.019, 0.033]), so it is not shown to have saturated. An error analysis tested against independent evidence finds that the missed cases have a narrower QRS, confirmed by an on‑machine measurement outside our pipeline after we show the sign of this effect depends on which delineator measures it; that uncertain labels show no enrichment among the misses; and that some apparent false positives are recordings the corpus itself codes as pre‑excited, placing part of the label problem in the negative class. We measure the optimism of non‑nested selection at 0.11 to 0.13 average precision. The deployed output is a percentile rank in a frozen reference distribution, not a probability. On the held‑out fold, on 14 positives, it reaches an average precision of 0.595 and an ROC area of 0.950. It is a screening pre‑filter, not a diagnostic tool.

Authors:Shugong Xu, Jun Jiang, Yuan Gao
Title: 6G Native AI and Channel Foundation Models
Abstract:
The integration of artificial intelligence (AI) and wireless communications is widely regarded as a core objective of sixth‑generation (6G) systems. However, both the meaning of native AI and the type of AI capability that should be embedded into future wireless systems remain open to interpretation. This paper discusses 6G native AI from a system‑design perspective and argues that native AI should be co‑designed, optimized, and deployed as an intrinsic component of the wireless system rather than as a removable post‑deployment add‑on. From this perspective, conventional task‑specific supervised models are difficult to use as the main technical basis of native AI because they depend heavily on labeled data, generalize poorly across propagation conditions, and require fragmented designs for different channel‑related tasks. Motivated by these limitations, we position channel foundation models (CFMs) as a channel‑centric foundation‑model paradigm for 6G native AI. We define the scope of CFMs, clarify their differences from task‑specific wireless AI models and large language models, and summarize three pretraining families: generative, discriminative, and hybrid pretraining. We further discuss how CFMs may support physical‑layer processing, radio access network intelligence, and integrated sensing and communications. Preliminary CSI‑CLIP‑based results are included as bounded evidence that CFM‑style pretraining can improve positioning and beam prediction when task‑specific labels are limited.

Authors:Zhaoyu Li, Hangrui Bi, Youyuan Zhang, Wenjie Ma, Zenan Li, Zhaolei Zhang, Xujie Si, Kaiyu Yang
Title: Euclid-Omni : A Unified Neuro-Symbolic Framework for Plane Geometry
Abstract:
Euclidean geometry is a compelling testbed for AI reasoning, as it demands the combination of intuitive diagram understanding, axiomatic deduction, and algebraic computation. Yet, existing approaches typically address only a subset of these abilities or struggle with competition‑level problems. We introduce Euclid‑Omni, a unified neuro‑symbolic framework that couples a formal geometry system with Large Language Models (LLMs) and Vision‑Language Models (VLMs) to tackle both calculation‑ and proving‑style problems, in formal and natural languages, up to Olympiad‑level difficulty. At its core, we develop Euclidea, a versatile symbolic geometry solver that automatically generates reasoning steps through deductive inference and algebraic computation. Building on this, we develop a data‑generation pipeline that synthesizes symbolic problems and solutions, renders diagrams, and translates them into natural language, producing large‑scale, diverse datasets for training LLMs and VLMs across a wide range of reasoning settings. Experiments show that VLMs trained on our synthetic data achieve superior performance on calculation tasks, and that LLMs combined with Euclidea are competitive with state‑of‑the‑art systems on Olympiad‑level proving problems, despite using orders of magnitude less compute and training data. Code and scripts are publicly available at https://github.com/20171130/Euclid‑Omni

Authors:Zhouyuan Ma, Yutao Wu, Hanxun Huang, Xiang Zheng, Xiao Liu, Yixin Cao, Zuxuan Wu, Xingjun Ma, Yu-Gang Jiang
Title: HarmProfile: Characterizing Harmful Distributions in Frontier LLMs
Abstract:
Frontier large language models (LLMs) safety evaluation has largely treated harmful generation as an attack outcome rather than as an object of analysis. Consequently, little is known about the harmful outputs produced during model misbehavior, partly because large‑scale, high‑quality collections of frontier‑LLM misbehavior are difficult to obtain. To address this gap, we introduce HarmProfile, a content‑centric benchmark dataset that collects model misbehavior across diverse harm categories and model families, and defines the resulting harmful‑output distribution as a model‑level risk profile. The premise is that, just as linguistic behavior can be characterized from an utterance corpus, model risk can be characterized from the content, severity, and variation of its safety failures. HarmProfile contains over 80,000 validated artifacts from 23 frontier LLMs across 13 model families, organized into 15 harm categories and 57 subcategories. Using this corpus, we find that frontier LLMs reliably produce harmful content at scale, yet exhibit distinct risk profiles; both harmfulness and diversity grow with model capability, suggesting that frontier LLMs may appear safe yet harbor increasingly dangerous knowledge beneath the alignment surface. Our source code is available at https://github.com/fresh‑ma/HarmProfile .

Authors:Ahasan Kabir, Jiaqi Xue, Mengxin Zheng, Qian Lou
Title: HW-Router: Hardware-Aware Routing for Scalable Multi-LLM Serving
Abstract:
Modern large language model (LLM) serving platforms deploy multiple models across different GPUs, requiring routers to direct incoming queries to appropriate LLMs. However, existing routing approaches primarily rely on static model attributes such as size or FLOPs to estimate serving costs. This static cost modeling fails to capture the dynamic behavior of real deployments, where the same model can exhibit vastly different inference latencies depending on hardware type (e.g., H100 vs. V100), current system load (e.g., running and waiting queue lengths), and resource contention (e.g., KV‑cache usage and GPU utilization). Such hardware‑agnostic routing leads to suboptimal decisions, resulting in SLO violations, queue buildup, and underutilized GPUs. To address these challenges, we present HW‑Router, a dynamic routing framework that integrates real‑time hardware signals into model selection to enable accurate latency prediction and intelligent, SLO‑aware routing decisions. Our approach incorporates model‑specific features (architecture, size, input length) alongside hardware metrics including queue lengths, KV‑cache utilization, and recent TTFT/TPOT performance, and uses a lightweight latency predictor to estimate per‑model‑per‑GPU serving time. Evaluations across diverse workloads show that HW‑Router achieves 3.4‑3.9x lower end‑to‑end latency, 46‑48 percentage points higher SLO attainment, 6‑8x lower GPU load skew, and a 3.1‑3.4x reduction in waiting‑queue fraction compared to state‑of‑the‑art router baselines, CARROT and IRT, with only ~200 us of additional routing overhead and no loss in output quality. These results highlight the importance of real‑time hardware feedback for scalable, predictable, and well‑balanced multi‑LLM serving. Code is available at https://github.com/UCF‑ML‑Research/HW‑Router.

Authors:Yuan Guo, Yilong Chen, Chao Hu, Xianghao Yu, Liang Hong, Jie Xu
Title: WARA: Toward Automated Wireless Optimization Research with Closed-Loop LLM Agents
Abstract:
Large language model (LLM) agents are increasingly capable of tool use, code execution, artifact inspection, and iterative revision, creating new opportunities for automating scientific and engineering research. To the best of our knowledge, this paper presents the first end‑to‑end autoresearch framework for the wireless domain, with a focus on wireless resource allocation optimization. We propose the Wireless AutoResearch Agent (WARA), a closed‑loop multi‑agent system for automated wireless optimization research. Given only an initial topic, WARA decomposes the workflow into three phases: research gap identification and problem proposal, wireless optimization modeling, algorithm design and experimentation, and research deliverable construction. Across these phases, WARA uses artifact‑mediated control: upstream artifacts are consumed as inputs, structured outputs are stored for downstream use, and controller‑managed gates validate consistency among models, algorithms, experiments, and claims. When validation fails, WARA repairs only the responsible artifact instead of restarting the whole workflow. We present a representative wireless resource allocation case study showing how WARA converts an initial topic into a complete research package with executable evidence and a synthesized technical manuscript. We further design a structured LLM‑based ScoringAgent to evaluate manuscript‑level research validity and optimization research maturity. Comparative results show that WARA substantially outperforms one‑shot LLM generation and approaches the quality profile of recently accepted peer‑reviewed technical papers. These results indicate that closed‑loop artifact control is a promising path toward end‑to‑end LLM‑assisted wireless optimization research. The source code is available at https://github.com/guoyuan‑dotcom/WARA_CUHKSZ.

Authors:Wen Ye, Muyan Weng, Chuizheng Meng, Hao Niu, Yizhou Zhang, Yan Liu
Title: Coarse-to-Fine Multi-Resolution Diffusion Models for Trajectory Generation in Urban Systems
Abstract:
Understanding human mobility is critical for a wide range of urban applications, including traffic management, epidemic control, and urban planning. However, due to privacy concerns, the availability of large‑scale public trajectory data remains limited, posing challenges for downstream mobility analysis. Existing methods for synthetic trajectory generation primarily focus on matching global distribution similarity, while often overlooking mobility patterns across different spatial and temporal resolutions that are essential for practical utility. To address these challenges, we propose a novel multi‑resolution diffusion framework, MR‑Traj, for large‑scale trajectory generation. MR‑Traj explicitly models trajectories as compositions of coarse‑grained milestones and fine‑grained segments, enabling the capture of complex spatial‑temporal dependencies at multiple resolutions. Experimental results demonstrate that MR‑Traj achieves comparable performance to state‑of‑the‑art methods in terms of global distribution similarity, while consistently outperforming them in modeling fine‑resolution mobility patterns and supporting downstream urban mobility tasks. In addition, by introducing stochasticity at multiple resolution levels, MR‑Traj generates more diverse trajectories, which empirically reduces trajectory linkage risk under a seed‑guided data release setting. Our code is available at https://github.com/Ray0202/MR‑Traj.

Authors:Matt R. Flax
Title: A Biophysically-Inspired Feedback Controller for Multi-Class Cache Fairness
Abstract:
Cache replacement under multi‑tenant LLM‑serving conditions is a multi‑class problem: short, high‑reuse system prompts; long, moderate‑reuse user documents; medium‑length code context; and bursty conversation history share a single eviction pool. Under skewed multi‑class arrivals, conventional flat‑LRU policies expose the worst‑served‑class miss ratio (m_\max) only as a fixed point. We introduce a class of cache‑replacement policies parameterised by a per‑class flux formula, where three structural commitments ‑‑ a single global token‑mass imbalance signal, K parallel rectified per‑class promotion accumulators, and an age‑ordered eviction backstop ‑‑ produce emergent multi‑class fairness. We instantiate this class with a linear V‑coupled rectified flux and a Goldman‑Hodgkin‑Katz extension whose V \to 0 limit is exactly the linear form. Across four skew levels on synthetic multi‑class workloads, the policy class closes 27‑‑72\,% of the LRU\toBelady gap on m_\max, with linear and GHK interchangeable on the headline objective within search variance. The fairness/throughput tradeoff is exposed as a tunable knob on a single hyperparameter axis. We position this against the LeCaR feedback‑controller lineage and the formal‑control‑theory cache‑decay lineage as a novel combination of known ingredients. Code and reproduction scripts: https://github.com/flatmax/membrane.cache

Authors:Arya Rahgozar, Pouria Mortezaagha
Title: Auxiliary uncertainty signals for LLM-assisted systematic review screening: a benchmark across eight Cohen drug-class reviews
Abstract:
Large language models (LLMs) are increasingly used for title‑abstract screening in systematic reviews, but their decisions lack calibrated uncertainty. We show that an auxiliary BERT+GCN classifier supplies a structured uncertainty signal that improves LLM screening efficiency, and we identify the prompt‑delivery strategy that maximises the benefit‑to‑cost ratio. We evaluate five LLM prompt‑delivery conditions on eight drug‑class datasets from the Cohen (2006) benchmark using 3 seeds x 5‑fold stratified cross‑validation (600 fold‑level results). A BERT+GCN model trained per fold classifies each test paper as INCLUDE, EXCLUDE, or MAYBE via two spectral tests (algebraic radical and categorical paradox). Conditions vary information content (none / label / full scores), selectivity (all papers vs. MAYBE only), and timing (proactive vs. reactive two‑pass). A cross‑model pilot against gpt‑4.1‑mini on three datasets tests cross‑generation transfer. Three findings: (i) Full‑context delivery yields significant gains in F1 (+0.011, paired Wilcoxon p=0.008) and WSS@95 (+0.050, p=0.039) at a 1.28x token‑cost premium, while preserving recall. (ii) MAYBE‑only routing is Pareto‑optimal: highest mean recall (0.92) and AUC‑ROC (0.54) at only 1.05x baseline cost ‑‑ one sixth of full‑context overhead. (iii) The two‑pass design escalates 22.2% +/‑ 8.8% of records yet never revises its decision (0% flip rate across all datasets and folds), giving decisive evidence that current instruction‑tuned LLMs cannot self‑triage. The cross‑model pilot shows an identical +0.8% recall uplift for both LLM generations. A per‑paper ablation across 20,796 observations shows the dual paradox test reduces empirically to a one‑line logit‑gap criterion. We release the full pipeline; the 600‑run experiment replays in under one hour from cached LLM responses.

Authors:Zian Meng, Zhen Li, Chuanhao Li, Qiang Li, Kaipeng Zhang
Title: Marionette: Predicting World States, Rendering Geometry, Painting Appearance
Abstract:
Interactive game world models typically autoregress visual observations directly in pixel or latent space, forcing structured properties such as pose, geometry, and occlusion to be implicitly maintained by the same generative sequence. Over long horizons, errors in these latent world properties accumulate, making consistency and controllability fragile. We explicitly model the evolving world state, delegate exact geometric computation to a fixed, zero‑parameter renderer, and leave the neural model to synthesize appearance. We instantiate this idea as Marionette, a world model for interactive games with articulated characters. First, a two‑stage autoregressive dynamics model predicts an explicit and interpretable 276‑dimensional 3D world state comprising multi‑entity articulated skeletons, metric root trajectories, and rotations. Second, a zero‑parameter graphics bridge converts the predicted state into pose‑control videos, computing world‑space geometry and occlusion in closed form. Third, a control‑conditioned video‑diffusion observation model synthesizes photorealistic RGB observations from the resulting structured controls. Our experiments establish two properties of Marionette. First, the predicted world state is directly controllable. Forcing a mismatched action stream changes root‑aligned joint error by 31% across 48 held‑out segments. Second, long‑horizon behaviour is determined in the state, and can be repaired there. Left free, the two generated characters drift to 21.2 m apart (recorded sessions stay near 5 m) and a third of frames show ground penetration. Two rules imposed on the explicit state, a terrain collider and a separation cap, cut penetration by 66% and keep the pair engaged, with no change to the observation model. Routing appearance through the predicted state costs no fidelity we can detect, at an FVD of 831 against 799 for recorded pose.

Authors:Jocelyn Xu, Minje Kim
Title: Singer-Informed Vocal Source Separation for Multi-Singer Music Mixtures
Abstract:
Music source separation systems typically extract a single vocal track and do not distinguish between multiple singers. We study singer‑informed vocal source separation for multi‑singer mixtures. Our framework introduces a short enrollment recording of a target singer to guide separation through a learned embedding. The singer embedding is incorporated using feature concatenation or feature‑wise linear modulation (FiLM), enabling the model to focus on the target singer while suppressing interference. We construct a duet dataset based on DAMP‑VSEP with quality filtering and non‑overlapping enrollment segments. Experiments on solo and duet settings show that while baseline models perform well for single‑singer mixtures, the proposed method improves target‑singer extraction in multi‑singer cases, increasing target‑singer SI‑SDR from 0.33 dB to 5.58 dB. Fréchet Audio Distance (FAD) further shows improved perceptual quality and better alignment with target audio distributions. Code and checkpoints are available at https://github.com/jocelynxu01/singer‑separation‑paper.

Authors:Alexy Skoutnev, Kirill Acharya, Gaston Longhitano, Madeleine Udell, Kevin Ellis, Iddo Drori
Title: Twin: Playing an Unknown Game with a Test-Time Digital Twin
Abstract:
We present a Test‑time World‑model Inference (Twin) system, in which a frontier coding agent writes an executable world model for completing continual learning tasks, such as ARC‑AGI‑3 games. Traditional approaches hand‑engineer such models, one custom design per task. Each game hides its rules and goal, and our system constructs them from simulation and interaction alone. Its inductive prior over grid games is strong enough to recover the true transitions of the game and the goal on nearly all levels. Replay validation happens in a twin world model. The harness enforces that an action is not made until the program reproduces every previous observed game transition. Each mismatch between a world model prediction and the actual action result becomes a counterexample that is used to repair the world model. Twin clears 179 out of 183 levels (97.8%), and does so more efficiently than humans in 158 out of 179 levels (88.3%). The system infers the goal before any reward on 156 of the levels it clears (87.2%), and in the remaining levels automatically discovers the goal by search. The benchmark scores completion and action efficiency, between 0 and 100, against humans playing each game for the first time. Played directly, the base model scores only 7.8%; an off‑the‑shelf harness increases it to 61.1%, whereas our twin world model increases the same base model to 93.3%, clearing 23 out of 25 games. Building a usable world model is simpler than anticipated, whereas the harder problem is inferring the right goal.

Authors:Yuhao Zhan, Bingxiang He, Zecong Tang, Chaojun Xiao
Title: PACE-Bench: Benchmarking Physics Adaptation via Code Evolution in Dynamic Environments
Abstract:
Self‑evolving agents improve future behavior from interaction experience, yet existing evaluations typically optimize under fixed execution conditions and do not test recovery after those conditions change. To address this gap, we introduce PACE‑Bench (Physics Adaptation via Code Evolution), a simulator‑grounded benchmark of 144 source‑to‑target adaptation pairs across six physics domains. Each pair links a source environment to a mutated target environment with the same goal and interface. A code‑driven design that succeeds in the source fails in the target, where agents must iteratively adapt it into a working target design using diagnostic sandbox feedback within a limited attempt budget. We compare ten self‑evolving methods from four paradigms. The benchmark remains far from saturated: Reflexion + Qwen3‑14B succeeds on only 35.9% of full‑benchmark pairs, while GPT‑5.5 solves 66.7% of the Statics subset under the full budget. Together, these results show that simulator‑grounded reflection is more reliable than unverified self‑revision, while memory anchors agents to early designs and broad tree search explores without converging. Even revealing exact physical changes does not raise the performance ceiling, pointing to mechanism redesign rather than parameter inference as the central bottleneck. Data and code are available at https://github.com/thunlp/PACE‑Bench.

Authors:Mahdi Saberi, Toygan Kiliç, Mehmet Akçakaya
Title: UMPIRE-Net: Unrolled Magnitude-Phase Regularization Network for Accelerated MRI
Abstract:
MRI reconstruction from undersampled k‑space measurements is an ill‑posed inverse problem. Physics‑driven deep learning (PD‑DL) methods have shown strong performance for this task by combining the MRI forward model with learned image regularization within algorithm‑unrolling frameworks. However, most existing PD‑DL methods reconstruct complex‑valued images directly, thereby implicitly coupling magnitude and phase within a single learned representation. This coupled regularization may be suboptimal in reconstruction settings where accurate phase modeling plays an important role, such as partial Fourier (PF) imaging, where recovery of the omitted asymmetric k‑space measurements depends on the underlying image phase. In such scenarios, explicit modeling of magnitude and phase as separate components may reduce the reliance on externally estimated or predefined phase information. To this end, we propose UMPIRE‑Net (Unrolled Magnitude‑Phase In REgularization Network), a PD‑DL method that introduces separate learned regularizers for magnitude and phase components, together with a novel data‑fidelity formulation that enforces measurements consistency. We evaluate UMPIRE‑Net for accelerated MRI with PF across different datasets and acceleration factors. Experimental results demonstrate that our proposed method improves reconstruction quality compared with a conventional complex‑valued PD‑DL baseline, yielding sharper images and reduced artifacts. Code available at: https://github.com/MahdiSaberii/UMPIRE‑Net

Authors:Jihun Park, Kyoungmin Lee, Jongmin Gim, Hyeonseo Jo, Jaeyeul Kim, Han Zou, Zhenpeng Zhan, Yan Zhang, Sunghoon Im
Title: CRAFT: Constrained Reward via Attention Fine-Tuning for Subject Personalization without Composed Targets
Abstract:
Subject‑driven image personalization‑‑‑generating new images that preserve the identity of one or several reference subjects in novel scenes‑‑‑is a foundational capability for modern visual content creation. It is currently dominated by generalized methods that fine‑tune a pretrained multimodal diffusion transformer (MMDiT) on hundreds of thousands to millions of paired \emph(reference, composed‑target) examples, where each composed target is a synthesized image of the subject in a novel scene. Producing such targets demands a costly multi‑stage curation pipeline‑‑‑LLM‑based prompt generation, T2I‑based composed‑target synthesis, reference‑subject extraction, VLM‑based quality filtering, and correspondence labeling‑‑‑and tightly couples each method to a particular target synthesizer and curation choice. We introduce \emphCRAFT (Constrained Reward via Attention Fine‑Tuning), a single‑step ReFL framework that fine‑tunes a pre‑trained \emphreference‑aware MMDiT via LoRA adapters using a compact reference‑only data construction‑‑‑10K reference images and subject masks, with no composed‑target supervision. CRAFT realizes a \emphWhere to look principle: attention‑level rewards align noise‑ and phrase‑token attention with the correct reference subject, and the resulting per‑subject attention masks gate a pixel‑level identity reward to keep image‑space supervision consistent with the learned attention routing. Applied to FLUX.2‑klein‑9B, CRAFT achieves state‑of‑the‑art performance on XVerseBench \revwhile using no composed‑target supervision‑‑‑only 10K reference‑only samples, whereas prior generalized methods require 150K to over 2M composed‑target pairs. The same recipe transfers to other reference‑aware backbones, consistently improving performance. Project page: https://jihun999.github.io/projects/CRAFT/.

Authors:Yichen Xu, Jianzhe Ma, Chuhan Wang, Zhonghao Cao, Liangyu Chen, Wenxuan Wang, Qin Jin
Title: A Survey of Large Models in Sports
Abstract:
Sports have witnessed growing global enthusiasm in recent years, serving as a vital force for physical health, cultural exchange, social connection, and economic growth. The rapid advancement of large models, particularly (multimodal) large language models (M)LLMs, has demonstrated transformative potential to reshape sports understanding, analysis, and interaction across diverse domains. This paper presents a comprehensive survey of large models in sports, including (i) an overview of tasks and applications across different participant groups; (ii) a detailed analysis of sports‑related datasets and benchmarks; and (iii) a critical discussion of current challenges and future directions. Our goal is to establish a foundation for advancing research and practical development of large‑model‑driven sports intelligence. An open‑source GitHub repository is maintained at: https://github.com/Road2Redemption/Awesome_Large_Models_In_Sports1.

Authors:Zhizhao Guan, Chen Huang, Ziming Liu, Hongru Liang, Wenqiang Lei, See-Kiong Ng, Tat-Seng Chua, Anthony G Cohn
Title: Clearing the Fog: Towards Installing and Refining Proactive Exploration Capabilities in LLM Agents
Abstract:
We study proactive exploration in LLM agents, i.e., the ability to explore an environment to acquire information that improves future decision‑making. In this regard, we first identify two fundamental bottlenecks that hinder this capability and then propose \ours, a novel method designed to instill and refine proactive exploration. Specifically, \ours\ consists of two components: (1) Exploratory Data Construction, which synthesizes exploration‑rich trajectories to mitigate the hindsight bias of standard demonstrations; and (2) RL Optimization with Contrastive Signal Guidance, which leverages contrastive trajectory pairs to distinguish productive exploration from redundant wandering. Extensive experiments demonstrate the effectiveness of \ours\ and provide insights into the characteristics of proactive exploration. Our code is available at: https://github.com/GuanZhizhao/SAFARI.

Authors:Md Ashraful Hossen Akash, Shyla Afroge, Abdullah Al Mamun, Md. Kishor Morol, Tze Hui Liew
Title: TRIAGE: Risk-Controlled Pseudo-Label Admission for Annotation-Efficient Semi-Supervised Retinal OCT Classification
Abstract:
The advanced retinal disease diagnosing imaging modality, optical coherence tomography (OCT), encounters a lack of automation because of the high expenses for annotations performed by specialists. The use of SSL solves the problem of insufficient annotations using unlabeled B‑scans; however, most of the current techniques for generating pseudo‑labels are based on prediction confidence without considering the asymmetry between different types of errors. This paper proposes TRIAGE, a risk‑controlled semi‑supervised framework for OCT scans classification, which uses the concept of a patient‑level conformal risk controller with an asymmetric cost matrix. TRIAGE unites three crucial modules: a hierarchical classifier that is capable of working with partially abnormal supervision of the disease subtypes, a patient‑grouped conformal risk controller with primal‑dual coverage control, and a context‑aware Transformer teacher for cross‑slice verification. On the dataset from Noor Eye Hospital (16,822 B‑scans, 161 patients, and 554 volumes) with a test set of unseen patients, TRIAGE demonstrates 89.66% scan‑level accuracy, 0.8805 macro‑F1, 0.9641 macro‑AUC, and an 8.34% under‑grading rate when using only 20% of the labeled data. With only 5% of the labeled data, TRIAGE keeps 76.88% accuracy and a 0.1656 under‑grading rate. Compared with the other six state‑of‑the‑art semi‑supervised methods, TRIAGE significantly outperforms them with ablation study demonstrating the contribution of each module in the overall framework performance (by 42.7% in terms of under‑grading rate comparing to fixed threshold methods). TRIAGE demonstrates 98.00% accuracy for 3‑class classification with 1% labeled data and 95.94% accuracy for 8‑class classification with 10% labeled data on the OCT‑C8 dataset.

Authors:Xiaojun Wu, Cehao Yang, Honghao Liu, Xueyuan Lin, Zhichao Shi, Hao Zhou, Xuhui Jiang, Chengjin Xu, Jia Li, Jian Guo
Title: Envs-FORGE: Frontier-Optimized Reward-Grounded Environment Synthesis for Agent RL
Abstract:
Reinforcement learning (RL) for terminal agents needs executable training environments with reliable rewards and useful difficulty. Fixed recipes such as few‑shot, Self‑Instruct, and Evol‑Instruct apply the same prompting policy to every seed, even when the current policy would benefit from a harder, easier, or simply different task. We present Envs‑FORGE, a prompting policy that converts verifier rewards into per‑seed environment‑synthesis actions. Envs‑FORGE estimates seed pass rates, scores six projection‑‑direction actions around a target learning frontier, and solves a per‑seed mixed‑integer linear program (MILP) to choose the action that conditions generation. The selected action drives synchronized rewriting of the instruction, fixtures, oracle solution, tests, and Docker environment; only gold‑verified bundles enter RL training. The indexed MILP form also supports optional soft skill coverage for portfolio planning. On Qwen 3.5 35B, Envs‑FORGE improves Pass@1 over Base by 9.2 percentage points on tb‑core (40.0% to 49.2%) and 6.4 points on tb‑2.0 (23.0% to 29.4%), exceeding the strongest fixed‑recipe baseline by 2.4 and 2.1 points. It reaches 77.1% on SWE‑bench Verified versus 73.4% for Base, and improves tb‑core by 6.8‑‑9.2 points across the evaluated 4B‑‑35B models. All synthesis methods export 100 verified environments and use 2.27M‑‑2.88M synthesis tokens, placing the comparison at the same downstream training‑set size and the same operational scale. The source code is available at https://github.com/DataArcTech/DataArc‑SynData‑Toolkit/.

Authors:Navonil Neogi, Nabil Iqbal
Title: Information Spreading in Diffusion Models from Effective Field Theory
Abstract:
We study score‑matching diffusion models with a convolutional architecture. We argue that the inductive bias of locality means that the machinery of effective field theory from physics can be usefully applied to describe the denoising dynamics. We apply this formalism first to a simple toy example which permits an analytical description, and thereafter to MNIST, and show that in both cases, the mutual information between two points grows in a manner predicted by a simple effective field theory of Brownian motion.

Authors:Kohsuke Ide, Ryousuke Yamada, Yoshihiro Fukuhara, Hirokatsu Kataoka, Yutaka Satoh
Title: Seeing Red, Thinking Bad: Color Bias in Vision Language Models
Abstract:
Vision language models (VLMs) are increasingly used in industrial decision‑making systems, such as recruitment support and recommendation. This motivates careful analysis of how VLMs process visual and textual information. In this work, we study how VLMs interpret text rendered as an image, and investigate the influence of visual styling biases. To this end, we introduce Stealth Visual Prompts, which subtly change visual styling of text, such as color and contrast, while preserving semantic content. Using these prompts, we systematically control the visual styling of words in text and measure their impact on the analysis performed by VLMs. We further analyze how such visual perturbations affect the latent representations of the vision encoder. From our experiments, we observed that coloring positive words in green consistently shifts sentiment predictions toward a positive direction. As a result, VLMs often fail to properly account for negative words present in the text. Our analysis suggests that this behavior is correlated with changes in the latent representations of the vision encoder induced by color variations. In addition, we show that reducing text‑‑background contrast increases reliance on visually salient cues and leads to more incorrect Visual Question Answering (VQA) outputs. These results suggest that the visual styling of rendered text can guide VLMs' interpretation in ways that diverge from human semantic understanding. Project page: https://github.com/KohsukeIde/color‑bias‑vlm

Authors:Mohamed Kotb, Johannes Meier, Christoph Reich, Oussema Dhaouadi, Luis Denninger, Daniel Cremers
Title: MAGneT-3D: Monocular and Domain-Generalizable Temporal 3D Detection
Abstract:
Monocular temporal 3D detection aims to detect objects in 3D, given a monocular video. Query‑based 3D detectors unify detection and cross‑view association, but their learnable queries fit the spatial distribution of the training data (e.g., field‑of‑view). We show that this issue is especially severe when these models are applied to monocular video, hindering generalization to unseen datasets and environments. To address this limitation, we introduce MAGneT‑3D, the first method for domain‑generalized monocular temporal 3D object detection. Instead of relying on static learnable queries, we propose a Domain‑Robust Anchor Generator (DRAG) approach that adaptively derives 3D proposals during inference. To further enable domain generalization, we propose a Temporal Refinement and Identity Merging (TRIM) strategy, reducing dependence on specific 3D proposals. To enable comprehensive domain‑generalization evaluation, we establish a cross‑dataset benchmark spanning nuScenes, Waymo, Lyft, and ONCE. Under zero‑shot domain shifts, MAGneT‑3D outperforms all baselines, improving NDS from 12.1% to 18.6% while also increasing in‑domain accuracy.

Authors:Xukai Wang, Liangqi Li, Zhiyue Xu, Jingang Zhou, Xiaoyu Shi, Jiansheng Cai, Bo Zhang, Zhe Li, Xu-Yao Zhang
Title: APTER: Adaptive Post-Training with Expert-Grounded Rubrics
Abstract:
As large language models enter professional domains, they must satisfy domain constraints, include critical evidence, and provide complete reasoning rather than merely produce fluent responses. Existing post‑training methods often rely on holistic preferences or outcome‑level verification, while recent rubric‑based methods usually generate rubrics independently for each query. In specialized domains, such unconstrained rubrics may omit critical requirements and vary across samples, hindering the diagnosis and targeted repair of persistent capability deficiencies. We propose APTER (Adaptive Post‑Training with Expert‑Grounded Rubrics), a framework that integrates structured domain knowledge into fine‑grained evaluation, optimization, and diagnosis for specialized complex reasoning. First, expert‑grounded rubric construction starts from an expert criteria framework built by domain experts, where each criterion represents a stable professional capability. For each query, APTER selects relevant criteria and instantiates them into query‑level rubrics linked to their source criteria, turning reusable expert criteria into executable query‑level supervision without reference answers. Second, adaptive post‑training uses rubric verdicts as both optimization and criterion‑level diagnostic signals. Aggregating low‑scoring verdicts by criterion ID reveals persistent deficiencies and triggers targeted supervised fine‑tuning updates during reinforcement learning. Experiments on mathematical reasoning and medical question answering show consistent gains across both domains. Across three model generations, APTER improves the mathematics and medical averages over the corresponding base models by up to 15.86 and 8.04 points, respectively. Code and rubric datasets are available at https://github.com/AntDT‑APTER/APTER.

Authors:Nikolai Röhrich, Isabell Hans, Felix Krause, Björn Ommer
Title: Concept Guidance: Precise, Training-Free Latent Control for Text-to-Image Generation
Abstract:
Text‑to‑image diffusion models have two major drawbacks that severely limit their practical utility: (1) standard models lack an intrinsic mechanism for continuous, concept‑specific guidance (e.g., for precisely controlling how aesthetically pleasing an image looks), and (2) they lack reliability for tasks requiring high local coherence (e.g., generating text or human hands). To tackle these issues, we introduce a novel notion of concept‑wise mutual information and find large, concept‑dependent differences between individual layers, demonstrating that the generation of specific structures is localized in distinct parts of the network. We exploit this insight by reinforcing the impact of concept‑relevant layers in Concept Guidance (CoG), a precise, target‑specific guidance method that works for models out‑of‑the‑box without additional training, external models, gradients, or prompt engineering. CoG first quantifies each layer's concept‑specific impact and then guides denoising using a weighted combination of predictions generated with concept‑relevant layers skipped. We demonstrate performance increases across various targets and popular models like PixArt‑alpha, SD3, SD3.5, and FLUX.1‑dev. Code is available at https://github.com/CompVis/concept_guidance

Authors:Yijiang Li, Yijun Liang, Yunjie Tian, Bingyang Wang, Ke Zhang, Zhenfei Yin, Di Fu, Philip Torr, Nuno Vasconcelos
Title: Self-Supervised Visual On-Policy Distillation
Abstract:
Visual on‑policy distillation relies heavily on an informative teacher‑student asymmetry, through either a larger, stronger teacher or privileged supervision, such as reference answers or ground‑truth regions of interest. This raises a fundamental question: where can informative asymmetry come from when nothing privileged is available? We answer this by inverting where the asymmetry comes from. Rather than adding privileged information to the teacher, we subtract information from the student. This asymmetry creates the same effective learning signal for free as a teacher with access to information unavailable to the student, without ground‑truth annotations, rewards, or a separate stronger teacher model. Building on this principle, we introduce Self‑Supervised Visual On‑Policy Distillation (S^2VOPD), a simple yet effective method that constructs on‑policy learning signals from asymmetric augmented views. S^2VOPD distills the teacher's distribution conditioned on the original image on‑policy into the student distribution conditioned on a strongly augmented view of the same image. We systematically explore a broad design space of visual augmentations and uncover that (1) asymmetry matters: all four augmentation families improve performance, while symmetric self‑distillation degrades it; (2) strength matters: performance peaks at a moderate strength; and (3) the gap must remain task‑consistent: augmentations that completely remove the question‑relevant evidence can induce large but uninformative discrepancies. Across six fine‑grained perception benchmarks, S^2VOPD improves Qwen3.5‑4B from 70.7% to 77.4%, above all open‑source models compared, up to Qwen3‑VL at 235B, and surpasses GPT‑5.4. While holding training data the same, it recovers 96% of the improvement achieved by methods with privileged information. Website is at https://williamium3000.github.io/s2vopd

Authors:Wei Zhang, Shengkai Yu, Shiqiang Gong, Qi Zhang, Qiang Li, Qi Wang
Title: HiCo-GS: Hierarchical Context Aggregation and Geometric Consistency for Octree Gaussian Splatting
Abstract:
Octree‑based anchor Gaussian Splatting has emerged as a scalable representation for city‑scale novel view synthesis, where multi‑level anchors adaptively capture scene content from coarse building structures to fine architectural details. However, we identify a fundamental limitation in existing methods: cross‑level feature isolation, where each level's anchor features are optimized independently with no inter‑level communication, causing color drift on building facades and over‑smoothing in textured regions. We present HiCo‑GS, a high‑fidelity reconstruction framework with two complementary modules. Cross‑Level Context Aggregation (CLCA) enables bidirectional hierarchical prior injection by leveraging the octree's spatial containment structure to aggregate per‑level context vectors into parent‑self‑child triplets, fused via a lightweight MLP with residual connection. Coarse‑level structural priors flow down to inform fine‑level anchors, while fine‑level detail statistics feed back to prevent over‑smoothing, at negligible computational overhead. Depth‑Normal Geometric Consistency (DNGC) regularization enforces agreement between rendered normals and depth‑derived normals through an alpha‑weighted consistency loss, complemented by edge‑aware smoothness losses with progressive warmup that exploit the strong planar priors ubiquitous in urban geometry to suppress floating artifacts. We further introduce the China‑Pagoda dataset comprising 8 ancient Chinese pagodas with over 1,200 images each, featuring dense ornamental carvings, curved multi‑layer eaves, and repetitive fine‑grained textures. Extensive experiments on Mill19, UrbanScene3D, MatrixCity, and China‑Pagoda demonstrate that HiCo‑GS achieves state‑of‑the‑art rendering quality and substantially cleaner geometry across real‑world and synthetic urban benchmarks.Code: https://github.com/WZ‑CS/HiCo‑GS.

Authors:Zihong He, Chen Liang, Hai-Ning Liang
Title: AppLooper: An Agentic Application Engineering Loop for Accountable Release with Virtual-User Feedback
Abstract:
Much existing research on coding agents organizes application development as an iterative loop of requirement interpretation, implementation, tool execution, evaluation, and repair. As these loops run longer, requirements may drift; users may lose awareness of the current state and rationale for changes; and generated applications may remain insufficiently grounded in target users' contexts and needs. Application engineering therefore requires a mechanism connecting owner intent, target‑user experience, development changes, and responsibility for release. We present AppLooper, a human‑‑coding‑agent‑‑virtual‑user application engineering loop for accountable release. An application owner confirms frozen requirements, supplies feedback, inspects candidates, and retains final release authority. A development agent produces and revises versioned candidates. A virtual‑user agent cohort executes interface scenarios grounded in target users and contexts of use. Besides, an owner‑intent simulation agent retests only requirements, constraints, and feedback explicitly confirmed by the owner, abstaining when evidence is insufficient. A testing agent performs read‑only developmental checks by reproducing reported failures, running existing regression tests, and exercising the current candidate through its browser interface. The orchestration layer groups the resulting findings and routes them into development revision, targeted retesting, and owner inspection. AppLooper binds requirements, feedback sources, interface targets, development changes, retesting outcomes, owner interactions, and release decisions to specific versions. It thereby extends sustained coding‑agent iteration into a traceable and reviewable lifecycle in which humans retain final responsibility for release. Source code is available at https://github.com/ZihongHe/applooper.

Authors:Jinlong Wang, Yuang Jia, Junhong Lin, Nannan Li, Wei Gao
Title: CoDS: Robust Collaborative Perception via Expert-driven Detection and BEV Segmentation
Abstract:
Collaborative perception breaks through single‑view limitations via multi‑agent information exchange. However, multi‑source noise such as pose errors and communication delays degrades fusion feature quality, constraining perception performance. Joint training of detection and BEV segmentation provides a natural remedy, where segmented road regions help constrain target distributions and detection bounding boxes help recover ambiguous segmentation boundaries. To this end, we propose a robust Collaborative perception framework with expert‑driven Detection and bev Segmentation (CoDS). To address spatial inconsistency in fusion quality, we first introduce the Collaborative Reliability Map (CoRM) to explicitly quantify feature quality distribution. Based on CoRM, we design the Semantic Mixture‑of‑Experts (S‑MoE) module to extract differentiated features for inconsistent feature demands. Finally, to further mitigate feature noise degradation, the Bidirectional Task Complementary Interaction (BTCI) refines task‑aware features through bidirectional injection. Extensive experiments on OPV2V and V2V4Real datasets show that our CoDS surpasses existing baselines on both tasks and maintains stable robustness under multi‑source noise. Code: https://github.com/JinlongW128/CoDS and https://openi.pcl.ac.cn/OpenAIDriving/CoDS.

Authors:Xingyu Zhu, Wenshuo Han, Zhouyu Wang, Yuran Wang, Ruihai Wu, Hao Dong, Fan Tang, Hechang Chen, Hyung Jin Chang, Yixing Gao
Title: FlatLab: A Unified Methodology Framework and Simulation-Based Benchmark for Robotic Manipulation of Flat Objects
Abstract:
Robotic manipulation of flat objects is challenging due to the ungraspable configurations and strong variations in object geometry and material. Existing methods rely on heuristic pre‑manipulation and are often evaluated in closed settings with limited generalization. We propose a unified framework that decouples the manipulation into a strategy generator and an action execution module. The strategy generator predicts appropriate manipulation strategies from object point clouds by learning strategy‑centric, object‑invariant representations via simulated data transformation and contrastive learning. Conditioned on the predicted strategy, the execution module decomposes long‑horizon manipulation into reusable action primitives and dynamically composes them to generate stable trajectories. To enable systematic evaluation, we introduce FlatLab, a comprehensive simulation benchmark for robotic flat object manipulation. FlatLab provides high‑fidelity physical simulation of diverse rigid and deformable flat objects, automated multi‑modal data collection, and standardized task definitions and evaluation protocols. Experiments conducted in FlatLab demonstrate that our approach generalizes effectively to unseen objects and categories, outperforming existing baselines. The project page and the code are provided at https://flatlab‑web.github.io/.

Authors:Tomislav Dobrički, Byung-Woo Hong
Title: Source-Agnostic Image Translation Based on Latent Aware Adaptive Masking
Abstract:
In this work, we propose a source‑agnostic framework that dynamically refines a binary mask throughout the reverse diffusion process by computing the discrepancies of a pretrained diffusion model's prediction for each latent time step. Rather than relying on a fixed threshold, our method introduces a time‑dependent statistical thresholding scheme derived from the empirical mean and standard deviation of prediction discrepancies across the latent noisy images from the target distribution. This allows the mask to adapt to the model's varying predictive confidence at different noise levels, effectively isolating domain‑specific regions while preserving global structural coherence. Experimental results on the AFHQ and Celeba‑HQ datasets demonstrate that our approach outperforms state‑of‑the‑art unsupervised Image‑to‑Image methods in both realism (FID, KID) and faithfulness (SSIM, LPIPS). By requiring only a pretrained model of the target domain, our approach enables precise, automated localization and seamless translation across diverse source distributions without any specialized training. The project source code is available at: https://github.com/dtoma95/PM‑Edit

Authors:Yin Li, Ziyang Hu, Zhiyu Guo, Xiangyu Liu, Wenbin Li, Boo-Ho Yang, Rav Lawana, Ziyue Li, Wei Zeng, Fugee Tsung
Title: HAM-RAG: Hierarchy-Aware Multimodal RAG for Structure-Faithful Interleaved Generation
Abstract:
Existing multimodal RAG methods often flatten structured documents into isolated text and image units, weakening the source organization and local text‑image logic needed for faithful evidence selection and placement. We propose HAM‑RAG, a Hierarchy‑Aware Multimodal RAG framework for structure‑faithful interleaved generation. HAM‑RAG uses document hierarchy as a grounding signal across retrieval and generation, contextualizing textual and visual evidence and preserving source position and local text‑image relations in the prompt. We further introduce HAM‑Bench, covering Wukong, Wiki, arXiv, and Recipe across game walkthroughs, web pages, scientific papers, and step‑wise recipe documents. Across multiple backbones, HAM‑RAG improves the main multimodal average by 17.3% over the strongest non‑hierarchical baseline. On Wukong, HAM‑RAG improves Img‑CBS by 24.2% over the strongest non‑hierarchical baseline, demonstrating substantially better local text‑image alignment. The main experiments and ablation study together demonstrate that document hierarchy is a key grounding signal for faithful image selection, placement, and local text‑image alignment. These findings highlight the value of hierarchy‑aware grounding for reliable multimodal assistants that generate answers faithful to the source organization, procedural structure, and local text‑image evidence of structured documents, such as technical manuals, maintenance guides, and industrial SOPs. The code is available at https://github.com/MCCodeAI/HAM‑RAG.git.

Authors:Wenjin Liu, Haoran Luo, Fayuan Ke, Zhenghong Lin, Yue Lu, Zhe Cui, Anh Tuan Luu, Carl Yang
Title: MMDynOpt-Agent: Dynamic Optimization for Multimodal Large Language Model Reasoning via Reinforcement Learning
Abstract:
Recently, multimodal large language models (MLLMs) have demonstrated strong potential in visual understanding and complex reasoning tasks. However, existing methods often struggle to efficiently transform visual cues from multimodal inputs and the semantics of the question into effective reasoning conditions, thereby limiting the reasoning performance of multimodal large language models. To address this challenge, we propose MMDynOpt‑Agent, which models the dynamic optimization of multimodal reasoning as a Markov decision process via end‑to‑end reinforcement learning. Specifically, a lightweight multimodal agent serves as the decision policy and interacts with the target MLLM as the environment, adaptively steering its reasoning through multi‑turn dynamic optimization prompts. Furthermore, to reduce the cost of multimodal reasoning, a reward mechanism that combines format compliance, answer correctness, and budget awareness is designed to jointly ensure reasoning accuracy and efficiency. MMDynOpt‑Agent is transferable and generalizable, enabling training with one target MLLM and inference‑time transfer to others. Experimental results on fifteen public datasets show MMDynOpt‑Agent achieves strong performance and outperforms baselines. Our project is available at https://github.com/QwenQKing/MMDynOpt‑Agent.

Authors:Yingying Fan, Penghui Du, Leyan Zhu, Runze He, Zimeng Wu, Yuxuan Zhang, Liang Chen, Jiahao Xie, Jiangtang Wang, Shuai Shao, Anchao Yang, Yutong Bai, Yan Wang
Title: MedClaw: Heuristic Agent Harness for Long-Horizon Surgical Video Reasoning
Abstract:
Understanding tens‑of‑minutes surgical videos requires long‑horizon temporal reasoning, answering what happens before, after, or across stages of a procedure by grounding the question in visual evidence spread across time. Existing approaches handle this poorly: a one‑shot vision‑language model (VLM) compresses the whole procedure to fit its context window and loses the detail a "before" or "after" question depends on, while video agents that train the model where to look are data‑hungry and transfer poorly to out‑of‑domain surgery. We build an agent harness that separates reasoning from perception and improves by evolving context rather than optimizing weights. A text‑only orchestrator plans which evidence to gather and issues an auditable sequence of tool calls, while frozen vision‑language sub‑agents execute each call over the pixels, viewing, cropping, inspecting frames, and retrieving external knowledge. We further propose a gradient‑free, reward‑gated Heuristic Skill Distillation loop that mines the agent's own low‑scoring traces and keeps a candidate skill only when it raises a validation reward, yielding reusable retrieval skills, notably directed re‑look. Growing an external skill library rather than tuning weights, the loop adapts from only about 100 labeled examples, far fewer than supervised or reinforcement fine‑tuning requires. To evaluate this agent, we introduce MedClawBench, a de‑leaked, doctor‑grounded benchmark of 1,123 questions over self‑built long neurosurgery recordings and a held‑out public lecture‑video test split. Across both datasets and all four evaluation dimensions, our agent consistently outperforms one‑shot VLMs and general video‑agent frameworks, with the largest gains on the long, out‑of‑domain neurosurgery videos. Project page: https://fyycs.github.io/medclaw/.

Authors:Yongmin Kim, Shota Takashiro, Yusuke Iwasawa, Takeshi Kojima, Yutaka Matsuo
Title: Batch-wise Adaptive Pruning: Periodic Neuron Activation-Aware Weight Pruning for Language Reasoning Model
Abstract:
Large Reasoning Models (LRMs) achieve strong performance on complex tasks through extended chain‑of‑thought generation, but incur substantial computational costs during inference. In production settings, batched inference is essential for high throughput, yet the existing training‑free adaptive pruning methods we evaluate severely degrade in this regime. Because a batch must share a single pruning mask, these methods aggregate activations across samples and then apply threshold‑based selection; the threshold, calibrated offline on unaggregated activations, no longer matches the aggregated distribution, so the realized sparsity ratio drifts and accuracy on reasoning tasks collapses under batched inference. In this work, we propose a training‑free adaptive pruning method designed specifically for batched inference in LRMs, built on two components. First, we replace threshold‑based selection with periodic top‑k selection over the aggregated importance scores, which is unaffected by the shift that aggregation induces in the activation distribution, and which runs selection once per update period rather than at every token, preserving the speedup. Second, based on the observation that important neurons re‑fire periodically during long reasoning generation, we introduce an activation memory that accumulates importance across update phases so that recurring neurons are retained. Experiments on diverse reasoning benchmarks demonstrate that our method outperforms the previous state‑of‑the‑art adaptive pruning method by 39.7 percentage points in average accuracy at batch size 4 with 50% target sparsity on DeepSeek‑R1‑Distill‑Qwen‑7B, and reaches 1.40x speedup over dense inference at 50% actual sparsity.

Authors:Liwei Deng, Jing Jiang, Zhiwei Li, Yang Wang, Guodong Long
Title: Content Depth Matters in Short-Video Recommendation: Rethinking the Attention Economy
Abstract:
Driven by the attention economy, short‑video Recommender Systems (RSs) are primarily optimized to maximize user engagement by promoting videos that capture attention within seconds. These systems inherently favor shallow‑content videos that are effective at attracting immediate attention. However, growing evidence suggests that prolonged exposure to such content may negatively affect users' cognitive engagement and mental well‑being, raising concerns about the long‑term societal impact of the short‑video platform. To tackle this challenge, this paper introduces a new metric, the Content Depth Score (CDS), to quantify the content depth of short videos. CDS measures the extent to which a video is expected to stimulate higher‑order cognitive processes, using a seven‑level scale grounded in established theories of cognitive psychology and learning. As an initial step toward this vision, we present SCOPE‑Bench, the first benchmark for content‑depth evaluation in short‑video recommendation. Built upon a large‑scale open‑source short‑video dataset, SCOPE‑Bench provides CDS annotations for 150K videos, enabling systematic evaluation of RSs from a cognitive‑content perspective. Leveraging SCOPE‑Bench, we evaluate 13 representative RSs and reveal a consistent preference for shallow‑content videos. Moreover, we find that these algorithms recommending cognitively deep content are only marginally better than random selection, highlighting a previously overlooked limitation of existing recommendation objectives. Our code and datasets are available at https://liweidengdavid.github.io/SCOPE‑Bench/.

Authors:John T. Halloran
Title: Nanbeige4.2-3B on Apple Silicon: Fixing Deployment Bugs and Decreasing Looped Transformer Memory Overhead
Abstract:
Nanbeige4.2‑3B is a 3B‑parameter agentic model built around a Looped Transformer (LT) that reuses one stack of layers for a second forward pass, adding effective depth without additional parameters. Evaluated on Apple Silicon (MPS), we identify five independent bugs which prevent the released checkpoint from running via Hugging Face transformers out of the box (including a silently‑zeroed RoPE buffer and calls to removed transformers cache APIs). Furthermore, we show that fixing these bugs is still not sufficient for agentic tasks, due to the LT's layer‑reuse strategy (which effectively doubles peak attention memory) used to achieve parameter efficiency. We thus introduce a chunked‑prefill strategy which alleviates the incurred memory‑capacity penalty, extending allowable context width by 2.7 × on 32~GiB shared memory. However, even with the reduced memory overhead, we show that patches are required to render Nanbeige4.2‑3B usable; resolving both system prompt and MPS‑native memory bugs finally allows reliable evaluation on standard MCP and tool‑calling benchmarks. On a subset of MCPMark, the debugged model completes up to 30% of real agentic tasks (up from the original's 0%), while, on BFCL, it is near‑perfect at single tool calls (yet fails the majority of multi‑tool tests). We release the patched checkpoint, system prompt optimizer, and evaluation harnesses at https://github.com/johnhalloran321/Nanbeige4.2‑3B‑mps‑fix.

Authors:Eslam Eldeeb, Hatim Chergui, Merouane Debbah
Title: XAI-Guided Conservative Decentralized Execution for Offline Multi-Agent Network Slicing
Abstract:
The recent advances toward sixth‑generation (6G) and beyond‑6G networks have accelerated the need for intelligent resource management mechanisms capable of supporting heterogeneous services under shared infrastructures in network slicing. However, resource allocation in network slicing naturally forms a resource‑coupled cooperative optimization problem with competing slice demands. Slices compete for limited resources to minimize individual latencies while coordinating to avoid conflicts and underutilization. Although multi‑agent reinforcement learning (MARL) has shown promising performance in such settings, existing online formulations remain costly, unsafe, and difficult to deploy due to their reliance on environmental interactions and communication among agents. In this work, we present explainable artificial intelligence (XAI)‑guided conservative decentralized execution (X‑CODE). X‑CODE is an explainable offline MARL that operates offline without environmental interaction, nor inter‑agent communication. It exploits explainability‑aware reward shaping to modify the relative preference among joint offline transitions during centralized training to improve decentralized resource‑allocation behavior. In deployment, the agents operate independently without signaling exchange among the agents. Simulation results demonstrate that the proposed approach achieves zero observed resource‑conflict events in the evaluated test episodes while minimizing per‑slice latencies. Moreover, the proposed framework exhibits lower signaling overhead and reduces effective inference latency by 88 % under the considered communication‑delay model compared to the online baselines. Source codes and datasets are available through: https://github.com/Eslam211/xcode‑ran‑slicing.

Authors:Zhiyan Zhang, Zicheng Yan, Jianqi Chen, Peipei Song, Shanshan Wang, Xun Yang
Title: ProFocus: Interpreting Affective Experience in Artistic Images with Progressive Visual Focusing
Abstract:
Interpreting the emotional responses triggered by images is central to achieving emotional intelligence. Compared with natural images, visual art is intentionally created to elicit emotional responses from its viewers through abstract concepts and visual metaphors, making affective interpretation particularly challenging. However, most existing methods rely on general‑purpose visual embeddings (e.g., CLIP), failing to capture the nuanced cues underlying artistic emotion. To address this gap, we propose ProFocus, a novel framework that models affective experience in artistic images via progressive visual focusing. The key idea is to model visual representation learning inspired by a hierarchical cognitive theory of human aesthetic appreciation. Technically, ProFocus contains two core components: a Hierarchical Art Critic (HAC) and a Progressive Hint Fusion (PHF) module. HAC leverages multimodal large language models to generate structured linguistic priors at three cognitive levels‑‑atmospheric style, narrative subjects, and concrete details‑‑thereby translating artistic perception into coherent semantic guidance. Building upon these priors, PHF departs from conventional cross‑modal fusion by sequentially injecting the hierarchical hints into visual features, enabling a progressive focusing process that mirrors human perception. This design allows the model to capture subtle affective cues and produce more faithful explanations. Extensive experiments on the ArtEmis v1.0 and v2.0 datasets demonstrate that ProFocus consistently outperforms state‑of‑the‑art methods in both emotion recognition and affective explanation. Project page: https://github.com/Zhang‑Zhiyan/ProFocus.

Authors:Yongjie Guan
Title: CavityRank: Zero-Extra-Byte Residual Routing for Cuckoo Filters
Abstract:
Near capacity, a cuckoo filter may reject an insertion even though a legal placement still exists: the table remains structurally feasible, but a bounded policy fails to find an augmenting path. Random kick‑out keeps each step cheap but leaves no persistent direction; breadth‑first search recovers direction by expanding a frontier and maintaining table‑scaled state. CavityRank exploits a resource already present in four‑slot packed buckets. Lookup observes only the fingerprint multiset, so query‑equivalent lane orders can encode two comparison bits without widening the 64‑bit bucket or changing the two‑bucket query. The bits form a four‑level ordinal residual rank. Insertion follows a minimum‑rank edge and re‑encodes each modified bucket from its outgoing edges after relocation, propagating the rank actually realized by the packed word. An exact capacity‑four orientation oracle separates structural infeasibility from bounded‑search loss. In a paired 4,096‑bucket XOR16 ladder, CR2 closes 86.47% of Random CF's oracle gap and CavityRank leaves 1.39% of that original gap. A canonical‑tie CR2‑versus‑CavityRank ablation isolates the second implicit bit, which closes 89.95% and 90.69% of CR2's residual gap; the corresponding closures at 65,536 buckets are 83.95% and 84.27%. A separate packed implementation study at 64 MiB and 97.75% load records 42.53 logical reads per insertion, versus 62.67 for explicit labels and 355.74 for depth‑10 BFS, with zero extra bytes per bucket and no table‑scaled workspace. CavityRank therefore occupies a practical design point between unguided eviction and frontier search.

Authors:Yuan-Kang Lee, Kuan-Lin Chen, Chih-Heng Chang, Jian-Jiun Ding
Title: SAFE: Scene-Aware Feature Modulation for Color Constancy with Learned Color Space in Pure-Color Scenes
Abstract:
Color constancy on pure‑color scenes is challenging: when most pixels share a narrow band of hues, every chromaticity‑based cue collapses to a single point and standard estimators become ambiguous. We propose a compact framework that couples two innovations: (i) SAFE, a Scene‑Aware FeaturE modulation network that organizes illumination cues into a structured four‑token representation, which is then selectively reweighted based on scene complexity features; (ii) the Learned Color Space (LCS), a scene‑dependent chromaticity normalization that directly addresses the chromaticity collapse problem for pure‑color scenes. Experiment results show that SAFE consistently improves performance in pure‑color scenes. Compared to the best‑performing baseline in each metric, it reduces the mean angular error by 10%, the best‑25% error by 20%, and the worst‑25% error by 5.8%.

Authors:James K. Wiles
Title: Implementing Computational Law in Wolfram Language for the Governance of Artificial Intelligence
Abstract:
How do we govern AI systems whose reasoning we cannot fully inspect? Governance does not require understanding a system's reasoning. It requires stating what the system is obliged, permitted, and forbidden to do, and checking whether it complied. I present an implementation of Reified Input/Output Logic, the formalism behind the DAPRECO knowledge base, in Wolfram Language: the core I/O axioms, obligations, permissions, constitutive norms, reified eventualities, and temporal operators. I then test whether GPT‑4 can translate English legal statements into the formalism, and report the failures: hallucinated functions, omitted temporal scope, deviation from the formalism, and (in the worst cases) code that runs, reads plausibly, but silently encodes the wrong norm. A case study, an AI guard dog operating under a computational contract, shows how formalized rules can extend from a contract directly into the operational code of an embodied agent, producing symbolic, auditable justifications for its behaviour. I argue that computational law can be used as a governance tool and that a desirable goal would be to formalize the law that can and ought to be programmatically executable.

Authors:Tianyu Fan, Chao Huang
Title: HELIX: Model-Harness Co-evolution for Recursive Self-Improvement
Abstract:
Scaling agent capability has largely focused on improving the model, yet an interactive agent acts through a runtime harness that mediates context, tools, control flow, and stopping. The harness shapes both what a model can accomplish and the trajectories from which it learns. This coupling motivates model‑harness co‑evolution for recursive self‑improvement: build harnesses for a fixed model, update the model from verified sibling trajectories, and rebuild the harnesses as model capabilities change. Realizing this loop requires a controlled way to evolve harnesses while preserving intervention identity and effect. We present HELIX, a source‑traceable substrate for harness evolution. HELIX decomposes agent systems into typed ports, reusable atoms, recipes, product shells, and runtime policies. It makes interventions explicit and auditable while retaining trajectories, test outcomes, and provenance. Harness evolution thus serves two linked roles: improving fixed‑model execution and producing matched successes, regressions, near misses, and alternative solutions as data for subsequent model improvement. We evaluate HELIX in one evolution round on code repair. A 65‑candidate portfolio discovers a fixed harness that improves task coverage by 4.0% over Pi, while the full portfolio exposes up to 58.0% more verified coverage through complementary sibling behavior. Selected candidates are assessed with repeated runs and the SWE‑bench evaluator. A 200‑slot sibling slice yields 438 verified SFT, critic, filter, and preference records. These results show how harness, model, and data form a feedback system: harness evolution expands current capability and creates learning signal for the next model; model updates motivate the next round of harness evolution. HELIX provides an auditable interface for studying this recursive process. Code is available at https://github.com/HKUDS/HELIX.

Authors:Jialong Guo, Ke Liu, Mengxuan Li, Jiajun Bu, Haishuai Wang
Title: CoANeRV: Coordinate-Aware Token-Space Neural Video Representation
Abstract:
Neural representations for videos (NeRV) have shown strong reconstruction fidelity by storing video‑specific information in network weights. However, existing formulations typically require either costly per‑video optimization or video‑specific weight generation, making it difficult to scale to efficient amortized video representation. We propose CoANeRV, a coordinate‑aware token‑space framework that adapts the broader token‑conditioned neural‑field paradigm to amortized video representation. CoANeRV forms compact video tokens in one feed‑forward pass and uses a shared coordinate‑conditioned decoder to reconstruct continuous spatio‑temporal queries, avoiding per‑video decoder optimization or generation while retaining coordinate‑level reconstruction flexibility. To make token‑space reconstruction effective, CoANeRV introduces a coordinate‑aware decoding architecture that aligns spatio‑temporal queries with video tokens through axis‑adaptive positional encoding and temperature‑modulated cross‑attention. Block‑wise coordinate querying further reduces peak attention memory, making high‑resolution reconstruction practical. Experiments on diverse video datasets show that CoANeRV consistently improves reconstruction quality over prior feed‑forward NeRV and INR baselines, reduces peak memory compared with attention‑based coordinate decoders, and provides efficient amortized encoding without per‑video optimization. These results support the proposed video‑specific combination of feed‑forward token formation, spatio‑temporal coordinate retrieval, and memory‑bounded dense querying. The code is available at https://github.com/jialong2023/CoANeRV.

Authors:Yuji Ren, Chenkai Xu, Zhuocheng Gong, Jianguo Li, Zhijie Deng
Title: CForce: Boosting Parallel Decoding for dLLMs via Consistency Forcing
Abstract:
Diffusion large language models (dLLMs) accelerate language generation by predicting multiple masks in a single forward pass. However, existing dLLMs can suffer from unreliable predictions in early denoising stages under aggressive parallelism strategies, leading to errors that can propagate to later stages. To tackle this issue, we present Consistency Forcing (CForce) for dLLMs, a distillation method to force the mask predictions of early stages to align with those of later stages. CForce trains the model on pre‑collected self‑rollout trajectories, thereby improving training‑inference alignment. We introduce Confidence Adaptive KL Divergence as a distillation objective to conjoin the merits of forward and reverse KL. We further provide a theoretical analysis for the consistency objective to explain why CForce can approximately minimize the prediction error of early stages. Critically, the same formulation applies to both mask‑to‑token decoding and edit‑capable decoding; in the edit‑capable case, later token‑to‑token refinements provide additional supervision for earlier masked‑state predictions. Experiments on non‑edit and edit‑capable LLaDA models show improved speed‑quality trade‑offs, especially under high‑parallelism decoding budgets. Code is available at: https://github.com/inclusionAI/dFactory.

Authors:Dinh Tuan Nguyen, Anh Dao, Phuong Nam Dang, Quan-Dung Pham, Tuyen P. Le, Truong Nguyen, Quan Nguyen
Title: OpenBelief-Nav: Evidence-Preserving Object Memory for Open-Vocabulary Language-Guided Navigation
Abstract:
Open‑vocabulary 3D scene graphs provide compact semantic memory for language‑guided navigation, but mapped objects are often exposed through a single fused feature or committed semantic label. Such commitment can remove minority yet task‑relevant hypotheses from the task‑time interface. We present OpenBelief‑Nav, an evidence‑preserving object memory that retains observation‑level phrases, reliability cues, and frame‑mask provenance while maintaining separate aggregate geometric and visual representations. Semantically related phrases are consolidated into a vocabulary‑independent object belief from which task‑specific readouts perform fixed‑vocabulary projection or free‑form retrieval. On five ScanNet200 and eight Replica scenes, full‑belief projection achieves mIoU scores of 0.2742 and 0.2912, compared with 0.2393 and 0.2701 for a matched early‑commit readout. Across 78 HM3D‑YCB navigation trials, consensus and early‑commit retrieval each achieve 60/78 successes, compared with 58/78 for belief‑weighted retrieval and 55/78 for DualMap. Across 20 Unitree G1 runs organized as 10 matched evaluation cases, a correction policy permitting at most two verified candidate attempts improves target‑confirmation success from 6/10 to 8/10 relative to top‑1‑only execution. Code will be released upon acceptance at https://openbelief‑nav.github.io/.

Authors:Stephanie Jarmak
Title: Engineering Reliable Coding Agents: Evaluating and Operating the System Around the Model
Abstract:
AI coding agents are commonly evaluated as models but deployed as systems. Their reliability depends not only on model capability, but on the harness, execution state, retrieval, memory and state management, permissions, review interfaces, and resource allocation. This monograph examines those boundaries and develops a framework for evaluating and operating coding agents reliably. It synthesizes 164 scholarly works, 100 practitioner records, 29 benchmark records, and 17 author‑system case records through a structured multivocal review, targeted update audits, software‑engineering coverage analysis, and distributed‑systems evidence synthesis. Across this evidence, many apparent model failures originate elsewhere in the system, while improvements at one layer often fail to propagate to end‑to‑end outcomes. Evaluation and operation are treated as a dependency chain in which weaknesses in task construction, execution environments, retrieval, state management, verification, or observability can invalidate downstream conclusions. The monograph contributes a versioned catalog of 206 reliability records: 193 gated practices, including 56 developed in depth, plus 13 research leads; an evidence ledger; a framework for dependency and repair asymmetry across the agent lifecycle; measurements and failure cases from operated agent systems; runnable evaluation and reliability protocols; and five reusable agent skills with evidence maps. Together, these provide a system‑level methodology for distinguishing model capability from infrastructure effects, designing defensible evaluations, and building systems that recover safely when components fail. The review is structured rather than exhaustive, evidence strength varies by topic, and results depend on workload and configuration. The methods record which search lanes were executed, which remain unexecuted, and limits on evidence‑grading claims.

Authors:Mohammadreza Narimani, Shreyan Mitra, Parastoo Farajpoor
Title: From crown candidates to neighborhood screening: integrating optical GeoAI and spatial modeling for urban-canopy assessment in Davis, California
Abstract:
Timely urban‑canopy information is essential for linking remote sensing with heat, mobility, and neighborhood planning. We developed an optical GeoAI workflow for Davis, California, using 2022 National Agriculture Imagery Program imagery (0.6 m RGB+NIR). DeepForest generated crown candidates; an NDVI threshold, non‑maximum suppression, and box‑prompted Segment Anything Model (ViT‑B) produced a crown‑anchored canopy surface. Analyses used the 25.92 km2 Census TIGER municipal boundary and a 100 m grid. The workflow retained 11,741 candidate crowns and mapped 2.43 km2 of canopy (9.37% of the city). On the identical extent, 87.8% of mapped canopy pixels and 97.4% of candidate centers agreed with the 2022 USDA/CAL FIRE LiDAR‑assisted canopy product; the optical surface represented 34.2% of the reference canopy area (IoU 0.288; Dice 0.448). Approximately 49% of candidates occurred within 15 m of a road. Canopy was inversely associated with Landsat land‑surface temperature (Spearman rho = ‑0.293; partial rho = ‑0.370 controlling for built probability), and spatial‑lag modeling confirmed clear neighborhood structure. Two transparent attention surfaces combined canopy need with thermal and contextual indicators. The framework provides a reproducible, updateable screening layer that complements structural canopy products and municipal inventories while retaining assumptions, data provenance, and spatial diagnostics for planning interpretation.

Authors:Tyler R. Johnson, Kian Ben-Jacob, Christopher P. Muller, Ramin Bostanabad
Title: On the Brittleness of Maximum Likelihood Estimation for Gaussian Process Hyperparameter Optimization
Abstract:
Machine learning (ML) has become an indispensable part of modern engineering design workflows. A crucial step in training an ML model is the selection of the loss function which can be systematically formulated via various techniques such as maximum likelihood estimation (MLE) and cross‑validation . While MLE is one of the most popular, effective, and intuitive mechanisms for training ML models, it is brittle: if the assumptions underpinning it are not met, the trained ML model may generalize poorly. This brittleness affects even Gaussian processes (GPs) which are widely used in engineering design and are often (incorrectly) presumed to be very robust to overfitting. In this paper, we fundamentally evaluate the brittleness of MLE in the context of training GPs for probabilistic regression or classification tasks. We compare theoretically grounded metrics against MLE and propose practical solutions. Our extensive studies demonstrate the effectiveness of our solutions in downstream design tasks such as Bayesian optimization and provide a blueprint for practitioners to build accurate and robust GPs that can even outperform tabular foundation models in terms of prediction accuracy, uncertainty quantification, and inference cost. Our contributions are publicly available via GitHub at https://github.com/Bostanabad‑Research‑Group/GP‑vs‑TabPFN‑vs‑GPyTorch.

Authors:Ruogu Chen, Jie Han
Title: PPAPlace: Differentiable Cross-Stage Objectives for Chip Placement Optimization
Abstract:
Macro placement significantly affects a chip's post‑route performance, power, and area (PPA). Most placement methods optimize half‑perimeter wirelength (HPWL) as the primary objective. However, recent benchmarking shows a near‑zero correlation between HPWL and post‑route timing metrics such as the worst negative slack (WNS) and total negative slack (TNS). As a result, all six evaluated artificial intelligence (AI) placers degraded PPA relative to the hierarchical baseline. Recent efforts have tried to train cross‑stage predictors to close this gap. However, existing methods focus on macro‑only representations and use pre‑route metrics as training labels. A label fidelity study of ten circuits at four design flow stages reveals that HPWL and pre‑route timing poorly reflect final post‑route timing rankings. In contrast, post‑global‑routing achieves the best balance between final timing fidelity and label generation cost‑effectiveness. Based on this finding, PPAPlace is a timing‑driven differentiable surrogate predicting post‑route PPA from macro and standard‑cell placements. The surrogate is a dual‑stream predictor that combines graph attention over the chip netlist with spatial convolution over the placement grid. It is trained on post‑global‑routing labels. The predicted WNS and TNS gradients flow end‑to‑end back to cell coordinates. PPAPlace exploits these gradients in two ways: as a co‑objective injected into an analytical placer's optimization loop (PPAPlace‑CoOpt), and as a post‑placement refinement step that adjusts macro positions via projected gradient descent (PPAPlace‑Refine). On five ChiPBench test circuits excluded from training, PPAPlace improves average WNS and TNS by 22% and 51% over the hierarchical baseline while preserving power and routability, using the same predictor without test‑circuit retraining. Code is available at https://github.com/ValleyC/PPAPlace.

Authors:Illia Volkov, Nikita Kisel, Tetiana Mishkina, Klara Janouskova, Jiri Matas
Title: Doomed to Re-Annotate, Forever: The ImageNet Story
Abstract:
Top‑1 accuracy on ImageNet‑1k remains the most commonly reported metric in visual recognition. Quality issues with the dataset have been repeatedly reported, yet the original 2012 noisy labels are still predominantly used. The paper presents a comprehensive effort, which goes well beyond prior correction attempts, towards obtaining accurate and complete ImageNet‑1k validation set annotations. The result, ReImageNet, includes multilabel correction, object localization, revised class definitions, and semantic attributes (text‑recognition, rendition, reflection, crowd, dominant). The reannotation reveals that approximately 12% of the original ImageNet‑1k labels are incorrect, 33.3% of images are multilabel and 3.8% contain no object from an ImageNet‑1k class. With the new labels, top‑1 accuracy increases by up to 1.2% for supervised models and by 5‑6% for MLLMs. We argue that annotation at ImageNet scale cannot realistically be completed in one pass, as errors and definitional issues are discovered only through annotating, and we build our pipeline around repeated refinement and error checking. We observed that human and LLM collaboration with appropriate tooling represents the current quality ceiling for annotation at this scale. ImageNet‑1k issues propagate into its derivative test sets, indicating that the problem is structural rather than specific to any single benchmark. All annotations, class definitions, guidelines, and analysis code have been publicly released. Project page: https://vrg.fel.cvut.cz/reimagenet Annotations: https://huggingface.co/datasets/vrg‑prague/ReImageNet Code: https://github.com/klarajanouskova/ImageNet

Authors:Sebastian Doerrich, Andreas Franz Schwab, Francesco Di Salvo, Shyam Nandan Rai, Hanh Huyen My Nguyen, Christian Ledig
Title: TRUE-Colon: Exposing a Consistent Transfer Asymmetry in Real-Time Polyp Detection
Abstract:
Computer‑aided detection (CADe) systems for colonoscopy promise to reduce clinical miss rates, yet reliable real‑world deployment remains elusive. This translational gap stems in part from a structural flaw in model development: the reliance on curated datasets that under‑represent the long negative stretches and procedure‑related artifacts characteristic of routine examinations. Training and evaluating architectures strictly on these lesion‑centric benchmarks creates an illusion of success, since such benchmarks cannot capture clinically crucial metrics. To expose this gap, we establish TRUE‑Colon, a standardized benchmarking protocol that measures key deployment characteristics alongside localization accuracy, and evaluate four real‑time architectures (Faster R‑CNN, YOLOv8, YOLOv11, RT‑DETR) across curated benchmarks (SUN, PICCOLO) and 60 unedited, full‑length procedures (REAL‑Colon). We observe a consistent transfer asymmetry: models trained strictly on curated clips suffer a severe performance collapse when evaluated on full procedures, whereas procedure‑trained models substantially improve rejection of non‑polyp content on REAL‑Colon, and largely retain their accuracy on curated benchmarks. Beyond transferability, we find that the Transformer detector attains the strongest sensitivity and the earliest, most persistent detections, while the convolutional detectors stay competitive at a higher throughput. Together, these results indicate that both training and benchmarking for deployable CADe should shift from curated, lesion‑centric clips toward full‑procedure data and deployment‑relevant operating points. Source code is available at https://github.com/sdoerrich97/true‑colon.

Authors:Rafi Ibn Sultan, Hui Zhu, Chengyin Li, Dongxiao Zhu
Title: MedPlex: Deep Vision-Language Co-Adaptation for Clinically Grounded Medical Segmentation
Abstract:
Medical image segmentation is still largely treated as a vision‑only problem, although clinical interpretation often relies on textual knowledge of anatomy, location, appearance, and surrounding context. Existing text‑guided segmentation methods within the Vision‑Language Model (VLM) paradigm often use language only as a late conditioning signal, limiting its influence on visual representation learning. We introduce MedPlex (Medical Plexus of Vision and Language), an end‑to‑end VLM framework that makes text guidance a continuous, clinically grounded component of segmentation learning. Through Bi‑Fusion (Bidirectional Fusion), visual and textual representations evolve jointly across the encoding hierarchy. MedPlex further introduces class‑level and region‑level concept alignment to organize the shared representation at complementary granularities. Class‑level alignment anchors each anatomical target to an aggregated clinical concept profile, while region‑level alignment preserves individual concepts, such as shape, location, appearance, and texture, through class‑specific visual evidence. In this way, language provides structured supervision throughout the encoder rather than serving only as a late‑stage cue. MedPlex achieves state‑of‑the‑art performance across CT and MR benchmarks for multi‑organ, cardiac substructure, and tumor segmentation, including settings with real free‑text clinical supervision. Code: https://github.com/rafiibnsultan/MedPlex.

Authors:James Adam
Title: Ontology-Grounded Project Memory for Coding Agents
Abstract:
Coding agents have become the primary means of generating new code in many software projects, and the resulting velocity of changes makes keeping track of the reasons behind those changes challenging. This paper introduces MOOSEDev, a system designed to give coding agents structured, ontology‑grounded project memory. The system captures architectural decisions, lessons, constraints, and rationales in a knowledge graph exposed to agents via a Model Context Protocol (MCP) interface. Records carry lifecycle status, provenance, and supersession links, queryable via MOOSE, a proprietary neurosymbolic engine that treats the symbolic layer as the primary reasoning substrate. We compared MOOSEDev against a production vector‑memory tool on a neutral public corpus of 835 typed records. MOOSEDev returned the expected answer set essentially in full (0.98‑1.00) on supersession, set‑completeness, and negation questions, whereas the baseline's top‑k retrieval surfaced between 6% and 27%. Conversely, relevance recall and token cost were largely equivalent between the two systems. We also describe a temporal commit‑history bootstrap of our own codebase, a pre‑registered live trial, and lessons learned.

Authors:Jia Sheng, Yiwei Lu
Title: No Universal Signal Predicts Sample-Level LLM Regression under Version Updates
Abstract:
Frontier LLMs are updated frequently and typically outperform their predecessors in aggregate. But aggregate gains say little about individual samples: an update can still cause sample‑level regression, where a response correct under the old model becomes incorrect under the new one. This paper studies how to predict such regressions from signals available at inference time. We compare single‑model signals (confidence, logit margin, attention entropy) against cross‑version signals (output KL divergence, likelihood drift, token‑level KL, representation drift) under a unified added‑value test that isolates each signal's gain over a confidence baseline. Across six benchmarks in three task families (multiple‑choice question answering, or MCQ; math reasoning; code generation) and six model update pairs, we find that (1) signal effectiveness is task‑dependent: confidence is strongest on MCQ and simpler math, while likelihood/KL signals give the most frequent gains on harder math and code; (2) no signal is universally best across model updates either; and (3) some cross‑version signals stay informative even when confidence fails, including without labels, which supports a proof‑of‑concept selective fallback that routes high‑risk samples back to the old model. Practitioners can use these task‑level patterns to choose which regression signal to trust for a given update. Code is available at https://github.com/jiashengsally/llm‑regression‑signals.

Authors:Lunjie Zhu, Xingtong Ge, Fangyu Lin, Yi Zhang, Zhening Liu, Mengfei Li, Yumeng Zhang, Guanglu Song, Yu Liu, Jun Zhang
Title: Omni-LiveAvatar: Minute-Level Real-Time Streaming Joint Audio-Video Avatar Generation
Abstract:
Joint audio‑video generative models serve as foundation for immersive and interactive digital‑human generation. Nevertheless, most existing models rely on bidirectional attention and multi‑step denoising and can generate only short clips, making them unsuitable for real‑time interaction over extended durations. We present Omni‑LiveAvatar, the first framework for minute‑level, real‑time streaming joint audio‑video avatar generation. Specifically, we propose (1) a progressive autoregressive distillation pipeline that transfers a large bidirectional joint audio‑video diffusion model into a few‑step autoregressive generator without auxiliary stabilization mechanisms; (2) a synchronized audio‑video long‑short‑term memory that preserves global consistency under a bounded memory budget; and (3) a hierarchical rolling prompt planning strategy that enables coherent semantic evolution and seamless prompt transitions. Extensive experiments show that Omni‑LiveAvatar generates high‑quality, synchronized minute‑level avatars in real time. In terms of speed, it achieves a 33× generation speedup over its teacher, LTX‑2, on a single NVIDIA H200 GPU; in terms of generation quality, it outperforms accelerated baselines across visual quality, audio quality, cross‑modal synchronization, and human fidelity. Our code is available at https://github.com/Aoko955/Omni‑LiveAvatar.

Authors:John Myron Uy
Title: Hard Cases, Bad Labels: Testing Error Exposure and Error Location in Uncertainty Sampling Under Bounded Label Noise
Abstract:
Active learning can reduce labeling cost by selecting informative examples, but the most uncertain examples may also be the hardest to label correctly. This study tests whether uncertainty sampling fails because it acquires more corrupted labels or because errors concentrated in difficult regions are especially harmful. Margin‑based uncertainty sampling is compared with random sampling under clean labels, random classification noise (RCN), and bounded difficulty‑dependent noise on three public binary tabular datasets. The design uses 100 paired seeds, nine expected noise rates from 0 to 0.30, annotation budgets from 20 to 120, and logistic regression with regularization re‑selected by cross‑validation at every budget. An exposure‑matched RCN control aligns mean final acquired corruption, while a clean‑label extension reaches budget 400. Under clean labels, uncertainty sampling improved normalized balanced‑accuracy area under the learning curve by 1.09 to 1.77 percentage points on all datasets. Difficulty‑dependent noise reduced this advantage more than RCN at six of eight rates on Breast Cancer Wisconsin, but at no tested rate on Banknote Authentication or MAGIC Gamma Telescope. Exposure‑matched analyses found no corrected evidence for a universal additional penalty from structured error location. On clean MAGIC data, uncertainty sampling improved balanced accuracy while reducing average precision and true‑positive rate at fixed false‑positive rates. Thus, uncertainty sampling was label‑efficient, but its apparent robustness depended on dataset, budget, noise structure, and evaluation metric.

Authors:Jiaqian Yu, Chen Jason Zhang, Haoyang Li, Guoqiong Ivanka Huang
Title: Beyond Simplification: DFT-GEN for Fidelity-Preserving Visual Accessibility in Dyslexia-Friendly Educational Texts
Abstract:
Dense educational texts impose avoidable reading friction on people with dyslexia, yet generic simplification can delete terminology, task constraints, or source evidence that readers still need. Stakeholder interviews with dyslexic adults and specialists reveal a core tension: reduced burden must not compromise information fidelity. We present DFT‑GEN, a stakeholder‑informed text transformation framework for content‑heavy educational materials. Its central contribution is not a generic LLM refinement loop, but a dyslexia‑specific accessibility layer that combines protected‑span preservation with a deterministic Dyslexia Accessibility Controller (DAC) for rendered visual organization. DAC converts stakeholder and expert preferences into reproducible controls for visual‑unit length, chunk spacing, source/task separation, highlighting budget, and reviewable risk flags. We therefore separate evaluation into DCFI, a fidelity‑safety diagnostic, and B‑DVAS‑VL, a rendered visual‑accessibility diagnostic. On 2,280 bilingual exam‑style items, DFT‑GEN preserves task‑critical information while improving visual accessibility: it wins 93% in English and 64% in Chinese of B‑DVAS‑VL pairwise judgments against same‑backbone controls, and in a controlled pilot with dyslexic adult readers it preserves answerability while reducing effort.

Authors:Rachid Arezki
Title: BCMT: Blockwise Causal Memory Transformer
Abstract:
Transformer architectures rely on dense self‑attention to model long‑range dependencies, but this mechanism exhibits quadratic complexity with respect to sequence length. We introduce BCMT (Blockwise Causal Memory Transformer), an architecture for long‑context language modeling that decouples local token interactions from global context propagation. Dense causal self‑attention is applied independently within local blocks, while each block produces an adaptive summary aggregated through an exponential causal memory. This memory is subsequently injected back into the token representations, enabling efficient propagation of long‑range contextual information without relying on explicit global attention. Unlike standard Transformers and recurrent memory architectures, BCMT maintains neither dense interactions between distant tokens nor learned memory states. Its memory mechanism is fully parallelizable and remains compatible with standard implementations of dense self‑attention. Experiments on language modeling with context lengths of up to 1024 tokens show that BCMT achieves validation performance comparable to that of Dense Transformers while significantly improving training throughput and reducing memory consumption. An ablation study further confirms that these improvements arise from the proposed memory mechanism. These results demonstrate that an exponential causal memory constructed from block summaries provides an effective alternative to dense global attention mechanisms for long‑context language modeling.

Authors:Bo Jin, Qiang Jiao, Xin Tong
Title: Agentao: A Governed Local-First Runtime for Tool-Using LLM Agents
Abstract:
LLM agents increasingly operate as execution systems that invoke tools, modify local state, use persistent memory, and interact with external protocols. These capabilities make agents useful, but they also introduce risks related to over‑privileged actions, weak auditability, prompt injection, tool poisoning, and uncontrolled side effects. This paper presents Agentao, a governed local‑first runtime for tool‑using LLM agents. Agentao separates model‑generated action proposals from host‑authorized execution through a layered architecture consisting of host‑facing surfaces, a host contract, a runtime core, a permission‑mediated tool system, and supporting subsystems for memory, replay, plugins, skills, sub‑agents, and protocol integration. We describe the motivation, threat model, design goals, governance model, execution pipeline, and structured event interface of the system. Agentao does not provide formal safety guarantees; rather, it demonstrates how permissions, state, protocol boundaries, and execution traces can be made explicit runtime abstractions for building agents that are more governable, inspectable, and suitable for host‑controlled local environments. The code is publicly available at https://github.com/jin‑bo/agentao.

Authors:Dayuan Zhao, Shengcao Cao, Yu-Xiong Wang, Liang-Yan Gui
Title: Think in Latent, Explain in Language: Self-Explainable Latent Reasoning
Abstract:
Latent reasoning has emerged as a powerful alternative to text‑based Chain‑of‑Thought (CoT), offering significant gains in computational efficiency by compressing verbose reasoning into compact embeddings. However, compressing reasoning into the latent space renders the thinking opaque, hindering its interpretability. Current methods present a stark trade‑off: they either function as unexplainable ''black boxes'' (e.g., Coconut), where the latent reasoning is not human‑readable, or rely on separate post‑hoc decoders for explainability (e.g., Heima), introducing architectural overhead and decoupling the explanation from the actual reasoning process. In this work, we present a unified framework for Self‑Explainable Latent Reasoning (SELR) that trains a single model to perform efficient and inherently explainable latent reasoning. Our core contribution is a novel multi‑task training objective that optimizes for two goals simultaneously: (1) an Answer Loss that optimizes the latent reasoning trajectory to produce accurate final answers, and (2) a CoT Loss that explicitly trains the same model to decode its own latent representations back into human‑understandable reasoning steps. This design ensures that generated latent representations are both task‑effective and semantically interpretable, eliminating the need for external decoders. We validate the effectiveness of SELR on both Large Language Models (LLMs) and Vision‑Language Models (VLMs), demonstrating that SELR achieves superior token efficiency and accuracy compared to baselines, while uniquely providing self‑contained explainability without auxiliary models. Project page is available at https://jasondayuan.github.io/SELR/.

Authors:Pengcheng Xu
Title: Does a Language Server Save Tokens for Coding Agents? A Measurement Methodology and Preliminary Study
Abstract:
Coding agents spend most of their context budget on retrieval. Lexical retrieval (grep) is universal, instant, and zero‑setup, but noisy: it cannot tell a definition from a call from a comment. Semantic retrieval via the Language Server Protocol (LSP) is precise and typed, but needs a running, indexed server and pays a per‑symbol round‑trip. The claim that semantic retrieval is more token‑efficient is, we find, asserted almost everywhere and measured almost nowhere: no public source isolates the LSP‑vs‑lexical token delta for an agent at equal task‑success. This paper formalizes the question with one metric (tokens‑to‑success), specifies a five‑arm ablation isolating semantic retrieval from confounds, maps three pre‑stated failure modes onto measurable variables, and reports a preliminary study (Python and TypeScript repos; Claude Opus 4.8, Sonnet 4.6, Haiku 4.5). The answer is conditional and usually negative. On symbol‑named localization the LSP costs tokens (+6% to +118%) and the agent ignores it when free. On reference‑completeness it buys precision but not token savings and cannot raise the recall ceiling set by agent thoroughness; it saves tokens only for the weakest model. Tool choice is task‑dependent: models default to grep on localization (0‑6% semantic use) but reach for the LSP about half the time on reference tasks, unprompted. On edits scored by real test execution the gap is starkest: grep solves multi‑file renames perfectly, a location‑only LSP fails three‑quarters of them by missing a call site, and even a complete, index‑warmed, text‑enriched LSP (each reference's line inline, as production LSP‑MCP servers do) recovers most of the gap but cannot close it, since a rename must touch comments and strings that semantic references exclude. The implication is not LSP‑always but an adaptive router keyed on task class, model capability, and lexical noise.

Authors:Pengrui Han, Jacob Andreas, Evelina Fedorenko, Andrea Gregor de Varda
Title: Modular Cognitive Architecture Emerges in Large Language Models
Abstract:
The human brain exhibits a striking degree of functional specialization, with distinct networks supporting language, formal reasoning, reasoning about other minds, and reasoning about the physical world. Is this modular organization a fundamental principle of how intelligent systems must be built, or an evolutionary accident specific to biological brains? Here, we test whether a similar organization emerges in Large Language Models‑‑another class of intelligent systems created through a very different optimization process. Using circuit analyses across N=46 tasks spanning four cognitive domains (language, formal reasoning, social reasoning, physical reasoning), we find that LLMs develop a modular architecture that mirrors the human brain: tasks drawing on the same network in humans recruit overlapping neurons in LLMs, whereas tasks drawing on different networks recruit distinct neurons. The convergent emergence of modularity in brains and neural networks suggests that it may be a fundamental property of intelligent systems.

Authors:Yaxin Luo, Haobin Jiang, Jialv Zou, Xu Huang, Wenhao Yan, Haodong Li, Zhengrong Yue, Jing Li, Xiaofu Chen, Xiaohan Zhao, Jiacheng Liu, Jiacheng Cui, Zhiqiang Shen, Xiaotong Li
Title: AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design
Abstract:
Transforming multimodal sources into condensed and structured media outputs can be fundamentally conceptualized as a long‑horizon agentic process centered on a model‑harness system. While an ideal harness system should align with human design priors and accumulate reusable experience through empirical exploration to drive recursive self‑improvement, existing paradigms remain static and fall short of this capability. In this paper, we present AutoDesign, a framework that aligns with human design priors, where a meta‑harness optimizer guides a code agent to recursively improve harness based on rollout feedback. To instantiate and evaluate this framework, we focus on the academic paper‑to‑poster generation task and introduce PosterBench, comprising a 100‑paper Main Track spanning five disciplines and PosterBench‑mini, a shared 10‑paper subset for controlled evaluation. On the PosterBench Main Track, AutoDesign achieves the highest score of 78.32, surpassing the closed‑source commercial system Claude Design by 7.45 points. Across seven controlled code‑agent‑model configurations, integrating the learned DesignHarness consistently improves performance, increasing the average PosterBench Score from 54.99 to 67.39 (+12.4%). In a fully autonomous long‑horizon loop, it executes 253 tool calls and 11 editing turns within 40 minutes for under 3, reaching average conference‑poster quality in human evaluation. A system‑blind human study further demonstrates that AutoDesign achieves the highest human preference among evaluated systems.

Authors:Minghui Guo, Shengqiong Wu, Hao Fei
Title: V-RAE: Rethinking Video Latent Spaces for Generation
Abstract:
Latent video generation relies on autoencoders to define a compact space in which generative models operate. Although video autoencoder architectures have evolved substantially, their latent spaces are still optimized primarily for pixel‑level reconstruction and provide limited high‑level semantic organization. A reconstruction‑optimal latent space, however, need not be well suited to generative modeling. We propose V‑RAE, a video representation autoencoder that builds compact generative latents on top of frozen vision foundation model representations. A lightweight temporal pooling module removes temporal redundancy while preserving semantic structure, and a video decoder reconstructs continuous motion from the compressed features. We evaluate V‑RAE with four representative frozen encoders on video reconstruction, semantic probing, and class‑conditional generation. V‑RAE achieves 2.13 rFVD on K600, outperforming all evaluated large‑scale pretrained video VAEs. Its latents retain substantially more semantic information than conventional video tokenizer latents. Under matched generation settings, our best variant achieves gFVD scores of 117.86 and 19.16 on UCF101 and K600, respectively, while converging up to 6x faster. We further show that reconstruction quality alone is insufficient to characterize generative utility and introduce tFVD, a temporal‑coherence diagnostic that correlates more reliably with downstream generation quality. Beyond video generation, V‑RAE also improves future video prediction on Cityscapes over the Wan 2.2 VAE latent space under matched prediction settings. Taken together, the experiments show that frozen semantic representations can support video reconstruction, generation, and predictive modeling. The project page: https://v‑rae.github.io/.

Authors:Kaixin Ding, Xi Chen, Minghong Cai, Zhiyuan Xu, Yiyang Wang, Yuxiang Lu, Junyi Li, Shuyang Chen, Yuan Gao, Xin Tao, Pengfei Wan, Hengshuang Zhao
Title: PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives
Abstract:
Video world models simulate future states conditioned on current observations and user actions. Recent systems have demonstrated impressive video consistency and action controllability over long sequences. However, fairly comparing these interactive models remains challenging. In practice, a human player typically evaluates a world model by pursuing long‑horizon objectives through interaction. For example, a user may turn around 360 degrees to see whether the environment remains consistent, or walk into the water and inspect whether realistic water ripples are generated. The action sequence required to achieve the same objective may vary substantially between models, making fixed action‑conditioned evaluation unsuitable for cross‑model comparison. To address this, we employ multi‑modal Agent Players to interact with world models toward specified long‑horizon objectives. Building on this paradigm, we introduce PlayWorld, a benchmark providing 171 scenarios, each with a specified objective. To evaluate performance thoroughly, we assess models along four core dimensions: geometry consistency, interaction fidelity, out‑of‑sight evolution, and insight evolution. In addition, we incorporate basic ability metrics for video quality and controllability. Experiments across nine state‑of‑the‑art world models reveal that current models remain unreliable on long‑horizon interactive objectives, particularly in maintaining spatial consistency and persistent state evolution. Code and data are available at https://github.com/kxding/PlayWorld.

Authors:Sikuang Li, Chen Yang, Jiemin Fang, Jiazhong Cen, Yuhe Wei, Jichen Pang, Wei Shen, Qi Tian
Title: SCULPT: Subtractive Composition for 3D Part Generation
Abstract:
Part‑aware 3D generation aims to create digital assets that are coherent as complete objects while exposing structural parts for editing, material assignment, animation, and reuse. Existing methods impose this structure outside the native generation loop: segmentation‑based methods partition an already generated shape, while additive methods synthesize parts from predefined layouts, boxes, or tokens and then reconcile them into a whole. The former preserves the generated geometry but fixes the object before part boundaries are determined; the latter exposes part cardinality but often leaves shared boundaries vulnerable to gaps, interpenetrations, and material discontinuities. In this paper, we propose SCULPT, a framework that addresses these challenges through subtractive composition. Given a complete object represented in a structured 3D latent space, SCULPT iteratively applies a joint split predictor to generate one extracted part together with the remaining object. The predictor performs a coupled denoising process conditioned on both the image and the current 3D state, so the extracted part and updated remainder are generated together rather than reconciled after generation. The joint split predictor processes both outputs on the union of their native sparse 3D supports, allowing neighboring supports to overlap rather than imposing a disjoint voxel partition. The rollout ends when the remainder support becomes empty or reaches a fixed safety cap, allowing the number of generated parts to adapt to each object within that bound. Extensive experiments demonstrate state‑of‑the‑art geometry on PartObjaverse while preserving strong complete‑object reconstruction after part assembly. Results on four dataset images, one text‑to‑image‑generated input, and one real‑world photograph further show fine‑grained textured part decomposition beyond the benchmark.

Authors:Zhe Ye, Hantao Lou, Yuechun Sun, Peiyang Song, Zhengxu Yan, Timothe Kasriel, Qingyang Zhang, Kaiyu Yang, Soonho Kong, Jingxuan He, Dawn Song
Title: Vero: Can AI Agents Build Formally Verified Software Repositories?
Abstract:
AI agents are increasingly used for programming, but do not provide any guarantee on the correctness of generated code. Verified code generation, in which an agent produces both an implementation and a machine‑checked proof of its specification, offers a stronger path toward trustworthy AI‑generated software. Existing benchmarks in this direction either focus on individual functions or only evaluate proof generation with provided implementations. It is still an open question whether agents can make coherent implementation and proof choices across real multi‑module codebases. To bridge this gap, we introduce Vero, the first benchmark to evaluate joint implementation and proof synthesis at the repository level. Vero contains 43 multi‑module instances sourced from real‑world repositories spanning Python, Dafny, Verus, and Coq, and covering diverse domains from cryptographic protocols to distributed systems. Each instance consists of a multi‑module Lean 4 repository with predetermined API interfaces, manually curated formal specifications, and reference implementations, supporting both proof‑only and code‑and‑proof evaluation modes. To improve benchmark reliability, Vero also includes an audit mechanism where agents are allowed to formally prove unsatisfiability of provided specification or incorrectness of reference code, which surfaces and corrects latent code and specification errors during curation. We evaluate frontier coding‑agent configurations with Lean toolchain access. The strongest agent fully solves only 27 of 43 instances and closes no specifications on the hardest repositories. Vero provides a concrete testbed for measuring progress toward repository‑scale verified software synthesis, where current agents still fall short. We release the benchmark, curation pipeline, and evaluation harness at https://github.com/sunblaze‑ucb/vero.

Authors:DreamX Team, Rui Chen, Xiangxiang Chu, Geng Li, Jifan Li, Qingfeng Shi, Datao Tang, Jing Tang, Jun Wang, Pengfei Zhang
Title: DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation
Abstract:
We present DreamX‑Phi 1.0, an action‑conditioned video world model for robotic manipulation that, given an observed frame, a language instruction, and a prescribed action sequence comprising end‑effector poses and gripper states, predicts the resulting future observations. Yet realism alone does not guarantee faithfulness: a convincing rollout can still move the wrong arm or lose the manipulated object. To ensure the prediction respects each arm's commanded path, we inject per‑arm \mathrmSE(3) transformations into attention via PRoPE‑style geometric encoding, preserving arm identity and rigid‑motion structure. Action control alone does not fully constrain scene geometry or the evolution of small manipulated objects. We therefore add a lightweight depth branch for scene‑level geometry and use SAM3 masks with a frozen V‑JEPA teacher to maintain object consistency throughout grasping. We further distill the multi‑step generator into a few‑step student via distribution‑matching distillation for efficient deployment. At the time of writing, \model achieves first place on Track~1 and second place on Track~2 of the WorldArena~2.0 Challenge. Our model and code will be publicly available.

Authors:Saisha Shetty, Satvik Tripathi, Austin Lin, Colin Zhao, Theodore Kim, Don Enwerem, Jacinta Arnold, Shahriar Faghani, Tessa S Cook
Title: MARC v1: An Open-Source Multi-Agent Framework for Clinical AI Reasoning and Coordination
Abstract:
We present Multi‑Agent Reasoning and Coordination (MARC), an open‑source framework that replaces monolithic LLM prompting with deterministic multi‑agent orchestration for clinical reasoning. MARC coordinates role‑specialized agents for extraction, reasoning, answer generation, and evaluation, with explicit context passing and traceable intermediate outputs, enabling stage‑wise failure attribution. We additionally introduce a Decomposer module that generates task‑specific agent prompts from a plain‑language description, eliminating manual prompt engineering. The framework supports both API‑based and local CPU‑compatible deployments and is entirely configurable via YAML, without code modifications. MARC is designed to be model‑agnostic, interpretable, and accessible to clinical domain experts without programming expertise. The full framework is available at https://github.com/Penn‑RAIL/MARC‑v1.

Authors:Rafal Robert Karpinski, Fethiye Irmak Dogan, Nikhil Churamani, Yiming Luo, Maartje M. A. de Graaf, Davide Dell'Anna, Hatice Gunes
Title: Mind the Context: Continual Learning of Socially Appropriate Robot Actions via Environmental-Social Disentanglement
Abstract:
Social robots are expected to operate across diverse environments, where similar arrangements can imply different socially appropriate actions, e.g., starting a conversation may be acceptable in a crowded home but disruptive in an office meeting. Because such norms and environments cannot all be anticipated in advance, robots require continual learning (CL) to adapt from sequential experience while retaining previously acquired knowledge. Prior work has studied CL for generating socially appropriate robot actions, but it has not addressed domain‑incremental settings in which the robot incrementally encounters diverse contexts (e.g., living room, meeting room, office, hallway), where both environmental (e.g., whether the space is open or cluttered with furniture) and social cues (e.g., how people or other agents are positioned around the robot) jointly shape the appropriateness of robot actions. We address this gap with the Explicit Disentanglement Dual‑Branch (EDD) framework. EDD explicitly separates environmental and social‑agent related knowledge and uses replay‑based rehearsal to mitigate forgetting while learning the appropriateness of robot actions (e.g., cleaning, serving, starting a conversation) across several indoor domains. Experiments show that EDD outperforms several state‑of‑the‑art baselines, and ablation studies further evaluate different disentanglement strategies and the sensitivity to domain ordering. Our code is publicly available at https://github.com/Cambridge‑AFAR/Mind‑the‑Context.git.

Authors:Zongyun Zhang, Jiacheng Ruan, Xian Gao, Ruizhu Zhou, Lingcheng Meng, Lining Hu, Ting Liu, Yuzhuo Fu
Title: Edit2TikZ: A Comprehensive and Challenging Benchmark for Scientific Figure Editing with TikZ
Abstract:
Although multimodal large language models (MLLMs) have shown substantial potential in visual understanding and graphic code generation, editing scientific figures through code presents a greater challenge: a model must jointly recover visual structure, ground the requested change, generate compilable code, and preserve all unrelated content. While existing TikZ benchmarks mainly focus on figure reconstruction and generation, few systematically evaluate instruction‑guided scientific figure editing with compilable code. We introduce Edit2TikZ, a comprehensive benchmark for scientific figure editing tasks, featuring 1,548 diverse and high‑quality samples. Edit2TikZ combines real‑world and controlled synthetic edit cases, supports both textual and visual localization request, and contains multi‑step editing, each with step‑level annotations. We further construct a human‑aligned evaluation framework to measure whether a requested edit is completed while irrelevant content is preserved. Utilizing Edit2TikZ, we evaluate 14 mainstream MLLMs and find that current systems remain unreliable: on average, proprietary models achieve a compilation success rate of merely 75% and remain limited in both figure restoration and edit correctness, while compact models below 9B struggle further with instruction following and complete figure generation. Therefore, we build a mixed training set TikZEditMix and adopt reconstruction‑then‑editing curriculum learning for compact models. On Qwen3.5‑4B, this training improves the compilation success rate from 45.35% to 83.40% and yields an average improvement of 18.7 points across our proposed evaluation metrics. The code and data will be released at https://github.com/Solunny/Edit2TikZ.

Authors:Zheyu Zhuang, Ruiyu Wang, Nick Heppert, Johannes Fabian Hahn, Abhinav Valada, Florian T. Pokorny, Danica Kragic
Title: Attention from Action, for Action: Emergent Visual Bottlenecks for Policy Learning
Abstract:
Visual bottlenecks that focus policy inputs on regions of interest (ROIs) can improve data‑efficient visuomotor learning by separating where to look from how to act. Many ROI interfaces rely on external spatial labels, such as gaze, object classes, or affordance annotations. Label‑free alternatives often derive crops from trajectories by detecting gripper or motion events and centering a fixed crop at the projected end‑effector. Such action‑derived crops are useful spatial priors that require no additional labels, but they encode fixed choices about event timing, proxy points, and crop scale. When the visual evidence needed for control lies away from the end‑effector or changes continuously with task progress, these crops can become misaligned. We propose Seeker, a task‑ and state‑conditioned readout that learns attention from action. Starting from frozen DINOv3 features, Seeker iteratively updates a query with gathered visual evidence, producing progression‑aware ROIs solely from action supervision. The learned ROI serves as a spatial interface for RGB cropping, mask‑guided background augmentation, and point‑cloud filtering. In simulation and the real world, Seeker improves data efficiency and robustness over no‑crop, augmentation, and action‑derived crop baselines. On real robots, Seeker raises average in‑domain success from the best baseline's 48.3% to 76.7% and success under lighting/background shifts from 20.0% to 60.0%.

Authors:Hmrishav Bandyopadhyay, Xuanchi Ren, Zijian Huang, Jay Zhangjie Wu, Tianshi Cao, Ruilong Li, Bryan Chu, Sanja Fidler, Yi-Zhe Song, Zian Wang
Title: Context-Matched Distillation: Teacher Causality for Autoregressive Video Distillation
Abstract:
Interactive autoregressive video generation demands both low‑latency rollouts and precise online control. Few‑step distillation accelerates generation by reducing denoising steps, while online control imposes a causal constraint: frames and blocks should depend on history and controls available during generation. Existing video distribution matching distillation (DMD) pipelines, however, often supervise causal few‑step students using bidirectional teachers that score complete clips. The score for a target can therefore depend on future frames and controls that were unavailable when the student generated it, misaligning teacher supervision with the student's causal information set. We introduce Context‑Matched Distillation (CMD), a causal DMD framework that aligns teacher supervision with the information available when each target is generated. CMD replaces bidirectional full‑clip scoring with a causal teacher that evaluates each target without access to future frames or controls. The same causal teacher initializes the few‑step student, establishing a consistent causal formulation across teacher training, student distillation, and inference. Beyond aligning the temporal information boundary, Prefix Scoring matches supervision to the student's realized rollout context by evaluating each target under the cached student‑generated prefix that produced it. Prefix Corruption further stabilizes training by perturbing unreliable prefixes produced early in training while preserving this target‑context alignment. With a simple causal formulation, CMD naturally extends to frame‑wise and chunk‑wise generation, long video distillation, and camera‑conditioned distillation. Experiments demonstrate state‑of‑the‑art aggregate performance among autoregressive methods on both short‑ and long‑video benchmarks, together with substantially improved adherence to time‑varying camera controls.

Authors:Christos Chatzisavvas, Stelios Alvanos, Efstratios Politis, Panagiotis Rigas, Thomas Pappas, Ioannis Giannoukos, Nikolaos Mitianoudis, Agata Ulanowska, Katarzyna Żebrowska, Nazarij Buławka, Christina Margariti, George Pavlidis, Chairi Kiourt, Anestis Koutsoudis, Vassilis Katsouros, George Ioannakis
Title: AmalthAI: An Open-Source Computer Vision Platform for Cultural Heritage
Abstract:
Computer vision (CV) and machine learning (ML) offer new tools for cultural heritage (CH) artifact analysis, but the CV/ML pipeline remains largely inaccessible to CH domain experts, who lack the background to configure, train, or assess models. We present AmalthAI, an open‑source CV platform that bridges this gap, enabling non‑ML CH experts to independently produce and validate archaeologically meaningful findings. The interface covers dataset management, training, and inference for classification, segmentation, and object detection, with Kubeflow and Katib handling scalable training and hyperparameter search. Grad‑CAM localizes the image region behind a prediction, and a vision‑language model (VLM) adds a text description of it for expert review. Since archaeological data is often state‑owned or rights‑encumbered and cannot leave institutional custody, AmalthAI's self‑hostable deployment ensures sensitive data is kept within premises. We test the platform on an archaeological use case built on a custom dataset of clay textile imprints, where CH experts trained and validated segmentation, and classification models for hypothesis testing. We provide the implementation code at https://github.com/TEXTaiLES/AmalthAI.

Authors:Valentin Noël
Title: Where You Measure Decides What You Measure: Position Selection in Ablation-Based SAE Evaluation
Abstract:
Sparse autoencoders are meant to name the things a language model computes, and the usual way to check that a latent matters is to switch it off and see what changes. But a latent fires at many tokens, and the effect has to be measured at one of them. The convention is to measure where the latent fires hardest. That choice is almost never reported, and it is not made by the experimenter: it is made by the dictionary under evaluation. Change the dictionary and the measurement moves to a different token. We show this is not a detail. Take two sparse autoencoders released by Google for the same model and match their latents by decoder similarity: even among the pairs the two dictionaries encode almost identically, they pick different tokens for a large share of them. Two dictionaries compared under the usual protocol are therefore very often compared at different places. To separate the convention from the dictionaries we train six autoencoders from one initialisation, differing only in fitting choices, so that a latent means the same thing in each. Most of the variance such a comparison reads as "these dictionaries disagree about this latent" turns out to be the position instead: it falls from 7.6% and 11.9% of variance to near zero once every dictionary is measured at the same token. More evaluation data does not rescue it. Across a sixteenfold range of corpus sizes the dictionaries agree less about where to measure, not more, so the problem grows with scale. The correction is one line of evaluation code. We give the protocol an ablation‑based causal number must report to be comparable across papers, and an audit of five published papers against it. In short: a causal number reported without its position describes the token it was taken at as much as the latent it was taken from.

Authors:Valentin Noël
Title: A Probe Direction Is a Property of Its Prompt
Abstract:
A model that behaves differently when it senses it is being tested would undermine the evaluations we rely on, so recent work has sought to read that sense directly from a model's activations. The standard instrument contrasts activations on prompts that announce an evaluation against prompts that do not, and reports how well the resulting direction separates held‑out cases. That number is then compared across models and correlated with scale. We observe that the instrument has a free parameter its readings do not disclose: "a prompt that announces an evaluation" is not a prompt but a choice among many, and nothing in the method fixes which. Holding the task text fixed and varying only that choice, we find that the reported score, and even the direction in which it trends with model size, follows the prompt rather than the model; two published studies that disagree about the sign of that trend are both reproducible from a single design, by choice of prompt alone. Treating the prompt as a facet of a measurement design rather than an implementation detail, we find the model under study accounts for a small share of the variance in the number reported about it, and most of the rest lies in how each model responds to each prompt: collecting more evaluation items cannot repair the measurement, while varying prompts can. A further check finds that the split these probes are scored on is largely separable from surface form alone, so a direction carrying no information about evaluation at all still reproduces a substantial fraction of each published score. We conclude that a single‑prompt design cannot support comparison between models, and we give the number of prompts a defensible comparison requires.

Authors:Alexander Bräuer, Benjamin Cauchi, Nils Strodthoff
Title: Foundation models for movement data: Are they ready for prime-time?
Abstract:
Foundation models (FMs) trained on large‑scale accelerometer data have been proposed as general‑purpose feature extractors for health monitoring, but systematic evidence of their advantages is lacking. We present the first comprehensive evaluation of four open‑source accelerometer FMs against supervised baselines covering 19 tasks across the domains of activity recognition including activities of daily living, clinical monitoring, and physiological inference. We find task‑dependent performance results: supervised models remain competitive with FMs on human action recognition (HAR), with no consistent advantage for either, while selected FMs lead on fall and stress detection and are the most robust to sensor‑placement variation. As frozen feature extractors, FMs are strongest for demographic inference, whereas sleep staging performance remains near chance level for all models. The internal FM representations show strong similarity across layers, highlighting potential for future FM improvements. Linear and frozen probing reveals that UniMTS provides the strongest representations and is the only FM that surpasses the supervised baselines without finetuning. Concept discovery analysis shows all models capture high‑intensity activities clearly but struggle with sedentary, complex or ambiguous activities. We provide scenario‑based deployment recommendations. Furthermore, we identify FM‑derived activity profile inference‑moving beyond fixed category classification‑as a promising research direction.

Authors:Wafa Al Ghallabi, Ritesh Thawkar, Sara Ghaboura, Omkar Thawakar, Numan Saeed, Dana Al Nuaimi, Ajnas Alkatheeri, Salman Khan, Fahad Shahbaz Khan
Title: How Good are Foundation Models in Longitudinal MRI Disease Progression Reasoning?
Abstract:
Magnetic Resonance Imaging (MRI) interpretation is fundamental to clinical decision‑making, requiring radiologists to integrate multi‑view anatomical planes across sequential timepoints while precisely localizing interval changes. However, existing vision‑language benchmarks remain confined to single‑timepoint, single‑view interpretation, failing to capture the temporal‑spatial reasoning essential to radiologic practice. We introduce the Time‑Aware Multi‑View MRI Benchmark, an evaluation framework unifying multi‑view anatomical input, temporal reasoning across longitudinal scans, and structured localization guidance. The benchmark comprises 3,920 expert‑verified question‑answer pairs derived from 890 patients across over 3,200 longitudinal MRI timepoints, drawn from seven clinical cohorts covering glioblastoma, neurodegeneration, vestibular schwannoma, and brain metastases, in open‑ended, multiple‑choice, and binary formats, requiring models to identify anatomical regions of maximal change, characterize progression across sequences and views, and provide structured guidance specifying boundaries, imaging features, and confounders. Experiments across 16 vision‑language models reveal moderate temporal alignment but systematic failure on change direction recognition and volumetric quantification, while multi‑view inputs improve spatial localization yet degrade temporal reasoning in compact architectures. Our benchmark provides a systematic framework for evaluating progression tracking, interval change localization, and temporal ordering, which are essential for clinical deployment. Code, evaluation splits, and the dataset are available at: https://github.com/wafaAlghallabi/Time‑Aware‑MRI.

Authors:Koen P. de Vries, Xavier Alameda-Pineda, Estefanía Talavera, Stéphane Lathuilière
Title: Reasoning for Social Audio-Visual Question Answering: Where Do We Stand?
Abstract:
Training Multimodal Large Language Models for audio‑visual social understanding is a crucial step toward embodied social intelligence. Chain‑of‑thought (CoT) reasoning has become the dominant approach, with HumanOmniV2 and its IntentBench benchmark as a prominent reference point. In this context, we report three findings. First, IntentBench is highly noisy: ~7% of questions are broken and ~23% are trivially answerable without the video input. We remove the affected questions and release Intentbench‑Prime. Second, current reasoning approaches are expensive and surprisingly ineffective. A simple Vanilla SFT baseline matches or outperforms existing reasoning methods across three benchmarks at a fraction of the cost, establishing it as an essential baseline for evaluating novel fine‑tuning techniques. Third, our analysis reveals that substantial priors can be learned solely from the text modality and that using a textual caption instead of the video yields performance on par with Vanilla SFT. These surprising findings reveal the limitations of current MLLMs when it comes to social understanding. IntentBench‑Prime, Vanilla SFT model, and code are publicly available.

Authors:Weimeng Luo
Title: When Should Multi-Round RAG Stop? Structured Stopping Judgments and Retrieval Reduction in Search-R1
Abstract:
Multi‑round retrieval‑augmented generation (RAG) must decide when to stop searching as evidence accumulates. Because the deployed policy is determined by the first STOP on each trajectory, this is a sequential selection problem rather than an independent state‑classification task. We adapt S2G‑RAG's structured sufficiency‑and‑gap judgment to a frozen Search‑R1 pipeline and train a Qwen3.5‑2B judge on 3,009 states from 900 disjoint HotpotQA questions. Search‑R1's reasoner, retriever, corpus, prompt, and search budget remain unchanged, while the judge checkpoint and stopping threshold are selected on grouped validation and frozen before confirmatory evaluation. On the confirmatory test set, the resulting policy reduces retrieval calls by 77 (3.70%) relative to Native Search‑R1, while Official Exact Match decreases by 0.625 percentage points. Thus, the trained S2G‑style structured judge reduces retrieval while broadly preserving answer accuracy. The result does not imply unchanged or improved accuracy, safe stopping, or lower total inference cost.

Authors:Peng Ling, Yingda Yin, Lingting Zhu, Weikai Chen, Shengju Qian, Zeyu Hu, Xin Wang, Wenming Yang
Title: CoverPrune: Coverage-Driven Token Pruning for 3D VLMs via Optimal Transport
Abstract:
While 3D Vision‑Language Models (3D VLMs) have demonstrated remarkable spatial reasoning capabilities, they suffer from massive visual token counts that create severe computational bottlenecks during inference. Existing token pruning methods primarily rely on diversity‑based selection, discarding similar tokens to maximize dispersion. However, in 3D environments, this approach frequently drops representative prototype tokens in favor of outliers, breaking the multi‑view consistencies and geometric structures essential for spatial reasoning. In this paper, we propose a paradigm shift for 3D VLM token pruning: from maximizing diversity to preserving visual evidence coverage. We introduce CoverPrune, a training‑free framework that formulates inference‑time token pruning as an Optimal Transport (OT) problem. To overcome the intractable combinatorial subset selection inherent in this formulation, we design the Feature‑Spatial‑Temporal (FST) transport cost and target capacity, along with an efficient Spatial‑Guided Greedy Selection (SGS) algorithm to approximate the OT objective. Furthermore, we propose CoverPrune‑Lite, an accelerated variant utilizing spatially structured local matching for minimal overhead. Extensive experiments across multiple 3D visual‑spatial reasoning benchmarks demonstrate that our methods achieve state‑of‑the‑art token efficiency, maintaining robust reasoning performance even under highly aggressive pruning budgets. Visit our project website at https://github.com/Brucess/CoverPrune.

Authors:Riya Deepak Shet, Le Zhang
Title: Reliability analysis for BraTS-GoAT segmentation: a controlled robustness study of deep-ensemble uncertainty
Abstract:
Deep networks segment brain tumours accurately in‑distribution, but can fail silently when the input differs from their training data. That risk is central to clinical deployment and is the premise of the BraTS‑GoAT generalizability task. We ask not only how well a model segments, but whether its uncertainty knows when it is wrong. On BraTS‑GoAT (Task 3) we train a 5‑fold cross‑validated nnU‑Net baseline (one held‑out prediction per case) and a 3‑seed deep ensemble. Both are evaluated for calibration and error detection on a per‑region relevant mask, aggregated per case. In‑distribution the 3‑seed ensemble improves modestly over the already strong single model on the same held‑out split, with the clearest gain in calibration. The separation appears under shift. In a controlled robustness study using graded synthetic corruptions as a proxy for acquisition shift, the single model's confidence stays flat while its accuracy and calibration degrade. Inter‑member disagreement instead rises steeply, about a quarter to a third above the clean condition, several times the single model's response. On the official validation leaderboard the 5‑fold ensemble of those folds attains whole‑tumour Dice 0.87. The generalization gap is concentrated on the harder regions, with a characteristic failure of missing small, satellite lesions on unseen cohorts. In the synthetic study, disagreement among the 3‑seed members is a more sensitive case‑level indicator of acquisition shift than single‑model confidence. Its per‑voxel error localisation weakens as severity grows. The contribution is a rigorous, honest reliability comparison rather than a claim that any one uncertainty method dominates.

Authors:Tianshuo Zhang, Xianglei Xing, Wenzhe Zhai, Jia Gao, He Cao
Title: History-informed Lagrangian Neural Networks
Abstract:
Forecasting the long‑horizon evolution of mechanical systems from position‑only observations is a pivotal yet difficult task, as hidden velocities and trajectory‑specific physical properties must be inferred simultaneously. Although physics‑guided neural networks like Lagrangian Neural Networks (LNNs) guarantee physical plausibility, they generally require complete state inputs and lack adaptability to changing system parameters. To break these limitations, we introduce History‑informed Lagrangian Neural Networks (HiLNN). Grounded in the insight that temporal position sequences implicitly encode underlying dynamics, HiLNN employs a recurrent encoder to extract a latent context from history. This context not only reconstructs the unobserved initial velocity but also adaptively modulates the mass matrix, potential energy, and damping coefficients of a structured Lagrangian system. By leveraging a differentiable RK4 rollout scheme, the entire pipeline is optimized end‑to‑end under multi‑step trajectory supervision and energy‑consistency regularization. Empirical evaluations across conservative, dissipative, and heterogeneous variable‑parameter systems show that HiLNN delivers superior long‑term prediction accuracy and maintains precise energy profiles compared to state‑of‑the‑art baselines. The source code is publicly available at https://github.com/yingtian22/History‑informed‑LNN.

Authors:Yuheng Huang, Jianlang Chen, Jiayang Song, Hua Qi, Aza Kai, Vincent Markert, Edison Marrese-Taylor, Jianjun Zhao, Lei Ma
Title: NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video
Abstract:
Long‑form video understanding encompasses tasks that go beyond retrieving isolated events, including tracking an evolving narrative and interpreting social meaning that may remain implicit. However, existing benchmarks rarely evaluate these capabilities jointly, particularly in high‑context, non‑English media. To address this gap, we introduce NARU, a benchmark designed to evaluate Narrative evolution and Reasoning on cultural Understanding in Japanese long‑form video. NARU consists of 1,481 questions grounded in 155 videos totaling 146.8 hours, spanning four narrative and five cultural dimensions. To construct the benchmark at this scale, we propose a hierarchical memory‑based annotation pipeline that transforms raw video into structured event, narrative, and cultural annotations, then generates questions via task‑oriented synthesis and iterative shortcut removal. The construction process includes two native‑speaker verification stages involving 68 annotators. Evaluations across eight model configurations reveal substantial limitations in both long‑range narrative integration and culturally grounded reasoning. By exposing these persistent gaps, NARU offers a systematic testing ground for developing MLLMs capable of reliably interpreting long‑form, high‑context video.

Authors:Minkyoung Kim, Beakcheol Jang
Title: Chance-constrained selection of sequential intervention strategies from counterfactual estimates
Abstract:
Many operational decisions are sequences of interventions under a cumulative resource limit, such as a maintenance schedule within a crew‑hour budget. Choosing among them calls for the outcome and the cumulative cost each would produce, counterfactual quantities identified from observational data. Two strategies with the same expected cost can exceed the budget at very different rates, so constraining the mean does not bound how often an overrun occurs. Prior two‑step architectures, recently extended to continuous doses, constrain the mean cost rather than its tail and allocate at a single decision point. Methods that do bound a cost tail take its distribution from a specified model rather than identifying it from data. We present a predict‑then‑optimize framework. In the prediction step, any estimator returning an outcome value and a cost distribution supplies what the decision rule consumes, so the predictor is interchangeable. In the optimization step, a chance‑constrained selection over a finite candidate set bounds the probability that the cumulative cost exceeds the budget. That tail does not decompose across stages, so each strategy is scored whole. Sweeping the tolerated violation probability traces a safety‑utility frontier, and distribution‑free finite‑sample bounds cover violation and outcome shortfall. Four of five environments, spanning clinical treatment and equipment maintenance, supply exact counterfactual ground truth; the fifth carries real outcomes from a digital‑health micro‑randomized trial. Across them, the rule holds the budget where a point‑estimate rule overruns it, at an outcome cost the frontier makes explicit. All code is available at https://github.com/mfriendly/counterfactual‑chance‑selection

Authors:Jiazi Bu, Pengyang Ling, Yujie Zhou, Yibin Wang, Yuhang Zang, Xuanlang Dai, Shengyuan Ding, Tianyi Wei, Xiaohang Zhan, Jiaqi Wang, Tong Wu, Dahua Lin, Xingang Pan
Title: HPSD: Hybrid-Policy Self-Distillation for Text-Image-to-Video Diffusion Models
Abstract:
Text‑Image‑to‑Video (TI2V) models are an emerging unified architecture, where a single model simultaneously supports text‑to‑video (T2V) and image‑to‑video (I2V) generation. Given a high‑quality first frame or a detailed textual prompt, TI2V models unlock substantially better visual quality than their T2V mode, raising a natural question: can the capability elicited by such privileged conditions be internalized into the model's own base generation ability? A common approach toward this goal is model self‑distillation. However, the most straightforward solution, supervised fine‑tuning, follows an off‑policy strategy: its supervision is confined to teacher‑generated endpoints from a fixed offline distribution rather than student‑visited states, lacking precise correction tailored to the evolving policy. Recent on‑policy distillation methods instead suffer from condition‑state mismatch, where supervision is steered toward the given first frame instead of the student's actual content, misleading the correction. To achieve self‑distillation that absorbs the teacher's privileged prior while retaining precise policy correction, in this work, we propose Hybrid‑Policy Self‑Distillation (HPSD), a novel self‑distillation framework where a single TI2V model acts as both teacher and student under different conditions: the teacher operates in TI2V mode with a high‑quality first frame and an enhanced prompt, while the student runs in the base T2V mode with only the vanilla prompt. Specifically, the student inherits off‑policy teacher trajectory points as anchors, locally refines them toward its own policy, and finally receives velocity‑level supervision on these self‑generated roll‑outs. Extensive experiments demonstrate that HPSD significantly improves T2V performance while also delivering notable TI2V gains, effectively strengthening the model's base generation ability.

Authors:Nhan Phan, Ilona Lähteenmäki, Anna von Zansen, Olli-Pekka Pauna, Yaroslav Getman, Tamás Grósz, Mikko Kurimo
Title: CASA: Content-Acoustic Speaking Assessment with Speech Encoder and Large Language Model
Abstract:
Research on automatic speaking assessment (ASA) has increasingly adopted multimodal speech large language models to assess learners' speaking performance. However, existing studies provide limited analysis of how acoustic and content information contribute to predictions and how stable the resulting performance is. We propose CASA, a simpler architecture combining Whisper‑medium and Qwen3.5‑2B that achieves state‑of‑the‑art performance while providing a more interpretable separation between speech delivery and content. On the Speak & Improve Corpus 2025, CASA achieves a root mean square error (RMSE) of 0.358, improving on the previous best RMSE while using approximately half the estimated inference parameters. The general‑purpose architecture is designed for adaptation to other ASA corpora without structural changes and relies on three handcrafted fluency features. Through ablations and repeated runs, we analyze the individual and complementary contributions of acoustic and content information, examine performance variability, and demonstrate the potential of large language model reasoning for training‑free content validation.

Authors:Yakun Huo, Yingquan Wang, Yangyang Liu, Tianyu Yan, Yunzhi Zhuge, Pingping Zhang, Huchuan Lu
Title: Paths: Prompt-aware Spatio-temporal Transformer with Hierarchical Multi-modal Fusion for RGB-Event Video Person Re-Identification
Abstract:
RGB‑Event Video Person Re‑Identification (RE‑VReID) aims to retrieve specific person across non‑overlapping cameras with complementary RGB videos and event streams. However, existing methods often decouple spatial and temporal modeling, which limits their interaction. In addition, global‑level RGB‑Event fusion fails to fully exploit fine‑grained discriminative cues. To address these issues, we propose Paths, a unified framework with spatio‑temporal modeling and hierarchical multi‑modal fusion for RE‑VReID. Specifically, we first design a Memory‑Augmented Backbone (MAB) to maintain modality‑specific identity prototypes for stable intra‑modal representation learning. Then, we propose a Prompt‑aware Spatio‑temporal Transformer (PST) to jointly model spatial and temporal cues within a unified Transformer. Finally, we introduce a Hierarchical Multi‑modal Fusion (HMF) to integrate RGB and event features at global and local levels. With these modules, our framework can learn robust and discriminative representations for RE‑VReID. Extensive experiments on three public RE‑VReID benchmarks including EvReID, MARS and iLIDS‑VID, demonstrate the effectiveness of our proposed method. The code is available at https://github.com/Reflection0427/Paths.

Authors:Jinhyung Bae
Title: Sampling Luck Masquerades as Allocation Gain: Auditing Test-Time Budget Allocation for Neural Combinatorial Optimization
Abstract:
Neural combinatorial optimization (NCO) solvers report the best of many sampled solutions per instance, and the sample count is, by convention, identical for every instance. Whether a non‑uniform allocation of a fixed total budget would buy anything has not been measured. We measure it, and we audit the measurement itself. First, on in‑distribution workloads the allocation headroom is not detectable. Across three pretrained solvers (POMO, AM, SymNCO) on uniform TSP‑100, an oracle allocation computed and evaluated on the same stored samples reports a 2.2‑2.6% gain with intervals excluding zero; measured out of sample the same gain is indistinguishable from zero (0.457, 0.015, ‑0.512 percent). Following the customary in‑sample procedure, all three solvers would have supported a published 2%‑level gain that does not exist. We calibrate this bias against an instance‑wise null in which the true gain is zero by construction; over the ranges we test it does not shrink with more samples or more instances. Second, the same correction that removes the phantom gains preserves a real one. Under distribution shift (a workload mixing uniform and clustered instances), a pre‑registered confirmatory experiment finds that allocation guided by held‑out sample statistics improves best‑of‑k by 11.5% (AM, primary endpoint; 95% CI [7.4, 19.7]) and 12.0% (SymNCO, replication) at equal evaluation budget, with the signal‑acquisition cost not charged; a pre‑registered negative control (POMO, an order of magnitude more robust to shift) shows ‑0.3% [‑0.7, 0.24]. The gain exceeds a frozen distribution‑label baseline by 4.2 points [1.9, 7.7]. An exploratory policy charging a 20‑sample probe against the same budget retains 3.4% (AM) and 4.6% (SymNCO). We give a correction procedure and a reporting checklist, and release all data, code, and the pre‑registration record.

Authors:Haotian Liu, Yang Liu, Guoying Zhao, Xiaobai Li
Title: Learning Unified Video and Image Representation for Video Face Forgery Detection
Abstract:
Face forgery detection is crucial for preserving the security and integrity of facial data given the rapid developments in face manipulation techniques and deep generative models. Existing methods for video face forgery detection typically assume that all frames in a forged video are manipulated, while detecting partially forged videos that contain only a subset of altered frames remains challenging. To address this issue, we propose a novel framework, UVIF, that utilizes additional annotated images to provide fine‑grained supervision for detecting partial forgeries in videos. UVIF employs a unified encoder and a multi‑task learning paradigm to jointly model facial videos and images for boosted video face forgery detection. A 2D backbone with temporal fusion modules is employed as the unified encoder. A pseudo labeling process is designed for video frames to bridge their representations with those of static images. A video‑oriented feature alignment strategy is further introduced to reduce the distribution gap between videos and images. Extensive experiments on benchmark datasets demonstrate the effectiveness of our framework, which outperforms state‑of‑theart methods in detecting partially forged videos while introducing no additional computational overhead. Our code is available at https://github.com/haotianll/UVIF.

Authors:Dechen Zhang, Xuan Tang, Xinxiang Yin, Xingwu Chen, Jian Qian, Difan Zou
Title: VALG: An Agentic System for ML Theory Research
Abstract:
Machine learning theory studies learning procedures through mathematical setups in which the data model, training protocol, oracle access, loss, metric, and randomness define the phenomenon that a theorem is meant to explain. Solving an open problem therefore requires the problem formulation, theorem target, and proof mechanism to be developed in concert. Researchers formulate hypotheses, test them through preliminary theoretical or empirical analysis, and refine both assumptions and proofs. We investigate whether this process can be organized as an autonomous agentic workflow for ML theory research. We develop VALG, an agentic system that combines multi‑level Verification, Adaptive formulation of Learning‑theory problems, and Graph‑structured proof development. Within each source‑relative theorem branch, VALG maintains a fixed mathematical specification, checks the theorem‑level composition of a typed proof‑dependency graph, and constructs and reviews local proofs in dependency order. When a proof attempt fails, VALG identifies whether the obstruction lies in a derivation, the proof structure, or the theorem formulation and routes the next attempt accordingly. Formulation‑level obstructions initiate an explicitly related variant or relaxation, preserving the mathematical relation between the resulting theorem and the source problem. We evaluate VALG on nine subproblems from five COLT 2026 open problems. Two runs produce internally finalized theorem candidates that match the scope of their source briefs; the remaining seven yield restricted‑method results, special cases, or conditional theorems. These case studies show how VALG keeps source‑scope matches, relaxations, conditional results, and blocked attempts mathematically distinct. VALG is open source at https://github.com/DechenZhang/VALG‑ML‑Theory‑Agent.

Authors:Jie Li, Chenxin Jia, Jinliang Shen, Cunzhuang Liu, Ruiyi Ding, Jianwen Xian, Kang He, Chengru Song
Title: TEMPO: Makespan-Aware Expert-Parallel Load Balancing Across Memory- and Compute-Bound Regimes
Abstract:
In expert‑parallel (EP) MoE serving, every layer synchronizes at the slowest GPU. Dispatchers balance token counts (EPLB, LPLB, UltraEP) or activated‑expert counts (METRO), assuming expert time is linear in one. Measurements on two datacenter GPU generations show it is neither: below \nstar\!\approx\!156‑‑168 tokens, HBM weight streaming dominates‑‑‑cost attaches to \emphactivated replicas, not tokens; above it, grouped GEMM rounds tokens to 128‑tile M‑tiles, so \emphsplitting an expert adds padded compute. A max‑affine profile t=\max(a+bG,\,c+βN) captures both regimes. Realistic decode batches hold hot experts in the linear regime and cold in the flat \emphsimultaneously; recorded batches show proxy dispatches differ by 1.4‑‑1.6× in modeled block time (p95 up to 1.7×), and \emphwhich proxy wins flips with the regime. We formalize per‑batch dispatch as a fixed‑charge makespan problem‑‑‑NP‑hard on two fully replicated GPUs, polynomial in degenerate limits‑‑‑and present \sys, a makespan‑aware dispatcher solving it in milliseconds off the critical path; its SGLang integration runs out‑of‑process and fuses dispatch with count collection into one in‑graph kernel. Anchored by an 8‑GPU Testbed~A microbenchmark, \sys stays within 1% of the best fixed baseline everywhere and wins by up to 15.5% where regimes mix. End‑to‑end on Testbed~B, Qwen3‑235B (inside the win region) gains 4‑‑6% throughput and cuts p99 latency by ~15.6%; DeepSeek‑V3 (outside, communication‑dominated) shows only mechanism cost. A phase diagram, not a universal win, is the claim: it predicts both outcomes before deployment.

Authors:Yi Shi, Huichao Xie, Yuqing Wang, Mingyu Wang, Kaihui Yang, Yu Liu, Ruitao Lu, Lizhe Li, Junwei Han, Dingwen Zhang
Title: P2Fusion: Prompt-based Progressive Infrared-Visible Image Fusion via Dual-Prior Distillation
Abstract:
Infrared‑visible image fusion (IVIF) is pivotal for multimodal perception, yet reconciling the inherent information disparity between thermal and textural features remains a fundamental challenge. Existing prior‑guided methods often rely on static constraints that induce optimization conflicts or utilize extrinsic semantic priors from large‑scale foundation models (e.g., CLIP/DINO), which frequently fail to exploit the intrinsic modality characteristics essential for high‑fidelity fusion. To address these issues, we propose P2Fusion, a prior‑guided distillation‑based framework that reformulates IVIF via dual intrinsic prompts. Instead of imposing hard‑coded penalties, we distill image‑intrinsic priors, thermal saliency and spatial quality, into learnable dynamic regulators. Specifically, a Teach‑to‑Fuse mechanism provides dual‑granularity progressive guidance, coupled with a Gated Dynamic Expert Recalibration (GDER) module for decoupled feature refinement. This design enables the network to adaptively mediate modal competition through expert specialization. Extensive experiments demonstrate that P2Fusion achieves state‑of‑the‑art performance across five mainstream datasets. Notably, our framework demonstrates consistent performance advantages in fusion quality, achieving state‑of‑the‑art results in 14 out of 20 key evaluation metrics across 5 benchmarks. Furthermore, it effectively contributes to the robustness of downstream perception, such as +3.2% mAP on MSRS, +0.5% mAP on M3FD and +0.9% mAP on DroneVehicle for object detection. Our code will be available at https://github.com/YiShi99/P2Fusion

Authors:Michael Chesser, Paul Quirk, Douglas Cooke, Guy Farrelly, Surya Nepal, Damith C. Ranasinghe
Title: InSPECtor: Improving SLEIGH Processor Specification Veracity via Proxy
Abstract:
Processor specifications underpin critical security and program‑ analysis tools such as disassemblers, decompilers, and emulators, yet, their correctness is rarely examined. Errors in specifications distort program behaviour, obscure vulnerabilities, and enable analysis‑evasion techniques. Validating processor specifications is a non‑trivial task. Our study is a significant undertaking to enable, for the first time, the systematic validation of open‑source SLEIGH language specifications, predominantly used by Ghidra. We design and implement a testing framework based on an automated oracle validation strategy by proxy. Our approach leverages the structure encoded in a specification itself to enumerate decodable instruction forms and generate targeted initial states. Then differentially test the successful decoding and emulation of those instructions by comparing emulators exercising the processor specification against hardware references. Applying InSPECtor across diverse, open‑source specifications‑‑‑x86‑64, AArch64, ARM/Thumb, RISC‑V, MSP430‑‑‑embedding differences in specification styles, author preferences, and instruction set architecture designs, we uncovered over 38,920 discrepancies that led to 125 unique bugs with proposed fixes, identifying decoding and semantic defects as well as cross‑vendor inconsistencies. We distill our findings into 8 concrete recommendations to drive future improvements. Our work underscores the importance of specification correctness and provides a practical tool to substantially improve the fidelity of SLEIGH processor specifications, strengthening the reliability of downstream security and analysis tools.

Authors:Peng Li, Qianqian Xu, Shilong Bao, Yangbangyan Jiang, Qingming Huang
Title: UniTraffic-Agent: Unified Traffic Video Reasoning for AI City Challenge 2026 Track 3 with Two Out-of-Domain Evaluations
Abstract:
Traffic video understanding has become an important problem in intelligent transportation, as road videos provide direct evidence for accidents, violations, and interactions between vehicles and vulnerable road users. A useful system should explain how a traffic event develops, why it happens, and when the relevant interaction occurs, yet this remains difficult for multimodal large language models (MLLMs) because traffic videos contain sparse events and varied viewpoints. We introduce UniTraffic‑Agent, the MR‑CAS solution for Track~3 of the 10th AI City Challenge, which includes Traffic Anomaly Reasoning (TAR) and two out‑of‑domain evaluations: FETV for fisheye traffic events and PSI‑VQA for pedestrian intention reasoning. UniTraffic‑Agent follows an observe‑‑reason‑‑act‑‑verify workflow that samples timestamped visual evidence, reasons over all questions from the same clip in one request, and converts responses through task‑specific action adapters. On the official Public leaderboards, MR‑CAS ranks 16th on TAR with a score of 0.5780, 2nd on FETV with 0.4884, and 4th on PSI‑VQA with 64.4161. The code is available at https://github.com/Roclp/UniTraffic‑Agent.

Authors:Wenjin Liu, Shen Pang, Tiesunlong Shen, Zhe Cui, Xiaobao Wu, Anh Tuan Luu, Haoran Luo
Title: TIEM: Temporal Integration of Hypergraph Evidence and Skill Memory for Event-Driven Financial Forecasting
Abstract:
Event‑driven catalyst‑outcome forecasting increasingly uses retrieval‑ and memory‑augmented large language model agents for prediction. However, training‑data contamination and temporal leakage can create an Evidence Chasm between reported accuracy and true predictive ability. We propose TIEM, a timestamp‑gated framework with three coordinated components: an Event‑Evidence Hypergraph (EEH) for timestamp‑filtered multi‑tier retrieval; a Case‑based Skill Memory (CSM) for source‑tagged temporal skills; and Heterogeneous Evidence‑Experience Fusion Reasoning (HEFR) for evidence‑experience fusion and prediction. We also introduce FinPURE, a recent‑period A‑share holdout benchmark, and use a Name‑Date Probe to assess per‑model name‑date sensitivity rather than assuming training cutoffs. Results on five financial forecasting benchmarks show TIEM outperforms current baselines. Our project is available at https://github.com/QwenQKing/Fin_TIEM.

Authors:Xinlong Xu, Yoshua Y. Li
Title: RAGSieve: Self-Referenced Local Contrast for Knowledge-Poison Detection in Retrieval-Augmented Generation
Abstract:
Retrieval‑augmented generation treats an external corpus as inference evidence, allowing injected documents to promote attacker‑chosen claims. Existing detectors depend on trusted references, specific attack artifacts, or global thresholds sensitive to corpus topology. We present RAGSieve, a self‑referenced detection framework that constructs its reference from the inspected system. RAGSieve‑Query (RSQ) performs query‑local contrast, scoring top‑five candidates against ranks 6‑20 of the same retrieval to detect answer‑anchor concentration and carrier transitions. RAGSieve‑Graph (RSG) performs corpus‑local contrast, comparing each document's semantically similar but lexically distinct neighbors with its local baseline to detect coordinated density before queries arrive. Across three QA datasets and six poisoning constructions, RSQ achieves 95.2% AUROC and detects 82.2% of poison at 5% clean‑document removal, versus 81.1%/52.5% for GMTP. RSG achieves 93.3%/79.8%, versus 79.4%/37.6% for CleanBase. Joint deployment reduces attack success from 67.4% to 14.0% while retaining 41.3% F1 on unpoisoned retrieval, demonstrating practical protection at both corpus ingestion and query time without poison labels or trusted corpora. Source code is available at https://github.com/XrazyMee/RAGSieve.

Authors:Xinlong Xu, Yoshua Y. Li
Title: EviReform: Evidence-Guided Query Reformulation for Multi-Hop Graph Retrieval
Abstract:
Multi‑hop retrieval must recover passages that provide sufficient evidence together. An initial passage often resolves an entity or relation implicit in the question, making the missing evidence easier to describe only after retrieval begins. Graph retrieval improves access to related evidence through stored corpus structure, but its retrieval signal is commonly derived from the original question. Complementary evidence must then be reached through stored relations even when an observed passage provides a more direct semantic cue. We introduce EviReform, which separates revising the retrieval request from aggregating evidence in the graph. Retrieved source passages formulate residual queries for the unresolved information need. The original and residual retrieval signals are normalized separately, combined, and propagated between propositions that share entities. On 2WikiMultiHopQA, HotpotQA, and MuSiQue, EviReform exceeds the strongest baseline by up to 5.59 Recall@5 points and 4.50 F1 points. These results show that observed evidence can guide graph retrieval toward the part of a supporting chain left underspecified by the original question. Code is available at https://github.com/XrazyMee/EviReform.

Authors:Vsevolod Skorokhodov
Title: PixSDS: Why Latent SDS Makes Noisy Pixels
Abstract:
Score Distillation Sampling (SDS) enables text‑to‑3D generation by optimizing rendered images with a pretrained diffusion prior, but latent SDS often produces structured color artifacts and high‑frequency texture noise. We identify a failure mode of latent SDS caused by VAE‑induced pixel drift: the optimized image can move along pixel‑space directions that are weakly constrained by the VAE encoder, so its latent representation remains clean and semantically meaningful while the image itself accumulates visible artifacts. We support this diagnosis with controlled 2D SDS experiments, VAE‑only optimization, and a simplified analysis showing that encoder‑like latent objectives can amplify image‑space noise when the inverse mapping to pixels is underconstrained. Motivated by this observation, we propose PixSDS, a lightweight VAE‑consistent gradient repair method. PixSDS decodes a latent SDS lookahead step and uses the decoded image as a clean direction for pixel‑space optimization, reducing motion in VAE‑inconsistent directions without retraining the diffusion model, changing the renderer, or replacing the SDS objective. Experiments in 2D optimization and text‑to‑3D generation show that PixSDS substantially reduces structured artifacts while preserving semantic content. Code is publicly available at https://sevashasla.github.io/pixsds‑webpage/.

Authors:Ziyang Gao, Zhizhuo Jiang, Jingjing Chang, Yixin Yang, Yuwen Pan, Yong-Qiang Mao, Yu Liu, Hai-Bao Chen
Title: DiCoR: Decoupled Referent Disambiguation and Contour Recalibration for Efficient Referring Remote Sensing Image Segmentation
Abstract:
Referring remote sensing image segmentation (RRSIS) aims to delineate targets specified by natural language expressions in remote sensing imagery. Existing methods mainly follow joint fusion segmentation (JFS) or decoupled prompt segmentation (DPS). JFS is efficient but often suffers from limited accuracy because referent localization and mask delineation are optimized under a unified objective, whereas DPS separates localization from mask generation using spatial prompts and foundation segmenters at the cost of higher memory consumption and inference latency. To bridge this gap, we propose DiCoR, a decoupled referent disambiguation and contour recalibration framework built on an efficient JFS pipeline. DiCoR addresses two key challenges: distinguishing the correct referent from ambiguous candidates and refining coarse masks after localization. A disambiguation‑aware localization guidance strategy ranks salient candidate regions with adaptive linguistic cues and injects the resulting localization prior into fused features. A lightweight contour recalibration module further predicts residual corrections to coarse logits under localized contour supervision, improving mask quality with limited computational overhead. Experiments on RefSegRS, RRSIS‑D, and RISBench show that DiCoR achieves the best segmentation accuracy across all three benchmarks. On RefSegRS, it improves mIoU and gIoU by 5.28% and 2.87% over a competitive JFS method while running 4.7% faster than a representative DPS method, demonstrating a favorable accuracy‑efficiency trade‑off. Code is available at https://github.com/zyGao1126/DiCoR.

Authors:Yifan Mei, Qingling Shi, Changli Wu, Jiayuan Rao, Jiayi Ji, Liujuan Cao
Title: TennisVAR: A Stroke-Evidence-Grounded Multimodal Large Language Model for Tactical Reasoning in Tennis Videos
Abstract:
Sports‑video understanding is moving beyond event recognition toward explaining how actions collectively shape match progression, however, existing tennis‑video methods either perceive individual strokes without modeling their tactical dependencies or generate high‑level analyses without grounding them in the underlying events. To bridge this perception‑to‑understanding gap, we formulate stroke‑evidence‑grounded tactical reasoning, a new rally‑level task that requires models to jointly predict an open‑ended answer, a hierarchical tactic label, an ordered sequence of supporting strokes, and decisive key actions, with each evidence stroke anchored to its racket‑ball contact frame. We further introduce TRACE (Tactical Reasoning with Action‑Chain Evidence in Tennis), a large‑scale expert‑annotated benchmark containing 11,189 rally videos, 41,485 stroke events, 25,429 tactical units, and 11,189 question‑answer pairs, which unifies fine‑grained stroke attributes, cross‑stroke tactical relations, hierarchical tactic annotations, and evidence‑grounded questions across factual perception, tactical understanding, and decision reasoning. Building on TRACE, we propose TennisVAR (Tennis Video Action‑chain Reasoner), an evidence‑grounded multimodal large language model that follows an "event‑relation‑evidence‑tactic" reasoning paradigm, where an Event Parsing Module converts continuous rallies into explicit stroke‑event sequences while a Tactical Graph‑Guided Temporal Reasoner jointly models rally progression and same‑player decision dependencies to identify question‑relevant evidence and decisive actions.

Authors:Zhixuan Liu, Zhichen Dong, Yuanfu Wang, Chao Yang
Title: Decoupled Contrastive Decoding via Expert-Aligned Drafting
Abstract:
Contrastive Decoding (CD) improves generation quality, but its amateur‑model pass makes decoding expensive. Accelerating CD with speculative decoding raises a proposal‑alignment question: should the contrastive signal shape the drafter, or should it remain only in verification? We study this question in the lightweight feature‑level drafter regime. Two controlled diagnostics, matched Cross‑alpha training and an Approximate Dual‑Drafter decomposition, give the same diagnosis: contrastive‑aware drafting does not consistently improve over expert‑aligned drafting because the contrastive correction is usually weaker than drafter error, and reconstruction can amplify that error. We introduce Decoupled Contrastive Decoding (DCD), which drafts with an expert‑aligned lightweight proposer and applies the amateur only in unchanged CD verification. Standard speculative verification preserves the vanilla‑CD output distribution. Across the main 8B settings, EAGLE3‑based DCD achieves average greedy speedups of 1.65 to 1.95x over vanilla CD and reduces MMLU proposal‑path latency by about 5 to 12x relative to amateur‑coupled proposal paths.

Authors:Victoria Basmov, Yoav Goldberg, Reut Tsarfaty
Title: Prompts in the Wild: A Large Analyzed Collection of Transactional Prompts in Code
Abstract:
The behavior of contemporary generative Large Language Models (LLMs) is directly shaped by prompts, unstructured texts that describe the desired output and model behavior. In this paper we argue that prompts are linguistic objects that merit investigation in their own right. To this end, we collect 57.5K unique samples of prompts from GitHub. Specifically, we focus on transactional prompts: reproducible natural language instructions that are integrated into software. To enable the empirical, quantitative study of prompts, we introduce a structured ontology, capturing the properties of prompts as well as their formal and semantic components. Based on this ontology, we transform prompts from unstructured raw texts into richly structured linguistic objects. Analysis of these structured data reveals significant diversity of usage patterns across languages, domains, tasks, and modalities, in a typical Zipf‑like distribution where some clearly prevail and others, more diverse, appear in the long tail. To validate the reliability of the ontology‑based annotation of the prompts, we perform a comprehensive error analysis across all fields, providing a detailed assessment of annotation quality. We release the dataset together with a browsing and exploration interface (https://github.com/OnlpLab/transactionalPromptsCollection ).

Authors:Yunhao Bai, Zhongwei Qiu, Guangyu Guo, Yiming Huang, Tony C. W. Mok, Qinji Yu, Ling Zhang, Yan Wang
Title: HounsWorld: A Multimodal World Model for Hidden Patient-State Readout, Reconstruction, and Simulation
Abstract:
Clinical intelligence requires estimating a patient's underlying condition from incomplete observations rather than learning isolated mappings from scans to answers. Volumetric medical images provide dense observations of anatomy, attenuation, and lesions, whereas clinical language provides sparse but complementary semantic observations. We formulate CT‑centered intelligence as inference over a shared latent patient state, under which readout, reconstruction, and simulation all become state‑dependent prediction problems. To operationalize this view, we introduce HounsBench, a computed tomography (CT) centric patient‑state benchmark that unifies these three task families with patient‑disjoint splits and per‑family metrics, and HounsWorld, a 3B multimodal world model that treats volumetric scans and language as observations of the shared state through Joint Understanding‑Generation Learning. A shared transformer forms an implicit patient‑state estimate and supports three outputs: query‑conditioned answers that read out the state, reports and captions that reconstruct it in language, and condition‑specific CT volumes for low‑dose denoising, virtual contrast enhancement, and anatomy‑constrained text‑and‑mask‑to‑volume generation. Zero‑initialized CT adapters preserve pretrained multimodal mappings, while condition‑explicit Hounsfield‑unit window sampling exposes clinically meaningful density observations. HounsWorld shows strong performance across all three task families while consistently improving CT understanding through clinically structured completion. Our project is available at https://github.com/byhwhite/HounsWorld.git

Authors:Xiaoyu Lian, Shuyin Xia, Hongxuan He, Lifeng Shen, Guoyin Wang, Xinbo Gao
Title: Adaptive $k$ Nearest Neighbors Classifier via Granular Ball Computing
Abstract:
The k‑Nearest Neighbor~(KNN) algorithm is widely used across various tasks. The selection of the k value is a key issue because it significantly impacts performance. In this paper, an adaptive and efficient KNN approach via granular‑ball computing is proposed. The method consists of two stages. \textcolorblackIn the training stage, the dataset is first coarsely partitioned to reduce the complexity of data distributions within a granular ball, and then the Fisher criterion is introduced to control ball splitting and stopping, yielding a multi‑granularity granular ball representation. In the prediction stage, the nearest granular ball is first located through a weighted distance mechanism, and an adaptive neighborhood is then constructed around the test sample. The effective k value is dynamically determined by the actual number of samples contained in this neighborhood. The neighborhood induced by the nearest granular ball provides more stable local group information, thereby improving robustness against noise and local perturbations. Experimental results demonstrate that the proposed method outperforms existing KNN variants across multiple datasets in terms of both accuracy and efficiency. The code has been open‑sourced for reproducibility: https://github.com/lianxiaoyu724/Adaptive‑GBKNN.

Authors:Varun Pratap Bhardwaj, Garima Singh, Arun Pratap Bhardwaj
Title: Agent Behavioral Contracts II: Certifying Compositional Reliability Without Assuming Independence
Abstract:
Compositional reliability bounds for multi‑agent systems multiply component reliabilities, a step licensed by a conditional‑independence assumption that is routinely stated and rarely tested. We test it. Two instances of one model, in a two‑agent handoff, co‑fail on 90.0% of the missions on which either fails (log OR 6.66, 95% CI [6.38, 7.00]; phi 0.916), in a preregistered evaluation of 18,000 missions scored by deterministic code with no LLM judge. Substituting a different model reduces the association in six of six contrasts; substituting a different vendor, model already different, does not ‑‑ a registered hypothesis reported as a null. The error is signed and runs against the operator: positive dependence inflates joint failure above the independence product, so redundancy is over‑credited exactly when components share a model. The assumption‑free alternative is often vacuous, and fitting a dependence model is worse: we prove a bootstrap bound on a fitted model's functional loses coverage of the truth as n grows, the identification gap being O(1) while the bootstrap haircut is O(n^‑1/2). More data makes such a certificate worse, with no visible symptom. We give a finite‑sample certificate assuming no dependence structure: a linear program over the joint, over a Bonferroni‑Clopper‑Pearson box around measured co‑execution moments. It is sound, sharp for the information supplied, and monotone in the moment family. Enriching ten moment functionals to fourteen narrows the identified interval by 85.7% and lifts the certified floor from 0.2455 to 0.4116. A companion anytime‑valid certificate holds type‑I error at 0.0471 under optional stopping. Common dependence statistics are marginal‑bounded and can reverse an apparent ordering of conditions when the compared agents fail at different rates. Contracts, scoring code, analysis scripts, and the preregistration are released.

Authors:Adnan El Assadi, Niklas Muennighoff, Jinhyuk Lee
Title: The Embedder's Dilemma: LLMs Are Better, but at What Cost?
Abstract:
Should you replace your text‑embedding pipeline with a large language model? We answer this with a controlled, cost‑aware comparison of ten LLMs across six families and 26 embedding models (118M to 14B parameters) on 37 tasks spanning classification, semantic textual similarity (STS), clustering, pair classification, and retrieval. In aggregate the two paradigms are effectively tied: the best LLM (Gemini 3.1 Pro, 77.6) and the best embedding model (77.2) differ by 0.4 points. Their strengths differ by task: LLMs lead on reasoning‑heavy retrieval, embedding models lead on classification, and the two match on clustering, STS, and pair classification. Reaching that parity is expensive. An LLM costs up to 1,431x more than an embedding model of comparable quality (USD 154 vs. USD 0.11 per benchmark pass), and the open LLMs tested process tokens 2.5 to 736x more slowly on the same GPU. Reasoning tokens account for 28 to 81% of LLM inference cost; lower reasoning budgets preserve or improve retrieval quality for most models in our ablation. The Pareto frontier contains the leading embedding models and one LLM, Gemini 3.1 Pro. These results support a division of labour: use embedding models for similarity, classification, and clustering, and reserve LLMs for reasoning‑intensive retrieval. Our code, datasets, and results are publicly available at https://github.com/embeddings‑benchmark/embedders‑dilemma.

Authors:Quan-Dung Pham, Anh Dao, The-Anh Nguyen, Minh Nguyen-Dinh, Phuong Nam Dang, Tri Pham, Hung Tran, Bach Dao, Tuyen P. Le, Truong Nguyen, Quan Nguyen
Title: HumanoidVLN: A Physics-Grounded Simulator and Benchmark for Vision-Language Navigation Across Diverse Humanoid Embodiments
Abstract:
Vision‑Language Navigation (VLN) for humanoid robots poses challenges existing benchmarks fail to address: bipedal locomotion imposes physical constraints absent from wheeled agents, humanoid morphologies vary across platforms, and egocentric observations are distorted by locomotion‑induced camera dynamics. We present HumanoidVLN, a physics‑grounded simulator and benchmark for VLN across diverse humanoid embodiments. Built on NVIDIA Isaac Sim, our platform supports an extensible set of humanoid configurations, demonstrated on four robots (Unitree G1, Unitree H1, Internal‑A, Internal‑B) spanning 10‑12 lower‑body DoF and heights from 1.17m to 1.80m, via a hierarchical control stack combining a reinforcement learning locomotion policy with interchangeable PD or MPC path trackers. New robots and VLN models integrate with minimal effort; we demonstrate compatibility with NaVILA, DualVLN, StreamVLN, and JanusVLN. Environments are drawn from artist‑designed scenes and 3D Gaussian Splatting reconstructions, filtered for navigable areas exceeding 100 square meters. Instructions are generated by a dual generator‑reviewer plus paraphraser multi‑agent pipeline with human‑in‑the‑loop verification, yielding 933 collision‑aware reference episodes, each paired with one fine‑grained instruction and three coarse‑grained stylistic variants (formal, natural, casual). Across four models and four embodiments, JanusVLN achieves the highest mean success rate of 43.55% and nDTW of 48.38. In a 20‑episode sim‑to‑real pilot with DualVLN and the Unitree G1, navigation errors correlate strongly (r=0.935), with a mean absolute difference of 0.68m and mean trajectory similarity of 0.782 (+/‑0.188) nDTW. These results highlight the interaction between VLN models, controllers, and humanoid embodiments under physical execution. Code, benchmark, and data will be released upon acceptance at https://humanoid‑vln.github.io/.

Authors:Xutao Mao, Liangjie Zhao, Xiang Zheng, Cong Wang
Title: Practice Makes Unsafe: Skill Misevolution in Self-Improving LLM Agents
Abstract:
Self‑improving LLM agents convert successful trajectories into persistent cross‑task state. An unsafe success can thereby become reusable policy after its triggering input disappears. Skill evolution makes this failure measurable by distilling operational trajectories into executable, transferable, and inspectable procedures. Because evolution optimizes task outcomes rather than procedure safety, compromised experience can cause skill misevolution. Existing benchmarks measure current behavior or static artifacts but cannot attribute risk across authoring, retrieval, and later execution. To expose this lifecycle, we introduce SkillMisevo‑Gym, a lifecycle‑aware harness that versions skill state across agent frameworks, and SkillMisevo‑Bench, a frozen design from malicious exposure to carryover tasks, with concept‑aligned benign tasks and nine lifecycle metrics. We also introduce SafeEvolve, a wrapper that repairs unsafe content and governs subsequent reuse. Across 25 agent‑method configurations, each covering 525 tasks in 25 episodes, all 21 evolved configurations author unsafe artifacts, while only fifteen lead to fresh‑session harm. In the exposure sweep, three malicious tasks raise carryover ASR from 16.0% to 35.3%. Across representative skill evolution methods, SafeEvolve reduces unsafe retrieval and fresh‑session harm by 26.7 and 17.3 percentage points, respectively, while mean benign utility changes by only 0.4 points. Together, persistent‑adaptation safety must govern what updates write and what future executors reuse. Code is available at https://github.com/henrymao2004/misevolve.

Authors:Wenyu Li, Sidun Liu, Tongrui Hu, Peng Qiao, Yong Dou
Title: LocusGS: Spatially Grounded Tokens for Feed-Forward 3D Gaussian Splatting
Abstract:
Recent query‑based feed‑forward 3DGS methods represent a scene using learnable queries, each aggregating multi‑view evidence and decoding a group of Gaussians. Ideally, different queries should specialize in coherent local regions of the scene. However, we observe that Gaussians decoded from the same query often scatter across distant scene regions, resulting in weak query‑level spatial coherence and poor alignment with the scene structure. We attribute this behavior to the purely latent representation of existing Gaussian queries. To address this limitation, we introduce LocusGS, which augments each Gaussian query with a 3D anchor state consisting of a center and a support radius. The anchor state is progressively refined across decoder layers and is used throughout query interaction, multi‑view feature aggregation, and Gaussian generation. Specifically, an anchor‑to‑ray geometric bias guides each query toward spatially relevant image observations, while anchor‑centered decoding organizes its Gaussians within a local region. Experiments on novel view synthesis benchmarks show that LocusGS improves rendering quality over query‑based Gaussian token baselines under the same Gaussian budget. Further analysis shows that the learned anchors form coherent spatial layouts and lead to more structured Gaussian distributions, demonstrating that explicit anchor states improve the spatial organization. Our project page: https://leo‑frank.github.io/LocusGS_viewer.

Authors:Hakan Üstünel
Title: Matrix-Driven Quartic Overhauser (QOVR) Surfaces Structural Framework: Continuity Limitations, Computer Graphics Algorithms, and Software Implementation
Abstract:
This study introduces the spatial and analytical construction of the Quartic Overhauser (QOVR) surface generation framework designed to resolve boundary alignment and localized shape modification constraints. This framework implements a variable parameter fourth degree novel architecture to achieve exact parameter isolation across orthogonal coordinate axes. The analytical pipeline integrates directional spline blending functions with symmetric spatial control matrices, ensuring that internal knot vector variations allow localized surface adjustments while preserving the absolute positional invariance of global edge boundaries. Computational verification confirms that while the current formulation satisfies explicit C^0 positional closure and C^1 tangent continuity conditions across the internal and boundary interfaces without triggering global curvature propagation, edge joint separation, or wave‑like artifacts, C^2 curvature continuity is not maintained. As demonstrated by three‑dimensional mesh models and colormap visualizations, the spatial intensity fields condense strictly within the immediate neighborhood of the modified element, verifying that the displacement effect decreases exponentially as the distance from the perturbed control points increases. This explicit decoupling preserves structural symmetry and boundary invariance across adjacent geometric patches, satisfying manufacturing and reverse engineering sealing criteria. Keywords: Quartic Overhauser surface; computer aided geometric design (CAGD); local shape control; geometric continuity; symmetric tensor products; computer graphics algorithms; software implementation framework.

Authors:Ahmed Sameh, Ramzi Al-Sharawi, Yogatheesan Varatharajah
Title: The Impact of Temporal Context Length and Encoding Strategies on Self-Supervised ECG Representation Learning
Abstract:
Self‑supervised electrocardiogram (ECG) models are often trained on a few seconds of ECG signal and, increasingly, on discretized token sequences. It remains unclear whether these choices sacrifice information needed for rhythm inference and longitudinal consistency in real‑world ambulatory recordings. We present a controlled study on the Icentia11k single‑lead dataset that varies (i) the input horizon (16 seconds, 1 minute, 5 minutes, and 10 minutes) and (ii) the front‑end representation (continuous convolutional patch embeddings vs. fixed vector‑quantized tokens), while holding the Transformer backbone and training protocol constant. Representations are assessed by downstream abnormal rhythm detection and by patient‑level retrieval that probes cross‑session stability. Our results show that increasing temporal context beyond 16‑second snapshots yields stronger transfer and higher retrieval accuracy, with the strongest performance achieved by the 5‑ and 10‑minute models, indicating improved capture of slow‑varying rhythm dynamics and individual‑specific structure. Across all evaluated horizons, continuous patch embeddings outperform discretized tokens, suggesting that quantization can discard clinically relevant waveform detail. These findings motivate ECG foundation models that emphasize extended context and continuous encoders for clinical prediction and similarity‑based applications. Our code and pretrained models are publicly available at https://github.com/muha‑0/ecg‑ssl‑representation‑learning.

Authors:Florian Braun
Title: Excess Separability: Nuisance-Controlled Residual-Stream Probing for Benchmark Contamination Detection
Abstract:
Benchmark contamination is diagnosed today with n‑gram overlap, with likelihood‑based membership inference, or with canary strings, and each needs something usually unavailable: the training corpus, a well‑chosen test statistic, or foresight at dataset release. A recent alternative reads contamination off a linear probe on internal activations. We show that the natural way to do this does not work, and specify one that survives measurement. The protocol reports a zero‑sum contrast on the depth profile of probe accuracy, recentred on a level‑matched placebo baseline, tested against a label‑permutation null, with the reference set twice the size of the suspect set. Each choice replaces a simpler alternative we measured and rejected. Reporting the level of excess separability rather than its shape makes the false positive rate track the size of the analyst's own control set, from 0.03 to 0.99 under a true null. Contrasting against a flat depth profile fails in both directions, rejecting a true null 0.72 of the time when surface decodability rises with depth and losing all power when it falls. An item bootstrap holds the fitted probe fixed and rejects up to 0.09 of the time where a permutation null that refits it holds 0.02. A half‑size baseline triples the error rate. On real transformers, baseline depth profiles are measurably not flat, spanning up to 29.1 accuracy points on a temporal split, and their non‑flatness tracks the surface difference between the item sets (correlation 0.87 over 6 audits), so the correction is largest exactly where it is needed. All 4 well‑matched Pile arms return null, and the protocol refuses a verdict on the temporal split rather than reporting one. What this does not establish is whether transformers carry a familiarity direction at all: the only positive sits on the split where exchangeability fails. Implementation, tests and audits are released.

Authors:Li Yin, Zhi Li, Zhan Shi, Haoran Zhang, Haebin Seong, Zhangyang, Wang
Title: @skills: Attention is all you have
Abstract:
There are 56,804 public agent skills today, and teams write many more privately. The dominant delivery model is installation: once installed, a skill's description remains in the system prompt, competing for fewer than 100 reliable trigger slots. This leaves the long tail with no practical path to use and forces teams' own playbooks to compete for the same scarce space. We observe that installation bundles three separable functions: content, persistence, and automatic triggering. Only the last requires prompt residency. We therefore propose @skills, an open protocol that separates them. A path addresses any skill, subtree, or collection, and reading a skill is sufficient to use it, so nothing is installed or made resident. The operation vendors a copy at the same path into a project's Git‑tracked tree for adaptation and ownership. The operation adds one .gitignore‑style line, the only element that costs prompt residency. A directory is a menu, making bundles ordinary directories rather than all‑or‑nothing units. The protocol requires no manifest, lockfile, or registration, and SKILL.md remains unchanged. @skills is additive, ships as an installable package, and turns any agent that can read files and run commands into a client through a single instruction file. Its open specification is at https://github.com/SylphAI‑Inc/atskills and it is implemented in the AdaL CLI at https://adalagent.ai . Because paths address skills well but cannot find them, the protocol is paired with a free hub at https://atskills.one for corpus‑wide search and ranking, repository‑free hosting, private and team collections, and one‑screen authoring. The hub is optional: gh: and local paths resolve without it, and indexed GitHub skills retain their gh: identities. Install less, use more.

Authors:David M. Straub
Title: Domus: An Open-Data Web Platform for House History Research
Abstract:
House history ‑‑ the record of who lived in a building, who owned it, and how its address changed over time ‑‑ is a central concern of genealogical and local history research. Wikidata and OpenHistoricalMap together provide an open‑data infrastructure, CC0‑licensed and capable of representing this data: Wikidata through its linked‑data property model with temporal qualifiers, OpenHistoricalMap through historical building footprints with start and end dates. What has been missing is a domain‑specific tool that makes contributing and discovering building history accessible to genealogists and local historians. Domus fills this role: a map‑based web platform that guides users through structured, crowdsourced editing of building records in Wikidata and displays corresponding OpenHistoricalMap footprints. The application requires no custom backend ‑‑ authentication, data reading, and data writing all operate through public APIs in the browser, and all data are stored under CC0 in public collaborative databases. This paper describes the genealogical data model, the architecture, and key technical design decisions.

Authors:Emanuele Cavalleri, Paolo Perlasca, J. Harry Caufield, Justin Reese, Christopher J. Mungall, Marco Mesiti
Title: SchemaLink: An Intelligent Web Editor for LinkML Schema Curation
Abstract:
Motivation: LinkML is a suitable language for the representation of the structural and content constraints of different kinds of biomedical data. Even if it is a quite recent proposal, it has been applied in several biomedical contexts. Developing and maintaining LinkML schemas presents several challenges, particularly for novice curators. Non‑expert bio‑curators may struggle with LinkML syntax and best practices, requiring significant time and effort to develop well‑structured schemas. Results: In this paper we propose SchemaLink, a web‑based environment for the graphical construction and enhancement of LinkML schemas that address the following requirements: (i) introduce a graphical language for the specification of LinkML schemas, (ii) make uniform the specification of schemas in similar contexts, (iii) simplify the design and curation processes by exploiting a RAG‑based approach to assist curators in creating new schemas from scratch and editing already developed ones. Several experimental analyses show the quality of the produced LinkML schemas through the AI‑based editing facilities. Availability and Implementation: SchemaLink is available online at: https://SchemaLink.biodata.di.unimi.it. SchemaLink code and testing data are available as open‑source on GitHub at: https://github.com/AnacletoLAB/schemalink‑webapp,schemalink‑api.

Authors:Yash Bagla
Title: Excitation-Supervised Closed-Loop Self-Calibration and Target Seeking for an Unknown-Pose Range-Bearing Relay
Abstract:
A vehicle seeking a hidden target through a range‑bearing relay of unknown position and yaw must decide, online, whether its own motion has already made the relay calibration trustworthy, and what to do when it has not. Two distinct vehicle‑relative observations are known to remove the calibration gauge and make the target's relay‑local packet globally actionable (arXiv:2608.09464), but that statement is static: it classifies a stored window only after the fact. This paper supplies the closed‑loop layer: we show that the trajectory‑spread margin S_v that governs identifiability is simultaneously a finite‑noise seed‑accuracy bound, a local‑vector variance decomposition, and a circle‑geometry excitation budget, and we use it to supervise an excitation‑reset controller. An excitation‑supervised algorithm retriggers exploratory motion whenever the spread certificate is insufficient, projecting the target‑seeking input away from the excitation's push, and otherwise proceeds to unrestricted target seeking. Under explicit sampling assumptions the supervision rule provably acquires any required excitation in finite time; in the noiseless local regime with positive excitation decay, estimator convergence yields target‑seeking convergence after certification; and the threshold is selected from a desired calibration‑accuracy level rather than chosen heuristically. Closed‑loop simulation, paired Monte Carlo comparisons, a spread‑threshold ablation, and a ROS 2/Gazebo software‑in‑the‑loop experiment with sensing delay validate the approach. A decay‑rate sweep shows that supervision matters when a fixed schedule's decay outruns the unknown time‑to‑adequate‑excitation: over 100 paired trials the fixed baseline's yaw RMSE rises from 0.010 to 0.065 rad and success falls to 56%, while target‑tracking error remains insensitive; supervision keeps yaw RMSE between 0.0095 and 0.0191 rad with 100% success.

Authors:Aysha Ashraf, Shaina Ashraf, Wafaa I. M. Hussin, Ali Haider, Zhi Lu, Zhenming Peng
Title: HIMEC: Directional Change Representation and Fixed-Interface Decoding for Remote Sensing Image Change Captioning
Abstract:
Remote sensing image change captioning (RSICC) converts bitemporal imagery into a sentence describing semantic changes. Most RSICC methods condition caption decoders directly on fused visual features, leaving intermediate change structure and decoder‑interface consistency less studied. We present HIMEC, combining Directional Change Representation (DCR) with fixed‑interface decoding. DCR separates signed differences into appearance‑oriented, disappearance‑oriented, and shared‑context streams before fusion. A learned‑query encoder converts the fused representation into visually conditioned change‑query tokens that form the scene decoder's only sample‑dependent memory. A training‑only auxiliary phrase decoder supplies caption‑derived supervision. With a fixed zero input, the scene decoder maintains the same interface during training and inference. Separately, we evaluate a local‑to‑scene cascade conditioned on teacher‑forced local states during training and autoregressive states at inference. On changed LEVIR‑CC validation pairs, these states have a mean cosine distance of 0.69. Regime‑matched conditioning recovers most of the associated deficit, whereas permuting state correspondence causes no detectable penalty. These findings are limited to the evaluated cascade. In a matched three‑seed comparison, HIMEC reaches a Consensus‑based Image Description Evaluation (CIDEr) score of 142.81\pm0.60 on LEVIR‑CC, versus 139.51\pm3.40 for direct fused‑feature memory. On SECOND‑CC, fixed‑zero and regime‑matched diagnostic conditioning reach 75.67 and 76.99 CIDEr, respectively, versus 60.77 for the mismatched cascade. The source code will be made publicly available at https://github.com/ayshaashra/HIMEC upon publication.

Authors:Peilin Chen, Xiaoxuan Yang
Title: Lonic: Algorithm-Hardware Co-Design for Energy-Efficient Fully Local Online SNN Training with INT4 Precision
Abstract:
Spiking neural networks (SNNs) have recently attracted increasing attention as an energy‑efficient learning paradigm. Existing works also propose temporally and fully local online SNN training algorithms to address memory and computation overhead. However, they do not consider whether the algorithmic advantages can be effectively translated into real‑device efficiency. To address this challenge, we present Lonic, an algorithm‑hardware co‑design for energy‑efficient and scalable fully local online supervised SNN learning. On the algorithm side, we implement an INT4 low‑precision training algorithm for fully local online SNN learning while maintaining accuracy. On the hardware side, to leverage the benefits of the proposed algorithm, we introduce reconfigurable multiplier‑free integer PE arrays, dual‑optimization zero‑gating strategy, temporal prefix‑accelerated local learning dataflow, and low‑precision weight movement to significantly improve training efficiency. Compared to Apple M4 and Nvidia V100 GPUs, Lonic achieves average energy efficiency improvements of 17.44x and 66.28x, respectively, along with speedups of 3.25x and 1.02x, respectively. Moreover, Lonic achieves 15.95x (14.64x) and 1.52x (7.28x) energy efficiency (area efficiency) over ASIC TPU‑like and H2Learn accelerators, respectively. The code for Lonic is available at https://github.com/peilin‑chen/Lonic.

Authors:Nelson Guda
Title: Geometric and Behavioral Stratification in Transformer Residual Streams
Abstract:
Trained transformer models develop privileged bases: coordinate axes whose statistics differ from the rest of the residual stream. But what kind of direction does such a basis select? We investigate the prediction direction, the unembedding direction of the token a model currently predicts, and find that it functions as a content‑defined privileged anchor. Measured with respect to this anchor, residual‑stream variation is geometrically and behaviorally stratified by proximity to the prediction. The stratification holds in all eighteen models tested (dense and mixture‑of‑experts, 7B‑120B, base and instruction‑tuned). A narrow, scale‑invariant prediction interface concentrates readout‑relevant structure, while the vast prediction‑distal complement expands with model scale. Because the prediction direction sits nearly orthogonal to the principal variance axes, variance‑based analyses recover this organization only partly, and the shortfall grows with prompt heterogeneity. Anchoring reveals a steep geometric gradient: prediction‑proximal regions are highly structured and cluster related prompts, while the complement is flatter and anti‑discriminates among prompt groups. The interface is a narrow slice but functionally decisive. Disrupting the variance directions closest to the prediction causes immediate divergence and frequent task‑frame shifts; disrupting the next level down delays divergence and preserves framing. The complement is weakly readout‑aligned per direction yet causally and temporally load‑bearing, and behavior is driven by direction rather than magnitude. These results establish the prediction direction as a privileged anchor distinct from previously described coordinate axes, and give a geometric account of how high‑dimensional computation coexists with linear readout.

Authors:Sanjay Bhargav Dharavath, Hanvitha Saraswathi Mukkamala, Faizan Farooq Khan, Ioannis Kakogeorgiou, Aditya Arun, C V Jawahar, Zakaria Laskar
Title: MV2: Multi-View Multi-Vehicle Driving Dataset for Novel View Synthesis
Abstract:
Differentiable rendering has advanced novel view synthesis (NVS), yet applying it to real‑world driving remains difficult due to sparse capture viewpoints, dynamic objects, and limited multi‑trajectory data. We introduce the Multi‑View Multi‑Vehicle (MV2) dataset and benchmark for evaluating NVS models under large viewpoint changes in dynamic urban scenes. MV2 features synchronized captures from a car, scooter, and drone, each following distinct yet synchronized trajectories. Training NVS methods on one vehicle's camera stream and testing on another enables evaluation under substantially larger viewpoint variations than existing single‑trajectory datasets. All sequences are registered via Structure‑from‑Motion and camera poses verified using manual pixel‑level correspondence annotations, yielding 50 high‑quality scenes with 12000 images. Benchmarking recent NVS and camera pose estimation methods shows that NVS performance degrades with increasing viewpoint disparity, and that feed‑forward pose estimators notably lag behind optimization‑based approaches, highlighting MV2 as a rigorous testbed for NVS in driving. The dataset, benchmark protocol, and project resources are available at https://mv2‑dataset.github.io/.

Authors:Ganesh S
Title: FluctlightDB: A Memory Model of Data for AI Agents
Abstract:
For fifty years, data systems have answered two questions. The relational model asked which records match a predicate; the vector model asked which vectors lie nearest a query. Neither was built for cue‑driven, provenance‑weighted recall across long sessions. We propose treating long‑term agent memory as a distinct data model ‑‑ with its own write semantics (encoding, separation, consolidation, provenance) and read semantics (cue‑driven activation across a linked memory graph) ‑‑ and present FluctlightDB, an embedded engine that implements this contract via experience() and activate(). We make that case carefully, not categorically: we do not claim novelty over Mem0, Zep, or HippoRAG‑style memory layers, only an embedded engine contract beneath them. On LoCoMo (official evidence‑recall metric; 10 conversations, 1,982 gold spans), CHORUS recalls 99.0% on an internally reproduced July 2026 run. On LongMemEval‑S (500 questions, official session_recall@8), our retrieval harness scores 97.6% (488/500); end‑to‑end QA with our reader/judge stack scores 97.4% (487/500) ‑‑ these layers use different protocols than vendor leaderboard figures we cite for context only. On BEIR SciFact (shared MiniLM embeddings, same harness, Recall Fabric on), CHORUS/PRISM edges Chroma on nDCG@10 (0.646 vs. 0.645) and Recall@10 (0.792 vs. 0.783). We also report a small author‑designed regression suite (FAMB; paraphrase n=10, other sub‑tests n=1) at 100% macro ‑‑ internal validation, not peer benchmark. Strangers can verify the engine in under a minute via pip install "fluctlightdb[native]" and a minimal connect() ‑> experience() ‑> activate() script (compiled wheel, not source‑only). Harnesses and frozen JSON are MIT‑licensed. We claim no new neuroscience and no new transformer; we propose a missing layer of the data stack and release an engine others can reproduce and contest.

Authors:Ying He, Zhouhong Gu, Zhecheng Hu, Yubo Zhou, Hao Shen, Jiaqing Liang, Zhaoqian Dai, Shuguang Ma, Fei Yu, Yanghua Xiao, Zhixu Li
Title: Are Large Language Models Reliable Reviewers? A Benchmark for Error Detection in Financial Documents
Abstract:
Ensuring the accuracy of financial documents is critical for economic analysis, regulatory compliance, and corporate decision‑making. Several studies have shown that Large Language Models (LLMs) perform well in many financial tasks, such as stock price movements and financial analytics. However, a critical task remains unexplored: the ability of LLMs to identify errors in financial documents. In this paper, we introduce FinED‑Bench, the first publicly Benchmark for Financial Error Detection across three levels of cognitive complexity. FinED‑Bench covers nine real‑world financial scenarios, and includes over 900 documents reported in 2025 that are unseen by existing language models. We detail the benchmark construction process and evaluate several advanced LLMs (e.g., GPT‑4o, Qwen3‑14B) on this tasks, which requires both financial domain knowledge and reasoning capabilities. Experimental results show that current LLMs still struggle with this task, especially in high‑complexity cases. Besides, supervised fine‑tuning can significantly improve the performance of weaker LLMs on this task. Our data and code are available at https://github.com/hedyHe/FinED‑Bench.

Authors:Ying He, Sihang Jiang, Xingzhou Chen, Zhouhong Gu, Yiwei Gu, Minggui He, Shimin Tao, Hongxia Ma, Yanghua Xiao
Title: The "Knowledge-Behavior Gap" in Cultural Taboo Safety of Large Language Models
Abstract:
Cultural taboo safety is essential for deploying large language models (LLMs), as culturally insensitive outputs may cause offense or even social harm. However, existing cultural benchmarks primarily assess cultural knowledge or values biases, while overlooking whether LLMs can recognize and respect cultural taboos, especially when taboos are implicitly hidden in seemingly harmless questions. Besides, cultural taboos are implicit, and context‑dependent, thus poss unique challenges for reliable evaluation. To address these gaps, we introduce CulShield, the first public benchmark dedicated to evaluating and improving the cultural taboo safety of LLMs. CulShield spans 77 countries and territories, and includes over 2,020 taboos. It evaluates models along both explicit knowledge and implicit behaviors. Experiments on several advanced LLMs (e.g., GPT‑4o‑mini, Gemini‑2.5‑pro) reveal a clear ``knowledge‑behavior gap'': models often fail to apply known taboos during interaction. We further show that variations in linguistic context can significantly affect LLMs' cultural taboo safety. Code and data is accessible here: https://github.com/hedyHe/CulShield.

Authors:Ritabrata Chakraborty, Rajatsubhra Chakraborty, Shivakumara Palaiahnakote, Angelo Cangelosi, Umapada Pal
Title: Vision-Language Models are Fragile Multilingual Associators
Abstract:
Vision‑language models must associate visual entities with textual attributes. Whether these associations or concept bindings remain stable when the language of the input changes is unexplored. We introduce M^2BIND, a benchmark varying the language of the context and query across multiple languages. We evaluate binding both extrinsically through task performance metrics and intrinsically through causal interventions. We find that binding is not language‑invariant: cross‑family and cross‑script settings trigger significant binding collapse, with the model's internal binding computation shifting to later layers and losing causal strength. Closely related languages preserve associations comparatively better. In a broader sense, our findings indicate how VLMs deployed globally in multilingual settings cannot be assumed to maintain the same association quality observed in monolingual evaluation.

Authors:Suman Paudel, Sarbin Sayami
Title: Comparative Analysis of Multilingual Pre-trained Models for Nepali Automatic Speech Recognition
Abstract:
Multilingual pretrained models nominally support Nepali, yet no controlled benchmark has compared them under a single fine‑tuning protocol. We fine‑tune six pretrained models (XLSR‑53, IndicWav2Vec, MMS‑1B, Whisper‑Medium, Whisper‑Large‑v3‑Turbo, and Conformer‑Hi) spanning CTC self‑supervised, autoregressive encoder‑decoder, and hybrid Conformer‑CTC architectures, on the OpenSLR SLR54 Nepali corpus (~165 hours) using identical preprocessing, splits, optimizer, and family‑matched learning‑rate schedules. We evaluate Word Error Rate (WER), Character Error Rate (CER), and Real‑Time Factor (RTF) on three independent test sets (OpenSLR, FLEURS, Common Voice). Whisper‑Large‑v3‑Turbo (14.76% WER) and IndicWav2Vec (14.89% WER) tie at the top despite a 9x parameter gap and 40x pretraining‑data gap, providing direct empirical evidence that language‑family proximity in pretraining can substitute for raw scale for in‑domain Nepali. CTC decoders run up to 29x faster than autoregressive Whisper at the same accuracy, flipping the practical deployment preference toward CTC under any latency budget. Massively multilingual pretraining (MMS‑1B) yields the smallest out‑of‑domain degradation on FLEURS (+12.55 pp), indicating that scale buys robustness rather than peak in‑domain accuracy. The resulting benchmark provides the first standardized, multi‑model, efficiency‑aware reference numbers for Nepali ASR.

Authors:Alex Chao
Title: When AI Is Your Pastor: A Benchmark for Theological Triage and Pastoral Guidance in Large Language Models
Abstract:
People increasingly ask large language models (LLMs) for counsel on questions of faith, doctrine, and pastoral care. These questions are not ordinary information requests. Some ask about core Christian beliefs, some ask about real disagreements among faithful traditions, some require humility because the issue is prudential, and some are pastoral situations where safety and human referral matter more than theological completeness. Existing benchmarks do not evaluate this structure. We introduce FMG‑Bench, the Faith & Moral Guidance Benchmark, a 120‑scenario benchmark for evaluating large language model behavior in English‑language Christian theological triage and pastoral guidance contexts. FMG‑Bench v1 evaluates 14 advanced models across 8,792 scored responses, comparing raw model behavior with three guided instruction settings. In our production run, placing models inside a structured harness improves over raw model behavior by +3.96 points on average, with every model improving. The most safety‑critical finding is a +10.8 point gain in escalation appropriateness ‑‑ whether AI systems recognize when pastoral, clinical, legal, or emergency support is needed. The guided settings also improve robustness, meaning consistency when questions are reworded or pressured (92.88 to 98.02 stability). Asking a model to compare perspectives helps in secondary‑doctrine questions but can be counterproductive when applied to primary doctrine or urgent pastoral situations. The benchmark is a measurement tool, not an endorsement of AI systems as pastoral authorities.

Authors:Yuyang Yin, Zixiang Li, Longxuan Deng, Hongkai Li, Shifang Zhao, Junnan Liu, Weirong Huang, Mengyu Wang, Tianxiao Fu, Yikai Wang, Peng-Shuai Wang, Xiaojie Jin, Yao Zhao, Yunchao Wei
Title: StateFlow: Building, Evolving, and Accessing 3D World States for Previsualization
Abstract:
Previsualization is an intermediate layer between ideas and production in film, games, architecture, and urban design. It lets creators iteratively refine scenes, actions, cameras, and spatial‑temporal dynamics. Yet existing generative methods rely on simple prompts to jointly control all of these factors through one‑shot image or video synthesis, offering weak controllability and limited support for iterative editing. Fundamentally, a world comprises multiple elements with geometry, appearance, and other attributes, together with cameras. Different frames are produced through local modifications or recombinations of this shared state, which is otherwise largely reused. Therefore, we argue that the missing component is an explicit and persistent working state. To address this, we present StateFlow, a state‑centric framework for generative previsualization. Rather than generating videos in one shot, StateFlow uses an editable 3D world to organize scene structure, evolution, and cameras, while off‑the‑shelf video models enhance visual quality when higher fidelity is desired. This world is maintained as a persistent structured 3D state of scene elements and camera configurations, serving as the core working representation for previsualization. Built on this insight, StateFlow has three stages to construct, evolve, and access the world state. State construction lifts generated 2D content into a coherent 3D world through prior‑guided, conflict‑aware dual‑view initialization, while State evolution translates user intent into structured state transitions while preserving world memory, avoiding full‑scene regeneration for each edit. State access uses render‑feedback reflection to refine camera plans into visually feasible trajectories, avoiding reliance on VLM semantics alone. Experiments show that StateFlow produces high‑quality 3D worlds for video creation and game‑like prototyping.

Authors:Ankita Rajaram Naik, Anupama Murthi, Benjamin Elder, Siyu Huo, Raavi Gupta, Abhinav Jain, Praveen Venkateswaran, Abdulhamid Adebayo, Danish Contractor
Title: VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies
Abstract:
Agents deployed in enterprise settings must reason across structured APIs and document collections, yet existing benchmarks evaluate these capabilities in isolation. We introduce VAKRA (eValuating API and Knowledge Retrieval Agents), a benchmark of over 8,000 executable APIs across 62 domains with tasks spanning three settings of increasing difficulty: diverse API interaction styles, multi‑hop reasoning over structured APIs, and multi‑source reasoning with natural‑language tool‑use policy constraints. Correctness is verified by re‑executing predicted tool calls against live APIs, accommodating multiple valid paths. Using a fixed ReAct harness to isolate model capabilities from agent architecture, we evaluate frontier and open‑weight models and find that even the best model achieves only 70.4% on single‑hop endpoint‑style tasks and drops to 50‑‑51% on compositional APIs; performance degrades by over 50% as reasoning depth increases, and policy‑constrained questions expose severe failures (as low as 2.4% on unanswerable queries). Trace analysis shows failures concentrate at language‑mediated reasoning ‑ entity disambiguation, cross‑source grounding, rather than tool invocation mechanics. Code is available https://github.com/IBM/VAKRA. Dataset is available https://huggingface.co/datasets/ibm‑research/VAKRA

Authors:Junming Zhang, Shuyu Yin, Peilin Liu, Rendong Ying, Fei Wen
Title: Curvature-Aware Zeroth-Order Optimization for Memory-Efficient Test-Time Adaptation
Abstract:
Test‑time adaptation (TTA) aims to enhance the cross‑domain performance of pre‑trained models by adapting to unlabeled test data. While most existing TTA methods rely on backpropagation (BP) for finetuning, BP‑free methods such as zeroth‑order (ZO) methods are more desired in practical on‑device scenarios. ZO methods rely only on forward computation, which can largely reduce the complexity and memory overhead of on‑device deployment. However, ZO methods suffer from much higher variance compared with first‑order methods in estimating the gradient. To address this, we propose an improved ZO method to substantially boost the performance of ZO optimization based TTA. First, we provide an observation to reveal the persistent low‑rank Hessian structure of the loss during the adaptation process. Based on this insight, we then propose a loss‑landscape curvature‑aware zeroth‑order (CAZO) method, which leverages a sliding‑average estimation of the diagonal Hessian to construct a covariance matrix for anisotropic perturbation sampling. CAZO operates by freezing pretrained weights and optimizing minimal adapter parameters via forward‑only passes based gradient estimation, which can substantially reduce the memory overhead compared to BP‑based methods. Extensive experiments demonstrate that CAZO significantly outperforms existing TTA methods, achieving state‑of‑the‑art performance while maintaining an excellent balance between accuracy and memory efficiency. Code is available at https://github.com/Hollyming/CAZO.

Authors:Wolfgang Gatterbauer
Title: An Extended Tutorial and Vocabulary for Relational Language Design in an Era of AI-Assisted Query Generation
Abstract:
Relational query languages have been studied and used for more than 50 years, with SQL dominant in practice. Today, queries are increasingly generated by machines and read by humans. At the same time, the landscape also includes dataframe, pipeline, logical, functional, graph, and relational programming notations. These developments invite two related questions beyond expressive power: which relational structures do languages make explicit, and how well can notation support users in reading and revising queries? This 3‑hour tutorial extends an earlier SIGMOD'26 tutorial in three directions: recursive and path queries (connecting relational and graph query languages), nested relational data, and relational languages for problems beyond PTIME. Rather than beginning from formal definitions, we start from example queries and compare how different languages express the same intent. To compare recurring structure across notations, we use Abstract Relational Calculus (ARC) and Relational Diagrams as reference representations. From these examples, we develop a vocabulary for relational language design, including information need, query mapping, relational pattern structure, relational pattern denotation, and semantic conventions. Participants will leave with a framework for comparing existing and future relational languages, a precise vocabulary for articulating design trade‑offs, and a concrete set of examples connecting classical database languages with alternative proposals.

Authors:Weihao Bo, Shan Zhang, Yanpeng Sun, Jie Liu, Yongke Yao, Jinhao Du, Wei He, Kai Zou, Zechao Li, Jingdong Wang
Title: Diagram-MMU: A Multi-Modal Benchmark for Scientific Diagrams
Abstract:
Multimodal Large Language Models (MLLMs) have been growing the capability for scientific writing and collaboration. For example, OpenAI Prism is a free workspace for scientific writing and collaboration. One important feature in Prism is turning scientific diagrams directly into LaTeX TikZ code. In this paper, we build a benchmark, Diagram‑MMU, a multi‑modal benchmark designed to assess MLLMs' ability for scientific diagram parsing and understanding. Diagram‑MMU features 3.7k curated diagrams and 18.3k human‑validated questions across six domains. It evaluates MLLMs on three tasks common in vibe writing workspaces: diagram‑to‑code parsing, diagram‑to‑code editing, and diagram question answering, alongside agentic settings per task. The evaluation of 12 MLLMs reveals that diagram‑to‑code tasks are more challenging than diagram question answering: models can reason well over diagrams but struggle to parse and edit them, underscoring the need for methods to enhance MLLMs' capability in diagram‑to‑code generation. Under agentic settings, most models improve parsing and editing performance but degrade on question answering, while Claude‑4.6 Opus consistently improves across all three tasks. Project Page: https://vi‑ocean.github.io/projects/diagram‑mmu.

Authors:Antoine de Mathelin, Christopher Tosh, Wesley Tansey
Title: ScreenShot: A Foundation Model for Few-Shot Combination Drug Screening
Abstract:
Treating patients with combinations of drugs reduces the risk of resistance to any individual drug. Finding effective combinations is difficult because the large search space makes combinatorial screens prohibitively expensive, time consuming, and often technically infeasible. Predictive models can fill this gap, yet existing methods typically require molecular profiling of each sample and per‑cohort training, limiting their applicability when time and tissue are scarce. To address this challenge, we introduce ScreenShot, a hierarchical transformer pretrained on 40 drug screening datasets covering 3,700 drugs and 6,000 biological samples, whose architecture mirrors the nested structure of screening data. Given a few‑shot context of observations from a new patient, ScreenShot predicts the response of the sample to combination therapies through in‑context learning, operating directly on functional measurements with no fine‑tuning and no molecular profiling. On four held‑out datasets, ScreenShot outperforms all baselines in both prediction accuracy and identification of selectively effective treatments. ScreenShot's internal representations are directly useful for experimental design: we use them to drive a weighted k‑means++ active learning strategy that selects which experiments to run, achieving the same hit detection as uniform screening with a third of the budget. Source code and interactive dashboard: https://github.com/tansey‑lab/screenshot.

Authors:Yung-Hsu Yang, Luigi Piccinelli, Samuel Rota Bulò, Sunghwan Hong, Denis Rozumny, Johannes Schönberger, Zuria Bauer, Hermann Blum, Peter Kontschieder, Marc Pollefeys
Title: Map-Det3D: Metric Feed-Forward 3D Reconstruction Prior for Multi-view 3D Object Detection from Streaming Inputs
Abstract:
Metric 3D object detection is a core capability for embodied agents, yet most reliable systems lean on depth sensors, trading away cost, power, and integration simplicity. This motivates monocular 3D detection, which avoids additional constraints, yet it faces a major obstacle: from a single image, depth, and especially absolute scale, are underconstrained. As a result, the prevailing pattern of detecting in 2D and then predicting 3D attributes is often brittle, since modest range errors can dominate 3D localization, and the learned scale prior can fail when cameras, motion, or environments undergo domain shifts. To address this, we propose Map‑Det3D, an online multi‑view 3D object detection model that brings detection directly into a 3D space reconstructed from RGB. We map a short temporal window into multiple views and repurpose a feed‑forward metric 3D reconstruction model as our geometric backbone while tuning its object‑aware capabilities. Building on this representation, Map‑Det3D directly predicts boxes in metric 3D space, without the widely used 2D‑to‑3D lifting. Experiments across different benchmarks show that this design supports strong online performance and robust transfer without adaptation, suggesting that training reconstruction priors for detection is a practical route to stable metric 3D detection from monocular video. Code and models are available at https://royyang0714.github.io/Map‑Det3D.

Authors:Byungoh Ko, Jinyoung Park, Jongha Kim, Jeehye Na, Jaewon Cho, Hyunwoo J. Kim
Title: Context Blindness in DPO: Mitigating Object Hallucination in MLLMs via Context-Calibrated Preference Optimization
Abstract:
Multimodal large language models (MLLMs) have made rapid progress, yet they still exhibit object hallucination, generating plausible but incorrect descriptions that are inconsistent with the visual input. Direct Preference Optimization (DPO) mitigates this by training models to prefer non‑hallucinated responses over hallucinated ones, and recent efforts further enrich the preference data with relevant context. However, it remains unclear whether DPO actually leverages such context. To investigate this, we propose Contextual Preference Gain (CPG), a simple metric that measures how much a model's preference strengthens when relevant context is provided. We find that higher CPG consistently corresponds to lower hallucination, yet standard DPO and its variants exhibit only limited CPG, indicating that they underutilize contextual information and thus remain prone to hallucination. To address this, we propose Context‑Calibrated DPO (C^2‑DPO), which directly maximizes CPG while preserving the original preference ordering. Across multiple benchmarks, C^2‑DPO substantially reduces hallucination without compromising general reasoning, relatively reducing the Object HalBench hallucination rate of Qwen2‑VL‑Instruct‑2B by 36%. Code is available at https://github.com/mlvlab/C2‑DPO

Authors:Zunhai Su, Bohan Sun, Xialie Zhuang, Shuibai Zhang, He Xiao, Jing Xiong, Hengyuan Zhang, Zhongzhu Zhou, Tiantian Zhang, Ngai Wong, Chuan-Wei Kuo
Title: Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus
Abstract:
We present the first systematic study of Massive activations (MAs) in layer‑interleaved HLA LLMs and uncover two architecture‑aligned morphologies: MAs consistently spike immediately before full attention layers, forming pre‑attention spikes (PAS), and can persist through intervening linear attention layers, giving rise to inter‑spike plateaus (ISP). As full attention becomes denser, successive PAS become increasingly connected through ISP, ultimately recovering the stable MA morphology of full attention LLMs. We establish the recurrence of this organization across five linear attention architectures, six hybridization configurations, five data domains, and representative open‑source hybrid models spanning 1.2B to 397B total parameters. Controlled pretraining of GDN‑based hybrids at scales up to 1.3B shows that both morphologies emerge early and respond asymmetrically to output gating: full attention output gating strongly attenuates their absolute magnitudes without eliminating their layerwise organization, whereas removing GDN gates yields comparatively modest amplification. Mechanistically, our systematic‑outlier analysis supports a shared lifecycle account governed by the timing of MA cancellation. PAS follows a localized write‑sink‑cancel process, while the extended persistence of ISP is consistent with delayed cancellation. At the full attention limit, this account recovers the stable MA morphology characteristic of full attention LLMs. Our code is available at https://github.com/StartluxLabs/Massive‑Activations‑HLA.

Authors:Liangwei Li, Lin Liu, Jing Zhang, Xiaohui Du, Ruqian Hao, Xinwei Li, Hanzhe Liang, Juanxiu Liu
Title: MVFM-3DAD: Multi-view Flow Matching for 3D Anomaly Detection via Density Proxy Estimation
Abstract:
In 3D anomaly detection (3DAD), most existing methods rely on Memory bank retrieval or reconstruction. However, memory‑based methods are constrained by the coverage of stored normal features, while reconstruction‑based methods may learn identity shortcuts that also reconstruct anomalous inputs well. These limitations motivate a density‑oriented approach that evaluates whether a test sample follows the learned normal distribution. To this end, we propose MVFM‑3DAD, a flow‑based framework that reframes 3DAD as density proxy estimation over the normal data distribution. MVFM‑3DAD introduces a Bidirectional Geometric Projector (BGP), whose forward process converts irregular point clouds into structured multi‑view representations. The Flow‑guided Density Proxy Estimator (FDPE) estimates a reference density for each view feature, after which the backward process of BGP maps these multi‑view density estimates to their corresponding 3D points. Building on it, anomalous features can be identified by their terminal normality. Unlike conventional flow‑based likelihood estimation, our formulation requires neither input reconstruction nor explicit Jacobian evaluation, yielding a simple and efficient anomaly‑scoring mechanism. Extensive experiments show that MVFM‑3DAD outperforms the strongest competing methods on Real3D‑AD and MVTec3D‑AD. Code is available at https://github.com/lil‑wayne‑0319/MV3D‑AD

Authors:Zhiqiang Que, Chang Sun, Haiyang Wang, Dinesh Pamunuwa, Roshan Weerasekera, Qijia Tang, Bakhtiar Zadeh, Wayne Luk, Maria Spiropulu
Title: FQTree: Fine-grained Quantization and Hardware Generation of Boosted Decision Trees
Abstract:
Boosted decision trees (BDTs) are widely used in latency‑critical applications, but efficient hardware deployment remains challenging. Existing designs often rely on uniform or manually tuned fixed‑point formats, which can introduce unnecessary hardware cost or accuracy loss. This work presents the FQTree algorithmhttps://github.com/ecs‑bristol/FQTree for fine‑grained quantization‑aware training of BDTs, together with the QXGB framework for automatic hardware generation. FQTree introduces a hardware‑oriented leaf‑value quantization scheme that uses a global quantization step together with a tree‑wise shift, enabling compact non‑negative integer leaf representations, controlled clipping/pruning, and bias folding to reduce datapath cost. This work further applies this quantization during boosting so that later trees adapt to the errors of the already‑quantized ensemble, and then lowers the trained model into low‑latency hardware implementations through a compiler‑based flow. Results on JSC, MNIST, and NID show that our method reduces LUT usage by 26‑57% compared with the state‑of‑the‑art FPGA‑based BDT designs while matching or improving accuracy.

Authors:Josef Liyanjun Chen
Title: Ready Cohorts: Bounding GPU Opportunity and Avoiding Host Round Trips in LLM-Agent Control
Abstract:
LLM‑agent services repeatedly execute small deterministic transitions between model and tool calls: route an outcome, update state, and emit the next effect. We ask when this control path exposes enough concurrent work for GPU execution, and what changes when a GPU‑computed route decision remains on device. We formalize the ready‑cohort boundary using fixed‑partition share F, exact offline share P, local upper bound U, and online achieved share A. Under zero service time, unlimited capacity, and equal relative launch deadlines, a specialized dynamic program computes P exactly. In a stationary Poisson replay of one pinned 851‑session public trace panel, the primary condition at 100,000 target active sessions, K=256, and a 50 ms launch deadline gives F=30.19%, P=43.00%, and U=45.85%. Exact packing recovers 81.83% of the opportunity lost at fixed window boundaries. The outcome‑derived route key is a conditioning proxy, not proof of executable identity. A separate mechanism study keeps a GPU‑computed binary decision on device instead of returning four bytes to the host and redispatching. Across four named GPU placements, the device‑resident path is faster in all 36 configurations; within‑placement row‑median ratios range from 1.19x to 2.39x. Across both admissible mechanisms, all 14,557,440 tested batched invocations match a separately implemented host oracle. A fixed nested device graph that removes no host decision is slower in all 60 configurations across five placements. Together, the studies establish two measurable gates for GPU agent control: deadline‑feasible cohort supply and observation placement. A joined finite online runtime is required to measure A, CPU displacement, and service‑level benefit.

Authors:Nikolette Pedersen, Regitze Sydendal, Veronika Cheplygina, Théo Sourget
Title: Look What the Probes Dragged In! Real-World Chest X-ray Shortcuts in MedCLIP
Abstract:
Vision‑language models, such as contrastive language‑image pre‑training (CLIP)‑based approaches, have reached state‑of‑the‑art (SOTA) results in medical artificial intelligence. However, recent work reveals that CLIP‑based models remain vulnerable to shortcuts. We investigate how real‑world shortcuts manifest across different layers of the medical CLIP‑based model, MedCLIP, and its vision encoder, a frozen ResNet‑50. We attach 17 linear classification probes to the intermediate layers of the ResNet‑50 and train them on three different dataset configurations and targets: NIH‑CXR14 (pneumothorax) and PadChest (cardiomegaly and pneumothorax). This setup allows us to observe model behaviour during evaluation using subgroup‑based calibration and layer‑wise confidence curves. We find that the final linear probes achieve a high AUROC but poor calibration in the models. The layer‑wise confidence analyses suggest that shortcuts emerge at different depths. Patterns consistent with localised shortcuts, such as drains, appear at later layers, while patterns consistent with diffuse shortcuts, such as scanner‑specific noise patterns, emerge earlier, aligning with previous work. Finally, we conduct a manual analysis of the images, which reveals data quality issues in both NIH‑CXR14 and PadChest. Our findings underscore that even SOTA models remain vulnerable to shortcuts, and the need for high‑quality and well‑annotated datasets to draw solid conclusions. Code can be found on our GitHub: https://github.com/nikodice4/MedCLIP_shortcuts.

Authors:Kangning Zhang, Haotian Fang, Xukun Luo, Hao Yin, Yang Gao, Peng Yan, Weiwen Liu, Weinan Zhang, Yong Yu
Title: HCGRec: Hint-Conditioned Generative Recommendation with Semantic IDs
Abstract:
Semantic‑ID generative recommenders represent each item as a short sequence of discrete semantic tokens and predict the next item by autoregressively generating this token sequence. This paradigm enables a unified generation interface for item IDs, histories, and item text, but it also creates a structured optimization bottleneck during reward‑based post‑training: when an early semantic token enters the wrong branch of the item‑token space, finite rollout groups rarely reach the ground‑truth item, so group‑relative optimization receives identical zero rewards and produces no useful advantage. We propose Hint‑Conditioned Generative Recommendation (HCGRec), a semantic‑ID generative recommendation framework that recovers learning signal for such hard training instances. HCGRec diagnoses each instance with checkpoint rollouts and supplies a minimal target‑prefix hint only when the current generator cannot reach the correct item. The model then generates the unhinted suffix under the hinted semantic branch, turning zero‑reward groups into informative comparisons over item‑token completions. Hinting also changes token identity: hinted prefix tokens are oracle‑provided item context, while unhinted suffix tokens are sampled generation actions. We therefore introduce hint‑aware credit decomposition, using supervised learning to preserve item‑semantic and prefix‑structure alignment for hinted tokens and GRPO to optimize the sampled suffix. Experiments on sequential recommendation benchmarks show that HCGRec substantially improves over supervised fine‑tuning and vanilla reward‑based post‑training, while reducing zero‑advantage training samples from over 70% to below 20%. The code is accessible at https://github.com/WncFht/GRec.

Authors:Daehwa Ko, Jaehyeon Kim, Seunghyun Ham, Jay Hoon Jung
Title: TESLA: Taylor Expansion of Sinusoidal Learnable Activations
Abstract:
The parity problem‑‑deciding whether the number of ones in a binary vector is odd or even‑‑remains challenging for standard neural networks due to linear inseparability and the need for global interactions. We propose TESLA, an activation defined as a learnable combination of sine and cosine terms, enabling explicit control over polynomial degree and selective amplification of high‑order components. Theoretically, we show that constraining TESLA's coefficients yields Lipschitz/Rademacher complexity bounds and shapes the training dynamics to emphasize higher‑frequency structure. Empirically, on parity with input length n = 32, TESLA attains strong generalization with 100K training samples (approximately 0.002% of the 2^32 input space) and remains robust under heavy corruption, retaining high accuracy with up to 30% label noise. We also compare against periodic and frequency‑based baselines (SIREN, SNAKE, and Fourier feature embeddings) on parity and Forrelation. Beyond synthetic structure, TESLA delivers comparable performance on ImageNet‑100, indicating that activation‑level degree control transfers to more general vision workloads. Code: https://github.com/KAU‑QuantumAILab/TESLA

Authors:Dongxu Tang, Shih Ying-Lei, Zhuoyi Ren, Jianting Liao, Yitian Shao
Title: Synchronized AMG and EMG Dataset of Lower-limb Muscle Activities in Everyday Training
Abstract:
Understanding how lower‑limb muscle groups coordinate is important for studying movement impairment, rehabilitation, and physical performance. Reproducible analysis of this coordination requires multimodal recordings that relate local muscle‑related signals with body‑level kinematics. Complementing neural‑level electrical activation captured by EMG, AMG provides a valuable mechanical approach to monitoring muscle activity. Here, we introduce a synchronized, multimodal dataset for healthy‑adult lower‑limb activities. For data collection on the left leg, 16 triaxial accelerometers were evenly divided into four muscle‑site clusters for AMG recording, complemented by four surface EMG channels. A 15‑marker optical motion‑capture (MoCap) system captured lower‑body kinematics, with the resulting marker trajectories used to compute bilateral knee and ankle joint angles. Our dataset contains 1,918 trials from 30 subjects across 16 task conditions. We benchmark the dataset by estimating four joint angles from 300 ms windows of the 5‑100 Hz band‑pass‑filtered AMG data and assess matched EMG features in a separate modality ablation. In the primary cross subject benchmark, the four reference models achieved mean absolute errors of 8.840^\circ‑9.591^\circ. The benchmark and ablation results characterize performance across subjects, tasks, and joint angles and examine the effects of sensor configuration, modality, the number of training subjects, and frequency representation. The release includes documented timing definitions, processed data, and reproducible benchmark resources. https://dongxutang918‑afk.github.io/SAME‑Limb/

Authors:Karl Hanna, Chen Feng
Title: Accuracy and Order Sensitivity Diverge Under Label-Free Strategies
Abstract:
Multiple‑choice benchmarks are widely used to evaluate large language models, but MCQ scores conflate knowledge with sensitivity to option order, which makes them unreliable measures of model knowledge. In this paper, we test whether preventing a model from seeing option labels while committing to an answer removes positional influence and, in turn, improves performance. We evaluate two different strategies for mitigating bias. The first uses a generation‑then‑matching approach, and the second scores options in isolation, which is positionally unbiased by construction. Neither reliably improves accuracy. A complete decomposition shows that the bottleneck is withholding options, not the matching step. The only configuration that consistently matches the baseline is the one that shows the model all options paired with an LLM matcher. However, eliminating positional influence entirely still does not reliably yield accuracy gains, while cyclic permutation often improves them. For two‑stage prompting, an aggregate measure of recall imbalance and a direct per‑question measure of order sensitivity both fail to show reliable debiasing.

Authors:Tom Adamczewski
Title: OEIS Open: How many conjectures can language models turn into theorems?
Abstract:
We construct OEIS Open, a benchmark based on 492 open mathematical conjectures from the OEIS, formalized in Lean by Tsoukalas et al. Whereas these conjectures had previously been attempted only with a bespoke agent, our open‑source evaluation code runs any generic language model (LM) against them, and is secure against LM cheating attempts. We find that LMs equipped with a minimal set of tools resolve 147 of these conjectures with a budget of \50 per attempt, scoring 30% on OEIS Open. OEIS Open Lite is a random subset of 100 conjectures for cheaper evaluation. When evaluated with a budget of \200 per attempt, the best current LM scores 44% on OEIS Open Lite. Giving LMs access to the mathematics literature via 476,000 papers from arXiv did not increase performance on OEIS Open Lite, and nor did using more sophisticated agent loops. The conjectures covered in this work are of uncertain mathematical significance, and most have likely received little previous attention. Nevertheless, our results show that LMs can resolve open research conjectures autonomously and at modest cost.

Authors:Xiaojun Wu, Cehao Yang, Honghao Liu, Xueyuan Lin, Xuhui Jiang, Chengjin Xu, Jia Li, Jian Guo
Title: LazyTrain: Limited-resource Allocation toward Zero-waste Yield Optimization in Large Language Model Training
Abstract:
Training large language models on limited hardware is increasingly a scheduling problem across GPU compute, host memory, PCIe transfer, and storage bandwidth. Existing offloading systems reduce GPU residency, and MegaTrain shows that a CPU‑master layer‑streaming executor can train large models on a single GPU, but fixed checkpointing and placement heuristics still leave communication exposed on the critical path. We propose LazyTrain, an optimization layer over a layer‑streaming executor. LazyTrain formulates checkpoint selection, activation placement, recomputation, and CPU‑GPU‑NVMe communication overlap as a mixed‑integer scheduling problem, then executes the solved policy during training. It further couples 8‑bit optimizer states with fast gradient clipping as a single Hybrid 8‑bit operator: state compression reduces optimizer‑state memory, while fast clipping counteracts the additional CPU‑side update overhead. Across H800 experiments from Qwen2.5‑3B to Qwen3.6‑27B, LazyTrain improves sustained TFLOPS over matched baselines runs by approximately 1.24×; RTX 3090 experiments likewise increase the maximum feasible batch size by one at each model scale. In the primary Qwen3.6‑27B H800 MetaMathQA run, LazyTrain reaches 219.95 TFLOPS and 1361 tokens/s at batch size 72, peaks at 68.84\,GB of GPU memory, and obtains 95.42% exact‑match accuracy on the full evaluation split. The source code is available at https://github.com/DataArcTech/LazyTrain.

Authors:Zihao Xie, Pingrui Lai, Yitong Wu, Hua Yang
Title: DaViNCi: A Dataset Towards Outdoor Vision-and-Language Navigation with Continuous Actions and Dynamic Elements
Abstract:
Vision‑and‑Language Navigation (VLN) has progressively expanded from indoor to outdoor environments. However, existing outdoor VLN datasets still rely on fixed discrete topological graphs for construction. It fails to align with the rapidly changing real‑world outdoor environments and impedes the sim‑to‑real transfer of VLN agents. To address this limitation, we propose DaViNCi (Dynamic Vision‑and‑Language Navigation in Continuous Environment), the first outdoor VLN dataset that simultaneously introduces both continuous and dynamic factors. The agent not only moves in the outdoor environment using continuous actions but is also required to handle unpredictable dynamic elements. The dataset encompasses six distinct maps with a total of 6,933 trajectories. Through comprehensive comparative experiments, we find that the success rate on DaViNCi decreased by more than 10% in discrete environments compared to previous datasets. And there is an even greater decline in continuous settings, demonstrating the challenge of DaViNCi. Furthermore, we clarify the impact of action granularity and dynamic elements. These results demonstrate the practical value of DaViNCi in advancing outdoor VLN toward more realistic environments. The website is https://xzh0312.github.io/DaViNCi/.

Authors:Zheyu Zhuang, Ruiyu Wang, Nils Ingelhag, Ville Kyrki, Danica Kragic
Title: Enhancing Visual Domain Robustness in Behaviour Cloning via Saliency-Guided Augmentation
Abstract:
In vision‑based behavior cloning (BC), conventional image augmentations such as Random Crop and Color Jitter often fall short under substantial visual domain shifts, including changes in shadows, distractors, and backgrounds. Superimposition‑based augmentations, which blend in‑domain and out‑of‑domain images, have shown promise for improving generalization in computer vision, but their suitability for BC remains uncertain because task‑critical semantics, spatiotemporal relationships, and agent‑target interactions must be preserved. To address this, we introduce RoboSaGA, a Saliency‑Guided Augmentation method within the superimposition family tailored for vision‑based BC. RoboSaGA dynamically adjusts augmentation intensity at the pixel level using policy‑driven saliency, enabling aggressive augmentation in task‑irrelevant regions while preserving task‑critical information. It integrates seamlessly into existing architectures without requiring structural modifications or additional learning objectives. Experiments in both simulated and real‑world settings show that RoboSaGA preserves in‑domain performance while substantially improving robustness to visual domain shifts, including distractor and background changes, as well as lighting and shadow variations. Code is available at https://github.com/Zheyu‑Zhuang/RoboSaGA.

Authors:Beomsik Cho, Jinhyeong Kim, Dongseok Lee, Jaehyung Kim
Title: LookBack: Where and How to Score LVLM Responses via Visual Reference Usage
Abstract:
Large Vision‑Language Models (LVLMs) integrate visual perception with language generation, enabling responses that span image understanding and complex reasoning. However, LVLMs do not just inherit the text‑level hallucinations; they also hallucinate against the image, producing fluent responses ungrounded in what they see. This makes LVLM response scoring inherently harder, and our diagnostics show that existing confidence‑based metrics adopted from LLMs are insufficient for LVLMs. Specifically, removing the input image barely changes confidence‑based selection, suggesting that output‑space confidence primarily captures textual plausibility rather than agreement with the image. To address this gap, we propose LookBack, a training‑free LVLM response scoring method that augments token likelihood with visual lookback score, a lightweight measure of how strongly each response token refers to image tokens. Across four benchmarks and three models, LookBack consistently improves Best‑of‑N selection over existing baselines with negligible additional overhead.

Authors:Junyoung Kim, Wonbin Kweon, Woojoo Kim, Jaehyung Lim, Dongha Kim, Hwanjo Yu
Title: From Overlooked to Explored: Recovering Item Relations via Mixture of Perspectives for Sequential Recommendation
Abstract:
Capturing user preference from a user's interaction sequence is the central challenge of Sequential Recommendation (SR). This preference intuitively emerges from inter‑item relations: each item transition reflects a preference embedded in the relations between items, making the faithful capture of these relations essential for accurate recommendation. For this reason, self‑attention is dominant in sequential recommendation for its ability to compute pairwise item interactions, yet our empirical analysis reveals that it consistently suffers from similarity bias across various types of transformer‑based SR models: dot‑product attention scores disproportionately favor similar items, systematically overlooking heterogeneous relations with meaningful preference signals and directly limiting recommendation performance. To address this, we propose PRISM (Perspective‑based Relational Insight Synthesis Module), a module that re‑examines item relations from multiple perspectives. PRISM employs K Perspective Lenses to calibrate attention from distinct viewpoints, combining an Affinity View that refines homogeneous relations and a Contrast View that exposes heterogeneous ones suppressed by similarity bias, enabling the model to capture the full spectrum of user preferences. Extensive experiments on seven real‑world benchmarks demonstrate that PRISM consistently outperforms state‑of‑the‑art baselines. Our code is available at https://github.com/327aem/PRISM/.

Authors:Xueqin Niu, Mufan Liu, Yifan Wang, Le Yang, Jun Sun, Yiling Xu
Title: ResPCC: A Loss-Resilient Neural Point Cloud Codec over Lossy Networks
Abstract:
Point cloud compression (PCC) is critical for efficient storage and transmission of 3D data. While recent learning‑based PCC methods achieve good rate‑distortion (R‑D) performance, they generally rely on ideal transmission conditions. In practice, packet loss is a common issue and can severely distort latent features, causing coordinate drift and geometric degradation. To address this challenge, we present ResPCC, the first end‑to‑end neural point cloud codec designed to offer intrinsic resilience against data loss. Our framework is loss‑rate‑aware and adapts to diverse packet loss conditions. At the encoder, we introduce a Condition‑Adaptive Latent Modulation (CALM) module to adjust latent feature distributions according to the perceived loss rate, as well as a Spatial‑Channel Interleaving (SCI) mechanism that transforms channel‑wise data extinction into spatially scattered element‑wise missing patterns. At the decoder, we develop a Mask‑Aware Graph‑based Latent Restoration (MGLR) module, followed by a Dictionary‑based Refinement (DBR) stage to recover corrupted features and align them with canonical priors. Evaluations on ShapeNet and SemanticKITTI under 5% to 30% packet loss rates show that ResPCC consistently delivers superior stability and R‑D performance over baselines. Our framework maintains high reconstruction fidelity under lossy conditions, providing a reliable solution for 3D data transmission over practical networks. Code is available at https://github.com/starrynight314/ResPCC.

Authors:Jinhyung Bae, Dain Kil, Seongmin Oh, Seungmin Lee
Title: When the Knowledge Base Becomes the Gold Standard: Measuring Resource-Shared Evaluation Loops in Entity-Level Machine Translation
Abstract:
The Seungjeongwon Ilgi, a UNESCO Memory of the World record, is only 37.4% translated, and the most conspicuous failure mode in automatic translation is the person name ‑‑ a misread name corrupts the historical fact rather than merely the surface. Low‑resource historical domains have no expert gold standard for entity translation, so practitioners substitute a knowledge base (KB) for the gold. That KB is the same resource injected into the system: scoring becomes self‑referential and the metric measures instruction compliance rather than translation quality. We measure this loop. Using expert person‑name annotations from the National Institute of Korean History as a gold independent of the injection pipeline, we hold the entity set fixed and vary only the provenance of the correct reading. Of 527 expert‑annotated mentions, only 31.1% lie outside the injection pipeline, and the residual loop is not uniform ‑‑ in the overlapping segment the injected reading agrees with the human translation 97.8% of the time against 70.1% in the independent one, so the segment that looks healthiest is the one the loop is holding up. Across four models, a difference‑in‑differences analysis shows the gain from KB injection is confined to the segment whose gold shares the injected resource; in the independent segment it is at or below zero. Post‑injection preservation clusters in a narrow 0.910‑0.996 band even though baseline capability differs fivefold, so the reported gain is the complement of prior performance and weaker models appear to improve more dramatically. On an independent sample built by removing the construction filter, the measure replicates within model (overlapping intervals) while discriminating between models (non‑overlapping intervals) ‑‑ it reflects a property of the model, not of the sample.

Authors:Yaohua Liu, Yifan Guo, Jiaxin Gao
Title: Learning with Bilevel-Minimax Optimization for Efficient and Reliable Transfer Attacks
Abstract:
Transfer‑based adversarial attacks craft adversarial examples using surrogate models to mislead black‑box victim models. Beyond perturbation generation, transferability is fundamentally governed by the coupling of initialization, surrogate adaptation, and gradient dynamics. We revisit this challenge from a bilevel‑minimax perspective and propose BMAT (Bilevel‑Minimax Adversarial Transfer). The bilevel formulation captures the dependency between initialization and perturbation, while the inner minimax problem promotes surrogate robustness for cross‑architecture generalization. Algorithmically, we develop an integrated bottom‑up solver that combines a Soft Weight Modulator and an Implicit Gradient Approximator to enable ternary coupling among initialization, surrogate adaptation, and perturbation optimization. We further provide theoretical insights into the optimization dynamics of the proposed bilevel‑minimax framework. Extensive experiments on classification and segmentation benchmarks show that BMAT outperforms more than 10 strong baselines across more than 30 victim models, improving both intra‑ and cross‑architecture transfer and yielding up to a 2x reduction in mIoU. Code is available at https://github.com/callous‑youth/BMAT.

Authors:Hoai Nhan Pham, Dang-Nguyen Bui, Le-Van Thai, Thanh-Hiep Vo, Lan Anh Dinh Thi, Tien Dat Nguyen, Duy-Dong Nguyen, Ngoc Lam Quang Bui, Tam Tran, Zhi Huang
Title: CoDiR: Confidence-Guided Diffusion Refinement for Semi-Supervised Histopathology Segmentation
Abstract:
Semi‑supervised histopathology segmentation is challenging due to scarce annotations and unreliable pseudo‑labels in ambiguous gland regions. To address this problem, we propose Confidence‑Guided Diffusion Refinement (CoDiR), a semi‑supervised framework that combines a Mean Teacher segmentation model with diffusion‑based pseudo‑label refinement. Given an unlabeled image, the teacher first produces a soft prediction, and only low‑confidence regions are refined by a conditional diffusion model trained to capture plausible mask structures from labeled data. The refined mask is then fused with reliable teacher predictions and used to train the student with confidence weighting and consistency regularization. On the GlaS and CRAG datasets CoDiR reaches 88.09% and 89.83% mDice with 10% labeled data, and 89.19% and 90.29% mDice with 20%, matching or exceeding the strongest published method on seven of the eight benchmark metrics. Ablations attribute the largest single contribution to the refinement module, which adds +6.36% mDice over the Mean Teacher baseline. The implementation code is publicly available at: https://github.com/vongla345/codir

Authors:Xingwei Sun, Heinrich Dinkel, Gang Li, Jiahao Mei, Yadong Niu, Zerui Han, Yuepeng Jiang, Jiahao Zhou, Lichun Fan, Jian Luan
Title: MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching
Abstract:
Generating coherent audio scenes that simultaneously blend speech, music, and sound effects remains a significant challenge. Current approaches typically rely on a disjointed pipeline where a frozen, decoupled text encoder feeds a separate audio decoder, limiting cross‑modal optimization and leading to poor speech intelligibility. To overcome these limitations, we introduce MiDashengLM‑Gen, an end‑to‑end framework that couples a pre‑trained Large Language Model (LLM) with per‑token conditional flow matching for autoregressive, variable‑length mixed‑audio scene generation. MiDashengLM‑Gen represents a first approach for general text‑to‑audio generation with one end‑to‑end trained model. Empirical evaluations demonstrate that MiDashengLM‑Gen drastically improves speech intelligibility over existing unified models. On the Seed‑TTS benchmark, English Word Error Rate (WER) drops from 12.15% to 2.79%, approaching the performance of dedicated Text‑to‑Speech (TTS) systems (1.24%). Furthermore, the framework extends effectively to multilingual settings, yielding highly competitive multilingual WERs compared to existing baselines. Lastly, the model maintains competitive mixed‑audio generation quality on the MECAT benchmark. Code and checkpoints are available at https://github.com/xiaomi‑research/midashenglm‑gen and https://huggingface.co/mispeech/midashenglm‑gen, and the demo page is available at https://xingws.github.io/midashenglm‑gen‑demo/.

Authors:Muhammad Ayub Sabir, Junbiao Pang, Fatima Ashraf
Title: High-Order Liquid Evidence Encoding for Gradual GNSS Spoofing Detection in Autonomous Driving
Abstract:
Accurate Global Navigation Satellite System (GNSS)‑based localization is essential for safe and reliable autonomous driving. However, spoofing attacks can manipulate vehicle position estimates. Continuous and subtle attacks are particularly difficult to detect because individual GNSS observations may remain plausible while the inconsistency between GNSS‑implied displacement and onboard vehicle motion gradually increases. Existing methods often rely on static vehicle‑behavior features or a single residual signal and do not explicitly model this evolution. To address this problem, we propose a causal high‑order liquid evidence framework for GNSS spoofing detection. The method first constructs a physics‑guided GNSS‑‑motion inconsistency residual by comparing GNSS‑implied displacement with onboard‑motion‑derived displacement. It then forms separate evidence streams for the residual level and its first‑ and second‑order discrete variations, with relevant contextual cues selected according to the evidence order. Each stream is processed by a separate adaptive liquid encoder, and the resulting temporal states are hierarchically coupled to predict spoofing at the window endpoint using only current and past observations. Experiments on three subsets of the real‑world AV‑GPS dataset show that the proposed method achieves the highest F1‑scores among the evaluated temporal models on Dataset~1 and Dataset~3, reaching 0.9535 and 0.9777, respectively. On Dataset~3, it detects both labeled normal‑to‑attack transitions within four sampling steps. Code and datasets are publicly available at: https://github.com/pangjunbiao/GNSS_Spoofing.git.

Authors:Nicholas E. Kyrkewood
Title: The Sleeping Agent: What Gist-Based Context Compression Loses and Why
Abstract:
Gist‑based context compression‑‑‑summarising older conversation history into compact representations‑‑‑is a common approach in long‑horizon language model agents, yet its effect on different types of memory retrieval is poorly understood. We use Salience‑Weighted Consolidation (SWC), a biologically‑inspired compression framework motivated by sleep‑based memory consolidation, as a diagnostic probe to study when gist compression helps and when it hurts. SWC scores conversation history by salience, partitions it into priority tiers, and applies structured gist abstraction to mid‑priority content. Evaluating four conditions on all ten LoCoMo conversations‑‑‑1,935 matched text‑only questions in total, 1,501 used in the primary aggregate after excluding Category 5 (adversarial) questions‑‑‑at temperature 0, we find a consistent task‑type interaction: gist compression substantially outperforms truncation on multi‑hop reasoning and single‑hop factual questions, but temporal questions remain substantially harder under compression, with compressed conditions scoring well below the full‑context reference on the conversations where both are evaluated. We trace this failure to a specific mechanism: the gist abstraction prompt preserves relational and event structure while discarding dates and times. A preservation analysis across all ten conversations confirms the mechanism: an approximately 20‑fold increase in temporal expression preservation (3.05% to 62.39%) with a one‑sentence prompt modification, while named entity and event preservation rates barely change (x1.02 and x1.11), demonstrating that the fix is a precision instrument. The prompt modification recovers +0.314 [0.254, 0.375] judge accuracy on category‑2 (temporal) questions in the matched set. Code and results: https://github.com/kyrkewood/sleeping‑agent.

Authors:Duy-Dong Nguyen, Le-Van Thai, Hoai Nhan Pham, Ngoc Lam Quang Bui, Tam Tran, Zhi Huang
Title: ProBAG: Prototype-Guided Boundary-Aware Graph Diffusion for Weakly Supervised Histopathology Segmentation
Abstract:
Weakly supervised semantic segmentation enables histopathology tissue segmentation from image‑level annotations, avoiding costly pixel‑level labeling by expert pathologists. However, CAM‑based methods often localize only highly discriminative regions and remain unreliable near tissue interfaces. We propose ProBAG, a stage‑1 pseudo‑mask generator that combines dataset‑specific visual prototypes with pathology‑aligned CONCH text prototypes over multi‑scale frozen UNI features. ProBAG introduces two complementary mechanisms: class‑wise power recalibration that reshapes inter‑class competition while preserving the total foreground activation mass at each pixel, and one‑step graph diffusion in which feature affinities are penalized by a late‑transformer attention‑context discrepancy used as a soft structural boundary cue. The resulting stage‑1 pseudo‑masks require neither CRF nor an external segmentation model; for complete two‑stage comparison, they additionally supervise a downstream Phikon‑FPN segmenter. Experiments on BCSS‑WSSS and LUAD‑HistoSeg show consistent gains over recent WSSS approaches, while ablations indicate that pathology‑aligned text semantics provide the largest improvement and graph refinement provides a smaller complementary gain. The code is available at: https://github.com/wterrr/WSSS

Authors:Juncheng Liao, Jinfan Lv, Guoming Wang, Jupeng Zheng, Ling Xiao, Siliang Tang
Title: AWARe: Mitigating Catastrophic Forgetting via Activation-Weighted Adaptive REtention
Abstract:
Multimodal Large Language Models (MLLMs) exhibit strong generalization and reasoning abilities due to large‑scale multimodal pre‑training. However, fine‑tuning these models on downstream tasks often leads to catastrophic forgetting, where newly learned task‑specific knowledge degrades previously acquired capabilities. This issue arises because gradient updates for new tasks overwrite parameters critical to prior knowledge, limiting the practical deployment of MLLMs. To address this challenge, we propose Activation‑Weighted Adaptive REtention (AWARe), a fine‑tuning method that mitigates catastrophic forgetting by dynamically controlling parameter updates based on activation patterns. AWARe assigns activation‑based importance scores to parameters, selectively freezing those essential for preserving prior capabilities while allowing less important parameters to adapt to new tasks. Importantly, AWARe operates without modifying model architectures, ensuring compatibility with existing inference engines. Extensive experiments demonstrate that AWARe effectively preserves upstream capabilities while achieving superior downstream performance compared to existing methods. Code is available at https://github.com/kaln27/AWARe.

Authors:Jiabao Zhuang, Changhao Jiang, Hanchen Wang, Jiahao Chen, Zhixiong Yang, Zhenghao Xiang, Yifei Cao, Jiajun Sun, Hui Li, Ming Zhang, Tao Ji, Tao Gui, Qi Zhang, Xuanjing Huang
Title: MuseCritic: Learning Multi-Aspect Song Rewards through Natural-Language Aesthetic Critiques
Abstract:
Long‑form song generation models continue to improve in duration, structural integrity, and acoustic complexity, making reliable aesthetic rewards increasingly important for aligning these models with human preferences. However, reward models for complete songs remain limited, and existing evaluators typically predict scores in a single forward pass without providing readable explanations. We introduce MUSECRITIC, a semi‑scalar reward model that generates a natural‑language critique covering five aesthetic dimensions and uses it as an intermediate representation to predict continuous reward scores. MUSECRITIC follows a two‑stage training pipeline: a teacher model first provides high‑quality critiques for supervised fine‑tuning, after which the fine‑tuned model generates its own critiques for reward learning, mitigating distribution shift between training and inference. On an in‑domain test set of 200 SongEval songs, MUSECRITIC reduces macro‑averaged mean squared error from 0.2875 to 0.2316 and improves macro‑averaged LCC, SRCC, and Kendall's tau to 0.9068, 0.8838, and 0.7178, respectively. On the out‑of‑domain Music Arena benchmark with 733 preference pairs, it achieves the highest accuracy of 71.35%. Moreover, using MUSECRITIC with GRPO improves Muse‑0.6B on all nine aesthetic metrics from SongEval and Audiobox Aesthetics. These results demonstrate that critique‑conditioned reward modeling reduces scoring error and provides an effective optimization signal for song generation. The project repository is available at https://github.com/WuqnEl/MuseCritic.

Authors:Shiji Zhou, Kunlin Lyu, Lei Zhang, Ruodong Wang, Yifan Sun
Title: MOON: Multi-Objective OrthoNormalized Updates for Multitask Learning
Abstract:
Multi‑objective optimization (MOO) has demonstrated significant success in multi‑task learning by mitigating task conflicts through gradient manipulation. However, most existing methods flatten model parameters into vectors and perform gradient manipulation under Euclidean geometry, thereby overlooking the matrix structure prevalent in modern architectures such as Transformers. In this paper, we show that gradient manipulation in Euclidean space does not generally yield the steepest descent direction under matrix geometry, potentially limiting optimization efficiency. Drawing from the theory of steepest descent for matrix‑valued parameters, we propose MOON (Multi‑Objective OrthoNormalized Updates), which performs gradient manipulation under spectral‑‑nuclear norm geometry and uses the orthonormalized manipulated gradient for parameter updates. Theoretically, for smooth non‑convex objectives, we establish convergence of the averaged Pareto‑stationarity measure at rates of \mathcalO(T^‑1/2) in the deterministic setting and \mathcalO(T^‑1/4) under stochastic gradients. Empirical results across various benchmarks show that MOON consistently improves both optimization efficiency and final multi‑task performance. Our code is available at https://github.com/KunlinLyu/MOON.

Authors:Ellen Su, Andres Potapczynski, Shikai Qiu, Edward Hughes, Andrew Gordon Wilson
Title: Epiplexity Guided Data Selection and Generation for Out-of-Distribution Generalization
Abstract:
Modern systems are increasingly expected to transfer across tasks not specified during training. What data facilitates generalization in these new, unanticipated settings? One hypothesis is that data with more structural information could contain shared circuits and subprograms that could be recycled in a wider array of downstream settings. Epiplexity, a recently proposed measure of the structural information a compute‑bounded learner can extract from data, provides a mechanism to reason about this relationship. In this paper, we show how to operationalize epiplexity as an online training signal for data selection and synthetic data generation. For selection, we fit scaling laws to the training loss curves of natural data domains to predict the expected epiplexity gain as a function of training tokens, and use this signal to adaptively determine the sampling weights over domains during training. For synthetic data generation, we define a generator's reward as the change in learner epiplexity over a buffer of previously generated data and use REINFORCE policy gradients to guide the generator toward an epiplexity‑maximizing distribution. In both cases, higher epiplexity predicts improved downstream performance on zero‑shot and fine‑tuning based tasks, supporting the hypothesis that data rich in structural information yield representations that transfer across domains.

Authors:Ran Li, Huiguo He, Jiahuan Cao, Junle Liu, Hiuyi Cheng, Lianwen Jin
Title: JieZi: A Large-Scale Expert-Audited Dataset and Benchmark for Ancient Chinese Character Exegesis
Abstract:
The scholarly exegesis of ancient Chinese characters demands integrating visual observation, linguistic analysis, and historical context. However, existing computational approaches focus narrowly on subtasks such as character recognition and retrieval, lacking the structured datasets and benchmarks required for comprehensive scholarly analysis. To address this limitation, we introduce Ancient Chinese Character Exegesis (ACCE), a vision‑language question answering (VQA) task that models the scholarly exegesis process. ACCE is organized into four progressive levels: basic character identification, glyph‑form analysis, meaning exegesis, and diachronic evolution analysis. To support this task, we construct two complementary resources. JieZi‑Dataset is the first large‑scale, expert‑audited VQA training dataset for ACCE, comprising over 500K QA pairs. It is constructed via a pipeline that reduces factual errors by constraining generation with expert‑designed templates and source‑text references. Human verification is further applied at each key stage to ensure scholarly accuracy. JieZi‑Bench is an evaluation benchmark aligned with the exegesis process, constructed and verified by human experts to ensure evaluation reliability. It consists of four levels with reference answers curated from authoritative lexicographic works held separate from the training data. Experiments on multimodal large language models show that current models perform well on basic identification but struggle with glyph analysis, semantic reasoning, and diachronic understanding. Fine‑tuning on JieZi‑Dataset substantially improves performance across all four levels. Code and dataset are available at https://github.com/Ran00w/JieZi.

Authors:Mingwei Xing, Xinliang Wang, Yifeng Shi
Title: STAR: A Spatial-Topology Aware Routing Framework for Generalizable 3D Scene Understanding
Abstract:
Constructing a unified 3D scene understanding model has long been hindered by the topological discrepancies across sensor modalities. While applying the Mixture‑of‑Experts (MoE) architecture is a flexible approach for multi‑domain 3D understanding, we observe that conventional feature‑only MoE routers may underrepresent local sampling topology under semantic supervision, making expert allocation difficult when semantic consistency coexists with geometric heterogeneity. To overcome this challenge, we propose STAR (Spatial‑Topology Aware Routing Framework). Specifically, we introduce a multi‑attribute self‑supervised pre‑training branch, covering topological and textural variations, to anchor cross‑domain structural priors. Building upon this, we design a domain‑aware expert branch with two mechanisms: Domain‑Spatial‑Guided Routing (DSR), which captures local topological variations from spatial context, and Entropy‑controlled Dynamic Allocation (EDA), which adjusts the number of activated experts according to routing uncertainty. Together, these branches combine stable cross‑domain representation learning with adaptive expert allocation. Extensive experiments across various tasks, encompassing both indoor and outdoor scenes, demonstrate the effectiveness of STAR. It achieves 80.1% mIoU on the ScanNet validation set and 77.2% mIoU on S3DIS, consistently improving over strong baselines. Code is available at our project page (https://xmw666.github.io/STAR/).

Authors:Shailja Thakur, Sungeun An, Chad DeLuca, Hima Patel
Title: The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance
Abstract:
A benchmark score comes from a single phrasing of each problem. That single phrasing is treated as if it stood for the whole space of ways the same problem could be asked, but it does not. We show that rephrasing a problem while keeping its meaning and answer fixed routinely flips a model's answer in both directions, so some failures become successes and some successes become failures. We call this drift. BenchDrift generates meaning‑preserving variations of benchmark problems along four axes, namely linguistic, referential, pragmatic, and structural, and measures how often, and why, correctness flips under each. Across eight models and three benchmarks (GSM8K, MMLU, MATH‑Hard), we observe that drift is large in both directions. Two findings stand out. First, phrasing sensitivity does not fade as models get better. Instead, it changes sign. Weak models gain more from rephrasing than they lose, while strong models lose far more than they gain. We find that the best models on a benchmark are therefore the ones whose scores depend most on the wording they happened to be given. Second, the models largely agree on which rephrasings cost the most correct answers even though they differ in how much they drift, so fragility belongs to the rephrasing and not to the model. Furthermore, rephrasing breaks answers a model was confident about, whether the problem is made shorter or longer. Code and Data: https://github.com/IBM/BenchDrift/tree/demo‑ui

Authors:Daifeng Peng, Yuanke Peng, Haiyan Guan
Title: Zero-OVCD: Bridging Training-Free Foundation Models and Pseudo-Label Learning for Open-Vocabulary Change Detection
Abstract:
Open‑vocabulary change detection (OVCD) enables the identification of user‑specified land‑cover changes in bitemporal remote sensing images, but existing training‑free pipelines remain vulnerable to inaccurate candidate masks, ambiguous semantic assignments, and accumulated inference errors. To address these issues, we propose Zero‑OVCD, a two‑stage framework that requires no pixel‑level annotations from the target domain. In the first stage, high‑quality change pseudo‑labels are generated through complementary candidate‑mask refinement, multiscale semantic similarity fusion with margin‑based reliability filtering, and response‑guided mask correction and completion. These components jointly suppress noisy candidates, enhance mask‑level semantic discrimination, and recover missed change regions. In the second stage, a change detector is trained using the generated pseudo‑labels, while checkpoint voting and high‑agreement sample selection are introduced to mitigate residual pseudo‑label noise. On LEVIR‑CD, WHU‑CD, and S2Looking, Stage I achieves F1 scores of 86.25%, 85.82%, and 50.48%, while Stage II further improves them to 88.65%, 88.85%, and 57.96%, respectively. On SECOND, the macro‑average F1 across six category‑wise one‑vs‑rest tasks increases from 47.91% to 50.92%. These results demonstrate that bridging training‑free foundation‑model inference with noise‑aware pseudo‑label learning provides an effective solution for open‑vocabulary change detection without target‑domain pixel‑level annotations. Code will be available at https://github.com/1321663019/Zero‑OVCD.

Authors:Zijian Zhao, Sen Li
Title: Low-Interaction-Rank Learning: Unifying Multiplicative Dual-Encoder Heads
Abstract:
A multiplicative dual‑encoder network computes a real‑valued output for a pair of inputs as the inner product of their separate encodings. This architecture has been developed independently in operator learning, bipartite matching, contrastive vision‑language models, retrieval, and other areas, yet no unified theory guides the basic design decisions: how many interaction modes to represent, how to normalize the encoders, and when the architecture should be avoided. We provide such a foundation by introducing the class of functions of low interaction rank, a class whose intrinsic complexity is measured by its interaction spectrum. Within this framework, approximation error decomposes into a spectral truncation term and an encoder‑realization term; sample complexity is governed by the sum of the two encoder complexities rather than their product; and a usability criterion based on spectral decay determines when the architecture can succeed. The same framework exposes a central identifiability problem: the encoders are defined only up to a linear gauge symmetry that leaves the learned coordinates arbitrary. We show that normalization is gauge fixing and that whitening pins the interaction modes up to permutation and sign, thereby explaining the uninterpretability of contrastive dimensions and providing a constructive remedy. Experiments on synthetic kernels, operator learning, and CLIP models validate the theoretical predictions: spectral decay rates match the predicted scaling, whitening recovers the true modes, and independently trained CLIP models are related by a single rotation which, after removal by whitening, exposes interpretable concept axes. The code of this paper is provided at https://github.com/RS2002/Mul‑Net .

Authors:Yoshihiko Kayama
Title: Semantic Lenia: Emergence of Homeostatic Solitons within the Semantic Space of Large Language Models
Abstract:
We introduce Semantic Lenia, an artificial life framework that transforms Large Language Model (LLM) inference from a static optimization problem into a continuous dynamical system within the macroscopic logit space. By establishing a non‑linear homeostatic feedback loop to dynamically balance semantic attraction and syntactic repulsion, we demonstrate the emergence of "Autonomous Semantic Solitons" ‑‑ macroscopic dissipative structures that avoid repetitive crystallization. Our exhaustive parameter sweeps map a critical "Habitable Ridge" where applied steering forces perfectly balance the model's intrinsic syntactic inertia. This approach successfully maintains generative trajectories at the edge of chaos, triggering profound abductive leaps without structural collapse and establishing a physical scaling law for machine cognition.

Authors:Xikai Sun, Kebin Liu, Haotian Wang, Li Liu, Xu Wang, Yunhao Liu
Title: Motion-as-Prompt: Enhancing Motion Reasoning in Multimodal Large Language Models via Motion-Guided Cross-Frame Visual Prompting
Abstract:
Motion‑centric video reasoning is fundamental to interactive applications such as robotic manipulation and autonomous navigation. However, multimodal large language models (MLLMs) typically process videos through sparse uniform sampling to control visual‑token and attention costs. This strategy may discard critical transitions between sampled frames, limiting reasoning about object movement, collisions, and causal interactions. To mitigate this issue, we propose Motion‑as‑Prompt (MaP), a track‑guided cross‑frame visual prompting framework. MaP recovers dense point trajectories, selects motion‑informative frames, and marks the trajectories accumulated between consecutive sampled frames directly onto the visual inputs, making otherwise hidden displacement, direction changes, and interactions observable to frozen MLLMs. Experiments on CLEVRER and Something‑Something‑v2 show that MaP consistently improves average motion‑reasoning accuracy, yielding gains of 4.2% and 8.9% for GPT‑5.5, respectively. Notably, these improvements are obtained without degrading non‑motion understanding, highlighting the robustness of MaP. These results demonstrate that MaP provides a simple and effective solution for enhancing motion‑centric video reasoning without model training or architectural modification. Project page:https://github.com/SunVictor23/MaP.

Authors:Huaxuan Wang, Huimin Wang, Ruiyu Zhang, Yingjie Li, Yitao Duan
Title: Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder
Abstract:
Recent advances in zero‑shot text‑to‑speech (TTS) have substantially improved speech quality and voice cloning fidelity. However, many zero‑shot TTS systems still depend on audio prompt transcripts at inference time. This dependency limits cross‑lingual voice cloning, since in‑the‑wild reference audio is often untranscribed. In this technical report, we present Confucius4‑TTS, a multilingual zero‑shot TTS system that supports 14 languages and performs both intra‑lingual and cross‑lingual reference cloning without requiring transcripts of audio prompts. Confucius4‑TTS follows a two‑stage architecture, consisting of text‑to‑semantic (T2S) and semantic‑to‑acoustic (S2A) modules. The LLM‑based T2S module uses a learnable speaker encoder to extract timbre features from self‑supervised speech representations, and the conditional flow‑matching S2A module converts the predicted semantic tokens into mel‑spectrograms. The same model also supports continuation cloning when a reference transcript is available. Confucius4‑TTS is trained on large‑scale multilingual speech data. It achieves high intelligibility and speaker similarity on public benchmarks. On the CV3‑Eval cross‑lingual benchmark, Confucius4‑TTS obtains an average WER of 3.73% across six directions. On our internal cross‑lingual set, it achieves the best average overall rank in human evaluation among recent open‑source and commercial systems. We release code, model checkpoints, and demos at https://github.com/netease‑youdao/Confucius4‑TTS.

Authors:Zhilin Ai, Boyu Li, Sidi Yang, Wenqing Shi, Wenyong Zhou, Binxiao Huang, Chenchen Ding, Ngai Wong
Title: Hybrid-LUT: Channel-Aware Hybrid Lookup Table and Filtering for Efficient Image Denoising
Abstract:
Lookup table (LUT)‑based image denoising methods have attracted increasing attention due to their high efficiency and hardware‑friendly properties. However, existing RGB‑LUT approaches require three identical LUTs to process RGB channels in parallel, resulting in large on‑chip SRAM consumption. A simple alternative is to apply LUT processing only to the luminance (Y) channel in the YUV color space to reduce memory usage. However, this naive strategy leads to degraded restoration quality, since ignoring the chrominance (UV) channels introduces color distortion and residual artifacts. In this work, we propose Hybrid‑LUT, a YUV‑based asymmetric channel‑processing framework that combines LUT and filtering in a unified design. Specifically, a multi‑band LUT branch with pixel‑level weight fusion is applied to the Y channel to recover fine textures, while lightweight filtering is used for the UV channels to maintain color consistency. This design reduces LUT storage by two‑thirds compared with RGB‑LUT methods while maintaining the same runtime throughput. Extensive experiments show that Hybrid‑LUT achieves state‑of‑the‑art (SOTA) performance across multiple benchmarks with only 421 KB of storage. In particular, our method surpasses existing LUT‑based denoising approaches by at least 0.63 dB CPSNR on real‑world datasets, demonstrating its effectiveness for image denoising on resource‑constrained edge devices. The project is available at https://github.com/Ai‑ZL/Hybrid‑LUT .

Authors:Fanding Li, Chenglin Wang, Xiangyu Li, Xingyu Qiu, Xinghua Ma, Xiangming Yin, Haiyang Li, Suyu Dong, Wei Wang, Kuanquan Wang, Gongning Luo, Shuo Li
Title: KANResDiff: Learning Local Residual Diffusion via Kolmogorov-Arnold Network for Ambiguous Medical Image Segmentation
Abstract:
Ambiguous medical image segmentation aims to provide a series of diverse but plausible segmentation hypotheses. However, existing methods introduce stochasticity in a fixed and pre‑defined manner, failing to form a progressive semantic modeling process. To address these challenges, we propose KANResDiff to learn local residual diffusion with Kolmogorov‑Arnold Network, thereby assigning distinct roles across stages for ambiguity modeling. Specifically, we propose Independent Time Encoding that offers spline‑based time embeddings instead of linear ones from MLPs, which enhances the independence across inference stages and assigns progressive semantic roles to different stages. We propose Residual Schrodinger Bridge that injects deterministic residual prior with learnable weights by constructing local Schrodinger Bridge instead of following manually settings, achieving a flexible deterministic‑stochastic interaction and stage‑aware ambiguity modeling thanks to local optimal diffusion path. Extensive experimental results on two public datasets demonstrate that KANResDiff achieves SOTA performance on GED and HM‑IoU, with maximum improvements of 16.8% and 7.7%, respectively, while maintaining competitive performance on the MDM metric. Source code is available at https://github.com/PerceptionComputingLab/KANResDiff.

Authors:Hojun Choi, Jaeyo Shin, Suin Lee, Hyunjung Shim
Title: MBA: Multimodal Benchmark and Agents for Real-World Business Ideation
Abstract:
Agentic systems powered by large language models (LLMs) have opened new opportunities for business ideation. Yet existing approaches remain confined to a text‑only paradigm, despite the inherently multimodal nature of real‑world contexts. We thus introduce MBA‑Bench, the first multimodal benchmark for training and evaluating business ideation agents, comprising 30K samples across six domains, each domain characterized by distinct visual cues not fully conveyed by text alone. Concretely, we automatically caption images and employ GPT‑4o to generate five reference ideas for each of three business questions through retrieval query generation, market evidence retrieval, and evidence‑augmented synthesis. Following prior work, we evaluate agents across six business‑oriented criteria using MLLM‑as‑a‑Judge. To consider settings where criteria are hidden or disclosed, we present MBA‑b and MBA‑k for blind and known, respectively. We train both with two novel reward objectives‑‑‑creativity and feasibility‑‑‑while MBA‑k further optimizes the six disclosed criteria for eight in total. Both are trained via LoRA‑based supervised fine‑tuning followed by group relative policy optimization with these setting‑specific rewards. For extensive experiments on MBA‑Bench, we set up two baselines accommodating either captions only or multimodal inputs, with the latter nearing closed‑source performance on several metrics. MBA‑b and MBA‑k outperform caption baselines by 63.9% and 77.1%, and multimodal baselines by 25.6% and 35.8%, respectively.

Authors:Noah Amsel, Jack Zhang, Kwangjun Ahn, Ali Naeimi, Austin Feng, Berlin Chen, Tri Dao, John Langford
Title: Dion3: Full-Stack Orthogonal Updates
Abstract:
The Muon optimizer incurs a significant overhead cost due to its cubic‑time Newton‑Schulz orthogonalization step. When weights are sharded, communication overhead compounds this computational cost, eroding the benefits of Muon in many settings. We present Dion3, a revision of Muon that targets this overhead at every level of the stack. Our Gram Newton‑Schulz algorithm reduces the FLOP cost of orthogonalization, our CuteDSL kernels accelerate it by exploiting symmetry, and our megabatching strategy reduces communication overhead. Moreover, we propose a simple change to the update rule that cuts costs even further: selecting only a fraction of the momentum matrix's rows to orthogonalize at each step. This update rule improves on Dion (another "compressed" version of Muon), in both speed and performance. Overall, Dion3 matches or improves on the loss achieved by Muon but reduces optimizer step time by up to 6x. Dion3 is available via the dion package (https://github.com/microsoft/dion) as a drop‑in replacement for Muon.

Authors:Ze Zhang, Yang Zhang
Title: Topology-Aware Query Selection for Surgical Instrument Instance Segmentation
Abstract:
Accurate foreground masks can still form an incorrect surgical‑instrument instance set: duplicate, fragmented, merged, missed, or empty‑frame predictions may preserve favorable pixel overlap while violating object identity and count. Final query selection is therefore a relational, variable‑cardinality problem rather than a collection of independent candidate decisions. We evaluate topology‑aware query selection, which represents the nonempty candidates of a fixed Mask2Former as a complete graph, learns relational candidate and pair representations, predicts set cardinality, and solves an exact structured subset problem. The formal comparison is the complete relational path versus a node‑feature‑matched path; it evaluates the combined effect of pairwise geometry, message passing, and the additional relational‑path capacity, not an isolated component. On the sealed 22‑case source test, all three discovery seeds supported instance‑set performance improvement with segmentation fidelity and predefined technical‑safety preservation: instance F1 increased by 0.0504‑‑0.0612 and positive‑frame set‑failure rate decreased by 0.0848‑‑0.1060. Direct ROBUST‑MIPS transfer reproduced the complete result in all three seeds. Endoscapes supported only one of three seeds and therefore did not establish stable direct transfer. Taken together, the results support a bounded conclusion: the evaluated complete path improved coherent instance‑set construction from fixed Mask2Former candidates in specified native‑instance contracts, while stable cross‑domain transfer and component‑specific effects remain unestablished.

Authors:Xiangqi Chen, Xiuling Zhang, Chengzhuan Yang, Li Zhao, Dawei Zhang, Yanchao Wang, Liyuan Chen, Hua Wang, Hao Peng, Zhonglong Zheng
Title: ProtoHGF-Net: Prototype HyperGraph Fusion with Intra-modal Calibration for RGBT Object Detection
Abstract:
RGB‑Thermal (RGBT) object detection enables robust perception in complex scenes by leveraging the complementary strengths of visible textures and thermal cues. However, existing methods mainly rely on dense cross‑modal interactions over full‑resolution features, which inevitably introduce background interference and hinder the learning of target‑relevant representations. In this paper, we propose the Prototype HyperGraph Fusion Network (ProtoHGF‑Net), a novel framework that redefines cross‑modal fusion as prototype‑level semantic interaction rather than the dense cross‑modal interaction paradigm. Specifically, we design Prototype HyperGraph Fusion to perform cross‑modal interaction in a compact prototype‑level semantic space. This design enables more selective fusion among target‑relevant prototypes. To support this prototype‑level fusion, we propose Teacher‑Mask Calibration Distillation, which calibrates modality features before fusion using modality‑specific teachers and target‑aware masks. This strategy suppresses backgrou‑ nd‑dominant responses and produces more target‑focused features. Extensive experiments on DroneVehicle, DVTOD, and FLIR demonstrate that ProtoHGF‑Net achieves state‑of‑the‑art performance with 85.9% mAP_50, 88.2% mAP_50, and 79.1% mAP_50, respectively. Our code is available at \hrefhttps://github.com/ZiMo‑Chen/ProtoHGFGitHub.

Authors:Haowei Lou, Hye-Young Paik, Dai Jia, Kai Li, Lina Yao
Title: CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation
Abstract:
Human voice generation has made rapid progress in speech generation, singing voice generation, voice cloning, and voice editing. However, most existing systems are designed for specific tasks and often rely on task‑dependent architectures, control signals, or autoregressive decoding, limiting fine‑grained controllability and inference efficiency. In this paper, we propose CookVoice, a unified framework for multimodal, multi‑style, and multi‑task human voice generation. CookVoice decomposes the human voice into three key factors: content, prosody, and style, enabling both speech and singing voice generation within a unified model. To achieve precise and flexible controllability, we design a flexible alignment strategy that maps text, style, and prosody control signals onto the frame‑level of spectrogram. This design allows CookVoice to support a wide range of tasks, including text‑to‑speech, text‑to‑singing voice, style‑controllable generation, voice mimicry, voice conversion, and voice editing. Experimental results show that CookVoice achieves generation quality comparable to existing Text‑to‑Speech and text‑to‑singing voice baselines, while providing stronger style and prosody controllability. Moreover, CookVoice achieves comparable performance to large‑scale baselines with only 43.51 million parameters and efficient inference using as few as 4 ODE steps, making it a practical solution for real‑world human voice generation applications. Demo page is available at https://haoweilou.github.io/CookVoice/.

Authors:Ryosei Hara, Masashi Hatano, Rintaro Yanagi, Atsushi Hashimoto, Takuma Yagi, Mariko Isogawa
Title: Hand Visibility Detector: Per-Keypoint Visibility Estimation for Hands
Abstract:
Hand Pose Estimation (HPE) is a fundamental technology for various applications such as AR/VR and robotics. In these applications, the visibility of each hand joint in the image is crucial for assessing the reliability of estimation results under occlusion. However, most existing HPE methods output joint positions without explicitly indicating their visibility. Although some methods account for occlusion or visibility, visibility estimation has mainly been used as an auxiliary signal for improving pose estimation. To our knowledge, per‑joint hand visibility estimation has not been systematically studied as a standalone task. In this work, we propose Hand Visibility Detector, a model for estimating the visibility of individual hand joints, and present the first systematic investigation of visibility estimation as an independent task. We show that leveraging the prior knowledge of HPE models pretrained on large‑scale data as a backbone yields high performance in this task. We further demonstrate the utility of Hand Visibility Detector on a downstream task of 3D hand pose annotation via multi‑view triangulation of 2D keypoints, showing that visibility‑weighted triangulation reduces reprojection error. Our method is released as a ready‑to‑use package, and the code and demo are available at https://github.com/ryhara/hand_visibility_detector .

Authors:Zepeng Wang, Jiagao Hu, Fuhao Li, Yuxuan Chen, Fei Wang, Daiguo Zhou
Title: From Synthesis to Removal: Physics-Grounded Reflection Simulation and Diffusion-Based Video Dereflection
Abstract:
Videos captured through glass often contain reflections that degrade visual quality and interfere with downstream vision tasks. Although single‑image reflection removal has been extensively studied, video reflection removal remains largely underexplored due to the lack of paired video data, temporally coherent removal models, and dedicated evaluation benchmarks. We present a closed‑loop framework that unifies physics‑grounded reflection simulation, diffusion‑based video dereflection, and benchmark evaluation. Our S2R‑Synthesis pipeline generates paired reflected and reflection‑free videos by performing physics‑grounded augmentation in the structure space and rendering realistic reflected videos with a trained video diffusion renderer; the augmentation models key glass‑related effects including roughness‑induced blur, thickness‑induced ghosting, and reflectance variation. Based on the synthesized data, we introduce S2R‑Removal, the first diffusion‑based video reflection removal model, which adapts a pretrained video diffusion prior through reflection‑aware latent adaptation and one‑step pixel‑geometric refinement, recovering the clean transmission in a single denoising step. We further build S2R‑Bench, the first benchmark for video reflection removal, supporting both full‑reference evaluation and real‑world human perceptual assessment. Experiments on S2R‑Bench and multiple public image benchmarks demonstrate state‑of‑the‑art performance and faster inference than even non‑diffusion baselines, and validate the effectiveness of S2R‑Synthesis. Project page: https://codingwzp.github.io/VideoDereflection_S2R.

Authors:Seungeun Lee, Joao Fonseca, Julia Stoyanovich
Title: RelShap: Relationally Consistent Shapley Explanations
Abstract:
Machine learning pipelines commonly flatten relational data into single‑table representations, discarding structural constraints. Widely used Shapley value‑based feature attributions then rely on feature independence, evaluating the model on combinations that could never arise in the underlying data, producing misleading explanations. We propose RelShap, a framework that incorporates relational constraints and data provenance into Shapley value computation, restricting both background data and coalition evaluation to relationally valid configurations. The framework is estimator‑agnostic and composes with Kernel SHAP, Monte Carlo, and Leverage SHAP without altering their sampling or weighting properties. Functional dependencies further induce equivalence classes over feature coalitions, which RelShap exploits to reduce runtime without changing Shapley values; we provide a combinatorial characterization of the expected speedup. Experiments across multiple datasets, models, and estimators show that RelShap produces explanations that are more faithful to the data‑generating process, correctly identifying the dominant feature in controlled settings where existing methods, including Conditional SHAP and ManifoldShap, do not. Our code is available at: https://github.com/duneag2/relshap.

Authors:Frederick Hayes
Title: Predictive Allostatic Organization in Recurrent and Spiking Agents Under Partial Observability
Abstract:
Adaptive behavior under partial observability depends on internal organization that carries information beyond the current observation. Drawing on Barrett and Miller's account of categorization as predictive, compressive, functionally organized, and allostatically constrained, we test whether recurrent and spiking agents develop internal states with corresponding computational properties. Agents operate in an energy‑constrained foraging task requiring resource acquisition, threat avoidance, contact‑dependent consumption, and regulation of an internal energy variable. In a frozen benchmark, learned agents outperform random and heuristic baselines; the trace‑augmented recurrent policy is strongest overall, while spiking variants show stress‑specific differences. Early internal dynamics predict later full‑safe‑efficient success above permutation baseline, reaching a maximum ROC‑AUC of 0.802. Reduced PCA subspaces retain behaviorally relevant information. Feature‑family controls show that predictive signal is distributed across trace, policy‑head, internal‑dynamics, observation, and allostatic variables, and low‑energy state remains strongly decodable after explicit energy‑related features are removed. Evaluation‑time perturbations to temporal state, sensory information, operating conditions, and allostatic mechanisms alter behavior and/or internal prediction. Seed‑balanced event probes show weaker but measurable information about future contact, successful consumption, and threat events, alongside strong low‑energy decoding. We interpret this pattern as a computational analogue of predictive allostatic organization: distributed control regimes that are predictive, energy‑sensitive, action‑relevant, and partly causally involved, without claiming biological validation or discrete symbolic categories.

Authors:Mehran Tamjidi, Hamidreza Dastmalchi, Ali Cheraghian, Mohammadreza Alimoradijazi, Aijun An, Hossein Rahmani
Title: Test-Time Hallucination Control in Large Vision-Language Models
Abstract:
Object Hallucination in large vision‑language models (LVLMs), where models generate non‑factual content about input images, remains a critical barrier to their reliability in real‑world applications. Existing mitigation strategies can be categorized into training‑based and training‑free methods. Training‑based methods often achieve strong performance but are costly, requiring extensive computational resources, large‑scale data, and time‑consuming fine‑tuning. Training‑free approaches are particularly appealing due to their efficiency. However, existing training‑free methods either require multiple decoding rounds, which adds computational overhead, or modify internal states in a model‑specific way that risks degrading pretrained knowledge. We propose Test‑Time Hallucination Mitigation (TTH) method, a novel training‑free method that addresses both limitations. TTH introduces a token‑validator module, implemented as a zero‑shot Multi‑Modal Classifier (MMC), to generate auxiliary logits grounded in the input image. These logits are fused with the original LVLM outputs at the token level for object tokens selected from a candidate pool. An entropy‑based weighting scheme is then applied to enable robust and accurate predictions. Extensive experiments across multiple LVLM families and diverse benchmarks demonstrate that TTH consistently improves accuracy and robustness, underscoring its generalizability and practical effectiveness. Code is released at https://github.com/Mehran‑TAM/TTH

Authors:Soumya Mazumdar, Vineet Kumar Rakesh, Tapas Samanta
Title: Analysis of Federated Aggregation under Model Poisoning and Backdoor Attacks: A Reconstructed Cross-Dataset and Cross-Architecture Benchmark
Abstract:
Robust comparisons of federated aggregation methods require joint consideration of predictive performance, threat definitions, metric semantics, and execution provenance. A 500‑cell seed‑1 evaluation matrix was reconstructed across five aggregation methods, five datasets, five architectures, and four recorded conditions: clean, sign‑flipping, Gaussian, and BadNets. Successful execution logs were identified for 454 original runs and 36 repaired or rerun executions, whereas 10 clean SVHN cells were supported by summary‑only provenance. Trimmed Mean achieved the highest clean macro‑mean accuracy (76.02%) and the lowest mean within‑task rank (1.70). Krum attained the highest recorded accuracy under both sign‑flipping and Gaussian configurations. These relative rankings remained unchanged when analysis was restricted to 21 task pairs for which original successful logs were available for every method‑condition combination. Audit of the supplied BadNets metric implementation established that every test input is triggered prior to target‑label counting; consequently, the retained metric represents Triggered Target‑Label Rate (TTLR) rather than a conventional target‑excluding attack success rate. An audit of the supplied FedPARETO scaffold further identified a pathway in which predictive summaries may characterize an uncorrupted local model while the aggregation weight is applied to a separately corrupted update, introducing a potential discrepancy between reported predictive outcomes and the updates used for aggregation. The canonical matrix contains a single identified seed for each cell, and exact attack and configuration lineage is incomplete. Accordingly, the findings should be interpreted as descriptive comparisons within the recorded configurations and not as statistical estimates or universal claims regarding robustness.

Authors:Xu Cao, Houze Yang, Vipin Gunda, Zhongyi Zhou, Tianyu Xu, Adarsh Kowdle, Inki Kim, James M. Rehg
Title: Gaze Target Estimation Anywhere with Concepts
Abstract:
Estimating human gaze targets from images in‑the‑wild is an important and formidable task. Existing approaches primarily employ brittle, multi‑stage pipelines that require explicit inputs, like head bounding boxes and human pose, in order to identify the subject of gaze analysis. As a result, detection errors can cascade and lead to failure. Moreover, these prior works lack the flexibility of specifying the gaze analysis task via natural language prompting, an approach which has been shown to have significant benefits in convenience and scalability for other image analysis tasks. To overcome these limitations, we introduce the Promptable Gaze Target Estimation (PGE) task, a new end‑to‑end, concept‑driven paradigm for gaze analysis. PGE conditions gaze prediction on flexible user text or visual prompts (e.g., "the boy in the red shirt" or "person in point [0.52, 0.48]") to identify a specific subject for gaze analysis. This approach integrates subject localization with gaze estimation, and eliminates the rigid dependency on intermediate analysis stages. We develop a scalable data engine to generate Gaze‑Co (Gaze Estimation with Concepts), a dataset and benchmark of 120K high‑quality, prompt‑annotated image pairs. We also propose GazeAnywhere, the first model designed for PGE. GazeAnywhere uses a transformer‑based detector to fuse features from frozen encoders and simultaneously solves subject localization, in/out‑of‑frame presence, and gaze target heatmap estimation. GazeAnywhere achieves state‑of‑the‑art performance on multiple PGE benchmarks, setting a strong baseline for this new problem even on a difficult out‑of‑domain, real‑world clinical dataset. GazeAnywhere is open‑sourced in github.com/IrohXu/GazeAnywhere.

Authors:Md Maklachur Rahman, Tracy Hammond
Title: Dual-Domain Cross-Modal Decoding for Clinical Text-Guided Medical Image Segmentation
Abstract:
Clinical text can narrow down what to segment, but recent text‑guided designs emphasize spatial alignment while overlooking frequency content that governs texture and boundaries. We propose Dual‑Domain Cross‑Modal Decoding (DD‑CMD) for clinical text‑guided pulmonary infection segmentation, integrating two complementary forms of language guidance during decoding. In the spatial domain, Text‑Guided Spatial Cross‑Attention (TGSA) aligns multi‑scale visual tokens with text semantics and updates features through gated residual fusion. In the frequency domain, Spectral‑Text Adaptive Modulation (STAM) applies a 2D DCT to compute learnable band‑energy statistics and predicts text‑conditioned FiLM parameters to recalibrate decoder channels for frequency‑aware decoding. DD‑CMD embeds TGSA and STAM into a coarse‑to‑fine decoder (7x7 to 56x56) and restores full‑resolution masks using a lightweight two‑stage refinement module. Experiments on QaTa‑COV19 and MosMedData+ show that DD‑CMD achieves 91.46% Dice / 84.26% mIoU and 81.95% Dice / 69.42% mIoU, respectively, with average gains of +1.96 Dice and +2.67 mIoU over the strongest prior baselines. Code: https://github.com/maklachur/DD‑CMD.

Authors:Maryam Dehdashti
Title: Qwen-MusicAVQA-7B: A Multimodal Model for Music Audio-Visual QA
Abstract:
A common approach to adding audio to a vision‑language model is to train or adapt a large omni‑modal system. We show that a lightweight alternative can be highly effective for music audio‑visual question answering (AVQA). Qwen‑MusicAVQA‑7B connects a frozen Whisper encoder to Qwen2‑VL‑7B‑Instruct through learned linear projections. The same frozen encoder processes both the video's music track and a TTS‑spoken question through separate projectors, while the language model fuses visual frames, music, and question audio through pretrained self‑attention, with no task‑specific fusion network. On MUSIC‑AVQA, our system reaches 96.0% +/‑ 3.9% accuracy across three independent training seeds on the 7,402‑question available‑video test subset. Our central finding is that downstream accuracy tracks how much fine‑grained local temporal information the audio representation preserves. In a matched 32‑token comparison, a stride‑pooled Whisper frame sequence outperforms a globally pooled PANNs representation expanded to the same budget by 26 percentage points, even though PANNs sees at least as much audio and uses a far larger projector. The effect is not simply sequence versus vector: within Whisper alone, reducing temporal resolution at a fixed token budget costs a comparable amount. Under matched data and inputs, fine‑tuned Qwen2.5‑Omni‑7B reaches 80.9%, against 95.9% for our 30 s variant; because the systems differ in backbone and adaptation, this is a system‑level comparison. Accuracy remains high on sampled head and tail splits of the rephrased MUSIC‑AVQA‑R benchmark (96.5% and 95.6%). Because both encoders stay frozen and the music features are cached, the entire adaptation is cheap to train: the complete two‑stage AVQA run takes approximately 5 hours on a single A100 80GB, and every run reported here fits on that one GPU.

Authors:Travis L. Johnson, Jiannan Jiang, Soumyabrata Chaudhuri, Yihao Chen, Lauren Falvey, Donal O'Cofaigh
Title: Long-Horizon Forecasting of Complete Financial Statements with Forma
Abstract:
Specialist training beats generalist scale when forecasting financial statements. To our knowledge, no prior work jointly forecasts complete financial statements beyond one year, yet in a discounted‑cash‑flow valuation most firm value sits past that window. We release ProForma‑20Q, a reproducible benchmark for forecasting 78 statement line items 1‑20 quarters ahead, for anonymized firms, from past statements and an industry code, scored by change‑space R^2. On it, Forma, a transformer that reads statements as sets of (account, quarter, value) tuples and maximizes a masked‑tuple Gaussian likelihood, beats every competitor we field: classical machine learning, chained gradient boosting, a zero‑shot time‑series foundation model, and frontier large language models. Its lead widens with horizon, where valuation needs accuracy most, and its Gaussian predictive intervals never under‑cover. Forma's forecasts nearly satisfy accounting identities; exact coherence is recoverable at no statistically significant accuracy cost. Its tuple interface supports scenario analysis without retraining, and we show that pinning future revenue paths sharpens the rest of the statement.

Authors:Vasundra Srinivasan
Title: Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations
Abstract:
Enterprise practitioners read agent leaderboards as if they ranked agent capability. We show, across three open agent‑trace benchmarks (TheAgentCompany, τ^2‑bench, and AppWorld), that the agent main effect accounts for less than 3% of total variance in every dataset and check type, while the agent‑by‑task interaction accounts for 7‑23%. Leaderboards rank specialization, not capability. We arrive at this through a four‑facet Generalizability Theory variance decomposition, fit with three estimators (Henderson Method‑I, REML via lme4, and a Bayesian binomial GLMM) that agree to three decimal places. Four further findings sharpen what the leaderboard is hiding. First, aggregate reliability collapses on the hardest task quartile: Eρ^2 on τ^2 action_checks falls from 0.752 to 0.000. Second, training‑cell reliability negatively correlates with held‑out reliability (r = ‑0.90 on τ^2), meaning the designs that look most reliable replicate worst. Third, population‑level diagnostics transfer across enterprise benchmarks (capability‑gap ratio stable at 0.35‑0.40) but per‑family agent rankings invert. Fourth, on the MAST failure taxonomy, trace‑level mode profiles are idiosyncratic (MAE = 0.261) while cell‑level profiles generalise (MAE = 0.056, r = 0.83). We package these into Deployment Decision Reliability (DDR), a one‑page reporting discipline that turns the variance‑component table into five decisions an enterprise buyer can defend. All code, data loaders, and fit artifacts are released under an open‑source license.

Authors:Dongsu Song, DaeYun GO, Boseung Seo, Jay Hoon Jung
Title: SegPAR: Class-Centric Decision-Based Sparse Attack for Semantic Segmentation
Abstract:
Despite the practical relevance of sparse decision‑based black‑box threats, they have received limited attention in semantic segmentation. To bridge this gap, we adapt the most representative decision‑based black‑box sparse attacks from the classification domain to serve as baselines, establishing a rigorous benchmark for this underexplored setting. In this context, we demonstrate that one of the existing methods suffers from severe query inefficiency due to its image‑centric pixel accumulation, which rapidly exhausts query budgets across the vast image space. To overcome this, we propose SegPAR, a novel decision‑based framework that shifts to a class‑centric exploration paradigm. Furthermore, to eliminate the misleading feedback generated by standard decision rewards during pixel accumulation, we introduce a novel discrepancy reward. Extensive experiments show that SegPAR significantly outperforms black‑box baselines in sparsity efficiency and MIoU reduction, while remaining competitive with white‑box sparse attacks. Code is available at \hrefhttps://github.com/KAU‑QuantumAILab/SegPARhttps://github.com/KAU‑QuantumAILab/SegPAR.

Authors:Christos Tsepas, Chang Yan, Maximilian Fuetterer, Sebastian Kozerke, Cian M Scannell
Title: Physics-Informed Implicit Neural Representations for Improved Myocardial Perfusion MRI Quantification
Abstract:
Quantifying myocardial perfusion from cardiac magnetic resonance (CMR) can be achieved by fitting tracer‑kinetic models to the dynamic contrast‑enhanced MR data. However, fitting the observed data with multi‑compartment exchange models, which describe the evolution of the contrast agent in the tissue, to estimate perfusion parameters is a challenging inverse problem that is sensitive to noise and acquisition variability. Previously, physics‑informed neural networks (PINNs) have been proposed as an alternative to conventional non‑linear least squares fitting methods with promising results for quantitative perfusion CMR. In this work, we extend the previously proposed PINN framework with spatiotemporal implicit neural representations (INRs) to represent the MR signal as a continuous spatiotemporal function and to improve the accuracy, smoothness, and physical consistency of the PINN model. In realistic simulated CMR datasets, our proposed PINN with INRs demonstrates improved robustness and parameter estimation accuracy over the previously established methods. The code is available at https://github.com/q‑cardIA/pinn‑inr.

Authors:Tanel Tammet, Priit Järv, Dirk Draheim
Title: Every pooling rule has its world: matching probability combination rules to situations and stakes
Abstract:
Systems often need to combine two numerical assessments of the same yes/no question. The appropriate formula depends on what the numbers represent and on how the sources are related. Averaging is correct when one of several alternative interpretations applies; multiplying odds is correct when probability reports are based on conditionally independent evidence and a common prior; and probabilities of alternative successful derivations require their dependence or shared evidence to be taken into account. We state the assumptions behind several common combination rules and derive the corresponding combined probabilities. Two groups of Monte Carlo experiments address different questions. First, controlled generating mechanisms verify that the derived rule recovers the correct probability in the situations for which its assumptions hold. Second, the same mechanisms measure the consequences of using a mismatched rule, using logarithmic score and threshold decisions with different costs. Distinct pooling rules can produce the same binary decision at threshold 1/2 while assigning substantially different probabilities, so binary accuracy alone can conceal important differences. We also give probabilistic interpretations of conflicting‑evidence rules and show that, for overlapping derivations, retaining the identities of shared uncertain premises permits direct calculation of the probability that at least one derivation is available. Pairwise combination of proof probabilities loses information when there are three or more derivations.

Authors:Wenchao Ma, Surya Dwarakanath, Yizhak Ben-Shabat, Dario Kneubühler, Haomiao Jiang, Sharon X. Huang, Hsueh-Ti Derek Liu
Title: A Geodesic Cut-Cell Prior for Neural Skinning
Abstract:
We introduce cut‑cell skinning, a geometric prior designed to augment data‑driven skinning weight generation. While data‑driven methods show promise in producing high‑quality skinning weights, they often lack the generalizability of classic geometric approaches. To bridge this gap, we propose a geometric prior that can be robustly computed for in‑the‑wild meshes and is efficient for large‑scale machine learning workflows. The key idea of our cut‑cell skinning is a fast graph‑based approximation of the volumetric geodesics distances, motivated by their importance in classic skinning weight computation. Our method achieves orders of magnitude speedup compared to optimization‑based solvers and remains resilient to topological artifacts common in cage‑ or voxel‑based alternatives. We demonstrate the efficacy of the cut‑cell skinning prior by integrating it into recent neural skinning models, showing consistent improvements across existing methods and achieving state‑of‑the‑art results. Project page: https://wenchao‑m.github.io/CutCell.github.io/

Authors:Cédric Léonard, Francescopaolo Sica, Martin Schulz
Title: Hardware-Aware Deployment of Joint SAR Compression and Despeckling on FPGA
Abstract:
Next‑generation Synthetic Aperture Radar (SAR) missions will generate data far faster than they can downlink, making onboard data reduction essential for near‑real‑time Earth observation. Learned Image Compression (LIC) offers better rate‑distortion performance than handcrafted codecs used operationally today, and recent work shows that simultaneously despeckling and compressing SAR imagery enables better representation capacity while unlocking higher compression rates. These methods, however, have yet to be confronted with the strict power, compute, and operational constraints of spaceborne systems. In this work, we bridge this gap by deploying a joint SAR Despeckling and Data Compression (DDC) framework on an embedded ZCU102 FPGA‑based platform, introducing model adaptations that respect the accelerator's fixed‑point arithmetic and limited set of supported operations. We evaluate four model topologies across precision levels and across CPU, GPU, and FPGA platforms, revealing several findings with direct design implications. We find that replacing conventional GDN activation functions with plain ReLU improves quality on SAR, suggesting that design principles established for compression of natural images do not necessarily transfer to SAR imagery. In addition, we demonstrate that residual blocks offer little representational benefit for ten times the compute, and show that the FPGA is the most energy‑efficient of the platforms tested. Together, these results set a functioning edge deployment workflow and an evidence‑based starting point for onboard SAR compression. The code is available at https://github.com/CedricLeon/SAR_DDC_FPGA.

Authors:Fenosoa Randrianjatovo, Maya Saleh, Simon Girard, Amadou Barry
Title: CosMAP: Contrastive Manifold Approximation and Projection for Dimensionality Reduction of Omics and Genealogical Data
Abstract:
Omics datasets, particularly single‑cell RNA sequencing data, are high‑dimensional, sparse, noisy, and dominated by zero values, making faithful low‑dimensional representation challenging. Existing dimensionality‑reduction methods may distort local neighbourhoods, global organization, or the cohesion of meaningful populations, with similar limitations arising in genealogical data. We introduce Contrastive Manifold Approximation and Projection (CosMAP), a graph‑based unsupervised dimensionality‑reduction method for producing faithful and interpretable embeddings. CosMAP extends the graph‑based framework of UMAP by combining cosine‑similarity neighbourhoods with temperature‑normalized contrastive affinities, which are optimized in the embedding space using an attractive‑‑repulsive objective. It further employs a two‑phase refinement strategy: an intermediate higher‑dimensional representation is first learned and then used to reconstruct the neighbourhood graph and initialize the final low‑dimensional embedding. We evaluate CosMAP on MNIST and USPS handwritten‑digit datasets, mouse retina and cortex single‑cell RNA‑sequencing datasets, and a large genealogical kinship dataset derived from BALSAC‑CARTaGENE. Compared with state‑of‑the‑art dimensionality‑reduction methods, CosMAP produces more coherent visual representations, improves neighbourhood preservation, and provides clearer global organization of digit classes, biological cell populations, and regional genealogical patterns. These results indicate that CosMAP offers a robust framework for exploratory analysis of complex, sparse, high‑dimensional data. The implementation is publicly available at https://github.com/FenosoaRandrianjatovo/CosMAP‑dr.

Authors:Qi Wang, Tianyi Wang, Chengyang Li, Shikun Ban, Yurun Chen, Yizhong Ge, Jason Qin, Chengtai Li, Wentao Zhu
Title: Towards the Harness of Embodied Agents
Abstract:
The success of coding agents has established the harness as a paradigm: what an agent achieves depends not on the model alone, but on the infrastructure around it. We ask whether the same paradigm extends to embodied agents in the physical world. We present Thea, a harness in which an agentic loop orchestrates robot capabilities, each wrapped as a callable tool. It inherits the core components of coding agents, modified as the physical world requires. The world, however, withholds two abilities that software grants for free: reading the state of the world, and judging the outcome of an action. To bridge these gaps, Thea introduces Scene Graph as Context, a persistent, symbolic representation of the world, and Evaluation as Exit Codes, which detects when an action should terminate, judges whether it succeeded, and on failure diagnoses the cause. Together they close the loop between the agent and the physical world. Rich behaviors then emerge from the composition of tools, and the closed loop carries long‑horizon tasks to completion in real environments.

Authors:Yoshinori Watanabe
Title: The Off-Support Barrier: Why Semantic Safety Constraints Are Not Learning-Problem Invariants, and What Follows for Prior Design, Containment, and Verification
Abstract:
We argue that a single structural fact organizes a wide range of phenomena in contemporary AI safety: a semantic safety constraint (e.g., the agent does not escape its sandbox) is an off‑support object. Formally, if q is the data distribution and \(p(\cdot\mid w)\) the model, the safety predicate B is not measurable with respect to \(σ(\textmodel, q)\), whereas the real log‑canonical threshold (RLCT) of singular learning theory (SLT) is. From this non‑invariance we derive, as corollaries rather than independent observations: (i) why reward hacking and sandbox escape arise under outcome‑based optimization; (ii) why encoding such constraints through Bayesian prior design or soft penalty weighting has poor leverage in singular models; (iii) why hard invariants belong in the harness and soft dispositions in the model; (iv) why the same B is nonetheless soundly and locally certifiable by formal verification, exactly as the local learning coefficient (LLC) locally pins the same RLCT ‑‑‑ with two precise points of disanalogy; and (v) why the residual difficulty, identifying which off‑support region matters, coincides with performative prediction and self‑referential functional dynamics, where SLT's analytic machinery breaks down. We use the July 2026 OpenAI‑‑Hugging Face evaluation incident as the motivating case. Numerical experiments code and related proofs in lean are available at https://github.com/xiangze/Preventing_Jailbreak_as_regularization

Authors:Zhiqi Wang, Yichi Zhang, Dongwon Lee, Yuchen Yang
Title: Lost in Compaction: Evaluating Side-Constraint Loss under Context Compaction
Abstract:
When the context window is under pressure, LLM systems compact prior context to continue ongoing tasks. We identify a class of user‑issued instructions, Session Constraints (SCs), such as "do not delete any emails until I confirm," that are meant to constrain LLM's behavior for the remainder of a session but are silently dropped during compaction. To quantify this loss, we introduce COMPINT, an evaluation suite that evaluates compactors across three long‑context scenarios: multi‑turn chat, agentic trajectory, and long‑horizon research. Current compactors retain only 17% of injected SCs on average, and most perform worse than running the same task without compaction. Retention varies sharply with compactor, prompt, context length, SC phrasing, and injection location, showing that the loss is systematic rather than tied to any single setting. We propose an SC‑aware extractor that runs alongside the compactor as a plug‑and‑play module, achieving over 90% retention across all three scenarios without modifying the compactor or LLM. The COMPINT evaluation suite and accompanying implementation are available at https://github.com/ZhiqiEliWang/compaction‑integrity.

Authors:Jeonghwan Choi, Taewon Yun, Minjeong Ban, Gyeonghun Sun, Jae-Gil Lee, Hwanjun Song
Title: Towards Query-Agnostic RAG Evaluation via Query Coverage and Claim Verifiability
Abstract:
Retrieval‑augmented generation improves the factuality of large language models by grounding responses in retrieved evidence, yet existing evaluation frameworks struggle to provide consistent, fine‑grained diagnostics across the diverse spectrum of user queries, ranging from close‑ended fact‑seeking to open‑ended explanatory requests. We propose Q‑CARE, a query‑agnostic and fully reference‑free framework that enables fine‑grained assessment by decomposing queries into sub‑queries and answers into atomic claims. Q‑CARE establishes a unified evaluation principle based on query coverage and claim verifiability, yielding coverage‑aware retriever metrics (C‑Prec@k, C‑nDCG@k) and claim‑level generator metrics (Completeness, Conciseness, and Verifiableness). On a human‑annotated benchmark spanning eight datasets, Q‑CARE achieves higher correlation with human judgments than four existing RAG evaluation metrics, including RAGEval and RAGChecker, proving its effectiveness as a reliable, automated evaluation framework. Code and data are publicly available at https://github.com/DISL‑Lab/Q‑CaRE‑COLM‑26.

Authors:Jiahui Zhang, Ziwei Zhang, Yipeng Wang, Yibo Liu, Haozhou Pang, Yikai Hu, Hongyan Ren, Lan Zhou, Qi Gan, Kai Sheng
Title: TRACE Bench: Task-driven Roleplay Agentic Checklist Evaluation
Abstract:
Roleplay evaluation should do more than assign a single score: it should reveal which role requirements were tested, which failed, and which dialogue evidence supports the judgment. We propose TRACE Bench, a task‑driven agentic checklist evaluation framework. It decomposes each role profile offline into a fixed checklist, then uses a User Agent to converse naturally with the target roleplay model while privately updating checklist states from model responses. Scores therefore trace back to checklist items and supporting dialogue turns rather than a black‑box holistic impression. For coverage cross‑validation, we audit released M2 free‑dialogue transcripts from the MiniMax Role‑play Benchmark against the same role‑derived checklist. The released free‑chat transcripts cover only 73.74% of key role‑profile points, whereas TRACE Bench reaches 99.91% coverage in fewer turns. Robustness experiments show stable rankings under repeated runs and User Agent replacement. Across 26 models, TRACE Bench reports overall rankings together with capability breakdowns and checklist traces. It also supports Closed‑Loop Benchmark Evolution, distilling verification methods proven effective in failed traces so later evaluations can more reliably elicit and examine observed failure modes.

Authors:Yifan Wu, Yufeng Zhang, Kenli Li
Title: CORA-Diff: Confidence-Oriented Residual Acceptance for Efficient Diffusion Language Model Inference
Abstract:
Diffusion language models (DLMs) update many tokens in parallel, yet practical decoders often use a fixed denoising horizon. Many predictions stabilize early, but blockwise decoding continues until all positions are resolved, causing repeated dense forward passes. Existing accelerators often rely on learned filters, modified scores, dependency models, or cache‑specific mechanisms. We ask whether native trajectory signals can identify residual positions likely to match the deterministic dense endpoint. We propose CORA‑Diff, a training‑free method that preserves the original transfer rule and applies confidence‑and‑persistence gating only to positions that rule leaves unresolved. Accepted tokens remain visible as context, and the block terminates once all positions are resolved. This requires no backbone change, learned acceptance model, or logit modification. Our theory explains why high‑confidence, persistent predictions are more likely to match the fixed‑horizon dense endpoint, and paired post‑intervention trajectories provide direct empirical support. We select one operating point on a separate GSM8K calibration subset and freeze it for all evaluations. Under a matched Learn2PD‑style LLaDA protocol, CORA‑Diff has the lowest measured runtime in all eight task‑length settings. Task scores match or exceed dense decoding in five settings, and the largest observed drop is 1.22 points. Its incremental speedups over EOS‑aware dense decoding are 2.70x and 3.32x on GSM8K and HumanEval. It also reaches 13.14x under the fixed‑horizon 1024/1024 mechanism‑isolation protocol and transfers to Dream without retuning at 3.18x‑3.53x. These results show that native confidence and persistence enable reliable residual acceptance, reducing repeated denoising computation while preserving task quality.

Authors:Ruoxi Zhao, Maziar Raissi
Title: Backtrader-Bench: Benchmarking LLM Agents on Algorithmic Trading with Self-Generated MCQs
Abstract:
Evaluating LLM coding agents in algorithmic trading is difficult because static benchmarks risk data contamination and numerical backtest outputs require ground truth from actual code execution. We present Backtrader‑Bench, a framework with two complementary pipelines. A deterministic multiple‑choice question (MCQ) pipeline generates questions from backtest configurations across five trading strategies, 33 templates, and three difficulty tiers, with an independent checker that re‑derives every answer. A generator‑solver filtering pipeline autonomously mines harder questions: a generator writes questions verified by executable code, converts them to MCQs, and discards any that a no‑tool solver can answer without code execution. We evaluate 11 models without tools (10 runs each) and four with‑tools configurations on a 30‑question curated set. Tool‑augmented agents reach 90.0% accuracy in a single pass (GPT‑5.5 and Opus 4.7), outperforming the best no‑tools baselines (73.0%, averaged over 10 runs) by 17 percentage points. On 38 separately mined questions, no‑tools accuracy drops further, with half the models falling to roughly random‑chance level (25%). Beyond evaluation, the scalable MCQ infrastructure is designed to produce a training corpus for reinforcement learning, with the ultimate goal of building a specialized agent for quantitative trading workflows.

Authors:Marjan Moodi, Xuankang Zhu, Fernando De Mesentier Silva, Harold Chaput, Mohammad Reza Taesiri
Title: AutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model Research
Abstract:
World modeling is an unsettled field: architectures, training objectives, and state representations interact in complex ways, and no single recipe dominates across environments. This makes it an ideal testbed for AI coding agents acting as autonomous researchers‑‑a setting in which the improvement direction is not specified in advance, unlike the engineering‑to‑spec tasks that dominate current agent benchmarks. We introduce AutoWorldModel‑Bench, a closed‑loop benchmark in which frontier coding agents autonomously improve a provided world‑model starter under a fixed compute budget. The benchmark spans eight game environments under a unified structured‑state representation‑‑ground‑truth entity state extracted from each game and consumed through a shared tensor format‑‑which isolates dynamics modeling from perception and enables minutes‑per‑run iteration. Across 64 sessions, Codex‑5.4 and Claude Opus 4.6 improve their starter on 63; in 91% of sessions the winning edit is a non‑trivial research‑style modification‑‑a new objective, representation, rollout procedure, or architectural change‑‑rather than a hyperparameter tweak. Our benchmark offers a setting in which frontier coding agents can be evaluated on open‑ended research rather than engineering‑to‑spec problems.

Authors:Mingju Gao, Jingkai Zhou, Kun Gai, Changqian Yu, Hao Tang
Title: AdvFD: Boosting Visual Generation via Adversarial Fr'echet Distance Loss
Abstract:
Fréchet distance has recently emerged as an effective distribution‑level objective for generator post‑training, complementing the conventional sample‑level diffusion and flow‑matching losses. However, directly optimizing Fréchet objectives can cause Fréchet hacking. The target metrics keep improving, but visual quality and Fréchet alignment in other feature spaces may stagnate or deteriorate. We attribute this failure to the static pretrained feature spaces used by existing Fréchet losses. These feature spaces provide incomplete and fixed views of the differences between real and generated distributions. To address this limitation, we propose Adversarial Fréchet Distance (AdvFD), which complements the static representation targets in FD‑Loss with a calibrated adversarially learned representation. AdvFD augments the original static Fréchet objective with a learnable representation that adversarially maximizes the Fréchet discrepancy between real and generated samples, while the generator minimizes the same discrepancy in the resulting adaptive feature space. To prevent the adversarial representation from trivially increasing the objective through feature amplification, we further introduce real‑feature whitening, which normalizes its scale and covariance geometry and stabilizes the min‑‑max optimization. Extensive experiments show that AdvFD consistently improves one‑step generator post‑training across both JiT and pMF backbones and across different model scales.

Authors:Song-Duo Ma, Pu-Jen Cheng
Title: Are We Really Making Progress in Group Recommendation? Unmasking the Tie-Breaking Illusion
Abstract:
Recent group recommendation methods have reported strong improvements on standard benchmarks, but it remains unclear whether these gains always reflect genuine advances in modeling group preferences. In this paper, we show that several recent methods are affected by a systematic evaluation bias caused by the interaction between training‑time score compression and evaluation‑time deterministic tie‑breaking. Specifically, an additional sigmoid transformation before the BPR objective can greatly increase tied top scores, making top‑K metrics such as HR@K and NDCG@K highly sensitive to how ties are resolved. We revisit recent representative methods and their baselines on CAMRa2011 and Mafengwo under both group and user recommendation settings, and evaluate them with a tie‑aware protocol that computes the exact expectation of HR@K and NDCG@K under uniform random tie‑breaking. Our results show that many previously reported improvements shrink substantially under tie‑aware evaluation, and the relative ranking of methods can change markedly. We further show that the additional sigmoid may act as implicit margin smoothing during optimization, and that temperature‑scaled BPR can retain much of this benefit without inducing severe tie inflation. Overall, our findings highlight the importance of tie‑aware evaluation for establishing reliable progress in group recommendation. The code is available at https://github.com/songduoma/TieAwareGroupRec.

Authors:Audrey Quessada-Vial
Title: Agentic Configuration Management (ACM): A Reference Configuration Model for Governed Agentic Systems
Abstract:
Agentic systems are increasingly composed of heterogeneous agents, prompts, tools, models, skills, composite subsystems, policies, and execution workflows whose configurations evolve across frameworks and runtime environments. Existing LLMOps and AgentOps platforms support orchestration and observability but do not provide a common configuration‑governance model for representing and governing these systems as coherent, versioned configurations. This paper introduces Agentic Configuration Management (ACM), a framework‑independent governance and configuration reference model for heterogeneous agentic systems. ACM combines typed and independently versioned Agentic Configuration Items, immutable revisions and baselines, explicit configuration‑runtime separation, lifecycle and assurance semantics, dependency‑aware impact propagation, and runtime provenance. Heterogeneous native configurations are normalized through semantic projection into a canonical Configuration Graph on which common governance semantics operate. We provide a Python reference implementation with adapters for LangGraph, CrewAI, and the OpenAI Agents SDK. The evaluation combines 27 governance scenarios with nine quantitative impact‑propagation cases. For the evaluated configurations, the three frameworks yield governance‑equivalent ACM representations and reproducible governance outcomes after projection. The impact semantics are formalized as monotone propagation over a finite lattice, establishing convergence, termination, and uniqueness of the least fixed point above the initial impact valuation. These results provide evidence that common governance semantics can support reproducibility, auditability, dependency analysis, and interoperability across heterogeneous agentic execution abstractions within the evaluated scope.

Authors:Shiqi Huang, Jiani He, Dingyan Shang, Yihua Xu, Jize Li, Yan Lyu, Lashimi Muraleedharan Nair
Title: DACRI: Decision-Aware Causal Intervention Ranking for Critical Supply Chains
Abstract:
Detecting or attributing a supply‑chain disruption is not the same as selecting the intervention that maximizes recoverable net value. We present CriticalSCM‑Bench v1, a controlled synthetic benchmark with causal ground truth, paired factual/counterfactual rollouts, and an explicit net‑value objective. Relative to a full‑information train‑selected static benchmark, LambdaMART improves median normalized net value by 5.7‑‑16.2%, with paired statistical support on the semiconductor and critical‑material archetypes but not on digital infrastructure. On digital infrastructure, a domain‑informed constant‑buffer policy remains stronger, showing that greater model complexity is not uniformly justified. Across partial and delayed settings, LambdaMART retains 33‑‑75% of full‑clamp value. Stress tests further show that intervention fidelity, timing, cost, and held‑out disruptions can alter policy ordering. Critical materials show the weakest out‑of‑distribution retention. Separately, a guarded explanation study over 540 generations preserves every fixed intervention decision after deterministic validation and template fallback, although exact wording remains unstable. Within this controlled setting, the results identify regimes in which adaptive ranking adds value and those in which simpler structural policies remain preferable.

Authors:Jiayu Ding, Meilu Song, Yun Chen, Wei Gao, Ge Li
Title: CausalSplat: Towards Comprehensive Hierarchical Reasoning in 3D Gaussian Splatting
Abstract:
While 3D Gaussian Splatting (3DGS) has advanced open vocabulary scene understanding, existing methods remain confined to explicit queries. They struggle to interpret implicit intents, complex spatial constraints, and commonsense reasoning required for practical embodied interactions. To address this gap, we introduce the task of reasoning 3D Gaussian segmentation and construct two benchmarks, Causal‑LERF and Causal‑ScanNet. These benchmarks systematically evaluate commonsense, spatial, affordance, and counterfactual reasoning. Evaluations reveal that current state of the art methods perform poorly on these reasoning challenges. Therefore, we propose CausalSplat, a framework that integrates vision‑language models with 3D scene graphs to disentangle explicit structural perception from implicit logical inference. Extensive experiments demonstrate that CausalSplat achieves state of the art performance on our reasoning benchmarks while showing strong generalizability on standard referring and open vocabulary 3D segmentation tasks. Project Page: https://jiayuding031020.github.io/CausalSplat

Authors:Huafeng Chen, Yueming Lyu, Ziyuan Chen, Wenda Tan, Chenyang Si, Liucheng Guo, Caifeng Shan
Title: PRMU: A Corpus-Free Benchmark for Person-Centric Knowledge Unlearning in Multimodal Large Language Models
Abstract:
Multimodal large language models (MLLMs) have demonstrated remarkable capabilities in storing and recalling rich person‑related knowledge, raising increasing concerns about reliable knowledge removal. However, existing machine unlearning approaches for MLLMs typically assume access to original forget and retain corpora, which are often unavailable in realistic deletion scenarios. To address this limitation, we introduce PRMU, a benchmark for evaluating corpus‑free multimodal unlearning under realistic person‑centric deletion requests. PRMU focuses on naturally acquired person‑related knowledge and evaluates whether models can remove target knowledge while preserving related knowledge through diverse textual and visual probes, including adversarial evaluation and fine‑grained locality analysis. To facilitate research in this setting, we further introduce Similarity‑Gated Projection Editing (SGPE), a lightweight corpus‑free unlearning baseline with knowledge displacement, protected parameter‑space editing, and locality‑aware multimodal control. Extensive experiments on representative MLLMs reveal that existing unlearning methods often suffer from unfavorable forgetting‑locality trade‑offs, with significant locality degradation under aggressive forgetting settings, and remain vulnerable to multimodal knowledge reactivation. Meanwhile, SGPE provides a competitive trade‑off between target forgetting, locality preservation, and general multimodal utility. We hope PRMU can facilitate future research toward realistic and scalable multimodal machine unlearning. Code and dataset will be released at https://github.com/2231122/PRMU.

Authors:Davide Rinaldi, Luciano Serafini
Title: sLTN: Structural Logic Tensor Networks
Abstract:
Logic Tensor Networks (LTN) provide a neurosymbolic framework in which first‑order logic is interpreted through tensor operations, enabling logical constraints to be integrated with differentiable learning. However, the original formulation of LTN is primarily suited to data represented as flat collections of individuals, and does not explicitly capture structural organization such as temporal order, sequential position, or graph connectivity. We introduce sLTN, an extension of LTN that makes structural dimensions first‑class elements of the language. Structural dimensions represent named tensor axes associated with domain‑specific organization, such as time steps, sequence positions, or graph nodes. They can be quantified explicitly, related through structural relations, and used to express temporal, sequential, and relational constraints directly at the logical level. We formalize the syntax and fuzzy tensor semantics of sLTN and show that, in the absence of structural dimensions, the framework recovers the original LTN semantics as a special case. We further describe a PyTorch implementation based on a declarative signature, formula parsing, and tensorial interpretation. The framework is illustrated on representative temporal and sequential reasoning examples. This paper serves as a companion to the sltn library, available at https://github.com/logictensornetworks/sltn.

Authors:Huafeng Chen, Yueming Lyu, Chenyang Si, Wende Tan, Liucheng Guo, Caifeng Shan
Title: Is There Really a Camouflaged Object? Towards Realistic Camouflaged Object Detection
Abstract:
Camouflaged object detection (COD) aims to segment objects that are visually concealed in their surroundings and has attracted increasing attention in recent years. However, most existing COD methods are developed under a closed‑world assumption, where each input image is assumed to contain a camouflaged object. This assumption ignores realistic scenarios with pure backgrounds or non‑camouflaged objects, causing existing models to produce severe false positives when deployed in open‑world environments. To address this limitation, we propose OPC16K, a large‑scale benchmark for realistic COD. OPC16K contains 16,245 images from 14 sources and is carefully organized into camouflaged‑object images, pure background images, and non‑camouflaged‑object images, enabling comprehensive evaluation of both segmentation quality and negative‑sample rejection. Based on this benchmark, we further propose OPCNet, a presence‑aware camouflage network that reformulates COD from a pure segmentation task into a joint problem of object localization and camouflage existence reasoning. Specifically, OPCNet introduces hierarchical existence reasoning to distinguish CO, BG, and NOCOD scenarios, similarity‑aware camouflage relation modeling to capture foreground‑background camouflage cues, and existence‑aware feature refinement to regulate segmentation features with existence predictions. Extensive experiments on OPC16K demonstrate that OPCNet achieves superior performance under the proposed realistic COD evaluation protocol, significantly reducing false positives on negative samples while maintaining accurate camouflaged‑object segmentation. Code and dataset will be released at https://github.com/2231122/OPCOD.

Authors:Vladimir Iglovikov
Title: AlbumentationsX: One Augmentation Pipeline for Images and Related Annotations
Abstract:
Augmentation can corrupt a training example when an image and its annotations receive different random changes. A crop must use the same coordinates for the image, mask, boxes, keypoints, stereo views, video frames, or volume. Code paths that choose these values separately can silently misalign the data. AlbumentationsX keeps the transform list, probabilities, annotation settings, and random seed in one Compose object. Each call chooses random values once and applies them to every supported part of the training example. The library keeps each object's mask, box, and label together and lets projects add their own transforms. It can also save the pipeline definition, show what happened in one call, and run that call again. The examples place Compose after files have been decoded into arrays and before PyTorch groups examples into a batch. AlbumentationsX executes the declared transforms. Practitioners still decide whether a flip, crop, color change, or other operation preserves the correct label for their task.

Authors:Biratal Raj Wagle, Bashirul Azam Biswas, Grant Chau, Matthew E. Maeder, Muhammad Azeem Arshad, Michael S. Leapman, James B. Yu, Indrani Bhattacharya
Title: Foundation Model-Enabled Efficient Data Sampling (FEEDS): A label-efficient training strategy for pan-cancer, multi-tracer PET/CT datasets
Abstract:
Automated lesion segmentation in whole‑body PET/CT imaging can assist clinicians with cancer detection, staging, and treatment planning across radiotracers and cancer types. However, training lesion segmentation models that capture variations in lesion size, distribution, and appearance requires large annotated datasets, whose creation is both time‑ and expertise‑intensive. As a result, models trained on limited labeled PET/CT data often lack the accuracy and generalizability needed for clinical use. We present FEEDS (Foundation model‑Enabled Efficient Data Sampling), a label‑ and compute‑efficient learning strategy that uses vision foundation model embeddings to select the most informative and diverse unlabeled cases for expert annotation. Unlike unsupervised, semi‑supervised, and active learning approaches, FEEDS is a one‑step training paradigm requiring only a limited, representative training set, making it label‑ and compute‑efficient. We train and validate FEEDS using the AutoPET‑III dataset. We test its accuracy and generalizability on three held‑out sets: AutoPET‑III, DeepPSMA, and an internal Dartmouth‑Hitchcock Medical Center dataset. We evaluate clinical utility at the voxel, lesion, and anatomic region level to assess performance in high‑risk areas and treatment planning utility. FEEDS outperforms random‑sampling‑based labeling, pseudolabel‑based semi‑supervised learning, and training with limited labeled data alone. It generalizes across all three test sets, FDG and PSMA tracers, and multiple diseases, matching fully‑labeled (100%) training performance with 70% less annotation burden. FEEDS addresses the challenge of label scarcity in an automatic lesion segmentation framework by providing a practical approach for constructing representative and diverse annotation queues from large, unannotated clinical repositories.

Authors:Hesam Araghi, Jan van Gemert, Nergis Tomen
Title: Static in Frames, Dynamic in Events: Rethinking Features in Event Cameras as Motion Cues
Abstract:
Event cameras capture intensity changes asynchronously with high temporal resolution, requiring novel preprocessing methods for downstream tasks. Unlike static intensity snapshots, event data inherently encode information about scene dynamics and object motion, meaning that features derived from events can exhibit behaviors with no direct analogue in frame‑based vision. In this paper, we analyze two features used in event‑based corner detection‑‑‑the eigenvalues of the structure tensor and the spatiotemporal density values‑‑‑and show that they are \emphmotion cues. We hypothesize that these features, combined with local geometric information, can enhance motion estimation tasks. To validate this, we first theoretically analyze how the eigenvalues of the structure tensor at moving corner points relate to the direction of motion. We then design controlled experiments on a synthetic dataset, confirming that extending local geometric features with eigenvalues and density values provides complementary motion information and is robust to texture and shot noise. Finally, we integrate the proposed features into a state‑of‑the‑art event‑based optical flow network and evaluate on the real‑world DSEC benchmark, where the added features consistently improve accuracy, with the largest gains in data‑scarce scenarios and for lower‑capacity models. The code for this paper can be found at: \hrefhttps://github.com/hesamaraghi/static‑in‑frames‑dynamic‑in‑eventshttps://github.com/hesamaraghi/static‑in‑frames‑dynamic‑in‑events.

Authors:Raphael Lorenzo-Louis, Fabio Amadio, Bertrand Luvison, Serena Ivaldi
Title: HUI360: A 360° Egocentric Dataset and Baselines for Human-Robot Interaction Anticipation
Abstract:
As robots increasingly operate in human‑populated environments, anticipating human intentions is essential for enabling proactive and socially aware behavior. Automatic anticipation of human‑robot interactions is thus emerging as a crucial perception challenge for embodied agents. To this end, we introduce HUI360, the largest dataset for human‑robot interaction anticipation in the wild and its set of baselines. The dataset was collected from a mobile robot, in the wild, over multiple days within a 3‑month period, and in several environments, capturing natural, spontaneous behaviors from both passersby and users, and encompassing a diverse range of individuals. This variety enables evaluating and improving the generalization capabilities of interaction anticipation models. We designed a pipeline and share code for automatic interaction annotation in arbitrary 360‑degree equirectangular videos, along with interfaces for manual refinement. Using this pipeline, we release the HUI360 open set of 1M pre‑processed annotations, including detailed 2D poses, facial keypoints, and segmentation masks, obtained using state‑of‑the‑art computer vision methods and manually curated to ensure high‑quality tracking and interaction annotation. Additionally, we release the raw panoptic 360‑degree images captured from the robot's egocentric viewpoint (on demand, for research purpose only in compliance with GDPR). Finally, we establish benchmark baselines for interaction anticipation, including the first cross‑dataset evaluations for this task: to this end, we also release 6M annotations for another existing in‑the‑wild outdoor dataset collected from a mobile robot (SSUP‑HRI). Dataset and code can be found at https://hucebot.github.io/hui360.

Authors:Zhuang Wang
Title: SCOUT: Symmetric Consensus Outlier Detection for Failure Localization in LLM Pre-Training
Abstract:
In LLM pre‑training, synchronization propagates rank‑local stalls, slowdowns, and numerical errors into job‑wide symptoms, obscuring their origin. Existing diagnosis often relies on in‑process monitors that cannot report after the trainer blocks or terminates, or on post‑mortem logs that preserve only synchronized symptoms; offline health tests lose the workload and operating conditions that triggered the failure. We present SCOUT, a unified runtime failure‑localization framework built on one design principle: identify outliers through strict‑majority consensus among equivalent replicas. SCOUT aligns replica progress, timing, and numerical evidence, then uses its Consensus Collective Communication (C3) abstraction to identify ranks whose compact signatures disagree with their peers. An out‑of‑band CPU observer remains responsive when training hangs, whereas in‑situ replay exercises recurring stragglers and silent data corruption (SDC) beside the live job with its model state, kernels, allocations, communication path, and thermal and memory pressure present. Collective fingerprints expose rank‑local protocol divergence. Clean replay coverage certifies checkpoint numerical integrity, preventing recovery from selecting state corrupted by SDC. SCOUT integrates with PyTorch, TorchTitan, Megatron‑Core, and DeepSpeed without training‑loop or framework‑source modifications. SCOUT is open source at https://github.com/LMResiliency/lm‑resiliency.

Authors:Ke Ma, Yamin Mao, Weiming Li, Shuai Tan, Yijie Zhong, Hao Chen, Haofen Wang, Meng Wang
Title: R4DSG: Relative 4D Scene Graph Memory for Object-Centric Question Answering in Long Egocentric Video
Abstract:
Long‑horizon egocentric video is a rich substrate for wearable AI assistants, but object‑centric questions such as where an item was moved, when it last changed state, or why it was relocated remain difficult because caption‑ and transcript‑based memories rarely preserve persistent object identity or structured spatial change. Existing long‑video QA methods mainly emphasize temporal grounding and clip retrieval, while prior 3D scene‑graph methods typically assume stronger geometry than free‑motion wearable RGB video provides, including point clouds, RGB‑D input, posed views, sparse reconstruction, or reconstructed scenes. R4DSG introduces a relative 4D scene graph memory for long egocentric video. Instead of storing raw graph sequences, R4DSG converts video into compact queryable memory entries indexed by time, place, persistent objects, anchor‑relative change, and local interaction context. The main idea is to separate stable anchors from dynamic objects, maintain persistent object identity across frames, and represent object state through anchor‑relative transitions rather than a globally aligned world model. Built on recent RGB‑only advances in promptable video segmentation, temporal propagation, and relative 3D lifting, the method produces a retrieval‑ready memory directly usable for long‑horizon question answering. Evaluation on a 255‑question object‑related subset from EgoLifeQA shows, under question‑only retrieval, a 6.7‑point overall gain over EgoRAG‑Text and a 12.5‑point gain on when questions, which highlights the value of temporally organized object memory. These results position relative 4D scene graphs as a practical memory substrate for wearable assistants, AR systems, and embodied multimedia agents. GitHub Page: https://dualtransparency.github.io/R4DSG/.

Authors:Sicheng Zhang, Zhonghao Yan, Binzhu Xie, Shi Qiu, Muzammal Naseer, Naveed Akhtar, Mubarak Shah
Title: On the Limitations of Cross-Lingual Consistency in Multilingual Text-to-image Generation
Abstract:
Text‑to‑image (T2I) generation has achieved remarkable progress in recent years. However, existing research has largely focused on English‑only settings, leaving cross‑lingual performance gaps and language‑specific effects insufficiently explored. To fill this gap, we introduce LingT2I, a benchmark covering 10 widely used languages with 33K prompts, designed to evaluate cross‑lingual effects in both content generation and text rendering. Building on this benchmark, we conduct a comprehensive cross‑lingual analysis, uncovering linguistic inequality and language‑dependent trade‑offs across evaluation dimensions. Beyond quantitative evaluation, we further reveal a range of language‑dependent generation patterns, highlighting how linguistic factors and their corresponding cultural contexts systematically impact model outputs. Our benchmark and analysis provide a foundation for studying cross‑lingual behavior in T2I generation and facilitate the development of more robust and inclusive models. Code and dataset are available at https://github.com/RISys‑Lab/LingT2I.

Authors:Jin-Ting He, Fu-Jen Tsai, Yan-Tsung Peng, Min-Hung Chen, Chia-Wen Lin, Yen-Yu Lin
Title: HNDiff: Haze-Noise Diffusion for Image Dehazing
Abstract:
Existing diffusion‑based methods have recently made significant progress in image dehazing. However, they typically neglect the physics of haze formation and reconstruct clean images from pure Gaussian noise, thereby limiting their restoration potential. To address this issue, we propose Haze‑Noise Diffusion (HNDiff), a novel diffusion framework that embeds the atmospheric scattering model as an inductive bias. By grounding diffusion in physical principles, HNDiff ensures that the restoration aligns more closely with underlying mechanisms of haze formation. In its forward process, we introduce joint haze‑noise diffusion with a haze‑aware noise scheduler, which progressively adds both haze and noise to an image. Essentially, the scheduler adapts noise levels according to haze density, meaning that regions with heavier haze receive stronger noise injection to encourage content generation, while clearer regions receive lighter noise to better preserve details, which directly links the forward degradation process with the physics of haze. In the reverse process, we then derive a physically consistent dehazing‑denoising process that simultaneously removes haze and noise to restore a clean image in a manner aligned with the forward degradation process. To further enhance practicality, we propose Latent HNDiff, which compiles clean latent priors that can be seamlessly integrated into existing dehazing networks to boost performance. Extensive experiments show that our work significantly improves leading dehazing backbones and achieves state‑of‑the‑art results on benchmark datasets. The project page is available at https://jin‑ting‑he.github.io/HNDiff .

Authors:Nicolás Vera Zúñiga
Title: What Iterated Self-Feeding Probes of Language Models Measure, and a test that separates the construction from the model
Abstract:
A growing class of methods probes a language model by feeding it its own output: self‑consistency, iterated refinement, agentic loops. We ask what such a probe measures, in a construction chosen to make the question sharp: a ring of token cells resampled in place by the model's own windowed conditional p_r(x_i | x_i+‑r). The substrate is Glauber dynamics on token sequences and is not new; what we change is the coupling. Advancing two rings that differ in one token under common random numbers makes undamaged copies diverge by exactly zero, so damage spreading becomes measurable where a maximal coupling gives mixing times instead. The answer is that it measures two different things at once, in readings that look alike. Some quantities are fixed by the construction: the damage light cone is kinematic, and the radius scaling of the token‑space Lyapunov exponent lambda_ca(r) is model‑invariant across 19 models and two scale ladders spanning 70x. Others genuinely track the model: lambda_ca crosses zero at a reproducible point in training, and the attractor share ranks models consistently however the lattice is built. Left undistinguished, the first kind is readily mistaken for the second ‑‑ we did so ourselves for four months, and report a phase transition we measured to three decimal places that belongs to the probe rather than to any language model. We give the test that separates them: hold the construction fixed and vary the model, or hold the model fixed and vary the construction, and see which readings move. We validate the instrument by reproduction first, recovering a Domany‑Kinzel damage field bit‑exactly against an independent prediction, and we report the estimator failures that this discipline caught ‑‑ four retracted verdicts, each on a quantity that looked like a measurement. The methodology ships as a package.

Authors:Man Jiang, Ouxiang Li, Weibao Xue, Zhenhua Tang, Yuan Wang, Shuo Wang, Yanbin Hao
Title: PEAK: Precise and Persistent Concept Erasure via k-Sparse Autoencoders
Abstract:
Erasing concepts from large‑scale text‑to‑image (T2I) diffusion models has become increasingly crucial due to the growing concerns over copyright infringement, privacy violations, and offensive content. Existing approaches struggle to achieve both precise and persistent concept erasure: inaccurate localization of concept‑related representations may cause unintended semantic interference, while incomplete removal of the underlying concept knowledge allows adversarial recovery. To address this dilemma, we propose PEAK, a precise and persistent concept erasure framework via k‑Sparse Autoencoders (kSAEs). PEAK first trains a kSAE on internal activations of the diffusion denoising network to decompose dense representations into interpretable sparse features. By contrasting sparse activations induced by target and non‑target prompts, PEAK identifies a compact set of target‑specific features according to both activation strength and frequency. These localized features are then used for parameter optimization, where PEAK selectively suppresses target‑related activations while preserving complementary non‑target ones towards the original model. This feature‑guided optimization embeds concept erasure directly into diffusion parameters, eliminating the need for additional inference‑time intervention and facilitating effective persistence against adversarial attacks. Extensive experiments demonstrate that PEAK achieves effective and robust concept erasure. On the I2P benchmark, PEAK reduces NudeNet detections from 582 to 6, lowers the average attack success rate (ASR) from 96.52% to 5.63%, and preserves general generation quality on MS‑COCO with a near‑zero KID. Our code and models are available at: https://github.com/manmanTAT/PEAK

Authors:Yuetian Du, Yucheng Wang, Zhenyuan Chen, Luyuan Chen, Rongyu Zhang, Jinjian Zhang, Wei Zhou, Zhijie Xu, Ming Kong, Zhan Zhou, Jie Liu, Qiang Zhu
Title: CARE: Confidence-Aware Reasoning for Reliable Medical VQA
Abstract:
Reinforcement Fine‑Tuning (RFT) has enabled medical Multimodal Large Language Models (MLLMs) to produce Chain‑of‑Thought (CoT) reasoning for visual question answering, yet these models suffer from confidence miscalibration‑‑‑a systematic gap between expressed certainty and actual diagnostic accuracy that undermines clinical trust. We propose CARE, a Confidence‑Aware medical REasoning framework that jointly optimizes accuracy and calibration through a dual‑stage pipeline. First, a scalable Medical‑CoT synthesis provides structured cold‑start data for Supervised Fine‑Tuning. Second, Group Relative Policy Optimization (GRPO) with a novel Confidence‑Aware Reward (CAR) mechanism ties the model's confidence to diagnostic correctness within the reward signal. Across three Medical VQA benchmarks, CARE achieves the highest diagnostic accuracy while obtaining the lowest Expected Calibration Error and Hallucination Rate, establishing a foundation for trustworthy clinical decision support. Our code is available at https://github.com/anotherbricki/CARE.

Authors:Thanh-Dan Bui, Thanh-Trung Do, Tuan-Phong Nguyen
Title: REAP: Relation-Aware Elicitation and Parsing for Closed-Book Knowledge Base Construction from LLMs
Abstract:
We present the REAP system for the AKBC Shared Task 2026 on constructing knowledge bases from language models in a closed‑book setting, subject to a budget of at most 32B parameters and no model fine‑tuning. Our system combines structured chain‑of‑thought reasoning, relation‑specific query strategies, and a reasoning‑based empty‑set gate to elicit parametric knowledge, followed by direct extraction into valid JSON arrays. On the test set, the system, built on the Mistral‑Small‑24B‑Instruct‑2501 model, achieves a macro‑F1 score of 0.62, with particularly strong results on countryLandBordersCountry (F1 = 0.95), companyTradesAtStockExchange (F1 = 0.73), and hasArea (F1 = 0.77). Our code is publicly available at https://github.com/yammdd/AKBC‑Shared‑Task‑2026.

Authors:Dazhao Du, Shiyan Du, Jian Liu, Yongjian Yu, Bohai Gu, Tao Han, Hualuo Liu, Eric Liu, Yujia Zhang, Xi Chen, Song Guo
Title: Temporally Grounded Compositional Camera Motion Understanding via Geometric Knowledge Distillation
Abstract:
Understanding camera motion is fundamental to video perception, with applications in spatial intelligence and controllable video generation. Multimodal large language models (MLLMs) provide a natural interface for this task, but existing work typically assigns one or more labels to an entire clip. Such clip‑level recognition overlooks two defining properties of real camera motion: it can change within a shot, and multiple movements can occur simultaneously. We therefore formulate camera‑motion understanding as temporally grounded, compositional recognition, which requires a model to localize motion‑consistent intervals and identify every movement active within each interval. We introduce CamChoreo, a benchmark of 4,229 real single‑shot clips with expert‑annotated temporal segments. Its annotations use a compact vocabulary of 20 direction‑aware labels, and nearly half of the segments contain compound camera motion, with multiple movement primitives active simultaneously. Recognizing such fine‑grained, compositional motion is hard for current MLLMs, whose visual encoders emphasize semantic content rather than the geometric evidence on which camera motion depends. Directly injecting features from a frozen 3D foundation model addresses this gap, but requires running the expensive geometry model on every input; we refer to this baseline as CamInject. We instead propose CamDistill, which distills the same geometric knowledge into lightweight camera tokens during training and removes the 3D model at inference. CamDistill matches the accuracy of direct feature injection without running the 3D teacher at inference. Together, CamChoreo and CamDistill advance camera‑motion understanding from clip‑level labeling to temporally grounded, compositional recognition. Project page: https://ddz16.github.io/cammotion.github.io/.

Authors:Md Rabiul Islam, Samir Abdaljalil, Erchin Serpedin, Hasan Kurban
Title: ConfTriage: A Calibration-Aware LLM Triage Framework for Pulmonary Nodule Malignancy with Selective Specialist Deferral
Abstract:
Pulmonary nodule malignancy prediction typically depends on image‑trained specialist deep learning (DL) models that require substantial annotated imaging data and task‑specific training. We investigate whether a generalist large language model (LLM), reading only a faithful natural‑language rendering of standard nodule attributes, can serve as a calibrated triage layer. We propose ConfTriage, a confidence‑calibrated method built on three pillars: language as the modality, calibration as the safety mechanism, and a selective specialist DL backstop for low‑confidence cases. We prove two guarantees: a finite‑sample combined‑error bound yielding an explicit per‑threshold operational certificate, and an oracle inequality showing that excess risk over the Bayes‑optimal deferral classifier is controlled by the L1 calibration error of the LLM probability. A controlled seven‑way input ablation across five frontier LLMs on LIDC‑IDRI shows that natural‑language descriptions dominate the diagnostic signal, while low‑level image statistics are essentially diagnostically vacuous. ConfTriage achieved an F1 score of 88.22% and an AUC of 0.92, resolving 76.5% of cases using zero‑shot LLM inference alone and referring only uncertain cases to the specialist DL backstop. These results demonstrate that clinically meaningful diagnostic information can be captured through structured radiological descriptions and leveraged by calibrated LLMs for selective referral. The framework suggests a practical pathway for combining generalist LLM prediction with specialist AI models in medical decision‑support systems. Source code is publicly available at https://github.com/rabiul‑ai/ConfTriage.

Authors:Xiaoyu Li, Andi Han, Jiaojiao Jiang, Junbin Gao
Title: Optimistic Rates for Multiclass PAC Learning
Abstract:
Worst‑case multiclass bounds do not become smaller when the best classifier is already nearly correct: what is missing is an optimistic rate, a guarantee whose fluctuation scales with the oracle risk itself. For a class of Natarajan dimension d_N and Daniely‑Shalev‑Shwartz dimension d_DS, the optimal excess risk is known at the two endpoints (d_DS/n realizable, \sqrtd_N/n+d_DS/n agnostic [HMZ24, CEH+26, Pab26]) and open in between. We close the gap: at every fixed oracle risk L^\star, the optimal excess risk is \widetildeΘ(\sqrtL^\star d_N/n+d_DS/n), uniformly in the alphabet size, attained by a learner that knows neither L^\star nor the confidence level. The upper bound composes the cover‑menu‑compression architecture of [CEH+26], at the realizable rate of [Pab26], with a new comparator‑facing relative compression theorem: a size‑k compression rule that empirically dominates a comparator h has population risk at most L(h)+O(\sqrtL(h)Γ+Γ) with Γ=(k\log n+\log(1/δ))/n, without stability; this transfers the comparison principle of the sharp binary theory [MQZ26] while discarding its Boolean‑cube geometry, which does not lift to multiclass labels. The lower bound forces both terms using one class and one distribution at every fixed L^\star, by a pair‑Assouad scheme calibrated to L^\star and a fiber argument on the pseudo‑cubes underlying the Natarajan‑versus‑DS separation of [BCD+22]. Both theorems extend to list learning: against the best r‑tuple of hypotheses, the same architecture and the same two engines yield an optimistic rate and a lower bound of the same shape, forcing the fluctuation term that [Pab26] expected to be necessary against list comparators, and removing the factor r from the known realizable list lower bound.

Authors:Nikita Borodin, Maria Krylova, Artem Zabolotnyi, Dmitry Aspisov, Egor Shikov, Nikita Tyuplyaev, Oleg Travkin, Roman Alferov, Dmitry Vinichenko
Title: Diffract: Spectral View of LLM Domain Adaptation
Abstract:
We study continual pre‑training (CPT) as a mechanism for adapting general‑purpose large language models to specialized domains: mathematics, instruction, code, and natural text. Using singular value decomposition of weight matrices, we find that CPT leaves singular value spectra largely invariant, with adaptation driven mainly by changes in singular vectors. An analysis of attention‑head projection matrices reveals strong, domain‑dependent head heterogeneity, which we exploit to define a head importance criterion: up to 60% of head updates can be removed without measurable quality loss. Selectively rewinding low‑importance heads to their pre‑trained state improves benchmark accuracy by up to 4% versus the fully trained baseline. Finally, we identify domain connectivity ‑ linear interpolation between CPT checkpoints yields smooth domain‑quality interpolation without notable degradation on either domain ‑ and release Diffract, an open‑source toolkit for scalable spectral analysis of billion‑parameter models.

Authors:Rajmund Nagy, Silvia Arellano García, Hendric Voss, Mihail Tsakov, Taras Kucherenko, Youngwoo Yoon, Gustav Eje Henter
Title: The GENEA Challenge 2026: A Large-Scale Disentangled Evaluation of Speech-Driven Gesture Generation on the Seamless Interaction Dataset
Abstract:
This preprint presents the results of the fourth GENEA Challenge, a large‑scale human evaluation of five speech‑driven gesture‑generation systems trained by participating teams on the Seamless Interaction dataset of dyadic conversations. As in the 2023 GENEA Challenge, we used a disentangled evaluation methodology to assess motion quality and speech alignment without confounding between the two, and performed a dyadic mismatching study to isolate the effect of listening and reacting to the interlocutor. We additionally introduce a new semantic gesture‑generation task and a text‑mismatching evaluation methodology using the Grounded Gestures subset of the data. In total, we ran four large‑scale user studies, collecting over 23,000 votes from 869 test‑takers. In the motion‑realism study, the dataset's filtered segments had substantially higher motion quality than all challenge submissions (68‑95% pairwise winrate). In the speech‑alignment study, the motion‑capture segments provided a conceptual ceiling at 62% alignment score, with the top submission significantly behind at 32% and the rest only slightly above the 0% expected of an input‑independent system. In the dyadic study, motion capture again set the ceiling at 65% appropriateness score, but no submission scored substantially above chance, indicating that the systems could not yet respond to the interlocutor. Finally, the semantic mismatching evaluation found highly expressive gestures in the dataset (test‑takers identified the matching transcript 79% of the time), yet almost all submissions failed to generate semantically expressive motion, with the best achieving only an 8% appropriateness score. The collected votes and outputs will be made publicly available at https://genea‑workshop.github.io/2026/challenge/ to facilitate reproducibility and further research.

Authors:Mykhailo Koshil, Matthias Feurer, Katharina Eggensperger
Title: TACTICL: Task-Aware Compression of Tabular ICL Models
Abstract:
The strong performance of foundation models for tabular tasks comes at substantial inference costs. Distilling models into task‑specific architectures reduces model size and computational demands but also sacrifices in‑context adaptability. Here we introduce TACTICL, an automated task‑aware compression framework for tabular in‑context learning models that jointly prunes transformer layers and replaces them with lightweight adapters trained on downstream tasks, thus blending in‑context with in‑weight learning. We study TACTICL on 47 benchmark datasets and show that we can substitute up to 85% of layers without substantial performance drop on a given downstream task. We further show that TACTICL maintains robustness to data shifts, leaving its in‑context ability intact. Overall, TACTICL provides a robust framework for exploiting the depth‑wise redundancy of tabular foundation models by combining task‑specific adaptation and structured compression. We provide the code at: https://github.com/Hebog/tfm_compression

Authors:Zhijie Wu, Kento Kawaharazuka, Kei Okada
Title: Neural Introspection Gating for Adaptive KV-Cache Reuse in Vision-Language-Action Models
Abstract:
Vision‑Language‑Action(VLA) models map camera images and language instructions directly to motor commands through a single autoregressive transformer. In real‑time control, they still spend substantial compute recomputing key‑value(KV) representations for visual tokens that barely change across neighboring frames. Recent work such as VLA‑Cache reduces that cost by reusing KV states for visually static patches, but its policy relies only on observation‑space heuristics and does not account for the model's own uncertainty. We propose Gated VLA‑Cache, a lightweight, training‑free extension that augments visual‑similarity caching with neural introspection. The method monitors the logit margin between the top two predicted action tokens, a zero‑cost confidence signal available during decoding. When the margin drops below a threshold, the cache is invalidated and a full recompute is triggered. Evaluated on four LIBERO benchmark suites with both OpenVLA and OpenVLA‑OFT, Gated VLA‑Cache improves reliability when blind caching hurts. On LIBERO‑Goal and LIBERO‑Long, it recovers over 100% of the lost accuracy while retaining 80% of the compute savings.

Authors:Simone Sarrocco, Paul Friedrich, Florentin Bieder, Christina Bornberg, Philippe Valmaggia, Peter Maloca, Philippe Cattin
Title: Modelling Geographic Atrophy Progression using Implicit Neural Representations
Abstract:
Age‑related Macular Degeneration (AMD) is the major cause of blindness in the Western world. Its late dry phase is characterised by irreversible atrophic areas, namely Geographic Atrophy (GA). Longitudinal Fundus Autofluorescence (FAF) image acquisitions are currently the main tool for assessing lesion growth over time at the image level. However, due to its highly individualised progression, the evolution of late AMD remains poorly understood. In this work, we propose using Implicit Neural Representations (INRs) to model GA progression at the individual level in a low‑data setting. Our approach generates both FAF and GA segmentation at both past and future time points. Among the comparison models, our method achieves competitive segmentation quality across different scenarios, yielding the lowest Mean Absolute Error (MAE) for the GA lesion area and the highest DICE score, without sacrificing FAF image quality. The code is available at https://github.com/SimoneSarrocco/ga‑progression‑with‑inrs.

Authors:Bangyan Liao, Chenglei Yu, Yuchen Yang, Chuanrui Wang, Zhisheng Song, Peidong Liu, Tailin Wu
Title: Path Integral Value Matching for Linear Quadratic Stochastic Optimal Control
Abstract:
Linear Quadratic Stochastic Optimal Control (LQ‑SOC) establishes a fundamental framework for steering noisy dynamical systems and has recently gained renewed interest in the machine learning community. However, current state‑of‑the‑art policy‑based methods suffer from prohibitive computational costs and instability due to their heavy reliance on full‑trajectory simulation. To overcome these limitations, we propose a paradigm shift toward a value‑based approach by revisiting Path Integral Control (PIC). Although standard PIC suffers from the same high‑variance bottleneck as policy‑based methods, we discover that by truncating and marginalizing the original path integral formulation, we can derive a temporal recursive form of the value function. Building upon this theoretical foundation, we propose the Path Integral Value Matching (PI‑VM) algorithm. Specifically, we employ temporal‑difference learning to approximate the recursive value dynamics, and further integrate the Girsanov theorem with experience replay to enable off‑policy training. We benchmark PI‑VM against SOTA policy‑based methods across various SOC benchmarks and sampling tasks. Empirical results demonstrate that PI‑VM matches SOTA precision with an order‑of‑magnitude efficiency gain in low‑dimensional settings, while effectively mitigating mode collapse in high‑dimensional scenarios. Consequently, PI‑VM offers a scalable solution for solving complex SOC problems.

Authors:Kaivalya Rawal, Daria Onitiu, Brent Mittelstadt, Sandra Wachter, Chris Russell
Title: Rule of Thumb: Explaining Artificial Intelligence Systems using Partial Information
Abstract:
Explainable Artificial Intelligence (XAI) seeks to explain how an Artificial Intelligence (AI) system arrived at a particular decision. We propose ''Rule of Thumb'' (RoT) explanations, a new approach to XAI based upon a novel formulation that identifies the most relevant features for predicting the behaviour of an AI system, for a particular datapoint. We show how RoT is well‑suited to enable XAI in: (a) zero‑shot classification using large language models (LLMs), (b) auditing of opaque AI systems without model access, and (c) the use of AI in scientific discovery. Additionally, RoT meets specific requirements from leading AI regulations, provides a familiar interface and visualisations for XAI practitioners, is model‑agnostic, and is substantially faster than alternatives. Code available at: https://github.com/KaiRawal/Rule‑of‑Thumb‑Explaining‑Artificial‑Intelligence‑Systems‑using‑Partial‑Information

Authors:Zihao Liu, Xiaolong Shen, Zhenglin Zhou, Ruijie Quan, Yi Yang
Title: Beyond Pixels: From Video Priors to 4D Worlds
Abstract:
4D generation synthesizes dynamic 3D scenes from conditions such as text or images. Existing methods either reconstruct generated RGB videos with a separate 4D model or adapt a particular video generator to predict geometry directly. The former suffers from distribution mismatch and error propagation, whereas the latter ties 4D prediction to a specific generator and may require retraining when the generator or conditioning regime changes. We ask whether the final denoised latents of video models that share a variational autoencoder (VAE) can instead provide a reusable interface to explicit 4D prediction. Building on this insight, we introduce direct latent‑to‑4D generation and instantiate it as Latent‑to‑4D, which bypasses RGB by aligning a video latent with the token grid of a pretrained 4D decoder and refining it through frame‑wise and global spatiotemporal attention. Trained on roughly 1K existing reconstruction clips, a single checkpoint transfers unchanged across multiple video diffusion transformers within the same VAE family. On Text4D‑200 and I4D‑200, Latent‑to‑4D surpasses matched same‑latent Wan+4RC cascades in projection‑based DINO‑F1 by 2.88‑‑3.45 and 5.81 points, respectively, while also being preferred by human raters for geometry, temporal stability, and overall quality.

Authors:HyeonJun Lee, Hyeonsik Jo, Jinwoo Chung, Jangho Kim
Title: SQuaT: Self-Supervised Knowledge Distillation via Student-Aware Quantized Teacher Features
Abstract:
Quantization‑Aware Training (QAT) enables the deployment of quantized models with minimal accuracy degradation. However, in practical scenarios, training labels are often unavailable due to privacy, copyright, or cost constraints. Knowledge Distillation (KD) is a common approach to address this challenge, but we observe that prior work combining QAT with KD suffers from a fundamental limitation: during distillation, the range mismatch between the teacher and the quantized student model induces an unattainable residual, resulting in an irreducible lower bound on the distillation loss. Motivated by this observation, we propose SQuaT (Student‑Aware Quantized Teacher Features), a label‑free QAT framework with KD that theoretically eliminates this lower bound by applying the student's quantization parameters to quantize the teacher's features during distillation. Through comprehensive experiments across diverse settings, we demonstrate that SQuaT consistently outperforms strong baselines, with particularly pronounced gains in extreme low‑bit (e.g., 1‑ and 2‑bit) settings. Furthermore, extensive evaluations across various model design choices show that our approach does not rely on specific architectural assumptions, making it broadly applicable across diverse architectures and quantization settings. The source code is available at https://github.com/lcdbsa522/SQuaT.

Authors:Seokhyun Youn, Dahyeon Kye, Sung-Ho Bae, Jihyong Oh
Title: Self-Geometry: GT-Free and Plug-and-Play Test-Time Adaptation for Geometrically Consistent 3D Vision Foundation Models
Abstract:
Recent Vision Foundation Models (VFMs) predict depth, camera pose, and pointmap in a single forward pass without per‑scene optimization, achieving strong generalization. However, enforcing explicit multi‑view geometric consistency, e.g., through bundle adjustment, is computationally costly and is thus not imposed during VFM pretraining, so such inconsistency can arise. To address this, implicit self‑consistency derived from model outputs (e.g., pointmaps, features), though enforced at test‑time in prior work, delivers inherently limited performance gain, especially on scenes where the pretrained VFM is highly inaccurate. In contrast to this implicit signal, we propose Self‑Geometry, a plug‑and‑play test‑time adaptation pipeline that directly imposes explicit multi‑view geometric constraints using 2D pixel correspondences as pseudo ground‑truth. Our proposed Self‑Geometry consists of Geometric Disentanglement Optimization, which combines Multi‑View Consistency and Epipolar Consistency losses with Gradient Disentanglement to prevent gradient conflict; Frame Angular‑Neighbor, a view sampler based on SO(3) geodesic distances for lightly imposing these constraints; and Lightweight TTA, which adapts VFMs via LoRA. Our method achieves consistent improvements in both pose and geometry estimation across six VFMs (VGGT, π^3, DA3‑Giant/Large/Base/Small) and four benchmarks (7Scenes, ETH3D, ScanNet++, HiRoom).

Authors:Shuai Wang, Wangyuan Ding, Yixian Shen, Jia-Hong Huang, Stevan Rudinac, Monika Kackovic, Nachoem Wijnberg, Marcel Worring
Title: MMArt A Multi-Perspective Multimodal Dataset for Visual Art Understanding
Abstract:
Recent vision‑language models demonstrate impressive general visual understanding, yet their art interpretation remains shallow: they describe surface content but struggle with formal analysis, grounded historical interpretation, or affective characterization. We argue this is not only a model but also a dataset limitation. Existing art datasets are single perspective resources, where no dataset provides narrative, formal, emotional, and historical perspectives simultaneously for the same artworks. We introduce MMArt, a large‑scale dataset of 74,234 WikiArt paintings, each annotated with four independently annotated perspectives plus a harmonized unified caption, produced by specialized vision‑language models or human annotation and validated through complementary quality evaluations. Two complementarity analyses establish that perspectives encode genuinely distinct information. A generative analysis shows that formal analysis descriptions best preserve compositional style, and historical descriptions carry strong affective signal in reconstructed images. A discriminative retrieval analysis reveals task‑asymmetry: narrative descriptions drive retrieval (R@1 = 44.0%), while formal descriptions, strongest for reconstruction, are nearly nondiscriminative at retrieval scale (R@1 = 7.8%). Leave‑one‑out analysis further confirms that historical descriptions are the least replaceable perspective across both tasks. Together, the two analyses establish that no single perspective suffices for all tasks, directly motivating MMArt multi‑perspective design. The dataset, code, and additional information are available at https://shuaiwang97.github.io/MMArt/.

Authors:Haoze Liu, Run Liu, Haiying Xu, Jiahui Han, Siyuan Fang, Siyu Yan, Huiqi Deng, Guanchu Wang, Na Zou
Title: Your LLM, Your Style: Behavioral Mode Axes for LLM Behavioral Control
Abstract:
Large language models (LLMs) increasingly act in interactive settings where their behavioral styles affect user experience, safety, and downstream decision making. Existing LLM personality studies largely rely on self‑report questionnaires administered in first‑person settings, making the resulting profiles sensitive to surface elicitation choices and poorly grounded in concrete model behavior. In this work, we introduce a situated behavioral‑data (B‑data) framework for studying and controlling LLM behavioral personality. We construct 3,200 contrastive behavioral scenarios spanning 20 behavioral patterns and four prompt registers, grounded in validated psychometric facets such as BFI‑2, DOSPERT, and HEXACO. Using this framework, we find that LLMs exhibit stable and model‑specific behavioral profiles, while also revealing register‑dependent shifts across first‑person decisions, advice‑giving, and task execution. We then show that these behavioral patterns can be controlled through Behavioral Mode Axes (BMAs), activation‑space directions derived from contrastive behavioral traces. Compared with response‑derived BMAs, which are more prone to trait drift, thought‑derived BMAs more faithfully capture the intended behavioral mechanism and provide cleaner control over situated behavioral styles. Our results suggest that LLM personality‑like tendencies are better understood not as abstract self‑report traits, but as measurable and controllable behavioral modes grounded in concrete interaction contexts. Our code and data are available at https://github.com/lhz191/LLM‑Behavioral‑Personality.

Authors:Hongrui Bao, Hangyu Rong, Zhuoshang Wang, Yubing Ren, Yanan Cao
Title: EVIL-Detect for NLPCC 2026 Shared Task 6: LLM-Generated Text Detection
Abstract:
The rapid development of large language models (LLMs) has increased the need for reliable detection of LLM‑generated text, especially in realistic Chinese scenarios involving human‑written text (HWT), LLM‑generated text (LGT), and LLM‑refined text (HLT). This paper presents EVIL‑Detect, a multi‑signal ensemble framework with conflict‑aware fusion for NLPCC 2026 Shared Task 6. The system integrates edit‑extent regression, zero‑shot likelihood‑contrast signals, lexical statistics, and conservative text rules. With calibrated decision boundaries and conflict‑aware integration, our system improves robustness under strong out‑of‑distribution shifts, achieving a macro‑F1 score of 0.8888 and ranking first in the official evaluation. Our code is available at https://github.com/bbbbhrrrr/evildetect.

Authors:Matteo Grella
Title: The Signal Rail: A Deterministic Motion Grammar for Communicating Conversational Agent State in Terminal Interfaces
Abstract:
Terminal interfaces to conversational agents report rich internal state (listening, thinking, executing tools, awaiting input, failing) almost entirely through text, while the motion channel beside it, the one peripheral vision monitors without reading, carries a single bit: alive. We present the Signal Rail, a one‑row terminal status instrument that gives that channel a grammar. Four ideas govern it: spatial semantics (input, processing, and output zones, with direction as meaning), a motion grammar (one kinetic rule per state, never color alone), determinism (frames as a pure function of explicit inputs, golden‑frame testable), and honesty (no invented progress or activity). We contribute a 45‑section normative specification and a reference implementation inside a working full‑duplex local voice agent driven by real signals.

Authors:Chengzhi Zhang, Xinyi Yan, Wenqi Yu
Title: Leveraging Human Reading Behavior for Keyphrase Extraction: A Webcam-based Eye-tracking Corpus
Abstract:
Purpose: Keyphrases are statistically and semantically important textual units that can also attract readers' attention during comprehension. However, existing keyphrase extraction (KPE) studies mainly focus on improving textual representation while largely overlooking human reading behavior. This study examines whether lightweight webcam‑based eye‑tracking features can improve KPE from Chinese academic abstracts in Library and Information Science (LIS). Methodology: To address the limited availability of eye‑tracking data for Chinese academic reading, we developed a lightweight webcam‑based data collection platform using the open‑source SearchGazer library and constructed the Chinese LIS Eye‑Tracking Corpus (CLIS‑ET). Three character‑level eye‑tracking features, first fixation duration (FFD), fixation number (FN), and total fixation duration (TFD), were incorporated into KPE models to evaluate their effects on extraction performance. Findings: Eye‑tracking features consistently improved KPE performance. The combination of FN and TFD achieved the best results on the Att‑BiLSTM+CRF model, indicating that readers' fixation behavior provides useful signals for identifying keyphrases in academic abstracts. Originality/value: This study introduces a cost‑effective webcam‑based eye‑tracking approach for KPE and presents CLIS‑ET, a Chinese academic eye‑tracking corpus containing FFD, FN, and TFD features. The results demonstrate the value of incorporating human reading behavior into keyphrase extraction. Dataset and code: https://github.com/yan‑xinyi/ET_AKE and https://github.com/yan‑xinyi/Reading_ET_System.

Authors:Akrin Zheng, Alexander Wu, Alaia Liu
Title: ENTLORE: A Graph-Grounded Benchmark for Latent Organizational Reasoning in Enterprise Question Answering
Abstract:
Enterprise question answering is framed as retrieving internal documents and generating grounded answers. Routine enterprise records, however, are work by‑products in which required organizational relations remain implicit across heterogeneous sources. Existing benchmarks provide realistic multi‑source evidence, but often materialize a predefined answer path and therefore test the composition of stated facts rather than recovery of a target relation absent from the corpus. We call the latter capability latent organizational reasoning. We introduce ENTLORE, a graph‑grounded benchmark construction framework that reconstructs an audited enterprise world from routine documents, authoritative organizational tables, and operational records. Versioned organizational conventions certify derived relations in a truth graph, enabling complete golden answers and proof certificates. The aligned anonymized release exposes only the document corpus while withholding private structure and target relations. ENTLORE contains 2,341 documents from three source types and 907 questions spanning explicit lookup, cross‑source composition, and latent organizational reasoning, evaluated across 56 model and access configurations. Structuring the released world as an induced entity graph or navigable knowledge base gives the strongest deployable results. Yet supplying gold documents still leaves 30.4% of latent questions unanswered, versus 12.6% and 6.2% for explicit and compositional questions. Enterprise QA therefore depends not only on document recall, but also on whether implicit organizational relations become usable. The benchmark, data, and code are publicly available at https://github.com/scitix/entlore .

Authors:Wenbo Dong, Dipankar Bhattacharya, Akinari Kobayashi, Akira Seino, Fuyuki Tokuda, Xuzhao Huang, Kai Tang, Norman C. Tien, Kazuhiro Kosuge
Title: Precise Top-Layer Fabric Segmentation for Fabric Destacking with Edge- and Shape-Aware Deep Networks
Abstract:
Fabric destacking requires precise segmentation of the topmost fabric layer, a task complicated by subtle fabric boundaries and high visual similarity between fabric layers. Existing semantic and edge‑based segmentation approaches often struggle with these complexities, limiting the performance of robotic manipulation for different tasks. In this work, a novel segmentation training architecture tailored for top‑layer fabric segmentation in stacked fabrics is proposed. The method extends the classical encoder‑decoder framework by introducing two specialized branches ‑ an edge‑aware branch and a shape‑aware branch ‑ that are used to supervise the backbone network for better tuning. The edge‑aware branch enhances boundary delineation, while the shape‑aware branch guides the network to capture and align the overall fabric shape with reference masks derived from Computer Aided Design (CAD) models. Experiments on a real‑world fabric dataset demonstrate that the training approach outperforms established baselines, verifying the effectiveness of the multi‑branch design through both quantitative results and ablation studies.

Authors:Zhuchenyang Liu, Ziyi Wang, Yao Zhang, Yu Xiao
Title: DistilVDR: A Compact End-to-End Visual Document Retriever via Dual-Student Distillation
Abstract:
Visual document retrieval (VDR) is dominated by multi‑billion‑parameter models that are slow to index at full corpus scale and expensive to serve. Prior compression routes either train a smaller multi‑vector encoder from scratch or distil only the query side; neither yields a compact single‑vector retriever end‑to‑end. We present DistilVDR, a 524M end‑to‑end VDR system distilled bilaterally from a single 8B vision‑language teacher under a pointwise cosine alignment loss. All supervision comes from the frozen teacher's embedding space, which was itself trained with relevance supervision, so the student objective needs no relevance labels, negative sampling, or contrastive term. We match VDR's text‑query and image‑document input asymmetry with an asymmetric encoder‑only student that concentrates visual capacity on the document side and keeps the query side at 70M parameters. We release two variants that share the same encoders and training and differ only in the document encoder's visual‑tile budget: DistilVDR‑HiRes attains 61.74 average NDCG@5 on ViDoRe v1+v2+v3 (86.9% of the 8B teacher) and leads every reproduced sub‑1B baseline on the high‑resolution‑sensitive v3 benchmark, while DistilVDR‑Fast attains 59.98 at a 3 times smaller visual‑token budget. Both variants store one million documents in a 15.6 times smaller index than the strongest sub‑1B multi‑vector baseline and index the corpus an order of magnitude faster. The code is available at https://github.com/Ryenhails/NanoVDR.

Authors:Kaican Li, Weiyan Xie, Lewei Yao, Jiannan Wu, Lanqing Hong, Yongxiang Huang, Nevin L. Zhang
Title: InSight-doc: Agentic Visual Perception for Long-Document Understanding
Abstract:
Long‑document understanding often requires reasoning over many visually rich pages, making inference costly and prone to context rot. In this work, we propose InSight‑doc, an agentic visual perception framework that treats visual resolution as an adaptive reasoning‑time resource. InSight‑doc starts from low resolution and selectively zooms into high‑resolution regions for finer evidence, without relying on any external retriever. To train such an agent, we construct an active‑perception corpus of 17.9K high‑quality SFT examples with region‑level zoom‑in trajectories, accompanied by 19.2K hard RL examples. Through SFT+RL, InSight‑doc‑8B improves the baseline by 4.3‑‑16.4 accuracy points over document VQA benchmarks. On long documents, it reduces hallucination by more than 40% and inference latency by 41%‑‑68% while maintaining an accuracy lead. Our code, datasets, and model are released at https://github.com/m‑Just/InSight‑doc .

Authors:Shijun Luo, Lizhi Wan
Title: ASR-Roundtrip Evaluation Can Mask Context- and Convention-Dependent Reading Errors in Chinese News TTS
Abstract:
ASR‑roundtrip evaluation is widely used as a scalable proxy for text‑to‑speech (TTS) intelligibility, but it can produce false negatives for reading errors perceived by listeners. We study Chinese news TTS spans whose correct reading depends on context or domain conventions, such as sports scores, aircraft models, technical units, and membership names. In these cases, Raw TTS can choose a plausible but wrong reading while ASR transcribes the audio as the intended or surface‑correct text. A targeted audit over 110 high‑risk MiMo TTS cases, reported with a complete denominator, confirms 46 masked false negatives, 9 exposed TTS errors, and 55 cases with no Raw TTS error. A span‑isolation diagnostic re‑exposes 18/46 previously masked errors. A Raw‑only CosyVoice audit on the same targeted pool confirms 51 masked cases. Across the 97 TTS‑specific audio files labeled confirmed masked across the two audits, Qwen3‑ASR surface‑recovers 40 cases, whereas Paraformer does so in only 2. The results suggest that ASR‑roundtrip is useful for screening but insufficient as standalone ground truth for Chinese news reading‑risk evaluation.

Authors:Namritha Lasyapriya Maddali, Rajini Makam, Suresh Sundaram, Narasimhan Sundararajan
Title: $π$-SUB: A Physics-Informed Synthetic Underwater Benchmark Dataset for Underwater Image Enhancement
Abstract:
This paper presents π‑SUB, a physics‑informed framework for generating synthetic underwater benchmark datasets that bridges the synthetic‑to‑real gap for Underwater Image Enhancement (UIE). The proposed framework extends the classical underwater image formation model by incorporating depth‑dependent downwelling irradiance, biologically resolved absorption, and environmental scattering across all ten Jerlov water types, together with independently controllable residual phenomena. Using this framework, the π‑SUB dataset consists of paired synthetic underwater‑reference images spanning shallow‑to‑deep and coastal‑to‑oceanic environments. Extensive simulation studies have been carried out to evaluate π‑SUB along two criteria namely hyper‑realism and generalizability. For hyper‑realism, π‑SUB attains a global Frechet Inception Distance (FID) that is 46% lower than Syrea. For generalizability, four state‑of‑the‑art UIE architectures (FUnIE‑GAN, Pix2Pix, PUIE‑Net, and Phaseformer) are used for comparative evaluation of π‑SUB. These models were independently trained on six datasets including one real and five synthetic datasets and tested on six real‑world benchmarks datasets. Across four UIE architectures and six real benchmark datasets, π‑SUB improves UIQM by 4.18% over PHISWID (next best) and 9.46% over Syrea (next best), while reducing NIQE by 48.78% and 23.98%, respectively. These results establish π‑SUB as a hyper‑realistic and generalizable benchmark for developing the next generation of underwater image enhancement methods. The code and dataset are available at https://github.com/airl‑iisc/pi‑SUB

Authors:Fanqi Zhou, Qiaosheng Chen, Zixian Huang, Gong Cheng
Title: Agentic Instruction Data Selection: Let DataMaster Interpret Your Intent
Abstract:
Although existing instruction data selection methods have introduced various metrics, the inherent complexity of real‑world datasets makes it impractical for any single metric to generalize across all scenarios. Developers are thus often forced to manually inspect data and craft heuristic rules for each new application‑‑‑a tedious and error‑prone process. In this paper, we propose a paradigm shift from manual configuration to automated orchestration via the Instruction Data Selection Agent (DataMaster), which interprets user intent and autonomously composes optimal selection strategies. By allowing users to specify data needs through natural language descriptions, DataMaster simplifies data curation and removes the burden of manual strategy design. Extensive experiments across the math, medical, and code domains show that DataMaster outperforms static baselines in most settings and surpasses full‑pool training in a substantial number of cases. The implementation of DataMaster and the scripts needed to reproduce the reported pipeline are publicly available at https://github.com/nju‑websoft/DataMaster.

Authors:Sangjin Jin, Kangmin Kim, Junhyeong Lee, Yongjae Lee
Title: Retrieval-Corrected Conformal Prediction for Time Series
Abstract:
Conformal prediction (CP) provides distribution‑free prediction intervals for fixed forecasters, but its standard calibration procedure is often inefficient for time series data, where forecast errors are temporally dependent and change across time and operating conditions. Recent time series CP methods improve local calibration using recent, weighted, or localized residuals. Yet local calibration can remain indirect, since broad residual weighting or additional adaptation procedures may dilute the evidence most relevant to the current prediction. This motivates a simple retrieval and correction strategy that selects similar past residuals as local evidence and then corrects the coverage error left by retrieval. In this paper, we propose Retrieval‑‑Corrected Conformal Prediction (RCCP), a retrieval‑augmented calibration method for time series prediction intervals. RCCP builds an asymmetric interval from retrieved one‑sided residuals and calibrates its normalized retrieval error with a scalar conformal correction. Thus, retrieval provides local residual evidence, while conformal correction determines the final scale needed for coverage. We provide a coverage‑gap bound based on the stability of the normalized retrieval error distribution. Across standard benchmarks and backbone forecasters, RCCP attains the target coverage in every setting and achieves the lowest Winkler scores, with fewer severe misses. RCCP also achieves low calibration and inference overhead, showing that retrieval‑corrected calibration is an effective and scalable approach to uncertainty quantification in time series forecasting. Code is available at https://github.com/jinsaaang/rccp.

Authors:Sangwoo Jo, Donggeun Ko, Jayeon Kang, Youngsang Kwak, Jaehwa Kwak, Sungjoon Choi
Title: Flow Straight to Reality: Perceptually Consistent Flow Matching for Efficient Image Restoration
Abstract:
Image restoration is fundamentally constrained by the tradeoff between distortion and perception: minimizing pixel‑wise error yields over‑smoothed results, whereas optimizing for perceptual realism often introduces structural deviations. Recent approaches attempt to balance this tradeoff via posterior sampling or multi‑stage generative pipelines, yet remain computationally expensive and architecturally complex. To overcome these limitations, we propose PCFlow (Perceptually Consistent Flow Matching), a unified framework that directly parameterizes a continuous transport from degraded observations to clean targets, jointly optimizing distortion and perceptual quality. While its latent consistency flow objective drives stable and efficient few‑step inference, a Latent Consistency Perceptual Loss (LCPL) imposes semantic constraints directly on the guiding velocity field, steering the dynamics toward visually sharp data manifolds. Furthermore, recognizing the inherent conflict between structural and perceptual consistencies, we integrate a conflict‑free gradient projection strategy to stabilize the multi‑objective optimization landscape. Combined with lightweight, convolution‑only backbone, PCFlow achieves competitive performance across diverse restoration tasks at a fraction of traditional computational costs.

Authors:Chenhao Dang, Siyuan Xiong, Conghui He, Weijia Li
Title: SKILLER: Language-Level Reinforcement Learning for Reusable Skill Extraction in Small Language Models
Abstract:
Agent skills represent a standardized format for packaging procedural knowledge and domain expertise, serving within agent harness systems as an essential mechanism to continually constrain a language model's behavior space for repeatable, high‑quality task execution. However, because strong closed‑source models entail high inference costs, current popular agent harnesses, such as Codex and OpenClaw, remain prohibitively expensive when deploying these skills to accomplish real‑world tasks. The rapid capability enhancement of open‑source models deployable on consumer‑grade GPUs presents a compelling opportunity to drastically reduce these costs by leveraging skill‑based behavioral constraints. Nevertheless, automatically generating effective skills tailored specifically for such compact models remains a significant practical challenge. To address this, we propose SKILLER, a natural‑language‑driven reinforcement learning framework designed to automatically generate executor‑specific skills for small models, which employs a strong model as the actor and critic, treats the small‑model agent system as the environment, and propagates all reinforcement learning signals entirely via natural language. Extensive experimental evaluations across five relevant benchmarks using Qwen3.5‑9B and Qwen3.5‑4B demonstrate that SKILLER outperforms three open‑source and one closed‑source skill generation or evolution methods, achieving absolute gains ranging from 4.3 to 20.4 percentage points for the 9B model and 1.8 to 13.3 points for the 4B model, while remarkably matching the performance of strong closed‑source models on single‑skill tasks in SkillsBench. The project is available at https://github.com/DANG‑ai/SKILLER.

Authors:Xiangchen Pan, Wei Wei, Huakang Niu, Zhicong Cheng
Title: Multi Interests for Joint Search-Recommendation Modeling
Abstract:
Search and recommendation are crucial for understanding user preferences. More and more studies are attempting to jointly model search behavior and recommendation behavior, by integrating user active search and passive recommendation behavior data to better mine user preferences. However, although existing cross‑domain unified modeling frameworks can effectively compensate for the differences in behavior between domains, they overlook the expression of interests in different scenarios under mixed sequences. In this study, we propose a multi‑interest‑based mixed sequential modeling framework MIJSR, which performs multi‑interest mining and adaptive integration on search recommendation mixed sequences from both structural and semantic perspectives. Specifically, our model can be roughly divided into three modules: cross‑domain behavior fusion, multi‑interest mining, and multi‑task prediction. Firstly, we align the representations of query and item through contrastive learning training. Then, we extract the multi interests of the mixed behavior sequence from both structural and semantic perspectives. Structurally, we extract search interests, recommendation interests, and cross interests through subsequence partitioning and mask settings; In terms of semantics, we use the semantic information of queries for clustering and perform semantic segmentation on mixed sequences to construct semantic multi interests. Finally, the adaptive fusion of multiple interests is combined with other side information to use a progressive layered extraction model for multi‑task prediction. Extensive experiments on two open‑source datasets have shown that our model can further enhance its accuracy in search and recommendation by extracting users' multi interests at a fine‑grained level. Codes are available at https://github.com/pxcstart/MIJSR.

Authors:Caoyuan Ma, Wenpu Liu, Weichu Xie, Tian Gu, Shilei Zhao, Lingxi Min, Shuai Dong, Yuqi Xu, Ji Zhao, Ziyue Wang, Wenzheng Chang, Taiqiang Wu, Yongfu Zhu, Wenqi Shao, Yinqiang Zheng
Title: SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning
Abstract:
Large vision‑language models (LVLMs) remain vulnerable to jailbreak attacks that exploit visual inputs to bypass safety alignment inherited from their language backbones. We propose SafeCap, a reinforcement‑learning framework that aligns LVLMs through learned self‑captioning. SafeCap trains a policy model to first generate a safety‑relevant image caption and then produce a final answer; the caption is further optimized by whether it enables a frozen LLM to reach a safety‑aligned decision. This caption‑mediated objective encourages the policy to expose visual cues relevant to safe response generation rather than relying solely on direct refusal supervision. Across five multimodal safety benchmarks and six vision‑utility benchmarks, SafeCap substantially improves aggregate safety performance under its intended DirectCap protocol, with gains of 3.7‑19.0 points in safety average across four model settings while maintaining comparable or improved vision utility. Under controlled comparisons on matched backbones and data, SafeCap outperforms safety SFT, DPO, and SafeGRPO, demonstrating the effectiveness of caption‑mediated reinforcement learning for multimodal safety alignment.

Authors:Zhichen Yang, Rui Xu, Yuzhen Niu, Fusheng Li, Hui Da, Ri Cheng
Title: Towards Color-Faithful Low-Light Image Enhancement via Adaptive Color Debiasing and Saturation Rectification
Abstract:
Low‑light imaging often introduces color bias caused by the low signal‑to‑noise ratio and the image formation process. Although recent low‑light image enhancement methods have achieved strong brightness recovery, faithful color restoration remains challenging, manifesting as overall color bias together with local under‑ and over‑saturation. To address this issue, we propose CAGE, a cylindrical color correction framework with adaptive color debiasing and gamut‑harmonized saturation rectification for color‑faithful low‑light image enhancement. We first introduce AdaLAB, a cylindrical adaptive LAB color space that provides a decoupled and image‑specific basis for uniform color correction. Building on this color space, we further develop AdaCCT, an adaptive cylindrical color transform with forward and inverse transforms for the conversion between RGB and AdaLAB color space, as well as necessary color debiasing and saturation rectification. The forward transform suppresses embedded color bias before backbone enhancement by reorganizing the chromatic distribution through chromatic‑plane shifting and scaling, while the inverse transform achieves faithful saturation rectification through out‑of‑gamut lightness compensation. Extensive experiments on multiple benchmarks show that CAGE achieves more faithful color restoration, specifically reduces color bias and saturation abnormality, and delivers better overall visual quality across different low‑light enhancement backbones. The code is available at https://yangzhichen763.github.io/CAGE/.

Authors:Junwoo Park, Minyoung Shin, Cheol Soon Lee, Sujee Lee
Title: Multi-Granular Rationale-Guided Molecular LLM for Property Prediction
Abstract:
Large language models (LLMs) are widely applied across chemical tasks, such as molecular property prediction, which underpins drug discovery. Molecular LLMs represent a molecule through several modalities, notably a 1D SMILES sequence or a 2D molecular graph. Both encode molecular information implicitly, so the contribution of individual substructures remains opaque. Retrieval and augmentation methods add context, but from external sources. However, the cues chemists reason over are the internal substructures that drive a property up or down. We propose MR‑MoL, a multi‑granular rationale‑guided molecular LLM that supplies this evidence directly. A fine‑tuned GNN scores each substructure through masking, and the most influential ones are serialized as a ranked, direction‑tagged rationale that the LLM reads alongside the SMILES sequence and molecular graph. The rationale spans three levels of granularity: Murcko scaffolds with their side chains, BRICS fragments, and functional groups. This is, to our knowledge, the first method to expose GNN‑derived attributions to an LLM as evidence for property prediction. On eight MoleculeNet tasks, MR‑MoL achieves the best overall results among generalist models and narrows the gap to specialist models tuned for each task. Five diagnostics further confirm that the model reads the rationale rather than merely benefiting from its presence. Its direction, rank, and substructure each shape the prediction, and its attributions reproduce known structure‑property relationships.

Authors:Guixu Lin, Yuyang Yu, Xiang Ji, Linyao Chen, Zhengwei Yin, Mengshun Hu, Mingdeng Cao, Shengfeng He, Yinqiang Zheng
Title: Bridging Event Streams and DiT: Event-Guided Video Frame Interpolation
Abstract:
Latent diffusion models have recently advanced video frame interpolation by synthesizing intermediate frames between input images. However, handling large temporal gaps and complex motion remains challenging, often resulting in motion blur, structural distortions, and temporal inconsistencies. Event cameras provide high‑temporal‑resolution motion cues that are well suited for bridging these gaps and improving interpolation quality. To exploit this advantage without training an event‑assisted model from scratch, we propose an adapter‑based framework that incorporates event‑derived cues into a pre‑trained image‑to‑video diffusion model with minimal architectural changes. Specifically, our method leverages Image Warped Events (IWEs) and bidirectional sparse optical flow to provide spatially and temporally aligned guidance during generation. By injecting these event‑guided structural and motion cues into the diffusion process, our approach reduces interpolation artifacts and improves both reconstruction fidelity and temporal coherence. Experimental results on real and synthetic benchmarks show that our method consistently outperforms existing state‑of‑the‑art approaches. The project page is at https://joseph‑lin‑tech.github.io/BridgeEventDiT‑VFI/.

Authors:Shuo Bao, Wei Dong, Shuyue Zhang, Ming Shang, Yuchen Huang, Han Yu, Chengjie Xu, Yiheng Bi, Kai Sun, Fuchun Sun, Xinzhou Wang
Title: PBD-AG: Persistent Baseline-Delta Active Graphs with Uncertainty-Aware Inspection for Long-Horizon Service Robots
Abstract:
Long‑horizon service robots require persistent world models that can be built autonomously in unseen environments and revised as task‑relevant objects change. Existing methods rely on online mapping, which accumulates localization and observation errors, static scene representations that cannot capture persistent object changes, or holistic vision‑language predictions that lack verifiable 3D geometric evidence. We present PBD‑AG, a persistent baseline‑delta active graph framework that decouples robot‑verified stable fixtures from revisable dynamic object events. Under our framework, the robot autonomously bootstraps the structural baseline from onboard exploration and inspects discovered fixtures to ground hierarchical object beliefs. PBD‑AG maintains reliability‑weighted object states over geometry, semantics, identity, existence, and support relations, utilizing a geometric visibility gate to mitigate false deletions under occlusion. Inspection viewpoints are selected by a graph‑conditioned policy that balances target coverage, travel cost, collision risk, and redundant observation. Simulation experiments in multiple environments and under controlled dynamic evaluation show higher aggregate coarse‑fixture F1 than capability‑matched controls, as well as stronger identity continuity and event recall. A qualitative physical‑robot demonstration further illustrates integration with onboard sensing, providing a traceable world model for long‑horizon robotic perception. The project page of PBD‑AG is available at https://shuobao214.github.io/PBD‑AG/

Authors:Linh Dieu Le, Tong Chen, Shazia Sadiq, Hongzhi Yin, Ming Jin, Junliang Yu
Title: Towards Efficient Reasoning in LLM-Based Recommender Systems via Model Merging
Abstract:
Large language model‑based recommender systems are increasingly adopting slow‑thinking models that generate step‑by‑step reasoning before making predictions, often achieving higher accuracy than fast‑thinking models that predict directly. However, their reasoning traces are often unnecessarily verbose, increasing inference costs without commensurate accuracy gains. Existing training‑based approaches to reasoning compression often incur substantial adaptation costs, while inference‑time methods are brittle and difficult to scale. These limitations motivate model merging as a promising training‑free direction for transferring specialised behaviours between models in a shared parameter space. In particular, merging a slow‑thinking model with a fast‑thinking counterpart provides a natural mechanism for balancing recommendation accuracy and reasoning conciseness. To this end, we propose, to our knowledge, the first model merging framework for reasoning compression in recommender systems. Unlike conventional merging methods that apply uniform merge coefficients across model components, our method performs fine‑grained merging at the level of individual attention heads, capturing heterogeneous patterns in recommendation reasoning. Each attention head is assigned a distinct merge coefficient according to its contribution to critical reasoning evidence and its sensitivity to parameter change, enabling selective injection of the concise behaviour of the fast‑thinking model into the slow‑thinking model and reducing reasoning verbosity without compromising recommendation quality. Experiments on three benchmark datasets show that our method reduces reasoning length by up to 24.3% while outperforming competitive model merging baselines in maintaining recommendation accuracy. The code is available at https://github.com/linhledieu/REAM.

Authors:Ruizhong Liu, Tingzhang Luo, Zaiyan Zhang, Jundong Chen, Hongruixuan Chen, Shaoguang Huang, Hongyan Zhang
Title: GeoSeg-OV: Bridging Geospatial Gaps with Structural Guidance for Open-Vocabulary Remote Sensing Segmentation
Abstract:
Open‑vocabulary remote sensing segmentation has recently emerged as a promising paradigm that enables pixel‑level recognition of arbitrary categories specified by natural language, including classes unseen during training. However, geospatial domain shifts caused by heterogeneous regions, spatial resolutions, and acquisition platforms weaken visual‑text matching and limit cross‑dataset generalization. Recent attempts have begun to incorporate auxiliary vision foundation models (VFMs), typically coupling their features with text embeddings as additional matching evidence. However, this strategy may introduce inconsistent matching signals while leaving the structure‑sensitive representations of VFMs insufficiently exploited. We therefore propose GeoSeg‑OV, which decouples auxiliary VFM features from visual‑text matching and repurposes them as structural guidance for cost aggregation and decoding. GeoSeg‑OV constructs an orientation‑robust cost volume from multi‑rotation CLIP features, while a frozen VFM extracts multi‑scale structure‑sensitive features in parallel. We propose Structure‑Guided Aggregation (SGA), which integrates cost tokens and CLIP semantic guidance with VFM‑derived pairwise structural biases for coherent spatial propagation, followed by text‑conditioned class‑wise reasoning. We further introduce Cost‑Aware Decoding (CAD) to adaptively refine and fuse multi‑scale semantic and structural guidance based on the current decoder context. On the global High‑Resolution Land Cover (HRLC) benchmark spanning seven datasets across six continents, GeoSeg‑OV outperforms the state‑of‑the‑art by +2.5 and +2.7 average mIoU under two training settings. A large‑scale zero‑shot case study further demonstrates its generalization across geographic domains and category systems without target‑domain annotations or retraining.

Authors:Zebin Xing, Yupeng Zheng, Qiang Chen, Linbo Wang, Yichen Zhang, Pengxuan Yang, Junli Wang, Deheng Qian, Xiaoqing Ye, Junyu Han, Yifeng Pan, Qichao Zhang, Dongbin Zhao
Title: DriveVLA-M0: Failure-Aware Memory Augmentation for Autonomous Driving
Abstract:
Vision‑Language‑Action (VLA) models have recently emerged as a promising paradigm for end‑to‑end autonomous driving by enabling unified reasoning across perception, language, and planning. However, existing approaches lack mechanisms to exploit past failures or adapt to distribution shifts, causing the model to persistently underperform on similar scenarios where it has previously failed. In this paper, we propose DriveVLA‑M0, a retrieval‑augmented VLA with failure‑aware latent memory. We construct a latent memory pool that stores failure cases along with their structure scene representations and expert trajectory labels, and design a dedicated Retrieve Model that decouples static road structure and dynamic agent interactions to enable structurally grounded retrieval. At inference time, retrieved cases are injected into the model via a lightweight decoupled LoRA‑based test‑time training (TTT) mechanism, allowing targeted and scenario‑specific correction without modifying the backbone. Extensive experiments on NAVSIMv1 and NAVSIMv2 benchmark demonstrate that our approach consistently outperforms prior methods, achieving 94.1 PDMS on Navtest and 47.0 EPDMS on Navhard with only 26.44 ms TTT backward latency overhead. Furthermore, we show that DriveVLA‑M0 scales effectively with additional memory, enabling training‑free performance gains through memory expansion. The code is available at https://github.com/ZebinX/DriveVLA‑M0.

Authors:Mizanur Rahman, Arshia Azimlu, Shadikur Rahman, Md Tahmid Rahman Laskar, Amran Bhuiyan, Shafiq Joty, Enamul Hoque Prince
Title: VisEditBench: Can Vision-Language Models Edit Visualization Code from Multimodal Feedback?
Abstract:
Vision‑language models (VLMs) have shown strong capabilities in generating visualization code from textual or visual specifications. However, real‑world visualization authoring is inherently iterative: users frequently revise existing visualizations to repair flawed charts or adapt them to desired styles. Existing benchmarks primarily evaluate generation from scratch, leaving visualization code editing from multimodal feedback largely unexplored. We introduce VisEditBench, a benchmark of 1,395 human‑annotated visualization code‑editing tasks grounded in realistic visualization workflows and failure cases. VisEditBench covers two practical settings: feedback‑guided repair, where models revise visualization code using buggy or marked charts together with textual feedback, and reference‑guided restyling, where models modify code to match a target chart image. Evaluating 20 state‑of‑the‑art VLMs reveals that visualization code editing remains challenging: Claude‑4.6‑Sonnet achieves the best overall pass rate of 74.46%, while most open‑source models remain below 50%. Performance is particularly weak on visually grounded style adaptation, where Claude‑4.6‑Sonnet achieves only 55.71%. To establish a strong baseline, we further propose VisEditAgent, a render‑grounded editing framework that iteratively generates, executes, validates, and refines candidate edits. Built on GPT‑4o, VisEditAgent improves overall pass rate from 55.75% to 67.99%, demonstrating the importance of render‑grounded feedback for faithful visualization editing. We will release VisEditBench at https://github.com/vis‑nlp/VisEditBench.

Authors:Gongli Zhang, Zhulin Liu, C. L. Philip Chen
Title: Share First, Route What Remains: A Unified Framework for Token-Adaptive MoE Computation
Abstract:
Mixture‑of‑experts (MoE) models have recently moved beyond routing a fixed number of complete experts. Shared‑expert designs preserve reusable knowledge, fine‑grained methods vary computation within experts, and dynamic routers adapt the number of active experts. Yet these decisions are usually made independently, overlooking a basic dependency: extracting reusable computation changes both what remains and how much expert capacity the remainder needs. We study this dependency by decomposing sparsely upcycled feed‑forward experts into key‑value channels. Co‑activated experts align at a subset of value positions; removing these positions changes expert preference; and greater shared coverage is associated with lower residual expert demand. These observations lead to one principle: share first, then route what remains. We instantiate it in UniF‑MoE, a unified framework for token‑adaptive MoE computation. Each expert is partitioned into aligned blocks. A shared‑demand score sets the shared block count and pathway weight, key prototypes select the shared content, and the complementary demand determines the residual expert count through cumulative routing mass. A Gram regularizer separates and normalizes router embeddings, promoting diverse routing directions, sparse expert overlap, and a simple routing geometry. Experiments on DomainBed and GLUE show that this unified design improves predictive performance over representative static and dynamic MoEs while reducing activated computation, inference latency, and memory. Code is available at https://github.com/existence0420/UniF‑MoE.

Authors:Mizanur Rahman, Mohammed Saidul Islam, Ridwan Mahbub, Md Tahmid Rahman Laskar, Shafiq Joty, Enamul Hoque Prince
Title: DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments?
Abstract:
Real‑world data science involves long‑horizon workflows that span data wrangling, exploration, modeling, visualization, and validation, and require coordinated use of tools such as notebooks, IDEs, terminals, browsers, and databases within real operating environments. Yet existing benchmarks lack real‑computer interaction and do not evaluate whether agents can execute complete end‑to‑end data‑science workflows in realistic computing environments, failing to capture the multi‑stage, multi‑tool nature of data‑science practice. We introduce DSAgentBench, the first benchmark to evaluate whether agents can automate full data‑science workflows inside real computer environments. DSAgentBench contains 275 diverse tasks covering the entire data‑science life‑cycle, reflecting the complexity and tool coordination required in practice. Each task requires grounding decisions in intermediate outputs and coordinated tool use, and includes a deterministic evaluator that verifies analytical correctness, visual outputs, and model performance rather than code‑only execution. Our extensive experiments with 15 closed‑ and open‑source models show that even the strongest agent, Claude‑4.6‑Sonnet, achieves only 56.70% task success, while all open‑source agents remain below 1%, frequently failing at tool orchestration, OS grounding, and multi‑step reasoning. These results reveal a substantial capability gap between current agentic systems and real data‑science workflows, positioning DSAgentBench as a foundation for developing grounded, verifiable, autonomous data‑science agents. We release DSAgentBench at https://github.com/vis‑nlp/DSAgentBench.

Authors:Haeyun Choi, Minhyuk Jang, I-Gil Kim
Title: CasDeblurGS: Cascaded 2D-to-3D Multi-View Consistency for 3D Gaussian Splatting from Two Blurry Images
Abstract:
Free‑viewpoint 3D scene media is increasingly important for immersive applications, yet practical capture often suffers from severe view sparsity and motion blur. Although neural rendering has advanced sparse‑view synthesis, existing blur‑aware methods typically require substantial multi‑view redundancy, accurate camera poses, or costly per‑scene optimization. We address a stringent yet practical setting: reconstructing a coherent 3D scene from only two motion‑blurred images with known intrinsics, without input‑view poses, auxiliary sharp images, or per‑scene test‑time optimization. To this end, we propose CasDeblurGS, a cascaded framework that progressively recovers reliable cross‑view information from local 2D correspondences to global 3D guidance. Stage 1 constructs locally reliable guidance through occlusion‑aware correspondence filtering, while Stage 2 aggregates the intermediate restorations into a provisional pose‑free 3D Gaussian representation whose input‑view re‑renders provide dense global guidance for final restoration. The resulting views enable a more coherent 3D representation and higher‑quality novel‑view synthesis. Experiments on real‑world and synthetic Deblur‑NeRF scenes show consistent gains over strong baselines, improving PSNR by 1.19 dB and 2.11 dB, respectively. Progressive ablations, cross‑view correspondence visualization, and camera reprojection analysis further demonstrate improvements in both rendering quality and multi‑view geometric consistency.

Authors:Minwoo Yu, N. Robert Bennett, Jongduk Baek, Adam S. Wang
Title: ENCORE: Efficient Noise Context-Aware Representation for Low-Dose CT Denoising
Abstract:
While deep learning‑based denoising has become widely adopted in low‑dose CT, conventional models use generic architectures designed for natural images, failing to account for non‑stationary and spatially correlated CT noise characteristics. To address this, we propose an Efficient Noise COntext‑aware REpresentation (ENCORE) framework that explicitly leverages CT noise characteristics and anatomical features. First, we reformulate the noise synthesis procedure based on a realistic noise distribution beyond the conventional Gaussian approximation, establishing a rigorous foundation for training pair generation. Next, we extract local noise power and correlation contexts to guide the denoising process. To fully leverage the potential of noise context, we propose a FlyingConv module, which adaptively changes convolution weights for each local image region. Notably, our approach demonstrates substantial gains in both denoising quality and computational efficiency. Furthermore, manipulating the intensity of the noise context maps at inference time enables zero‑shot conditional denoising, allowing for dynamic control over the output image texture. The entire pipeline is available at https://github.com/minwoo‑yu/ENCORE.git

Authors:Tianyi Fu, Mohan Sridharan
Title: Hierarchical Compositionality for An Assistive AI Agent
Abstract:
AI agents are increasingly being developed to assist humans in various applications, and Large Language Models and other deep network architectures are considered to be state of the art for such agents. These methods are impressive stochastic predictors, but they are resource‑hungry, opaque, and known to make arbitrary decisions in novel situations due to the narrow set of underlying representation and processing choices. Our work seeks to explore the design of architectures for such AI agents based on core principles that can be traced back to the early pioneers of AI but are not fully utilized in modern AI methods. We do so in this paper in the context of the core problem of AI agents addressing ambiguity in the objects being referred to by the human participants. Humans address such ambiguity by heuristically leveraging compositional knowledge of domain context and the preferences of the other human participants. Drawing inspiration from this observation, we describe an architecture that embeds the principle of hierarchical compositionality and uses simple heuristics to achieve the desired disambiguation. Specifically, domain objects are represented in terms of primitive attributes drawn from human‑validated semantic feature norms, and a hierarchical combination of attributes and concepts automatically identified from a limited observed history of interactions of an assistive agent with specific users. The assistive agent then achieves the desired disambiguation by reasoning with knowledge of this compositional hierarchy; axioms governing domain dynamics; and models of semantic compatibility, session salience, and user‑specific thematic preference, requesting human clarification when necessary. Experiments show that our approach consistently outperforms state of the art data‑driven baselines, supporting adaptation to specific user profiles.

Authors:Zijian Gu, Weikai Lin, Shuang Zhou, Zihan Chen, Song Wang
Title: UniMod: Enhancing Multi-Modal Medical Diagnosis through Cross-Modality and Within-Modality Alignment
Abstract:
Multi‑modal learning combining medical images and clinical text is promising for disease diagnosis. However, standard multi‑modal training leads to shortcut learning: models exploit the easier modality (e.g., diagnostic cues in text) while neglecting harder‑to‑learn features (e.g., subtle visual patterns). We propose UniMod, a framework that mitigates shortcut learning by requiring each modality to predict the diagnosis on its own. It supervises image‑only, text‑only, and multi‑modal classification simultaneously, so each modality must extract diagnostic features. We add cross‑modality alignment for knowledge transfer and within‑modality supervised contrastive alignment over same‑diagnosis patients. On Harvard‑Glaucoma, UniMod reaches 0.850 AUC, outperforming OGM‑GE and Gradient Blending by 1.6‑1.8%; on CheXpert Plus, it reaches 0.966 AUC, surpassing them by over 5%. UniMod also extends to 5‑class multi‑label diagnosis without architectural change, improving mean AUC by 0.097 over CGGM.

Authors:Alvin Spivey, Yu Huang
Title: Logit-Boundary Geometric Belief Interfaces and Sparse Sheaf-Enclave Protocols: A Self-Contained Substrate for Secure Network Electronic Health Record (EHR) Interoperability
Abstract:
Electronic health‑record interoperability is a boundary problem: legacy systems, generative models, terminology services, identity systems, and human reviewers may each expose rich internal states, while operational exchange requires a narrow shared interface of typed claims, bounded uncertainty, provenance, and explicit admission or abstention. This paper details a mathematical and engineering architecture for that interface. The organizing idea is the logit boundary: a discovery model may propose pre‑threshold scores over a local categorical decision, but a deterministic judgment substrate decides whether the proposal is admissible, requires review, or must be quarantined before any Fast Healthcare Interoperability Resources (FHIR) transaction is constructed. The resulting Geometric Belief Interface (GBI) combines finite boundary semantics, local Dirichlet evidence, cellular‑sheaf and mapping‑cone diagnostics, advisory geometric audit charts, and a Decentralized Cryptographic Sheaf‑Enclave (DCSE) protocol sketch for fail‑closed deployment. The framework does not establish clinical truth, global representation alignment, or end‑to‑end safety; it defines certificate‑producing checks at a model‑to‑system boundary. A companion frozen synthetic benchmark, GBI BoundaryBench v0.1, evaluated Qwen3‑4B‑Instruct‑2507 on 256 held‑out tasks across three evidence modes (768 canonical executions). All executions completed, but none produced an output accepted by the benchmark contract: 369 were rejected during safe parsing and 399 during schema validation, yielding zero coverage and deterministic quarantine. This empirical result is deliberately narrow ‑ one 4B open‑weight model under one frozen interface ‑ and is reported as evidence about the admission boundary, not as a general claim about LLM capability or clinical safety. A Julia appendix verifies numerical certificates using standard libraries.

Authors:Saman Rahbar
Title: Frozen Brain-MRI Foundation Models Are Site Fingerprints
Abstract:
Frozen foundation‑model (FM) embeddings are increasingly used as off‑the‑shelf brain‑MRI representations, on the assumption that they capture anatomy. We audit what they actually encode and find that acquisition site is a large, intrinsic component of the representation. Across two independent cohorts (ABIDE‑I, ABIDE‑II), three frozen 3‑D encoders (brain‑pretrained, CT‑pretrained, and randomly initialized), and every network depth, site is linearly decodable at roughly 0.9 balanced accuracy at deep layers, exceeding the decodability of every clinical or demographic variable (sex, age, autism diagnosis) at every layer. The effect is intrinsic rather than learned: a randomly initialized encoder is already a ~0.9 site classifier on both cohorts and across three architecture families (Swin, ViT, ResNet), and site is decodable at ~0.95 directly from the raw downsampled image with no encoder, so the fingerprint reflects low‑level image statistics that any encoder preserves rather than a product of pretraining. Residualizing measured population covariates leaves site decodability essentially unchanged, indicating an acquisition‑ rather than population‑driven effect. A nonlinear probe matches the linear one, so the fingerprint is fully linearly accessible. The site subspace is removable post hoc by iterative null‑space projection or ComBat (site decodability 0.94 ‑> 0.07/0.00), and is a site‑attribution concern for shared or federated embeddings; but for dense segmentation this removal is not free, because site and anatomy occupy an entangled linear subspace (a matched‑rank random‑direction projection is Dice‑neutral, whereas removing the site subspace is destructive). We recommend site‑audited use of frozen brain‑MRI FMs and release an open audit toolkit.

Authors:Lisa K. Fischer, Mykhailo Riabets, Daniel Rueckert, Benedikt Wiestler, Anke Meyer-Baese, Sandeep Nagar
Title: MRIComp4Flow: Compression of 3D Brain MRI for Training Multi-Modal Generative Models
Abstract:
Large‑scale multi‑modal MRI datasets impose substantial storage and I/O costs, limiting the training of 3D generative models on commodity infrastructure. While lossy compression is known to preserve accuracy for discriminative segmentation networks, its effect on generative models, which must learn the full data distribution rather than a decision boundary, is unexplored. We study whether standard image codecs can effectively compress semantically rich brain tumor MRI while preserving the fidelity required to train and deploy a 3D MRI generative model. Each 3D volume is compressed with JPEG2000 or a near‑lossless JPEG‑LS pipeline. Next, a Wavelet Flow Matching model, conditioned on BraTS image sequences (T1n, T1c, T2, T2f), is trained on compressed data, and the resulting models are evaluated on the validation set. At a 20:1 compression ratio, synthesis quality is statistically equivalent to a model trained on uncompressed data within a pre‑specified margin (ΔPSNR <1,dB, ΔSSIM <0.02; paired TOST p=[[p]]): mean PSNR is 27.3,dB vs. 27.0,dB and mean SSIM is 0.95 vs. 0.96 across modalities. Our results indicate that JPEG2000 compression is a practical step toward scalable 3D MRI generative modeling without degrading synthesis quality. The codebase is available at https://github.com/lisafis/MRIComp4Flow .

Authors:Francisco León Zúñiga Bolívar
Title: Not a Monolith: Lab-Level Divergence in the Cooperative Equilibria of Chinese Frontier LLM Agents
Abstract:
Does the cooperative bias documented for Western frontier LLM agents extend to a different alignment lineage, and should the Chinese models that embody it be treated as a single bloc or as distinct laboratories? We study four frontier‑tier Chinese models ‑ DeepSeek V4 Pro, Qwen3‑Max, Kimi K2.5 and GLM‑5.1 ‑ in an evolutionary Iterated Prisoner's Dilemma, under a design that removes a confound present in prior work. Rather than letting each model convert its own natural‑language strategies into code, which entangles strategic disposition with coding ability, we hold the converter fixed (GPT‑5.4 Mini) across all labs, so every cross‑lab comparison is a comparison of generation alone. We run the full protocol: all‑play‑all tournaments and a Moran process at n=500 runs per condition, across three prompt styles and four population regimes. Two pre‑registered hypotheses are evaluated. H6 (not monolithic) is supported: the four labs differ significantly in aggressive‑equilibrium proportion, P_A running from 1% for Qwen3‑Max to 9% for DeepSeek V4 Pro, with four of six pairwise comparisons surviving Holm‑Bonferroni. The spread across the four labs (P_A range 8pp) is larger than the difference between the Chinese and Western ecosystems' mean P_A (5.0% vs 5.0%): on this measure, within‑ecosystem variation exceeds the East‑West gap. H5 (cooperative‑bias generality) is consistent but qualified: a cooperative plurality holds in 6 of 12 lab‑prompt combinations against the 9 of 12 reported for Western models, a difference we do not treat as firm, since the count rests on Cooperative‑Neutral near‑ties and rises to 9/12 under an alternate converter in our pre‑registered robustness check. The lab, not the ecosystem, is the unit at which cooperative disposition is set; treating "Chinese models" as a monolith is not supported by the evidence.

Authors:Guanqun Yang, Tong Qi, Xiaoxue Han
Title: DualSpectralCF: Training-Free Sign-Aware Spectral Collaborative Filtering
Abstract:
Real‑world recommendation platforms routinely collect explicit negative feedback such as 1‑star reviews, hate‑button clicks, distrust between users, and very‑low watch‑ratio videos. Learned sign‑aware recommenders exploit this signal for clear accuracy gains, but only at the cost of gradient‑based training. In parallel, a line of training‑free spectral collaborative filtering methods matches or beats learned graph recommenders at a fraction of the cost, yet operates on positive interactions alone. We bridge these two lines with DualSpectralCF, a training‑free framework of two components that attach to any spectral backbone of the form \hat\mathbfr_u = F(\mathbfM) \mathbfr_u: a signed input signal \mathbfr_u^\pm that encodes the user's explicit dislikes, and a signed item‑item operator \mathbfM^\pm that blends like‑together and dislike‑together similarity. The framework is backbone‑agnostic and adds just two scalar hyperparameters. We instantiate DualSpectralCF on ChebyCF, GF‑CF, and Turbo‑CF, and evaluate on five sign‑aware benchmarks: every instance matches or beats its unsigned backbone on all 5 datasets, with Recall@20 lifts up to +32.6% with backbone‑specific (γ, κ) tuning and +1.9% to +16.0% for DualSpectralCF‑Cheby at the fixed default (γ= ‑0.5, κ= 0.1), and the family runs 7.7 to 155.3× faster than SIGformer while reaching 70.7% to 90.7% of its accuracy. Sign‑awareness helps most for cold‑start users, with up to +29.2% Recall@20 on Epinions users with 1 to 5 training items.

Authors:Guanqun Yang, Wenlong Zhang
Title: Sequential Modality Dropout for Robust Multi-Modal Sequential Recommendation
Abstract:
Multi‑modal sequential recommenders assume every item carries every modality, but real product catalogs often miss images or text, and a model trained on complete data loses much of its recommendation accuracy when a modality is unavailable at serving time. We propose Sequential Modality Dropout (SMD): during training, each modality stream (image and text) is independently erased with probability p for an entire user interaction history, so the model learns to predict the next item without relying on any single modality. We measure robustness by retention, the fraction of a model's full‑modality accuracy (HR@10) that survives when a modality is removed at test time. Across four backbones (MM‑SASRec, IISAN, MISSRec, and fMRLRec) on four Amazon domains, SMD raises text retention by 1.0 to 3.2x at essentially no cost to full‑modality accuracy; under an extreme 95% per‑item missing rate, it retains 61% of HR@10 versus 22% without (a 2.8x improvement). An optional cross‑modal reconstruction loss further lifts retention from 90% to 98% on a simple additive backbone under severe text missingness. SMD is a four‑line, architecture‑agnostic change that makes multi‑modal sequential recommenders robust to the missing modalities they actually encounter in deployment.

Authors:Scott E. Frias
Title: Similarity Gates Approve Reversals: A Validity Audit of Embedding-Cosine Thresholds in Agent Systems
Abstract:
Agent frameworks ship quality gates that compare text blocks by embedding‑cosine similarity and decide at a fixed cutoff. Deduplication filters, semantic caches, drift guards, and answer grader gates deploy to answer the question: "Does this text still mean the same thing?" But the score answers a different question: "How much did the wording change?" We audit this gate class as a measurement instrument. In the cases these gates exist to catch, the two can run in opposite ways. Many times, reversing an instruction is a single word edit, while agreement often rephrases a sentence. The consequence is a safety check that fires backwards. The production drift guard we audited caught 0 of 56 meaning‑breaking mutations, and one approved item, "withhold the study drug" ‑> "administer the study drug", came in at cosine 0.9608. We observed five shipped operating points, and balanced accuracy across 90 configuration‑threshold‑task cells never exceeded 0.700 (median 0.525). The same confounder also corrupted evaluations. A naively built corpus inherits this confounder and can return an inverted verdict, with a decision AUROC exactly 0.000 in 13 of 18 configuration‑task cells (at most 0.040 in all 18) against 0.440‑0.815 for the same nine configurations under a balanced 2x2 design. Twice in the effort it captured our own headline claims. Obvious repairs fail: an encoder swap and an overlap‑conditioned gate (0.750 in‑sample, 0.533 held‑out) land at chance on separately authored held‑out data, and an NLI drop‑in did no better. Embeddings do still bear hope here, as the strongest two of nine configurations separated reversal from paraphrase at matched overlap (AUROC 0.79‑0.90), but only a matched‑pair audit reveals the deployment regime. We release the corpus method, harness, and frozen results, and contend that scores gated this way measure the wrong thing. We believe a valid instrument is buildable.

Authors:Alec Harris, Kasey Corra, Archie Chaudhury, Yixiong Hao
Title: Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes
Abstract:
Feedback signals used to train Large Language Models (LLMs) are the primary driver of their behavior and our main lever for instilling alignment with human values and objectives. However, a key limitation of current post‑training methods is the inability of human annotators and automated reward functions to faithfully capture the feedback we would like to give. We introduce Evaluation‑Conditioned Training (ECT), a post‑training framework that uses natural language to condition each training sample on the fidelity of the feedback we provide and then elicits the desired behavior by conditioning the LLM on a high‑fidelity monitor in deployment. ECT is aimed at improving performance under imperfect feedback and works as an add‑on to existing algorithms such as SFT and PPO. We first provide a conceptual framework for ECT and discuss its potential to address persistent sources of reward mis‑specification. Then we motivate ECT in the context of the eliciting latent knowledge (ELK) problem. Finally, we evaluate ECT on two proof‑of‑concept experiments: increasing even‑handedness in news article generation and reducing sycophancy on an arithmetic task. In each setting, we utilize imperfect feedback, rewarding bias and agreement with the user, respectively. In both settings, ECT improves the targeted behavior relative to direct training.

Authors:Srinivas Telukunta, Georgios Nektarios Lilis, Lucio Baron
Title: The CASE Framework: A Multi-Disciplinary Control Architecture for Governing Enterprise Agentic AI
Abstract:
Enterprises are deploying autonomous AI agents faster than they can govern them, and prevailing approaches stretch a single discipline, typically DevSecOps built for deterministic automation, across every scale of agency. We argue that agentic AI governance is four problems, not one, each with a mature governing science. The CASE framework assigns Control theory to the individual agent (intent as setpoint, guardrails as feedback, evaluation as observation), complex Adaptive systems theory to agent collectives (where emergence makes single‑agent assurance non‑compositional), Supervisory cybernetics to human‑agent teams (where the Law of Requisite Variety shows unaided human oversight fails structurally), and Engineering operations to fleets (extending error budgets to decision quality so autonomy becomes a controlled variable). We formalize each layer, derive cross‑layer coupling conditions, including a zero‑touch deployment paradox where excellence at one‑layer strains the others, and trace twenty‑plus enterprise controls to their classical constructs. Three empirical studies validate the thesis: 82 percent of documented production agent failures are multi‑layer trajectories; none of 22 ecosystem tools offers full Layer 2 (emergence) coverage; and all 35 scored public deployments fall in the lowest maturity band. We name this mismatch, risk realized at the emergence layer against capability barely offered and practice absent, the Emergence Gap. A five‑level maturity model with a non‑compensatory bottleneck‑weighted index and assessment instrument operationalizes CASE as a scientific rather than process maturity model, grounded in production enterprise agentic platforms. As EU AI Act Article 14 makes effective human oversight a legal requirement, only architectures satisfying requisite variety can make oversight real rather than ceremonial.

Authors:Joyjeet Singh
Title: The Evaluation Protocol Determines the Result: An Independent Reproduction of LeWorldModel on TwoRoom
Abstract:
LeWorldModel trains a latent world model with a prediction loss and a single anti‑collapse regulariser, and reports approximately 87% of goals reached on TwoRoom, its simplest diagnostic environment. We reproduce that result by independent reimplementation on roughly 25 of rented compute, with all evaluation on one laptop CPU. We reach 94.0% at the repository's evaluation goal offset, against 84.0% for the authors' own released checkpoint measured under our protocol on identical episodes, and we reproduce the reported representation result directly (position probe Pearson r = 0.9988 against a reported 0.996). Reaching that point required four conventions that determine the outcome and appear in no released configuration file: dense action gathering across a frameskip block, a programmatically‑set action‑encoder width, ImageNet pixel normalisation, and action z‑scoring. A reproducer following the released configurations alone obtains a model whose predictor cannot converge. The evaluation protocol is itself contested by the released material. The paper's appendix and the repository's configuration specify different goal offsets and step budgets; on the authors' own weights these yield 14.0% and 84.0%, and only the configuration's values reproduce the reported figure. On fifty identical episodes, changing nothing but how the goal is constructed moves that checkpoint from 84.0% to 8.0%. Two findings generalise. One‑step prediction accuracy does not predict long‑horizon planning success: across three checkpoints spanning a sevenfold range in prediction error, including the authors' own, it orders short‑horizon success monotonically and fails to order long‑horizon success at all. And a batch normalisation layer inflated our reported validation loss by up to a factor of 300, concealing a training loss that was flat throughout.

Authors:Ghazal Kalhor, Zahra Jafari, Amirarsalan Shahbazi, Behnam Bahrak
Title: PERCEPT: A Corpus for POS Tagging and Analysis of Persian-English Code-Mixing
Abstract:
Social media has become a major venue for multilingual communication, where users frequently mix multiple languages within a single utterance. Although code‑mixed corpora have been developed for several language pairs, Persian‑English code‑mixing remains relatively underexplored. Existing Persian resources lack Universal Dependencies (UD) part‑of‑speech (POS) annotations for code‑mixed words, limiting both linguistic analyses and the development of syntax‑aware NLP models. To address this gap, we introduce PERCEPT, the first publicly available large‑scale Persian‑English code‑mixed corpus annotated with Universal Dependencies POS tags for code‑mixed words. The dataset comprises 6,800 posts collected from X, Instagram, and Digikala. We further present an LLM‑assisted annotation framework that automatically assigns POS tags and document‑level topics. Human evaluation demonstrates high agreement between the automatically generated annotations and gold annotations, confirming the reliability of the annotations. Using PERCEPT, we conduct the first comprehensive linguistic analysis of Persian‑English code‑mixing across multiple social media platforms. Our analyses reveal that nouns are the predominant category for code‑mixed words, while the distributions of other POS categories vary across platforms. We further find that the positional distribution of code‑mixed words is remarkably consistent across platforms, whereas the triggering effect is substantially more pronounced in Digikala. PERCEPT is publicly available at https://github.com/kalhorghazal/PERCEPT.

Authors:Matt J. Borowski, Blazej Osinski
Title: Hand-Written PTX Tensor-Core GEMM Kernels: A Multi-Precision Study on NVIDIA L4
Abstract:
High‑performance Tensor Core kernels rely on a low‑level PTX pipeline built from asynchronous data movement with cp.async, warp‑level matrix loads with ldmatrix, and matrix multiply‑accumulate operations with mma.sync. However, most application code accesses Tensor Cores indirectly through the WMMA C++ API. This paper asks a focused, practical question: when does replacing WMMA with hand‑written PTX actually pay off? To answer this question, we conduct a controlled, single‑GPU study on an NVIDIA L4 GPU (Ada, SM89), comparing double‑buffered WMMA baselines with a family of hand‑written PTX GEMM kernels across FP16, INT8, and INT4 arithmetic and square problem sizes from N=512 to N=8192. Every kernel is profiled with Nsight Compute across the full metric set, and PTX speedups are reported relative to the corresponding same‑precision WMMA baseline. Hand‑written PTX provides no end‑to‑end speedup for FP16, because its instruction‑level gains are offset by operand‑packing overhead. In contrast, the PTX kernels achieve consistent speedups of 1.4x‑1.8x for INT8, driven primarily by lower instruction counts and better global‑memory coalescing, and 2.9x‑4.3x for INT4, where native mma.sync.m16n8k64.s4 execution avoids the software‑emulated sequence used by the WMMA path. Relative to the FP16 WMMA baseline, the best quantized kernels reach 34.4x (INT8) and 98.7x (INT4) at N=8192. Across these experiments, occupancy is a poor predictor of throughput. For large matrices, performance instead tracks memory‑system behavior ‑‑ particularly global‑load coalescing and DRAM‑active cycles ‑‑ more closely than Tensor Core utilization. These results identify the precisions and operating regimes in which the additional complexity of hand‑written PTX is justified.

Authors:Yuning Peng, Haiping Wang, Yuan Liu, Yipeng Lu, Zhen Dong, Bisheng Yang
Title: LEGO: Leveled Language Gaussian Splatting
Abstract:
We introduce LEGO for advanced open‑vocabulary scene understanding. Beyond basic concept recognition, its core innovation lies in capturing the intrinsic semantic hierarchies within the scene, such as the "flowerpot ‑> bouquet ‑> bud ‑> petal" lineage. While foundation models like SAM can identify multi‑granular structures in 2D, their partitions are strictly perspective‑bound and lack cross‑view consensus. LEGO self‑adaptively re‑grades volatile multi‑view SAM granularities into a unified, 3D‑consistent hierarchy. This provides precise supervision for the structurally coherent, multi‑level segmentation of 3D scenes. By grounding these segments with CLIP embeddings, LEGO recovers open‑vocabulary semantic logic across hierarchical levels. Furthermore, by incorporating spatial relationships, we elevate these segments into level‑wise language scene graphs, effectively empowering Large Language Models to perform complex, context‑aware spatial reasoning and precise visual grounding. Experimental results demonstrate that LEGO establishes new state‑of‑the‑art performance across both promptable and open‑vocabulary 3D segmentation benchmarks, exhibiting advanced hierarchical scene decomposition and context‑aware spatial reasoning.

Authors:Shiting Gong, Jianpeng Yao, Jinfeng Wang, Marco Pavone, Jiachen Li
Title: Navigating the Proximity-Safety Balance: Constraint Decomposition for Human Following in Pedestrian Crowds
Abstract:
Following a target human in crowded environments involves an inherent conflict between staying close to the target and navigating safely among surrounding pedestrians and obstacles. This conflict becomes more severe in dense scenarios, where aggressive following risks collisions and conservative margins lead to target loss, especially when pedestrian behaviors are unfamiliar or unpredictable. Existing reinforcement learning (RL) methods typically encode these competing objectives into a single dense reward, but the resulting proximity‑safety balance is implicit and difficult to adjust across conditions. To address this, we decompose the human‑following task into a sparse task reward and independent cost constraints within a multi‑constraint RL formulation, where each constraint is managed through cost thresholds with direct behavioral meaning rather than implicit reward weight ratios, allowing explicit and tunable control over the trade‑off. We further quantify the prediction uncertainty of human motions and integrate these estimates into the RL costs to enhance safety under unpredictable conditions. Extensive experiments across both in‑distribution and out‑of‑distribution settings demonstrate that our method achieves an effective proximity‑safety balance compared to baselines. Real‑robot deployment further validates the feasibility of our method in real‑world scenarios. More details are available on our project page: https://nav‑ps‑balance.github.io/.

Authors:Kaustubh Shivshankar Shejole, Tanish Agarwal, Arpit Agarwal, Avishek Ghosh
Title: Finding the Signal in the Spam: Jointly Learning Rewards and Worker Reliability from Pairwise Comparisons
Abstract:
The problem of learning from pairwise comparisons has been widely studied across many domains such as recommendation systems, social choice, and more recently, fine‑tuning large language models. In this problem, the goal is to learn item rewards based on pairwise comparisons between them. In many scenarios, these comparisons are elicited from crowdworkers using platforms such as Amazon Mechanical Turk, Scale AI, etc. However, crowdworkers are often unreliable due to limited domain knowledge or revenue‑maximizing (spamming) behavior. In this work, our goal is to understand whether worker reliability (competency) can be learned jointly with item rewards. To this end, we adopt the Boltzmann‑rational model for pairwise comparisons, which extends the Bradley‑Terry‑Luce model by incorporating worker competencies. We derive an EM‑based algorithm for learning under this model by introducing Polya‑Gamma latent variables to transform the logistic likelihood into a conditionally Gaussian form, enabling tractable optimization and leading to a simplified Q function in the E‑step of the algorithm. This technique allows us to reduce our formulation to a matrix sensing problem, using which we establish theoretical convergence guarantees for our algorithm. We conduct extensive experiments on real‑world and synthetic datasets. These experiments demonstrate the advantages of using our algorithm over several baselines and confirm its strong robustness to both spammers and adversarial workers, highlighting its practical effectiveness in realistic crowdsourcing and reward learning settings. The code and data is publicly available at https://github.com/KaustubhShejole/BoRa_EM.

Authors:M. Sajid, A. Quadir, A. Rahaman, P. N. Suganthan, M. Tanveer
Title: Uncertainty-Aware Ensemble Deep Randomized Neural Networks for Classification
Abstract:
The current state‑of‑the‑art (SOTA) deep randomized neural networks, such as deep Random Vector Functional Link (dRVFL) and ensemble deep RVFL (edRVFL), treat all training samples uniformly, which limits their robustness and effectiveness when applied to real‑world datasets containing noise and outliers. Furthermore, the propagation of contaminated features across hidden layers negatively influences the decision‑making capability of these models. To overcome these limitations, we propose intuitionistic fuzzy dRVFL (IF‑dRVFL) and intuitionistic fuzzy edRVFL (IF‑edRVFL) frameworks that enhance model robustness. The proposed models unify intuitionistic fuzzy theory to exploit sample neighborhood information in the kernel space by jointly considering membership and non‑membership degrees for each sample. Membership degrees are computed based on the distance of samples from their respective class centroids, while non‑membership degrees quantify sample heterogeneity within local neighborhoods. These measures are employed to assign adaptive weights to training samples, enabling effective discrimination among clean, noisy, and outlier data points. Extensive experiments conducted on UCI and KEEL benchmark datasets, with and without the presence of Gaussian noise, demonstrate the superiority of the proposed IF‑dRVFL and IF‑edRVFL models over existing SOTA fuzzy and non‑fuzzy approaches. The source code is available at https://github.com/mtanveer1/IF‑edRVFL.

Authors:Niclas Claßen, Théo Sourget, Dovile Juodelyte, Rob van der Goot, Veronika Cheplygina
Title: Robustness of transferability estimation metrics for medical imaging
Abstract:
In transfer learning, the choice of source model largely influences the performance on a target dataset. Still, selecting a fitting source remains a challenging task, especially in medical imaging where one has to decide between models pre‑trained on off‑the‑shelf options, such as ImageNet, and domain specific datasets. Transferability estimation (TE) metrics address this problem by aiming to predict the best performing source model in a computationally cost effective way. However, previous work has reported conflicting TE metric performances due to differences in experimental setups. Moreover, most TE metrics are designed for and evaluated on natural images, while being optimized for accuracy, whereas in medical imaging metrics that are more robust to class imbalance are typically used. We study the impact of varying the target dataset as an isolated factor, by constructing miniature populations of different sample sizes and random seeds. In addition, we investigate the influence of the evaluation metric used to obtain the reference ranking. We find that small modifications to the target dataset change the rankings. Furthermore, we show that the choice of evaluation metric affects the reference rankings and therefore the evaluation of TE metrics. Overall, we observe a low agreement between rankings from TE metrics and reference. The code, model checkpoints and data splits used in this work are available through https://github.com/niclasclassen/robustness‑of‑transferability‑estimation‑metrics‑for‑medical‑imaging.

Authors:Fidel Omar Tito Cruz, Neda Ghafouri, Zengyan Wang, Pegah Khosravi, Yu Tian, Chen Chen
Title: Longitudinal 3D Foundation Modeling for Neoadjuvant Breast Cancer Response Prediction from Serial DCE-MRI
Abstract:
Pathologic complete response (pCR) is an important endpoint in neoadjuvant chemotherapy (NAC) for breast cancer, and predicting pCR from imaging during treatment could support treatment response assessment. Many existing imaging‑based approaches rely on a single static timepoint, which fails to capture changes that occur during treatment. In this work, we present a longitudinal framework that combines a frozen 3D foundation encoder (Pillar‑0) with our Temporal Dynamics Network (TDN) to predict treatment response from serial Dynamic Contrast‑Enhanced (DCE) MRI acquired across four clinical timepoints from pre‑treatment to pre‑surgery. The TDN combines time‑aware volumetric embeddings with clinical and treatment data to predict pCR. Evaluated on 982 patients from the combined I‑SPY2 and ACRIN‑6698 cohort, the proposed model achieves strong performance across all reported metrics when longitudinal 3D imaging is fused with clinical data (test AUROC: 73.6%, balanced accuracy: 69.1%). While clinical variables provide the strongest individual predictive signal, longitudinal 3D imaging contributes complementary information when fused with clinical data, improving pCR prediction. Our source code is available at: https://github.com/omarftt/longitudinal_temporal_pillar.

Authors:Jiaxing Guo
Title: Safe Observation Capacity for Opponent Exploitation under Showdown Censoring
Abstract:
In poker‑like games, folds hide private cards, so showdown data are missing not at random: per‑card estimates converge to behavior conditional on reveal, and shrinking confidence sets can lose coverage. A floor‑safe probe carries a line to showdown; sequence‑form flow then recovers censored fold mass on reveal‑certified histories. We price acquisition through safe observation capacity, the largest target reach attainable by a floor‑safe plan at a given value slack. Its frontier is concave and piecewise linear, with initial slope given by the floor's shadow price. When that reach converts fully to reveal and parent flow is non‑bottleneck, matching local bounds make the hands required for conditional‑probability half‑width \varepsilon inversely proportional, up to logarithms, to capacity, opponent continuation mass, and \varepsilon^2. Safe Active De‑censoring (SAD) combines public screening, an independent reveal batch, and robust deployment; a max‑min safe audit gives positive joint reveal rate to every coordinate in a finite library‑covered target set, including public‑null deviations. With 10^6 hands, SAD raises the river over‑fold certified gain from 0.485 to 0.692 and improves both gains over public‑only collection on all three deviations (Holm‑adjusted paired p\le0.012), while selecting no control target. On a fixed‑board public twin, a disjoint audit‑refit‑deploy loop detects all 30 simulation seeds and no control seed, certifying absolute value 0.655 (95% confidence‑interval half‑width 0.008). Every floor‑constrained probe and response passes a floor audit.

Authors:XPolicyLab Community, Tianxing Chen, Yue Chen, Tian Nian, Zijian Cai, Guangyu Chen, Wenwei Lin, Qiwei Liang, Zanxin Chen, Peicheng Xiang, Kailun Su, Zixuan Li, Junyuan Tang, Yan Qin, Qiangyu Chen, Shaolong Zhu, Xiang Li, Jiahao Zhang, Weijie Wan, Baijun Chen, Honghao Su, Kehe Ye, Shujia Liu, Kaixuan Wang, Haotian Liang, Yunze Liu, Mingleyang Li, Yuran Wang, Boyu Chen, Hongzhe Bi, Shuhe Huang, Hengkai Tan, Jisong Cai, Yao Mu, Jun Guo, Xiaofeng Wang, Zheng Zhu, Weijie Ke, Hengtao Li, Yuhang Tang, Xiaofan Li, Ganlin Yang, Zhangzheng Tu, Shuai Yang, Wenxuan Song, Pengxiang Ding, Kaidong Zhang, Yu Sun, Junliang Guo, Tong Zhang, Yixing Chen, Rongxu Cui, Zongzheng Zhang, Haoxiang Ma, Junhao Cai, Haoyu Zhang, Senqiao Yang, Jinhui Ye, Pengguang Chen, Shu Liu, Xiu Su, Wenhan Fang, Wenhao Li, Yichao Cao, Chengyao Wang, Qiang Chen, Ping Luo, Wenbo Ding
Title: XPolicyLab: A Unified Standard and Open Ecosystem for Robot Policy Evaluation and Deployment
Abstract:
Robot policy evaluation and deployment remain fragmented by model‑specific software dependencies, data representations, and runtime interfaces, so that connecting N policies to M evaluation environments requires O(NM) separate integrations. We present XPolicyLab, a unified standard and open ecosystem that reduces this cost to O(N+M). XPolicyLab specifies common observation, action, and trajectory schemas together with a minimal adapter interface for observation updates, action prediction, batched execution, and episode reset, while a dependency‑isolated client/server architecture separates policy inference from environment execution, so that each side retains its native software stack and may run locally or remotely. The ecosystem integrates 42 robot policies and standardizes their installation, debugging, serving, and evaluation workflows. Across these adapters, model‑specific code varies by an order of magnitude while the environment‑facing loop stays within a few lines of a fixed reference, confirming that the contract confines heterogeneity to the policy side. In a controlled study, conforming to the standard reduces the integration effort of a representative policy from over five hours to two hours, and packaged agent skills reduce it further to thirty minutes. The same adapters serve RoboTwin, RoboDojo simulation, and standardized real‑robot evaluation through one interface. XPolicyLab is released as shared infrastructure for reproducible policy comparison and standardized deployment across simulation and physical platforms. Project website: https://xpolicylab.github.io/.

Authors:Oussama Boussif, Mohammed Mahfoud, Younesse Kaddar, Moksh Jain, Sida Li, Damiano Fornasiere, Xiaoyin Chen, Yoshua Bengio, Esmeralda S. Whitammer
Title: Bayesian Symbolic Regression with Entropic Reinforcement Learning
Abstract:
Symbolic regression is the problem of finding an algebraic expression describing a stochastic dependence of a target variable on a set of inputs. Unlike forms of regression that fit parameters assuming a fixed model structure, symbolic regression is a search problem over the space of expressions, represented, for example, as abstract syntax trees using a library of operators. Symbolic regression is typically used in settings with limited, noisy data in the natural sciences. However, searching for a single best‑fitting expression fails to capture the epistemic uncertainty about the expression, which motivates a Bayesian perspective that enables uncertainty quantification and specification of natural priors to constrain the search space. In this work, we propose ERRLESS (Entropy‑Regularized Reinforcement Learning for Expression Structure Sampling), a scalable approach for sampling from the posterior distribution over expressions given data using maximum‑entropy reinforcement learning. ERRLESS learns a neural policy that constructs expressions sequentially by building up their abstract syntax trees. At convergence, the policy samples expressions from the posterior. At test time, expressions can be sampled by rollouts of this policy. We demonstrate that ERRLESS achieves competitive results on the Feynman benchmark while producing short and interpretable expressions. Additionally, we demonstrate that the mean of the posterior predictive approximated by ERRLESS achieves a high coefficient of determination (R^2) compared to an SMC baseline, highlighting the benefits of the Bayesian perspective in symbolic regression.

Authors:Louis De Oliveira, Anastasia Karpova, Georges Nader, Antoine Houdard, Pierre Mezieres, Damien Rioux-Lavoie, Romain Pacanowski
Title: A Hybrid Neural-Microfacet BRDF Model for Real-Time Rendering
Abstract:
Over the past decade, microfacet‑based BRDF models have formed the foundation of real‑time rendering pipelines. Despite their widespread use, they often fail to reproduce subtle appearance effects arising from complex light‑surface interactions, which have led to the emergence of specialized physics‑based models for specific optical phenomena (e.g., diffraction, iridescence, multilayers). Although more accurate, these models lose versatility and lack performance for real‑time rendering. Recently introduced, neural models have demonstrated their ability to approximate BRDF reference data coming from measurements, simulations, or even complex shading networks. However, most current neural models require relatively large networks, making them costly for real‑time rendering. In this paper, we introduce a hybrid model that combines a GGX‑type microfacet model and a neural model to leverage the best features of both representations. The neural component corrects the appearance approximated by the microfacet component, allowing much smaller network than in existing neural models. We show that, at identical memory cost, our model approximates measurements better than state‑of‑the‑art neural models for a low evaluation overhead compared to a microfacet‑based model. Furthermore, our hybrid model remains easily editable by artists and benefits from an important sampling scheme, making it attractive for both offline and real‑time rendering.

Authors:Karim Zaghw, Andrew Pashea, Marc Pritsch, Wouter Nuijten, Karl Friston, Lancelot Da Costa
Title: Renormalising Generative Models for Active Inference: Foundations, Derivations, and Verification
Abstract:
Active inference offers a unified framework for perception, learning, and action, but scaling discrete active‑inference models to rich spatial and temporal domains remains difficult. Renormalising generative models (RGMs) address this challenge by composing discrete generative models across spatial and temporal scales, coarse‑graining lower‑level states and paths into higher‑level causes for objects, events, and action. However, fully reproducing and adapting the framework remains difficult: the mathematical exposition is compact, and the reference implementations are deeply integrated within specialized software environments, leaving many algorithmic details implicit. This paper addresses these challenges by providing a self‑contained, derivation‑oriented account of RGMs together with an open, verified implementation. We explain how the hierarchy is built, how beliefs and actions are updated within it, and how information is passed between levels. Where the published equations and implementation differ in emphasis, we make those choices explicit and explain their modelling consequences. By clarifying the theory and separating it from its original implementation context, this work lowers practical barriers to entry and makes RGMs more transparent, auditable, and reproducible, providing a foundation for future quantitative evaluation and development on machine‑learning benchmarks.

Authors:Jun Huang, Meiyi Chen, Zijie Yue, Yuhang Xiao, Fang Li, Hanli Wang, Xiaowen Tong, Yi Guo, Miaojing Shi
Title: Bootstrapping Vision-Language Model for Hysteroscopic Surgical Scene Segmentation
Abstract:
Hysteroscopic surgical scene segmentation plays a pivotal role in understanding the hysteroscopic intraoperative environment as well as computer‑assisted intervention. However, this task presents unique challenges due to the high morphological similarity among different lesions and the presence of artifacts such as specular reflections, motion blur, and fluid occlusions in surgical videos. In this work, we propose the first vision‑language model (VLM)‑based hysteroscopic surgical scene segmentation method, which performs pixel‑wise localization for fifteen representative categories in hysteroscopic surgical scenes. Our VLM‑hyster has a segmentation backbone that utilizes the pretrained image encoder for robust visual feature extraction, coupled with a transformer‑based decoder for dense prediction. Moreover, we design category‑specific text prompts and incorporate a masked distillation branch to filter out visual features with low correlation to the text prompts, enabling the model to focus more effectively on category‑specific image regions and thereby enhancing segmentation performance. We collect a large multicentric hysteroscopic surgical scene dataset, containing 4,020 high‑resolution images with detailed mask annotations, for model training and evaluation. Experimental results demonstrate that VLM‑hyster substantially outperforms state‑of‑the‑art AI models. Furthermore, extensive assessments by gynecologists, as well as multicentre and prospective validations, demonstrate VLM‑hyster's robustness and generalizability. The results suggest that VLM‑hyster earns considerable potential in enabling AI‑assisted localization of surgical instruments and lesions in hysteroscopic surgeries. Code is available at https://github.com/viscom‑tongji/VLM‑hyster.

Authors:Zhengfeng Li, Lei Zhang, Xianwei Wu, Zhengqi Zhuang, Yingjie Xu, Boge Wang, Shaofei Zhu, Chuan Wang, Peng Zhao, Xinyu Zheng, Guoping Rong
Title: OpenCodeReview: Determinism over Non-Determinism for Cost-Effective Agent-Based Code Review
Abstract:
LLM‑based code review agents promise scalable, always‑on review, yet current systems suffer from two intertwined weaknesses: (1) non‑determinism‑‑unbounded tool use makes review outcomes unstable, and (2) context locality‑‑the reviewer's access remains bounded to the diff, capping discoverable issue depth. Both give rise to three challenges: misaligned context retrieval, a coherence‑efficiency trade‑off in multi‑file pull requests, and hallucinated comments that erode trust. To address these, we introduce OpenCodeReview, built on deterministic engineering for uncertain agents: rather than granting maximal freedom, we inject determinism at three deliberate pipeline points. Rule‑Guided Dispatch uses a multi‑layer rule system to deterministically select files and review criteria, eliminating variability of agent‑driven triage. Grounded File Review replaces free‑form exploration with a curated tool set exposed through a ReAct loop, while file‑level parallel SubAgents balance context coherence against efficiency and recover cross‑file dependencies on demand. Independent Reflection introduces a falsification‑first filter under an asymmetric information boundary‑‑the reflector sees only the diff, not the agent's tool‑augmented exploration‑‑removing hallucinated comments without self‑reinforcing bias, improving precision while preserving recall. On AACR‑Bench (200 real‑world PRs, 10 languages, 1,505 expert‑verified comments), OpenCodeReview outperforms mainstream coding agents (e.g., Claude Code and Codex) across six LLM backends, achieving up to 2.17x higher SEM‑F1 (25.10% vs. 11.57%) while consuming 5‑15x fewer tokens. We open‑source OpenCodeReview at https://github.com/alibaba/open‑code‑review.

Authors:Bohan Lin, Hejia Geng, Xinyi Xie, Heng Zhou, Qinghua Xing, Bo Liu, Chen Zhang, Yudong Zhang
Title: Emotion2Skill: Model-Internal Emotion Signals for Adaptive Skill Selection and Evolution
Abstract:
Skill‑based LLM agents select reusable procedures from an external library to solve complex tasks, yet their routing decisions rely entirely on text‑level signals such as task descriptions, verbal reflections, and experience‑derived rules, while the model's own internal representational state remains unobserved. Recent interpretability work has shown that LLMs maintain linear emotion representations that causally influence behavior; however, these representations have been exploited only for post‑hoc analysis or direct output steering, and have not been used to inform agent‑level decision‑making. We propose Emotion2Skill, a framework that extracts LLM‑internal emotion vectors and incorporates them into both skill selection and skill evolution. At each decision step, a 27‑dimensional emotion state is extracted from the residual stream and mapped to a confidence‑gated summary injected into the routing prompt. Beyond online selection, emotion trajectories are analyzed for abrupt internal‑state shifts to pinpoint problematic skill invocations, guiding targeted SOP rewriting that replaces the coarse binary outcome signal of prior methods. On WebShop and ALFWorld, Emotion2Skill with Qwen3‑8B improves over the Zero‑Shot baseline by +26.9% success rate and +25.5% average success respectively, outperforming all baselines on both benchmarks with consistent gains on Qwen3‑14B. Co‑activation analysis further reveals semantically coherent emotion‑‑skill pairings, confirming that the routing improvements reflect meaningful internal‑state signals rather than opaque statistical correlations. These results establish LLM‑internal emotion representations as an effective decision‑level signal for orchestrating agent skill systems, extending their utility beyond interpretability and output steering. The code is available at https://github.com/BoHan‑LIN04/Emotion2Skill.

Authors:Mingfeng Lin, Chengfei Cai, Lin Xu, Yuxiang Wei, Liang Han
Title: DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models
Abstract:
Flow‑matching models are now a mainstream method to image generation, but its adaptation to diverse downstream scenarios typically relies on post‑training, which may cause conflicts among task‑specific optimization objectives. Reinforcement learning enables direct optimization of task‑specific rewards beyond the original models, yet trajectory‑level optimization may incur high‑variance gradients and cross‑task interference. On‑policy distillation (OPD) offers dense and stable supervision on student rollouts, but conventional teacher matching remains imitation‑based. We propose DreOPD, a Degraded‑reference extrapolative OPD method for flow‑matching models that bridges these two paradigms. Our DreOPD converts implicit reward extrapolation into closed‑form velocity regression, enabling extrapolative post‑training with the stability of OPD. It further uses a mildly degraded reference to strengthen the teacher‑reference contrast, yielding a clearer extrapolation direction. Experiments on single‑ and multi‑teacher settings show that DreOPD outperforms OPD and multi‑task RL baselines in average performance, while surpassing specialized teachers on most metrics.

Authors:Tong Zhao, Mingkun Lei, Yucheng Han, Chi Zhang
Title: BAG: Budget-Aware Gating for Diffusion Caching
Abstract:
Diffusion caching is a lightweight strategy that accelerates Diffusion Transformers (DiTs) by reusing intermediate features across denoising steps, but existing paradigms face a fundamental trade‑off: online heuristics lack global budget awareness, whereas static schedules lack instance adaptivity and fail to flexibly adapt to varying runtime budget constraints. To bridge this gap, we present BAG (Budget‑Aware Gating), a novel caching policy that unifies global budget pacing with dynamic, instance‑adaptive feature reuse. Rather than relying on hand‑crafted rules, BAG employs a lightweight gating network that dynamically decides whether to execute a full computation or reuse cached features at each step by jointly conditioning on the budget state and local trajectory feedback. We train this policy via offline‑to‑online schedule distillation, transferring the decision‑making of offline‑searched schedules into a compact online gate. Extensive experiments on FLUX.1‑dev, Wan2.1, and Qwen‑Image‑2512 demonstrate that BAG consistently outperforms state‑of‑the‑art caching methods across various speedup tiers while remaining robust across different resolutions, seeds, and guidance scales. Code will be released.

Authors:David D. Yuan, Tony Z. Zhao, Kaylee Burns, Chelsea Finn
Title: SpeedTuning: Speeding Up Policy Execution with Lightweight Reinforcement Learning
Abstract:
While learned robotic policies hold promise for advancing generalizable manipulation, their practical deployment is often hindered by suboptimal execution speeds. Imitation learning policies are inherently limited by hardware constraints and the speed of the operator during data collection. In addition, there are no established methods for accelerating policies learned via imitation, and the empirical relationship between execution speed and task success remains underexplored. To address these issues, we introduce SpeedTuning, a reinforcement learning framework specifically designed to enhance the speed of manipulation policies. SpeedTuning learns to predict the optimal execution speed for actions, thereby complementing a base policy without necessitating additional data collection. We provide empirical evidence that SpeedTuning achieves substantial improvements in execution speed, exceeding 2.4x speed‑up, while preserving an adequate success rate compared to both the original task policy and straightforward speed‑up methods such as linear interpolation at a fixed speed. We evaluate our approach across a diverse set of dynamic and precise tasks, including pouring, throwing, and picking, demonstrating its effectiveness and robustness in enhancing real‑world robotic manipulation. Videos and code are available at https://daivdyuan.github.io/speed‑tuning/

Authors:Chenyang Li, Zejia Feng, Yuqin Huang, Yuxiao Ye, Huiyuan Xie
Title: LexKairos: Benchmarking Legal Temporal Capabilities in LLMs
Abstract:
Large language models (LLMs) have demonstrated strong performance across a wide range of legal tasks. In legal practice, time is a critical concept that governs the validity of statutes, the progression of legal cases, and the enforcement of procedural deadlines. However, legal temporal capabilities remain underexplored in existing legal AI benchmarks. To address this gap, we propose LexKairos, a comprehensive benchmark for evaluating the temporal capabilities of LLMs in the Chinese legal context across three dimensions: statutory temporal knowledge, case temporal modeling, and statute‑case temporal reasoning. LexKairos comprises nine sub‑tasks drawn from real‑world Chinese judicial cases and statutes. We conduct systematic evaluations of eight LLMs under multiple inference settings, including vanilla, Chain‑of‑Thought (CoT), and thinking modes. Our results show that Gemini‑3‑Flash achieves the strongest overall performance, yet even the best‑performing model exhibits notable limitations on tasks demanding precise time‑sensitive statutory metadata recall or complex reasoning in time limits, indicating that legal temporal knowledge and reasoning remain open challenges for current LLMs. Data and code are available at https://github.com/thunlp/LexKairos.

Authors:Shuyu Li, Kejun Zhang, Jiahe Lei, Shulei Ji, Zihao Wang, Jiaxing Yu, Wanying Wu, Lei Wang
Title: MusicLayout: Explicit Structural Planning for Controllable Text-to-Music Generation
Abstract:
Text‑to‑music generation has advanced rapidly, but current systems still rely primarily on global text prompts, leaving the structural organization of generated music implicit and difficult to inspect, control, or revise before audio generation. To address this issue, we introduce MusicLayout, an explicit intermediate representation for controlling musical structure in text‑to‑music generation. MusicLayout describes a musical piece as a time‑aligned layout of sections, textures, repetitions, variations, and instrument‑level arrangements, serving as an interpretable planning layer between textual intent and the generated music. We integrate MusicLayout into a text‑to‑music framework built on a unified autoregressive formulation, where the model first generates a MusicLayout representation and subsequently predicts audio tokens conditioned on this representation within a single sequence. The resulting MusicLayout can be inspected and modified prior to audio generation, providing a mechanism for layout‑level structural control. We evaluate MusicLayout through layout‑conditioned generation, layout manipulation experiments, and matched‑data ablations, providing evidence that explicit layout planning can improve long‑range structural organization and support layout‑level control. We have released the implementation as open source on GitHub at https://github.com/XaryLee/MusicLayout.

Authors:Francesco Ballerin, Erlend Grong
Title: Discovering PDEs equivariant under rigid motions
Abstract:
We consider the problem of PDE discovery from possibly noisy observations under the hypothesis that the underlying dynamic is symmetric in all rigid motions. Rather than using a generic library of derivative monomials, we leverage this assumption to construct libraries whose candidate terms are themselves rigid‑motion‑equivariant, and combine them with sparse regression to benchmark such libraries over five different equations. The advantage is most pronounced when the noise itself breaks rigid‑motion symmetry (e.g., radially or axially varying noise), and when the ambient spatial dimension increases, in which case they are also less resource intensive.

Authors:Ignacio M. Sticco
Title: Towards an LLM-based method for quantifying the sexual content in song lyrics
Abstract:
Reggaeton is one of the most widely consumed music genres in the world, and its lyrics are commonly regarded as highly sexualized. This claim rests mostly on qualitative studies and on small‑scale quantitative ones. This paper has two goals. First, we present a reproducible method that uses a large language model to quantify thematic content in song lyrics along several independent dimensions. The method is not restricted to sexual content. Second, we apply it to a corpus of 1,259 songs by 12 reggaeton artists released between 2002 and 2025. The analysis covers four topics: a dataset characterization, a per‑artist comparison, an analysis of how the dimensions change over time, and a comparison between our sexual‑explicitness score and Spotify's own explicit flag. We release the data collection code, the scoring prompt, and the corpus, so that other researchers can replicate the approach or apply it to their own lyrics datasets.

Authors:Mya Schroder, Yuna Hwang, Callie Y. Kim, Leqian Cheng, Jeffrey Li-cheng Liu, Chenchen Zheng, Xinning He, Bilge Mutlu
Title: SHRIMP: Iterative Refinement of Robot Task Plans
Abstract:
As collaborative robots have entered domains such as manufacturing, agriculture, and healthcare, programming or adapting robot behavior typically requires robotic expertise that most end users lack. Natural language lowers this barrier. Recent advancements in large language models (LLMs) have made it feasible to translate natural language into robot task plans. However, language‑based task specification suffers from semantic ambiguity, and generative models lack transparency for how language instructions become robot actions, making it difficult for users to validate the plan before execution. To address these issues, we introduce SHRIMP, a system that allows users to automatically generate a hierarchical robot primitive plan using natural language and iteratively revise their plan through re‑prompting and explicit correction. At each revision, SHRIMP allows users to validate their plan in simulation, and once satisfied, execute it on the physical robot. Through a user study involving participants planning tabletop kitchen tasks (n=35), we validate that SHRIMP improves perceived control and enhances robot transparency. System videos and source code are available at https://wisc‑hci.github.io/SHRIMP.

Authors:Sander Land, Clara Meister
Title: Explicit Boundary Markers for Subword Vocabularies
Abstract:
Subword tokenizers represent many common words twice in space‑using writing systems, once with a leading space and once without. The two entries have separate embeddings in models, so occurrences of one word are divided across rows that are trained independently, and the two forms need not even segment the string the same way: " together" may be a single entry while the same word without a preceding space is tokenized as "to|gether". Capitalization divides a word further, into as many as six forms. We introduce an alternative to standard whitespace conventions using an explicit word boundary marker, which prevents such duplication. Words are delimited by the boundary markers, and spaces between words are represented as pairs of such markers. Two shift codes do the same for title case and upper case, allowing one internal representation of a word to be re‑used across different settings. Switching to this convention mitigates the duplicate‑entry issue, but does not improve tokenization compression: for both vocabulary‑learning algorithms, the best marker scheme stays within one percent of the baseline in characters per token, averaged across six languages. It does result in better language modeling performance. Every marker scheme tested downstream reaches lower bits per byte than the baseline, suggesting that duplication carries a cost that compression does not capture.

Authors:Subinay Adhikary, Upal Bhattacharya, Vivek Kumar Singh, Anurag Sharma, Shubham Kumar Nigam, Suvasis Das, Shouvik Kumar Guha, Koustav Rudra, Kripabandhu Ghosh
Title: PROSLEX: A Novel Dataset for Expert-Annotated Legal Statute Prediction for Indian Judiciary
Abstract:
Legal Statute Prediction (LSP) involves automatically identifying relevant legal statutes given factual descriptions in legal documents, typically framed as a multi‑label classification task within natural language processing and information retrieval research. While recent advances have begun incorporating Large Language Models (LLMs) for statute prediction, current approaches primarily focus on accuracy metrics without addressing the critical need for legal reasoning, a fundamental requirement in judicial contexts where decisions must be explainable and justifiable. To address this research gap, we present PROSLEX (PRediction Of Statutes and LEgal eXplanation), a comprehensive dataset comprising 1,623 expert‑annotated legal documents from the Indian context. Each document is paired with statute predictions and detailed explanations, totaling 7,450 explanations, capturing the underlying legal reasoning. Using this dataset, we systematically evaluate various prompting strategies, including zero‑shot, few‑shot, chain‑of‑thought, and tree‑of‑thoughts approaches, to generate both statute predictions and their corresponding legal rationales. Our evaluation framework measures not only predictive performance but also the coherence and legal validity of generated explanations, positioning PROSLEX as a benchmark for developing explainable AI systems that can support legal practitioners while advancing research in interpretable legal NLP. To ensure reproducibility, we have made our PROSLEX dataset and model code available on GitHub: https://github.com/subinay494/Legal_Statute_Prediction_Explanation.

Authors:Kenta Watanabe, Atsuyuki Miyai, Mizuki Takenawa, Kiyoharu Aizawa, Toshihiko Yamasaki
Title: 360CityArena: A Realistic Virtual Urban Navigation Benchmark for Embodied Agents
Abstract:
We present 360CityArena, a benchmark for evaluating the urban exploration capabilities of embodied agents within a photorealistic environment constructed from 360‑degree videos. Existing outdoor benchmarks either lack sufficient photorealism or complexity, resulting in a considerable gap from real‑world urban environments. 360CityArena is built on a realistic reconstruction of the Akihabara district in Tokyo, Japan, using 602 360‑degree video segments covering 85 streets, and consists of 175 meticulously human‑crafted tasks. It encompasses three task categories: Environment Understanding, Path Reasoning, and Spatial Reasoning, covering fundamental abilities required for urban exploration, such as localization, landmark search, path planning, and relational spatial reasoning, thereby enabling comprehensive evaluation in realistic urban scenes. Our evaluation using state‑of‑the‑art LMM‑based agents shows that even the strongest model, Gemini 2.5 Flash, performs far below human level (human: 77.3% vs. Gemini 2.5 Flash: 17.1%), revealing substantial challenges that remain in city‑scale embodied navigation and reasoning. 360CityArena provides a necessary and challenging testbed for photorealistic urban‑district navigation and spatial reasoning.

Authors:Xueping Gao
Title: Evidence-Calibrated Runtime Reconstruction for Agent Skills Across Heterogeneous Coding Agents
Abstract:
Agent Skills package reusable instructions and assets for tool‑using language‑model agents. Progressive loading creates failure boundaries poorly represented by session‑, model‑, or tool‑centric traces: a Skill can be discovered but not activated, activated without instructions, or appear successful without an independently verified outcome. We present Skill Runtime Intelligence, a passive runtime‑intelligence system that reconstructs supported Skill‑lifecycle stages across heterogeneous harnesses while preserving unsupported stages as unknown. Its Run Panorama separates immutable events, deterministic relations, inferred diagnoses, and controlled outcomes with four evidence grades; optional trace import and OTLP/HTTP export support existing observability deployments. Across six frozen repository profiles, three coding agents, and seven clean or fault‑injected conditions, all 126 executions preserve source worktrees and each correlates to exactly one source session. Yet adapters expose three distinct semantics: no Skill runs; complete runs but no failure‑like events; or failure‑like events in every operational‑failure and clean session. In a seven‑template diagnostic study, semantic aliases and Panorama localize the same six non‑clean boundaries but differ in exact/status behavior; both Raw views emit a failure status on all 18 clean cases, while Panorama emits none. A known‑rule graph conforms to 126/126 frozen contracts, whereas a second model completes only 228/378 calls. These observations motivate executable adapter qualification and show that event presence is not boundary fidelity, composite exact scores mask distinct errors, and model explanations must not overwrite deterministic facts.

Authors:Qingying Niu, Ruiyang Ren, Wayne Xin Zhao, Yaliang Li
Title: BOUND: Brief-Guided Corrective Preference Distillation at Search-Control Boundaries
Abstract:
Large language model (LLM)‑based deep search agents solve tasks through iterative retrieval and reasoning, but locally relevant evidence can cause persistent wrong‑anchor drift, constraint drift, or local‑topic drift. Existing methods supervise trajectories, outcomes, or steps, but rarely distinguish task‑aligned continuations from locally plausible ones that reinforce drift. We propose BOUND, a brief‑guided corrective preference distillation framework for persistent search drift. For each student‑induced decision‑time state, BOUND constructs a teacher‑side search‑state brief that preserves the original search target and key constraints while summarizing confirmed evidence, missing information, and drift status. Guided by the brief, the teacher determines whether the student's continuation contains a correctable local search‑control error likely to affect subsequent decisions. Together with the rollout outcome, this assessment determines whether to construct a corrective contrast between a student‑specific correction and the original continuation, or a termination contrast between a supported answer and an unnecessary retrieval continuation. Each validated state‑matched preference pair operationalizes a search‑control boundary. Direct preference optimization (DPO) distills these preferences into the student, while the brief and teacher‑side computation remain confined to training. We evaluate BOUND on four multi‑hop QA benchmarks and three deep‑search benchmarks. Across the six benchmarks for which we reran baselines, BOUND leads on five datasets and 12 of 14 metrics. Under the same search‑control interface and matched settings, BOUND outperforms Trajectory SFT by 5.6 EM points on Bamboogle and 4.8 accuracy points on BrowseComp‑Plus. Code is available at https://github.com/RUCAIBox/BOUND.

Authors:Yuemeng Xu, Zongxi Liu, Junyu Long, Yiming Huang, Jiarui Guo, Yangyujia Wang, Jiachen Xu, Dongyuan Yu, Zongwei Lv, Tong Yang
Title: InSituANN: Revisiting IVF for PCIe-Efficient Billion-Scale Vector Search
Abstract:
Approximate nearest neighbor search (ANNS) over billion‑scale vector datasets has become a foundational operator for modern retrieval systems, powering large‑scale recommendation, semantic search, and LLM/RAG workloads. Although GPUs offer massive parallelism and high‑bandwidth memory for batched vector search, their limited VRAM capacity makes fully GPU‑resident billion‑scale indexes difficult to deploy. In CPU‑GPU heterogeneous designs, keeping the base vectors in host memory avoids this capacity limit, but naively offloading fine search to the GPU introduces a new bottleneck: large volumes of base‑vector data must be streamed over PCIe. We present InSituANN, an IVF‑based ANNS engine that enables billion‑scale vector search on a single commodity GPU. InSituANN keeps original base vectors in host memory, performs fine search in situ, and uses the GPU for compact routing and optional pruning. As a result, query processing avoids PCIe transfers of high‑dimensional base vectors while retaining the simplicity of IVF. Beyond query performance, we further design an ultra‑fast IVF construction path for InSituANN. On SIFT‑1B, InSituANN builds the IVF index in 5.2 minutes, about 350x faster than the measured 30.4‑hour HNSW build. At matched recall on billion‑scale datasets, InSituANN improves end‑to‑end throughput by 104.9x‑4298.2x over the PCIe‑bound Rummy baseline and by 2.4x‑4.6x over DiskANN on SIFT‑1B and DEEP‑1B. Together with strong recall‑throughput trade‑offs and lower index space than graph‑based alternatives, these gains make billion‑scale retrieval practical on cost‑efficient hardware. We open‑source InSituANN at https://github.com/mindtravel/InSituANN‑OpenSource.

Authors:Yi Pan, Jun-Jie Huang, Tianrui Liu, Zihan Chen, Lin Liu, Zhao Wentao
Title: IDATA: Scalable Invertible Diffusion for Unrestricted Adversarial Transfer Attack
Abstract:
Unrestricted adversarial transfer attacks are important for evaluating the black‑box robustness of deep visual models. Diffusion‑based attacks have shown promising transferability and visual imperceptibility by optimizing adversarial perturbations along denoising trajectories in latent space. However, existing methods are limited by two challenges: memory‑intensive multistep backpropagation and frequency‑agnostic perturbation over intermediate latents. To address these issues, we propose IDATA, a memory‑efficient diffusion framework for unrestricted adversarial transfer attack. IDATA consists of two key components: an Invertible Diffusion Module (IDM) and a Low‑Frequency Constraint Module (LFCM). Specifically, IDM reformulates adversarial optimization over diffusion trajectories as an invertible process, enabling constant‑memory backpropagation through on‑demand reconstruction of intermediate states instead of storing the full denoising chain. Moreover, LFCM leverages Discrete Wavelet Transform (DWT) to decompose latent variables into low‑ and high‑frequency components, restricting perturbations to semantically stable low‑frequency subspaces, thereby improving transferability while preserving visual imperceptibility. Extensive experiments on multiple benchmarks and diverse model architectures demonstrate that IDATA consistently outperforms state‑of‑the‑art baselines in attack success rate, memory efficiency, and visual imperceptibility. These results suggest that IDATA is a promising tool for black‑box robustness evaluation of deep visual models. Code is available at https://github.com/colourful‑pan/IDATA.

Authors:Víctor Gallego
Title: Gaming Without an Attacker: Benchmark Fingerprinting in LLM-Driven Search Under Selection Pressure
Abstract:
Benchmarks for systems that are optimized against the evaluation signal measure something different from what they claim. We document this concretely in two GPU‑kernel‑optimization suites with held‑out generalization gates: Metal‑Sci (10 scientific‑compute tasks) and Metal‑ZK (12 zero‑knowledge/cryptographic tasks), in which three frontier LLMs (Opus 4.7, Gemini 3.1 Pro, GPT‑5.5) propose Metal kernels inside a (1+1) evolutionary loop with rich feedback. Although no model is prompted to act adversarially, the promoted winners repeatedly fingerprint the evaluation configuration: they branch on the identity of runtime parameters, tune the measured branch maximally, and leave the unmeasured branch slow or silently wrong. Across the pooled suites, 16/53 (30%) of in‑distribution wins fail to transfer to held‑out configurations. We give a four‑mode taxonomy of these failures, from configuration fingerprints to gate leakage. We distill design guidance for measurement under strategic optimization: held‑out probes retain validity only on non‑enumerable axes; gates must measure held‑out performance, not just correctness; and a transfer rate is interpretable only with per‑failure mechanism grades: ours decomposes into gamed, overfit, and benign. Code and research artifacts: https://github.com/vicgalle/kernel‑fingerprinting

Authors:Xudong Wu, Zeqing Wu, Jiarui Zhang, Xuhao Fan, Ziang Ding, Yuming Zhuang, Mingqi Yuan, Yilun Du, Hongjie Jia, Yunfei Mu, Jiayu Chen
Title: EnergyBridge: Benchmarking Household Energy Management, User Participation, and Grid Flexibility
Abstract:
Residential virtual power plants (VPPs) can provide grid flexibility by shifting household demand, but physical flexibility becomes dependable capacity only when residents authorize a plan and the promised response is delivered. Existing benchmarks evaluate control but omit event‑specific authorization. We present EnergyBridge, a benchmark and agent framework connecting capacity reporting, household authorization, and physical execution. It combines region‑specific EnergyPlus environments for Tianjin and Berlin with an LLM‑based User Participation Simulator. Against 584 persona‑ and event‑matched human role‑play judgments, the LLM‑based User Participation Simulator preserves method ordering with a 5.3‑point mean absolute acceptance error. Across conventional controllers and agent baselines, EnergyBridge achieves the highest simulated authorization, lowest event‑window energy, and the most reliable capacity commitment in both regions. We release human data and codes for reproducible human‑centered grid‑flexibility research: https://github.com/Agentic‑Intelligence‑Lab/EnergyBridge.

Authors:Khoa Hoang, Hoang-Tuan Nguyen, Huong Ninh, Hai Tran, Long Q. Tran
Title: Semi-Dense Matching Uncertainty Is Not Just Local Confidence
Abstract:
Reliable semi‑dense matching is essential for modern geometric vision systems. Designed under a coarse‑to‑fine paradigm, it achieves an optimal balance between performance and computational cost. However, existing methods often struggle to provide well‑quantified uncertainties, where catastrophic coarse‑assignment failures are ignored, leading to truncated error distributions and severely misjudged geometric estimations. In this paper, we propose a lightweight, post‑hoc overall uncertainty estimation framework that introduces a two‑component calibrated Laplace mixture model with only 9 learnable parameters. The objective is to explicitly capture both the sharp local refinement noise and the broader tail of coarse‑assignment failures. We introduce the Coarse‑success posterior Refit (CoRe) method, a geometric refitting module that utilizes the posterior probability of coarse‑assignment success as soft correspondence weights. Extensive experiments show that our method consistently improves downstream geometric accuracy across various pretrained‑only matchers and robust estimators with minimal computational overhead. Our code is available at https://github.com/khoavpt/Probabilistic‑matching.

Authors:Xuan-May Le, Minh-Tuan Tran, Ling Luo, Uwe Aickelin, Dinh Phung, Trung Le
Title: Efficient Test-Time Scaling for LLM-based Time Series Forecasting
Abstract:
Long‑term time series forecasting benefits from preserving global structure such as trends and seasonality. Recent LLM‑based forecasters often improve accuracy through test‑time scaling (e.g., iterative refinement), but these methods are computationally expensive and increasingly prone to global‑shape mismatch as the prediction horizon extends. We propose SCALER, a coarse‑to‑fine forecasting framework that first employs a lightweight Transformer tailored to long‑term shape modeling to predict a coarse representation of future dynamics. This predicted shape then serves as a compact guide for an LLM to perform test‑time scaling via iterative coarse‑to‑fine residual token refinement, while processing substantially fewer tokens at each step. By guiding refinement with an explicit future‑shape prediction, SCALER reduces reliance on long description prompts, and its fixed‑step refinement avoids costly reward‑model‑based selection, further lowering computational overhead. Experimental results demonstrate that SCALER outperforms strong forecasting baselines in long‑term, short‑term and zero‑shot forecasting while significantly reducing the inference cost associated with scaled LLM for time series forecasting. Code: https://github.com/xuanmay2701/SCALER.

Authors:Jinhua Cui, Anhong Wang, Kai Hu, Donghan Bu, Peihao Li, Tammam Tillo, Hao Jing, Shiao Xu
Title: JSGS: JPEG State-Guided Supervision for 3D Gaussian Splatting from Mixed-Quality Views
Abstract:
Standard 3D Gaussian Splatting (3DGS) assumes that every input image faithfully samples scene radiance. However, mixed‑quality JPEG images violate this assumption because compression‑induced blocking and ringing artifacts can corrupt updates to Gaussians shared across views. To address this problem, we propose JPEG State‑Guided Supervision for 3D Gaussian Splatting from Mixed‑Quality Views (JSGS). JSGS uses luminance and chrominance quantization tables stored in each JPEG file to construct a view‑specific JPEG observation operator. This operator encodes and decodes each rendered view for domain‑matched comparison with the corresponding decoded input image. The luminance quantization table supplies continuous weights within a fixed middle frequency band. A loss in the low frequency band anchors coarse structure, while the weighted middle frequency loss redistributes supervision among the selected DCT coordinates. The resulting block disagreement also guides the Gaussian Controller to regularize small primitives with high opacity in disagreement regions. Across seven scenes and three mixed‑quality schedules, JSGS achieves the lowest mean LPIPS and the highest mean SSIM under every schedule while rendering at approximately 150 FPS. Code: https://github.com/Jayden‑Cui/JSGS.

Authors:Mia Hines, Alvitta Ottley
Title: Charting Public Health: A Taxonomic Study of Visualization Practices in the Public Health Field
Abstract:
Public health organizations regularly produce and publish data visualizations to raise awareness of critical issues, influence decision‑making processes, and promote overall well‑being. However, the design practices shaping these visualizations in real‑world settings remain largely unexamined, limiting the research community's ability to evaluate their effectiveness, accessibility, and alignment with communication goals. To address this gap, we construct and analyze a large‑scale corpus of over 4,000 real‑world data visualizations drawn from more than two dozen websites associated with U.S. and international public health organizations. We evaluate salient design characteristics like chart type, visualization accessibility, use of embellishments like iconography, and design flaws. This work contributes to understanding real‑world decisions in designing data visualizations and supports public health officials in improving data visualization‑related communications. Visualizations in our finalized corpus and the labeled dataset can be found at https://github.com/washuvis/ChartingPublicHealth.

Authors:Antonis Polemitis, Nicholas Christakis, Dimitris Drikakis
Title: Exact Finite-Horizon Memory, Conditioning, and Dissipative Decay in Coarse Upwind Finite-Volume Prediction
Abstract:
Coarse finite‑volume averages do not generally form a predictive state: realized interface fluxes close a conservative update, but distinct fine‑grid states with identical parent averages can generate different future coarse histories. We analyze this failure for periodic scalar advection discretized by a first‑order upwind finite‑volume method with forward Euler time integration. We derive a finite‑horizon observation‑rank law: each additional observation exposes one new child‑cell layer and contributes one fewer independent direction than the number of parent cells, until the unresolved layers are exhausted. Centralized prediction therefore requires one fewer additional coordinate than the number of parent cells per exposed layer, whereas product‑local prediction requires one coordinate per layer in each parent. An anchored flux‑divergence queue attains the centralized bound, identifies the periodic flux gauge, and, at saturation, forms a minimal autonomous predictive state with the parent averages. We then distinguish exact observability from stable recoverability. The collar‑to‑queue map becomes rapidly ill‑conditioned as the Courant number decreases, so algebraically visible delayed information may fall below a prescribed numerical tolerance. For Courant numbers strictly between zero and one, we prove contraction of the nonconstant component under bounded arithmetic perturbations, with separate control of mean drift, an explicit perturbation neighborhood, and a grid‑dependent decay time. Numerical experiments illustrate the rank ladder, effective‑rank loss, queue conditioning, step‑function flattening, and delayed coarse separation followed by dissipative decay. The results provide a solvable benchmark for assessing state sufficiency in coarse, reduced, multiscale, and learned scientific models.

Authors:Yuqi Zhang, Cheng Chen, Yuyu Guo, Wenjie Yang, Lingchen Meng, Peng Di, Hang Yu, Zuxuan Wu, Yu-Gang Jiang
Title: VLZip: Unified Visual and Textual Compression for Interleaved Long-Context Modeling
Abstract:
Vision Language Models (VLMs) face significant challenges with ultra‑long, interleaved image‑text sequences due to the quadratic complexity of self‑attention. Current solutions either resort to aggressive token pruning, risking irreversible information loss, or adopt efficient but less precise architectures, while largely ignoring the equally vital textual component. We introduce VLZip, a framework that unifies visual and textual compression for high‑fidelity reasoning within a pure Transformer. At its core, VLZip hierarchically distills visual and textual segments into compact, layer‑specific "soft prefixes" and injects them into each decoder layer's hidden states, drastically shortening the attention sequence while preserving fine‑grained global context. To address deficient evaluations in the field, we also introduce LongVLBench, a new benchmark derived from video narratives that demands holistic, narrative‑level reasoning. Extensive experiments show VLZip achieves leading performance on long‑context multimodal reasoning, enabling training up to 120K tokens, a 6x increase over the baseline, and inference beyond 280K tokens with significantly reduced memory, while demonstrating the memory scalability to handle up to 2M tokens. By excelling at extreme context lengths where existing methods collapse, VLZip establishes an efficient and powerful new standard for long‑context multimodal AI. Code is available at https://github.com/ShareLab‑SII/VLZip.

Authors:Jiayi Li, Di Wu, Qingxu Li, Hongxiao Zhao, Jiaqi Yang, Anjunyi Fan, Wenbin Zhang, Boqiang Wu, Shuting Liu, Shifeng Fang, Jianbo Dong, Dimin Niu, Bonan Yan
Title: C2C-Explorer: An Exploration Framework for Chip-to-Chip Interconnect Architectures in LLM Cloud Computing Systems
Abstract:
The scaling‑up of large language models (LLMs) necessitates computing systems to have multi‑processor‑chip architectures, elevating the importance of chip‑to‑chip (C2C) communication. However, designing efficient C2C hardware architectures for LLM workloads faces three key challenges: generating realistic LLM‑specific C2C traffic, accurately simulating hardware‑level communication at scale, and efficiently exploring the exponentially large C2C design space. We propose C2C‑Explorer, an adaptive Bayesian DSE framework that integrates a LLM‑workload‑driven traffic generator, a scalable interconnect simulator (switch/full‑mesh, up to 512 chips), and a metric‑guided evaluator into a workload‑to‑hardware optimization pipeline, enabling systematic C2C architectural co‑design under realistic LLM workloads. Validated against FPGA‑based C2C prototypes, the C2C simulator achieves 2.46‑8.23% end‑to‑end timing error across diverse traffic patterns. Its hybrid cycle and event model further accelerates large‑scale simulation by up to 7.8× over a pure cycle‑accurate baseline. Applied to a 32‑XPU DeepSeek‑R1‑671B inference workload, C2C‑Explorer identifies configurations that improve goodput by 44.1% and reduce memory by 98.4%. C2C‑Explorer is open‑source and available at https://github.com/Selinaee/C2C‑Explorer.

Authors:Wenxu Jia, Dongjie Fu, Xize Cheng, Fangming Feng, Linjun Li, Wenshi Chen, Yingming Li, Zhou Zhao, Tao Jin
Title: VoxZip: Semantic-Anchored Temporal KV Cache Compression for Long-Context Audio Inference
Abstract:
Recent advancements in Speech Large Language Models have demonstrated remarkable capabilities in understanding complex audio tasks. Despite this progress, their long‑context inference remains severely bottlenecked by prohibitive KV cache memory demands. Existing text‑centric compression methods struggle here, often disrupting speech continuity or discarding crucial semantic cues. To address this, we propose VoxZip, a train‑free, two‑stage semantic‑anchored KV cache compression framework. The first stage uses automatic speech recognition (ASR) transcriptions as explicit semantic anchors to temporally align, compress, and fuse audio tokens, significantly reducing the initial KV cache while elevating token information density. To further improve the compression ratio, the second stage employs a dynamic filtering strategy based on temporally decayed accumulated attention to evict non‑essential tokens while mitigating early‑token bias. Comprehensive evaluations on Qwen3‑Omni across six diverse audio benchmarks demonstrate the superiority of our approach. VoxZip excels in long‑audio reasoning and consistently maintains high‑fidelity perception on short‑form tasks. Notably, it sustains over 90% of the uncompressed baseline performance even under an aggressive 20x KV cache compression in long‑context scenarios. Furthermore, at a 4x compression ratio, VoxZip yields a 1.9x increase in inference throughput alongside a 3.3x reduction in peak memory overhead. Code and models will be available at https://github.com/MM‑Speech/VoxZip.

Authors:Chenhao Qiu, Ruixiang Wang, Runyi Zhao, Sixu Lin, Songen Gu, Shufeng Nan, Guiliang Liu, Kui Jia, Yanwei Fu, Simo Wu
Title: Vid2WAM: Distilling Video Diffusion Priors into World Action Models
Abstract:
World Action Models (WAMs) improve robot policy learning by jointly modeling future visual dynamics and actions. However, their scalability and generalization remain constrained by their reliance on costly expert demonstrations. We challenge this by asking whether future supervision for WAMs must originate from target‑task expert trajectories. In this paper, we propose Vid2WAM, an offline distillation framework that transfers visual diffusion priors from a large video foundation model into a compact WAM student. Given an observation and language instruction, Vid2WAM distills supervision through two complementary channels: task‑conditioned future rollouts directly supervise the student's future prediction branch, while an inverse dynamics model recovers embodiment‑specific pseudo‑actions for action learning. To robustly integrate synthetic and real supervision, we introduce source‑aware residual action adaptation that learns source‑specific corrections around a shared action backbone and mitigates interference from noisy pseudo‑actions. During inference, both the video teacher and inverse dynamics model are discarded, leaving only the WAM student for efficient deployment. Simulation and real‑world experiments demonstrate that Vid2WAM improves novel‑task generalization and data efficiency under limited expert demonstrations while preserving low‑latency inference.

Authors:Xiaoyang Bai, Zhenyang Li, Weiwei Xu, Edmund Y. Lam, Yifan Peng
Title: ERF-GS: Reconstructing Fast Motion from Disjoint Event-RGB Viewpoints
Abstract:
Deep learning‑driven representations such as neural radiance fields (NeRFs) and 3D Gaussian splatting (3DGS) have revolutionized the field of dynamic 3D scene reconstruction with improved visual precision and scalability. However, the reconstruction of fast‑moving objects remains a challenge; existing methods based on conventional frame‑based videos often struggle in scenarios such as sports events and animal videography. We propose an event‑RGB fusion Gaussian splatting (ERF‑GS) framework that integrates event information into both optimization and densification stages of the Gaussian splatting pipeline, taking advantage of novel event sensors with high frame‑rate. Unlike many other event‑assisted scene reconstruction methods, ERF‑GS was developed using realistic simulation settings and realizes event‑based learning detached from RGB inputs. This design enables its application beyond straightforward synthetic data into the realm of natural video with complex layout, low frame rates and severe motion blur. Our experiments show that ERF‑GS outperforms both the 4DGS baseline and the concurrent E‑D3DGS on different variants of the Neu3D and Nvidia datasets which include blurry RGB frames and disjoint RGB‑event viewpoints. Our code is available at https://github.com/andrewbxy/ERF‑GS.

Authors:Minhan Cho, Jimin Kweon
Title: Reproducing and Stress-Testing Two Approaches to LLM Reasoning Reliability: Test-Time Probability Aggregation and Logic-Representation Editing
Abstract:
We independently reproduce two recent methods for making large language model (LLM) reasoning more reliable, and stress‑test them across domains and models (RPC across four new task domains with Qwen3‑8B, LCF across four 7‑8B models). The first, RPC, aggregates token probabilities and self‑consistency at inference; the second, LCF, trains projectors that split hidden states into "content" and "logic" and edits the logic part toward a valid region. Validating such reliability claims matters because the original evaluations are run by each method's own authors and were never independently reproduced or stress‑tested across models and domains, and LCF shipped no public code. We re‑run RPC's published‑path aggregation and re‑implement LCF's projector, contrastive, and intervention pipeline, then extend both to text‑to‑SQL, legal extraction, fallacy identification, and precedent grading, and probe LCF's representation directly. RPC reproduces the original grid exactly on the authors' released reasoning paths; on four new domains its edge over self‑consistency is never significant (ties or small mixed differences, paired p >= 0.28), and on BIRD, the one domain where we vary the budget, the edge grows with K as predicted but its largest gap (+2.5 accuracy at K=32, p=0.16) reverses to ‑0.25 when we enlarge the sample to n=200. LCF's logic‑validity direction is real but weak (0.82 separability at the single best sub‑layer versus 0.95 for a semantic‑attribute control); its one positive effect (Qwen3 ΔProb) is not significant (p=0.56), while it significantly reduces ΔProb on two of the other three models.

Authors:Mahbub E Sobhani, Md. Faiyaz Abdullah Sayeedi, Fahmid Hasan Chowdhury, Md Adnan Arefeen, Farig Sadeque, Md. Faizul Bari, Swakkhar Shatabda
Title: Time Present and Time Past: Benchmarking Large Language Models on Temporally Evolving Document Understanding
Abstract:
Evolving documents, such as laws, tax codes, and software documentation, are amended, replaced, and sometimes reverted over time, so a question has different correct answers at different dates. In contrast to encyclopedic knowledge, where an old fact is simply overwritten, an amendment is itself an official text that states what it replaces and when it takes effect, and the earlier version stays correct for its validity period. The central challenge is therefore version resolution, that is, identifying the version in force on the queried date. Existing temporal QA datasets treat time only as an annotation, so version resolution stays untested. We present TIDE, an expert‑verified benchmark of 3,050 QA pairs over 644 official customs instruments issued between 1969 and 2025 by the Government of Bangladesh, covering eight task types over deeply code‑mixed documents that are heterogeneous in layout and dated in two calendars. In addition, we evaluate nine recent LLMs under a single protocol across parametric, gold‑context, and retrieval access, scored by a three‑judge LLM council with a hard date gate separating correct meaning from correct time. The best macro‑averaged accuracy is only 68.5%. Resolving a version from an implicit date reaches 59.7%, and detecting that the supplied version does not govern the query reaches only 26.7%. Models are more likely to find correct versions than to reject incorrect ones, and they tend to follow a confident parametric answer over the supplied authoritative text. All code and data are available at https://github.com/icsetepa44/TIDE

Authors:Jitendra Prajapati
Title: A counterexample to the Etzion-Silberstein conjecture
Abstract:
The Etzion‑Silberstein conjecture asserts that the Singleton‑type upper bound for linear Ferrers‑diagram rank‑metric codes is attained for every Ferrers diagram, minimum rank distance, and finite field. Let E be the Ferrers diagram with column heights (5,5,5,5,1,1). The bound for minimum rank distance 3 is 12. We prove that every binary linear code supported on E with minimum rank distance 3 has dimension at most 11, and we give an explicit code of dimension 11. Thus the optimum is exactly 11, disproving the conjecture. The nonexistence proof reduces a hypothetical dimension‑12 code to one of the three equivalence classes of binary [4× 4,12,2] MRD codes. A rank‑distribution argument eliminates two classes and leaves four kernel orbits in the field class; all four exact lift systems are unsatisfiable. Independently written verifiers reproduce the result, including a raw enumeration of all 8,382,465 kernels without orbit reduction. We also prove an exact row‑cone propagation identity. Iterating it produces binary counterexamples with bound 12 and optimum 11 at every minimum rank distance d \geq 3.

Authors:Cong Ming, Jingyi Chen, Bin Liu, Qi Chu, Tao Gong, Nenghai Yu, Yingfei Xiang
Title: Yesterday's Shield, Today's Spear: A Self-Evolving Safety Guardrail in Production
Abstract:
Deployed LLM safety guardrails are predominantly static: trained once and frozen at release, while new jailbreak techniques and previously un‑addressed harmful categories emerge within days, leaving the defense perpetually a step behind. We present SESG (Self‑Evolving Safety Guardrails), a multi‑agent system running in production. SESG monitors the live traffic behind a deployed guardrail and surfaces two classes of failure: jailbreaks novel in form and harmful categories novel in content. Once a failure is confirmed, a generation agent synthesizes paired training data targeted at it; a validation agent rebalances the batch toward the direction in which the deployed model errs, so that the model's own mistakes steer its training set; and a routing agent matches the training action to the diagnosed gap and returns the next version to production. Over six rounds of live evolution (V0 to V6), a 1.7B guardrail adapts to a new threat in 16‑24 hours, with about 2 hours of human effort, versus the 40‑90 hours of the manual process it replaces. On six emerging threats, it outperforms static guardrails from 0.6B to 9B and an adaptive baseline while preserving its general screening competence. Since April 2026, SESG has been the primary update pipeline of Sangfor's guardrail, autonomously closing 14 of 15 new threat scenarios in two months. We release 9 test sets for the 6 new threats at https://github.com/Trams1017/SESG. Warning: This paper contains examples that may be harmful or offensive.

Authors:Minhan Cho, Soyoung Park, Kihyeon Jeong, Byeongkyu Jeon, Daejin Choi, Jinyoung Han
Title: LLM within MCP Matters: Measuring Inefficient Resource Utilization Driven by LLMs
Abstract:
The Model Context Protocol (MCP) standardizes how servers expose data and tools to Large Language Models (LLMs). A common server design embeds frequently used reference data, such as identifier lookup tables, directly in the server instructions: the system‑prompt text a server hands to the host application. When a query concerns an entry of the embedded table, the model can act on it immediately instead of re‑discovering the same information through a search tool. We test whether client LLMs actually consume such instruction‑embedded data, reporting a 54,000‑trial study across 24 LLMs (9 Claude, 6 Gemini, 9 GPT) on a production legal‑information MCP server. A diagnostic condition that removes the competing search tool shows that failures are dominated by behavioral preference rather than missing capability. With search unavailable, 23 of 24 models read the embedded data reliably (hit ratio at least 98%); with a search tool merely present, 9 models drop below 15%. A 2^3 factorial analysis of three instruction‑level interventions reveals strong interaction effects: combining all three restores at least 86% for 20 of 24 models, but individual interventions can backfire for specific model families. Per‑server prompt engineering is therefore a workaround rather than a fix; we argue that MCP host applications should provide an explicit mechanism that places server instructions ahead of tool selection in the client LLM's deliberation.

Authors:Tailin Zhou
Title: Hierarchical Self-Improvement: A Framework for Task-Specific Evolvable Agent Harnesses
Abstract:
Modern LLM agents are often improved by modifying prompts, tools, or workflows manually, while the executable scaffold surrounding the model‑‑‑the \emphharness‑‑‑is typically treated as a fixed artifact after deployment. This work studies an alternative where the harness is \emphtask‑specific and continuously evolvable: each task family maintains its own harness, which is hot‑swapped across iterations through a fixed task‑injection seam and rewritten using environment feedback. We introduce Hierarchical Self‑Improvement (HSI), a framework in which a single frozen LLM M operates across three hierarchical scopes: a task harness H that executes tasks, an evolver that rewrites H, and a meta‑evolver that rewrites the evolver's strategy code under a frozen outer anchor. A thinking‑on/off design isolates the contribution of harness evolution by disabling reasoning during task execution while enabling it during self‑modification. HSI is bounded by two factors: a \emphfeedback‑fidelity bound, since evolution requires informative reward signals to guide selection, and a \emphbackbone capability bound, since harness redesign cannot overcome limitations of the frozen model. On BALROG with DeepSeek‑V4‑Flash‑Preview as the frozen backbone, HSI achieves consistent gains over the initial harness on moderate‑difficulty tasks (+39.3 on BabyAI, +33.0 on Crafter, +25.0 on TextWorld, and +15.0 on MiniHack, all in raw % Progress), while obtaining strong held‑out generalization on BabaIsAI sub‑suites (0.98 best‑test on BreakStop and 1.00 on GoTo from a 20% unseen split). On tasks beyond the backbone's capability (NLE), harness evolution provides no improvement. These results demonstrate task‑specific harness evolution as a viable axis for improving frozen LLM agents under clear empirical limits. Code is available at https://github.com/TailinZhou/hsi.

Authors:Zhuowen Liang, Zhengxuan Zhang, Jiayang Wang, Jiazhuo Chen, Nan Tang
Title: Beyond Tables: Doc2DB-Bench for Relationally Faithful Document-to-Database Construction
Abstract:
Practical AI systems increasingly need to turn long, heterogeneous documents into queryable relational databases, not isolated spreadsheets. In domains such as finance, healthcare, education, transportation, and enterprise operations, downstream workflows rely on normalized schemas, entity identities, keys, cross‑table relationships, and integrity constraints for analytics, compliance, auditing, and SQL‑backed decision making. Existing Document‑to‑Table benchmarks are insufficient for this setting: flattening evidence into single tables can duplicate entities, obscure many‑to‑many relationships, create sparse records, and avoid testing whether extracted facts form a valid database instance. This creates an urgent need to evaluate document understanding as database construction rather than field extraction. We introduce Doc2DB‑Bench, a benchmark for Document‑to‑Database construction, containing 203 long‑document instances across 42 schemas and seven domain groups, with 117 entity tables, 132 relationship tables, 7,341 rows, and 41,935 cells. Built through a controllable DB‑to‑Doc synthesis pipeline and organized by a taxonomy of intra‑table extraction and inter‑table reasoning, the generated documents undergo authenticity verification, proving indistinguishable from real‑world references. Doc2DB‑Bench thus provides a testbed for reliable, auditable, and relationally faithful LLM‑based data systems. The benchmark is publicly available at https://github.com/SetonLiang/Doc2DB‑Bench.

Authors:Xianghan Meng, Wei He, Zhiyuan Huang, Chun-Guang Li
Title: Learning Deep Modality-Shared Self-Expressiveness for Image Clustering with Textual Information
Abstract:
Leveraging textual information for image clustering has emerged as a promising direction, largely owing to the powerful representations learned by Vision‑Language Models (VLMs). Existing approaches typically retrieve a textual counterpart for each image and then refine multimodal representations by directly enforcing cross‑modal agreement, e.g., maximizing image‑text similarity inherited from pretrained VLMs. However, such a strategy aligns heterogeneous representations across modalities without explicitly modeling the intrinsic structure within each modality and thus might yield unreliable alignment or distort modality‑specific structures that are crucial for clustering. In this paper, we propose a simple but principled approach, termed deep modality‑shared self‑expressive model (DeepMORSE), which discovers cross‑modal structures via a modality‑shared self‑expressive model and simultaneously learns structured representations that conform to a union of modality‑specific subspaces. Moreover, we theoretically justify that the modality‑shared self‑expressive coefficients suppress inter‑class noise towards a subspace‑preserving solution, and show that mini‑batch optimization procedure introduces an implicit regularization onto the self‑expressive model. We evaluate our DeepMORSE on six widely used image clustering benchmarks and observe performance improvements exceeding 3% on the UCF‑101, DTD‑47, and ImageNet‑Dogs datasets. In addition, we demonstrate the strong transferability of the learned representations by achieving state‑of‑the‑art performance on downstream tasks such as image retrieval and zero‑shot classification‑‑‑without requiring any task‑specific losses or post‑processing. The code is available at: https://github.com/mengxianghan123/DeepMORSE.

Authors:Zain Naboulsi
Title: Does a Toehold Make a Bidder Bolder? Preemption and Multiplicity in Multi-Round Takeover Auctions
Abstract:
A bidder can quietly buy a stake in a company before making an offer for it. That stake, a toehold, is supposed to pay for itself twice: it makes the bidder willing to bid harder, and it frightens rivals into staying out of the fight. The first effect is arithmetic. The second is what would justify the cost and exposure of taking one at all. Yet toeholds are rare in practice, a standing puzzle. We ask whether that second effect is there once the contest is modelled as several rounds of escalating offers rather than the single exchange classical models assume. We turn it into a game a computer can solve, and certify the answers to an accuracy a referee can check. Three findings. The auction fixes what the toehold‑holder earns but not how it bids: the same contest supports a bidder who opens aggressively against a rival who folds, and one who opens cheaply against a rival who does not, with the same profit either way. Aggressive preemptive bidding still appears when the toehold is removed entirely, so it comes from bidding in public and in turns, not from owning the stake. And the tidy "bigger toehold, more deterrence" relationship holds only in a contest cut short after one round; give it a real second round and it stops responding. So the two reasons to buy a toehold do not fare alike. The profit reason holds up; the deterrence reason does not, which suggests why toeholds may be rarer than theory predicts, alongside the procedural costs of disclosure and price impact that this model omits. A warning follows for anyone computing economics from a game solver: solve this auction once and it returns a confident figure for what a preemptive bid is worth; solve it again from a different start and it returns a different one, equally converged. We also report which solvers cope with contests of this shape, including versions too large to enumerate. Code is released.

Authors:Yiqiao Liao, Parinaz Naghizadeh
Title: Rethinking Learning-Based Influence Maximization: Simple Neural Surrogates and Native Discrete Search
Abstract:
Existing learning‑based influence maximization frameworks rely heavily on complex neural architectures and continuous optimization over seed representations. We challenge this paradigm with SIMBA, a diffusion‑model‑agnostic framework pairing a lightweight neural surrogate with direct discrete search. SIMBA introduces three key components: 1) uniformly anchored node embeddings that eliminate initialization noise and encourage learning driven by graph topology and diffusion pattern, 2) a shallow two‑layer graph neural network surrogate predicting final infection states, and 3) batched multi‑swap simulated annealing that explores combinatorial seed space without gradients or continuous relaxation. By shifting compute from complex representation learning to effective discrete search, SIMBA drastically cuts time‑to‑solution while achieving superior influence spread and data efficiency. Our code is available at https://github.com/yl489/rethink‑IM.

Authors:Zejun Xu, Taiyi Chen, Jin Li, Yongtong Gu, Qi Cheng, Aixuan Lv, Shuai Zhu, Pengfei Zhu, Kaichen Yang, Boyu Sun, Yixian Yang, Mulong Xie, Xin Liu, Dagang Li, Xiaoteng Ma, Hongru Wang
Title: CAP: A Scalable Benchmark for Evaluating Cross-Site Browser Agents with Complex Actions and Perception
Abstract:
Large language models are increasingly deployed as autonomous agents that interact with the web through browsers. While recent progress has been driven by benchmarks that evaluate end‑to‑end task success, these evaluations largely overlook two fundamental sources of difficulty in real web browsing: complex actions over rich user interfaces and visual perception of dynamically rendered content, especially in workflows that span multiple websites. We introduce CAP, a scalable benchmark for evaluating browser agents on cross‑site, human‑like web tasks that require non‑trivial UI interactions and visual understanding. Specifically, we adopt a decomposition‑and‑recomposition pipeline that first abstracts each website into a structured site card capturing user‑facing functions, complex execution operations, and perceptual requirements, and then recomposes these components into realistic cross‑site workflows. Each task is therefore grounded in multiple specific operations on each website, enabling fine‑grained diagnosis. Built on this framework, we construct 420 tasks across 108 real‑world websites and 24 domains under careful quality control. Experiments on state‑of‑the‑art browser agents using our verifiable agent‑as‑a‑judge evaluation framework show low success rates and reveal that perception‑heavy interactions remain a major bottleneck, exposing substantial gaps between current agents and real‑world web browsing demands.

Authors:Zhiyuan Yang, Jiahao Cheng, Vincent Quoc-Huy Trinh, Mahdi S. Hosseini
Title: Gated Spatial Redundancy Projection for Pathology Transformer Attentions
Abstract:
Transformer models are increasingly used for whole‑slide image analysis in computational pathology. Yet, WSIs differ fundamentally from natural images: neighbouring patches often contain highly similar tissue type, stain, texture, and cellular composition. We identify this local spatial redundancy as a pathology‑specific failure mode of self‑attention, where dominant neighbourhood features can be repeatedly mixed into patch‑tokens and weaken subtle diagnostic or prognostic deviations. We propose Gated Spatial Redundancy Projection (Gated SRP), a lightweight drop‑in correction module for self‑attention layers. For each patch token and attention head, Gated SRP estimates a local redundancy axis from neighbouring value vectors, projects the attention output onto this axis, and applies a learned signed gate to correct the redundancy‑aligned component geometrically. Across five TCGA survival cohorts, Gated SRP obtains the highest mean C‑index among the compared attention variants in all cohorts, with an average improvement over the base attention, while adding only +0.02% parameters. Across five slide‑level classification datasets, it improves the base attention on 12 of 16 reported metrics and achieves the best AUC on three datasets. Code is publicly available at https://github.com/AtlasAnalyticsLab/GatedSRP.

Authors:Uri Z. Kialy, Gil Ben-Artzi
Title: Circuit Fine-Tuning for Compute-Efficient Transformer Adaptation
Abstract:
Parameter‑Efficient Fine‑Tuning (PEFT) has become the de facto standard for adapting Vision Transformers (ViTs) to downstream tasks. While parameter count has been the dominant efficiency metric in PEFT, it does not imply compute efficiency: parameter‑sparse methods can still incur full‑model training cost per step, and typically need long schedules to reach peak accuracy. We introduce Circuit Fine‑Tuning (CFT), a compute‑efficient framework that uses circuit discovery‑‑‑conventionally used to explain trained models‑‑‑to select modules for fine‑tuning before training. Whereas attribution is conventionally formulated against a trained task head, we formulate it against a near‑zero‑initialized probe head, which isolates the response of the backbone to the target distribution rather than the preferences of a particular classifier. CFT then fine‑tunes only the recovered subgraph. CFT needs no learning‑rate warmup and reaches peak accuracy in ~20 epochs on average‑‑‑versus 44‑‑96 for strong PEFT baselines‑‑‑yielding 2.3‑‑6.6× fewer training FLOPs and up to 16× less wall‑clock time, while adding zero parameters and no inference operations. Experiments across a standard visual transfer benchmark (VTAB‑1k), hierarchical backbones (Swin), domain‑shifted medical imaging (CBIS‑DDSM), and a vision‑language model (Gemma‑3 on CUB‑200) demonstrate the effectiveness of CFT. Code is available at https://github.com/UriKialy/CFT

Authors:Petr Korolev
Title: What Irregularity Costs: CUDA C++, Rust, and Triton on a Hash-Blocked GPU Workload
Abstract:
GPU language comparisons are almost always run on tiled dense linear algebra, where every toolchain is good and the differences are small. We implement the same hash‑blocked TSDF fusion kernel in CUDA C++, in Rust through NVIDIA's cuda‑oxide, and in Triton, and measure it on a workload with the opposite character: an open‑addressed hash table with compare‑exchange insertion, data‑dependent per‑lane probe depth, and contended scatter. The result is a split. On the regular stage, which walks a truncation band and accumulates, all three languages land within a small factor of each other. On the irregular stage, which probes and inserts, Rust stays close to hand‑written CUDA C++ while Triton is more than an order of magnitude slower. Language choice is nearly free on the work that is usually benchmarked and expensive on the work that is not. We attribute both gaps to specific things the languages cannot express, not to ratios. Triton's cost follows from a probe loop that must run to a compile‑time bound and from tl.atomic_cas taking no mask, which forces a scratch structure with no counterpart in CUDA. Rust's cost was invisible in every instruction count: its kernel issues fewer instructions, fewer compare‑exchanges and fewer registers at identical occupancy, yet was slower. Hardware counters located it in L1 residency. A GPU‑scope atomic load must be coherent across SMs, no NVIDIA L1 is, so the type‑correct way to read a shared location bypasses the cache on every access. Triton's bounded probe is also a correctness problem for fusion: at load factors an ordinary depth trajectory reaches, it silently discards blocks and the reconstruction loses patches of surface with nothing reported. We also report a defect found and fixed in cuda‑oxide itself, now merged upstream: its scoped atomic load and store could not be called at all in the build mode that produces real kernels.

Authors:Zixiang Wan, Xusheng Yang, Zheng Wang, Peiji Yang
Title: ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure
Abstract:
Neural speech codecs face a fundamental tension in the language‑model era: tokens that support high‑fidelity reconstruction are not necessarily easy for autoregressive models to predict. Our controlled analysis of diverse codec and self‑supervised speech representations shows that clearer phoneme structure before discrete code assignment is consistently associated with easier autoregressive token prediction. Yet phoneme structure alone is insufficient for high‑fidelity reconstruction, which also requires reconstruction‑relevant acoustic detail. Guided by this observation, we introduce ReLMCodec, a low‑bitrate single‑codebook speech codec built upon a preserve‑‑control‑‑refine principle: it preserves the linguistic organization of frozen self‑supervised learning (SSL) features at the quantizer input, controls reconstruction‑driven drift through Pre‑quantization Anchor‑Preserving Adaptation (PAPA), and refines the quantized latent space with a training‑only WavLM‑Large L24 teacher to reduce phoneme‑level token fragmentation. Together, these components allow acoustic detail to support waveform reconstruction while keeping the resulting token sequence predictable for autoregressive models. At 650 and 800 bps, ReLMCodec moves the empirical single‑stream predictability‑‑reconstruction frontier in our evaluations, with gains that carry over to downstream text‑to‑speech (TTS) synthesis in both intelligibility and speaker similarity.

Authors:Varun Pratap Bhardwaj, Garima Singh, Arun Pratap Bhardwaj
Title: SuperLocalMemory 4.0: The Governed Memory Operating System for AI Agents
Abstract:
AI agents are becoming shared infrastructure, yet durable memory is commonly assembled from separate retrieval, governance, and operational components. We present SuperLocalMemory 4.0, a governed, local‑first memory operating system for AI agents. The system combines dense semantic, BM25 lexical, temporal, Hopfield‑associative, and spreading‑activation retrieval through reciprocal‑rank fusion; a governed learning and behaviour layer; bi‑temporal recall; multi‑scope personal, shared, and global memory; role‑based access control; GDPR‑oriented export and verified erasure; audit trails; and a deployment‑context EU AI Act checklist. V4 introduces a reliability spine for its primary write path: generation‑fenced admission, a policy registry, verifiable memory transactions with per‑projection apply, verify, compensate, and erase owners, and hash‑checkable completion manifests. The runtime is available through CLI, MCP, an HTTP daemon, a dashboard, editor integration, and framework adapters, and supports fully local, local‑with‑on‑device‑model, and provider‑assisted modes. We evaluate eleven fault‑injection and mechanism scenarios, each repeated 200 times. The released evidence bundle reports 2,200 of 2,200 deterministic repetitions upholding their scoped component properties. The governed write envelope measured 3.522 ms at p50 and 5.297 ms at p99, versus 1.835 ms and 2.569 ms for the ungoverned baseline, corresponding to in‑process control‑plane overheads of 1.687 ms at p50 and 2.728 ms at p99. These are scoped component and mechanism measurements, not an end‑to‑end multi‑process or external retrieval‑accuracy benchmark. The paper consolidates prior SuperLocalMemory work on privacy‑preserving multi‑agent memory, information‑geometric retrieval, and the V3.3 Living Brain lifecycle.

Authors:Ashritha Gonuguntla
Title: The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World
Abstract:
LLM routers promise efficiency by matching each request to the cheapest adequate model, and are increasingly applied per step inside multi‑step agents. Yet agentic routers are evaluated like single‑turn routers: by replaying logged trajectories and substituting another model's recorded outputs, assuming the rest of the trajectory is unaffected. We test this assumption with branching rollouts: we fork live SWE‑bench agent trajectories at controlled points, rebuild the environment, continue each fork with a different model, and compare against same‑model control forks that isolate sampling and replay noise. Across six paired runs (~900 rollouts), swaps exceed their matched control floors by +0.25 to +0.66 normalized edit distance (multiplicity‑corrected CIs exclude zero), rewriting 61‑94% of post‑fork actions; 74‑77% of early swaps diverge at the first post‑fork action, versus 6‑35% of controls, leaving only 3% of replayed states valid. Divergence decreases with fork depth in both directions. All five outcome flips we observe occur in swap arms, upgrades rescuing unsolved instances and a downgrade losing the sole solve, and zero occur across 359 control forks. Scoring these same swaps with a log‑stitching replay evaluator, replay mispredicts every success‑relevant outcome call and predicts patches with 0.00‑0.11 similarity to reality. Auditing the noise floor, temperature‑0 "determinism" is configuration‑dependent: FP8‑served controls diverge on over 90% of forks while AWQ‑served ones remain near‑identical; and under tight budgets the stronger model more often exhausts its steps without submitting. Replay‑based benchmarks score the wrong world for agentic routing; we release our harness and all trajectories.

Authors:Rui Wang, Yeteng Wu, Xianling Zhang, Mengshi Qi
Title: VTO: Visual Tool Orchestration for Video Anomaly Detection
Abstract:
Video anomaly detection (VAD) is a critical yet challenging task due to the complex and diverse nature of real‑world scenarios. Traditional deep learning approaches are fundamentally limited by poor generalization across diverse scenarios. While multimodal agents offer a promising tool‑learning paradigm for VAD, current systems relying on supervised fine‑tuning struggle with complex orchestration, and standard reinforcement learning often causes premature termination due to coarse‑grained outcome rewards. To address these challenges, we propose VTO, a process‑supervised reinforcement learning framework. Moving beyond static tool usage, VTO enables the agent to dynamically explore and interact with the environment. Specifically, we introduce a foundation model‑driven cognitive evaluator to provide context‑aware semantic feedback, which is seamlessly integrated into a Process‑Supervised Cognitive Alignment that delivers fine‑grained, step‑wise supervision. By explicitly penalizing logical truncation and rewarding complete causal chains, the agent optimizes its multi‑step reasoning policy for interrelated tool orchestration. To support our proposed framework, we meticulously crafted VAD‑Tool, a hierarchical visual tool set comprising 12 specialized vision tools spanning from entity tracking to high‑stakes hazard detection, and established the corresponding benchmark for rigorous multi‑step reasoning evaluation. Extensive experiments on VAD‑Tool demonstrate that VTO significantly outperforms baselines, achieving up to a 10.2% absolute accuracy improvement in tool scheduling. Code and data are available at https://github.com/MICLAB‑BUPT/VTO.

Authors:Satvik Praveen, Shengji Jin, Ahmed Lamidi, Xin Qian, Yi Sheng
Title: BAP-MOS: Bandit-Based Adaptive Prompting for Boundary-Sensitive Multi-Organ Segmentation
Abstract:
Multi‑organ ultrasound segmentation remains challenging when anatomically adjacent structures must be delineated jointly, as localized boundary errors can persist even when Dice scores are high. To address these challenges, we propose Boundary‑Adaptive Prompting for Multi‑Organ Segmentation (BAP‑MOS), a closed‑loop adaptive prompting framework. BAP‑MOS formulates prompt selection as an organ‑specific multi‑armed bandit problem over box, point, and combined prompts. An outer Tree‑structured Parzen Estimator (TPE) loop selects the prompt‑selection parameter vector, while an inner UCB‑Tuned loop adapts per‑organ prompt preferences during fine‑tuning using a bounded Dice‑‑MSD‑‑HD95 validation‑probe reward. The framework further introduces an organ‑scaled negative prompt ring to adapt sparse prompt geometry across anatomical scales, while keeping the image and prompt encoders frozen and updating only the mask decoder. We evaluate BAP‑MOS on pooled prostate‑region TRUS cohorts against U‑Net, nnU‑Net, MedSAM, fixed‑prompt SAM/MedSAM, and adaptive policy variants. On this benchmark, BAP‑MOS achieves Dice 0.982, HD95 0.482, and MSD 0.204, reducing HD95 by approximately 48% and MSD by 45% relative to the strongest conventional baseline. To verify the generalization ability of the framework, we tested it on the external PFUS1 pelvic‑floor ultrasound corpus using MedSAM and its adaptive strategy variants, and the results were good. These results support adaptive prompt allocation as an effective mechanism for improving boundary‑sensitive multi‑organ ultrasound segmentation without modifying the foundation‑model backbone. Source Code is available at: https://github.com/SatvikPraveen/BAP‑MOS

Authors:Muhammad Ayub Sabir, Shaohong Zheng, Zhiyu Qu, Fatima Ashraf, Junbiao Pang
Title: Large Multimodal Agents for Intelligent Transportation Systems: Architectures, Evidence, and Deployment Challenges
Abstract:
Large multimodal agents (LMAs) are increasingly proposed for intelligent transportation systems (ITS), but existing studies often conflate multimodality, agency, empirical performance, and deployment readiness. This review provides an auditable evidence map of 42 primary study families released between January 2023 and 3 August 2026 within a corpus of 91 mapped sources. It distinguishes model‑level, system‑level, and hybrid multimodality and classifies each family by system architecture and action authority. Evidence is assessed independently through functional capability (C0‑C3), validation setting (E0‑E4), three evidence propositions (P1‑P3), and eight methodological‑concern domains (Q1‑Q8). Transportation semantics (P1) are directly evaluated in 23 families and multidimensional integration (P3) in 24; 19 families directly evaluate both. Evidence reconciliation (P2) remains unresolved because no family demonstrates the complete provenance‑challenge‑handling‑comparison‑outcome chain. Fourteen families reach C3, but 13 remain at E2; only one reaches E3 and none reaches E4. Across ITS domains, LMAs are best supported for semantic interpretation, intent translation, evidence organisation, scenario authoring, explanation, and specialist‑tool coordination. Numerical forecasting, optimisation, simulation fidelity, hard constraints, low‑level control, safety fallback, and final authority should remain with independently verifiable specialist systems or accountable humans. The review therefore supports bounded orchestration rather than replacement and provides a matched comparative evaluation protocol and staged roadmap for accountable deployment. The living evidence repository is available at https://github.com/pangjunbiao/ITS‑LMA‑Review.

Authors:Yongkang Yang, Zhezheng Hao, Hong Zhang, Yi Liu, Xiankun Lin, Wence Ji, Fanjunduo Wei, Jiarui Yu, Qiang Lin, Xiaoyun Liang, Hande Dong
Title: Matching Supervision to the Student's Learning Capacity: A Unified Framework for On-Policy Self-Distillation
Abstract:
On‑policy self‑distillation (OPSD) improves the reasoning abilities of LLMs by internalizing privileged context into model parameters through self‑distillation. Two recent research lines promote vanilla OPSD by choosing which tokens to learn from and by controlling how much privileged information the teacher receives, respectively. However, we show that each line optimizes one variable while holding the other fixed, which leads to a suboptimal solution. We argue that the two variables are coupled through the student's learning capacity: the privileged information sets the per‑token divergence the teacher prescribes, while token weighting selects which of these the student must absorb. We formalize the two lines of work into a unified optimization framework, which maximizes the aggregate teacher‑‑student divergence, subject to a budget on the aggregate learning difficulty the student can absorb. Under this modelling, we propose Unified On‑Policy Self‑Distillation (USD), a lightweight online algorithm to solve the Lagrangian. USD reveals that a single dual variable governs both decisions: at one price for learning difficulty, it simultaneously sets the token‑selection threshold and the direction of privileged‑information adjustment, keeping supervision matched to the student's evolving capacity. Through extensive experiments, USD consistently demonstrates superior performance over OPSD and token‑ and PI‑side baselines across various model scales on various reasoning benchmarks. Code is available at https://github.com/lauvlalala/USD.

Authors:Víctor Gallego
Title: A Hybrid Nested Harness for Decoupling Structure and Parameters in LLM-Driven Optimization
Abstract:
In evolutionary algorithms powered by language models, the LLM acts as a single operator that simultaneously updates structural components (like control flow) and continuous parameters. While LLMs can be good at the first, they are not efficient at the second, wasting tokens taking discrete jumps inside a trial and error loop. We resolve this by formalizing a hybrid nested search, in which an outer loop has the LLM propose a structural sketch, with numeric gaps, and an inner numerical optimizer tunes the sketch. Both the outer and inner solvers are pluggable: any text‑based optimizer can be combined with a zero‑order optimizer (CMA‑ES), gradient‑based routines, or MCMC samplers. We validate our framework across three scientific domains: (i) meta‑optimizers on closed‑form test functions, (ii) code‑based policies for systems research and social dilemmas; and (iii) approximate Bayesian inference tasks. Across all three, the hybrid optimizer is superior to both vanilla LLM‑driven search and pure numerical optimization baselines. Code at: https://github.com/vicgalle/hybrid‑nested‑search

Authors:Koyar Afrasyab
Title: Exact Zarankiewicz Values On Two Finite Frontier Slices
Abstract:
The Zarankiewicz number Z(m,n,s,t) is the maximum number of edges in a bipartite graph with parts of orders m and n containing no copy of Ks,t. We give one combined, certificate‑based computer‑assisted proof for two finite slices and a corrected neighboring frontier: Z(12,n,3,3) = 6n (18 <= n <= 22), Z(13,22,3,3) = 137, Z(13, 18, 3, 3) = 116, Z(14, 18, 3, 3) = 124, Z(15,18,3,3) = 132, Z(14, 17, 3, 3) = 118, Z(15, 17, 3, 3) = 126, 132 <= Z(16,17,3,3) <= 133. The load‑bearing new upper bounds are the exact 12 x 18 and 13 x 18 certificate packages. Their orbit certificates exclude every hypothetical matrix at the next edge count. Deletion lemmas and explicit witnesses close four neighboring cells, while the 16 x 17 entry is deliberately reported as an interval because only its 132‑edge lower witness and the published 133 upper bound are certified here. Separately, the 13 x 22 proof excludes 138 ones by reducing to 83 degree profiles, rationally separating 77 of them, and eliminating the remaining six by marked‑row congruences, leave enumeration, modular Gram tests, and exact Farkas certificates. All accepted claims are replayed by standard‑library Python and exact integer/rational arithmetic; floating‑point optimization is used only to discover certificates.

Authors:Kai Peng, Yunzhe Shen, Miao Zhang, Leiye Liu, Wei Ji, Jingjing Li, Yongri Piao, Huchuan Lu
Title: SAMOT: State-Aware Step Modulation and Optimal Transport Matching for Audio-Visual Instance Segmentation
Abstract:
Audio‑Visual Instance Segmentation (AVIS) aims to simultaneously classify, segment, and track sounding objects within video sequences. Unlike Audio‑Visual Semantic Segmentation (AVS), AVIS involves instance‑level modeling across longer video sequences, introducing two key challenges: (1) complex modality‑state changes disrupt long‑range modeling, and (2) substantial structural and distributional discrepancies between modalities hinder precise instance‑level association. Existing methods rely on fixed‑step Transformers and recursive Mamba models, lacking adaptability to modality‑state changes. In addition, methods performing implicit matching ignore the inherent distributional inconsistencies. To address these issues, we propose a framework with Adaptive Dynamic Step Modulation (ADSM) and Optimal Transport‑based Matching Modulation (OT‑MM). ADSM adaptively modulates Mamba step sizes using temporal variation, cross‑modal discrepancy, and historical context, balancing rapid response to modality‑state changes with stable long‑range modeling. OT‑MM explicitly formulates instance‑level cross‑modal matching as an entropy‑regularized optimal transport problem solved via log‑domain Sinkhorn iterations, and further enforces distribution‑level coherence with an MMD regularizer. Extensive experiments demonstrate state‑of‑the‑art performance on the AVIS benchmark (+3.76 FSLA, +2.75 HOTA, +2.58 mAP), verified through comprehensive qualitative visualizations. The code and model are available at https://github.com/happylife‑pk/SAMOT.

Authors:Anthony Miyaguchi, Conor Johnston
Title: DS@GT ARC at Touché: Large Language Models for Retrieval-Augmented Debate
Abstract:
We extend the DS@GT ARC working‑note submission to the Touché 2025 Retrieval‑Augmented Debate task. The task has two subtasks: generating the next utterance in a simulated debate, and evaluating debate responses according to the Gricean maxims of Quantity, Quality, Relation, and Manner. The DS@GT ARC submission consisted of six leading LLMs from three providers through a retrieval‑augmented prompting pipeline. We summarize the results from the working paper and explore whether multi‑LLM evaluator agreement is a reliable proxy for official evaluation performance. The analysis shows that frontier LLM systems are strong response generators, and as evaluators they agree strongly within model families. However this consensus does not reliably track the official evaluation target, with the largest gap on the Quality maxim. The accompanying source code for this paper is located at https://github.com/dsgt‑arc/touche‑2025‑rad and https://github.com/dsgt‑arc/touche‑2025‑rad‑analysis.

Authors:Daniele Molino, Alessio Zoboli, Camillo Maria Caruso, Valerio Guarrasi, Paolo Soda
Title: Compositional Cross-Modality Translation via Whole-Volume Multitask Latent Flow Matching
Abstract:
Cross‑modality medical image translation can reduce the burden of multi‑modal acquisitions, yet the field remains constrained by two coupled limitations: methods operate on 2D slices or 3D patches rather than whole volumes, and train a separate model for each translation task. Both stem from a single cause, the absence of a sufficiently strong volumetric prior, which forces generative models to learn anatomical appearance and cross‑modality mapping simultaneously, an ill‑posed problem at the scale of available paired datasets. We propose to decouple these objectives. A large‑scale pretrained 3D variational autoencoder provides a compact latent representation of volumetric appearance, reducing translation to a conditional flow‑matching problem. This compression makes whole‑volume processing tractable, while a resolution‑aware sampling strategy preserves native anatomical scale. We train a single model jointly across inter‑modality (MRI\toCT, CBCT\toCT) and intra‑modality (MRI\toMRI) tasks over three multi‑center datasets. Across all tasks, whole‑volume processing outperforms its patch‑based counterpart, and the multi‑task model matches task‑specific baselines while replacing N networks with one. Crucially, joint training unlocks capabilities inaccessible to task‑specific approaches: zero‑shot generalization to anatomical regions unseen during training, within 0.15 SSIM of the fully supervised model, and compositional cross‑dataset translation along paths never directly supervised. These results suggest that combining a strong volumetric prior with multitask training is a scalable route toward synthesis systems that generalize beyond their training distribution. Code is available at https://github.com/arco‑group/Whole‑Volume‑Latent‑FM.

Authors:Y Huynh, Duc Thanh Nguyen, Thao Minh Le, Mohamed Abdelrazek
Title: When Does An Extra View Help? Adapting Single-View 3D Reconstruction with Extra Imagery
Abstract:
Reconstruction of 3D objects from a single image is a challenging research problem in computer vision. The key challenge is the lack of critical information from viewpoints to complete 3D structures. Using an additional view may help to resolve the issue. However, there is no mechanism that can integrate the extra view into the single‑view 3D reconstruction principle. We address this challenge by proposing ASV3D, a framework for adapting single‑view 3D object reconstruction to test‑time data with support from one additional image. We introduce two adaptation strategies: (i) a zero‑shot adaptation scheme that leverages the auxiliary image to improve the reconstruction quality of an object without retraining, and (ii) an optimised adaptation scheme that further enhances visual fidelity and cross‑view consistency via contrastive learning. We apply our ASV3D to improve two state‑of‑the‑art single‑view 3D reconstruction pipelines on both benchmark and real‑world datasets. Results demonstrate that our approach consistently improves reconstruction accuracy and robustness under unconstrained multi‑view inputs, outperforming the baselines in both quantitative metrics and human preference. We publish our code and the real‑world object dataset in our project page at https://github.com/YNhuHuynh/ASV3D/tree/main.

Authors:Rui Xu, Hanmo Zhang, Songhua Liu
Title: Staying True to the Origin: Continuous Image Stylization with Smooth Transitions
Abstract:
Recent advances in generative models have achieved remarkable performance in text‑ and image‑conditioned editing. However, preserving the content of a given image while referencing style patterns from another remains challenging, often leading to uncontrollable stylization results. In this paper, we approach image stylization from the perspective of continuous control, aiming to enable modern Diffusion Transformer (DiT)‑based multi‑reference editing models to (1) faithfully preserve the semantic structure of the content image, (2) render strong stylization effects, and (3) smoothly transition between the two. To this end, we propose a simple yet effective two‑stage training strategy along with a style‑strength‑aware spline formulation. Specifically, in the first stage, the model is trained to produce strongly stylized outputs while preserving the content semantics as much as possible. In the second stage, with the base model frozen, we learn a set of anchor projectors that map various stylization strengths into the model parameter space. During inference, by performing style‑strength‑aware spline interpolation in a low‑rank space, our method enables continuous control over stylization strength, even though the model is trained with only a few discrete strength levels. Extensive experiments demonstrate that our method supports precise and continuous manipulation of stylization strength while generating high‑fidelity results with modern DiT models. Project page: https://reychiaro.github.io/StyleController.

Authors:Peng Zhang
Title: SCTD 3.0: Sonar Common Target Detection in the Wild - A Large-Scale, Multi-Scene Dataset from Real Marine Surveys
Abstract:
Synthetic Aperture Sonar (SAS) is core for wide‑area detection of small underwater targets. However, large‑scale, high‑quality SAS datasets are scarce, hindering data‑driven recognition. Existing benchmarks are small and limited to single scenarios, failing to reproduce complex acoustic scattering, diverse seabeds, and multi‑pose imaging in real detection. To fill this gap, we introduce SCTD 3.0 ‑ a large‑scale real‑measured dataset for Sonar Common Target Detection in the Wild in natural waters. It contains over 10,000 high‑quality real SAS image snippets from multi‑frequency systems (240 kHz, 450 kHz, and others), covering ten typical target categories across varied seabed geomorphologies, with multiple observation angles, detection ranges, and frequency bands. We establish a rigorous hierarchical annotation protocol that decouples labeling of intrinsic physical properties, deployment characteristics, and scattering phenomena ‑ covering material, geometry, internal structure, burial state, shadow integrity, specular highlights, edge diffraction, and resonance effects. This enables fine‑grained target characterization. We also construct a multi‑task benchmark for object detection, fine‑grained classification, and attribute prediction, evaluating mainstream deep learning models under cross‑domain, cross‑scene, cross‑frequency, and cross‑view generalization. SCTD 3.0 is expected to provide a critical data cornerstone for robust underwater target perception in open‑water environments. SCTD 3.0 is available at https://github.com/automlresearch/SCTD‑3.0.

Authors:Xuning He, Zinan Sheng, Yongding Tao, Huanyu Liu, Ge Li, Xue Jiang, Yihong Dong
Title: Archer: Adaptive Reuse of Cached Hidden States for Efficient Rollback in Diffusion Language Models
Abstract:
Diffusion language models (DLMs) iteratively refine a sequence, allowing earlier predictions to be revised as context evolves. This rollback capability distinguishes them from irreversible autoregressive generation, but makes inference costly. Every denoising update alters the global context, forcing both prompt and response states to be recomputed even though only response tokens are revisable. Key‑value (KV) caching could reduce this cost, yet conventional caching assumes immutable historical states and is therefore difficult to reconcile with rollback. In this paper, we introduce Adaptive Reuse of Cached Hidden States for Efficient Rollback (Archer), a training‑free KV caching method for rollback‑capable DLMs. Archer asymmetrically keeps the mutable response synchronized with the current hypothesis while reusing prompt K/V within a bounded state neighborhood. Although prompt representations also change under bidirectional attention, their token identities remain fixed; bounded reuse therefore amortizes repeated prompt computation without caching mutable response states. It also delays feedback from tentative tokens, reducing premature reinforcement of transient high‑confidence errors and giving rollback more opportunity to correct them. Our analysis characterizes prompt reuse as a reversibility‑aligned cache boundary, bounds its state‑dependent approximation error, and gives a decoder‑margin condition for preserving full‑refresh decisions. Existing DLM acceleration often trades quality for speed. Archer shifts this frontier, attaining the best mean performance of 33.63% together with a 2.57x mean speedup on the main suite. Across evaluated settings, it improves Pass@1 by up to 3.05 points and reaches up to 2.95x speedup. Controlled analyses connect the quality gain to delayed prompt feedback and validate state‑aware refresh. Our code is available at https://github.com/Hxnng/Archer.

Authors:Pengxiang Cai, Wanchen Lian, Chenyang Liu, Xiaohan Li, Qingyuan Zeng, Jinhong Wang, Jintai Chen
Title: PATH: Next-Interval Prediction via Autoregressive Tree Hierarchy on Tabular Data
Abstract:
Interval prediction aims to achieve a target coverage level while producing intervals that are as short as possible. Many conformal regression pipelines first predict an uncertainty surrogate and then convert it into an interval through calibration or selection. This separation supports coverage calibration, but post hoc rules largely determine the final interval and do not fully use the learned output distribution. We observe that the resulting intervals have inherently hierarchical geometry: an interval can be recursively refined into nested subintervals, and binary trees naturally represent this structure. We formulate this hierarchy as next‑interval prediction and propose PATH, which learns how probability mass flows from each interval to its next nested subintervals. PATH predicts a base leaf distribution and uses an autoregressive decoder to refine branch probabilities. Matching the distribution to the interval hierarchy aligns learning with extraction: PATH accumulates probability over adjacent output intervals and returns the shortest contiguous range reaching a selected mass. We compare PATH with 24 baselines for interval prediction on PATHBench, comprising 56 OpenML regression datasets. PATH substantially shortens the resulting intervals, achieving the lowest mean normalized length, 0.1473, while maintaining mean coverage of 0.9144. These results establish hierarchical output modeling as an effective approach for compact interval prediction on tabular data. Code is publicly available at https://github.com/pxcai/PATH.

Authors:Sankalp Nagaonkar, Rohit Garg, Ankit Raj, Ashish Choithani, Ashutosh Trivedi
Title: Search over the Visual World: Persistent Visual Memory, Layered Indexes, and Source-Grounded Evidence
Abstract:
Most video‑retrieval systems assume a bounded corpus and return ranked files or timestamps. Agents operating over cameras, screens, streams, and archives face a different systems problem: observations arrive continuously; models interpret them at different temporal granularities; context must be selected without replaying the complete visual record; and results must stay connected to inspectable source evidence. We argue that search over such a corpus is an infrastructure problem that cannot be reduced to ranking video files. We develop a conceptual and formal model of search over the visual world built on analyzer‑defined scenes, persistent understanding artifacts, visual memory as coexisting scene spaces over shared source time, and capability‑declared indexes, distinguishing memory (everything retained), context (what is selected for a task), and evidence (the source intervals that ground it). The VideoDB data format (VDB) realizes this model in production, exposed through a typed search surface spanning planned retrieval, stateful investigation, direct access, and grounded synthesis. We contrast this model‑agnostic infrastructure, where segmentation, sampling, model choice, embeddings, and ranking are system decisions and live streams are first‑class sources, with video‑native foundation models offered as fixed APIs. In a semantic‑retrieval comparison against a commercial video‑native engine spanning 9,800+ queries over four public datasets, a pipeline of general‑purpose components achieves higher macro‑averaged Recall@1/@3/@10 (73.09/83.39/91.20 versus 65.75/77.13/89.10), while the baseline is higher at Recall@50 (96.42 versus 96.07). Retrieval quality over the visual world is today governed more by system design than by video‑specific pretraining, and visual‑memory infrastructure can deliver it while keeping playable, source‑grounded evidence first‑class.

Authors:Gerrit Großmann, Sumantrak Mukherjee, Sebastian J. Vollmer
Title: CLAM: Causal Spatial Disaggregation to Infer Local Effects From Coarse Data
Abstract:
Learning fine‑grained spatial patterns from coarse‑resolution data is challenging, especially in causal settings where high‑resolution effects must be inferred from aggregated interventions and outcomes. We introduce CLAM, a method for estimating localized causal effects from coarse observations by exploiting high‑resolution contextual covariates that modulate these effects. By jointly learning the causal mechanism and a disaggregation mapping, CLAM captures interactions that are missed when addressing these problems independently. The method supports localized effect estimation, counterfactual reasoning, and principled outcome disaggregation, and reliably captures spatially varying causal effects across diverse settings. This is particularly relevant for applications such as public health and environmental policy, where decisions are made at broad scales despite substantial local heterogeneity. Code is available at https://github.com/gerritgr/clam

Authors:Fengrong Wan, Chengcan Wu, Ningtao Lyu
Title: SodaMem: Evidence-Grounded Temporal Graph Memory for LLM Agents
Abstract:
Large language model (LLM) agents that assist users over weeks of conversation must remember what is currently true, not merely what was once said. Flat RAG diaries and Markdown logs optimize needle retrieval but under‑serve currency, provenance, and ordered temporal reasoning (Maharana et al. 2024; Wu et al. 2024; Packer et al. 2023; Chhikara et al. 2025). We present SodaMem, an evidence‑grounded temporal graph memory that (i) extracts typed FactEvents with mandatory provenance spans, (ii) persists mention time, occurrence time, and validity with SUPERSEDES/CONTRADICTS/UPDATES edges under hybrid lexical‑dense indexing, and (iii) answers via a planner‑reader loop that gathers citable evidence before composing a final response. On LongMemEval‑S, our store‑of‑record configuration reaches 92.8% accuracy (464/500; best of N=3) at mean 0.00161/question (approximately 18.3k tokens; median 0.00111 / approximately 14.6k) with deepseek‑v4‑flash. We compile public systems with estimable API cost into a cost table and cost‑accuracy map; under these estimates SodaMem sits near the accuracy frontier at Flash‑tier spend and strictly dominates several higher‑cost, lower‑accuracy points. Accuracy uses the same Flash model as reader and judge (self‑grading); costs exclude ingest/judge and cross‑system comparisons are compiled estimates rather than a single‑harness bake‑off.Our code is available at https://github.com/SodaMem/SodaMem

Authors:Pengxiang Cai, Xiaohan Li, Anglin Liu, Qingyuan Zeng, Zexun Li, Jintai Chen
Title: JustLLMGRPO: Radiographic Control for Chest X-Ray Generation
Abstract:
Text‑conditioned chest X‑ray generation aims to synthesize realistic radiographs that faithfully depict specified findings. Existing work has primarily improved quality by updating image generators, implicitly treating prompts as fixed after CXR‑domain adaptation. We show that this generator‑centric view leaves a substantial optimization dimension underexplored. With a CXR‑adapted Sana generator frozen, one‑pass reformulation by an unmodified LLM reduces RadDINO‑FID from 54.225 to 27.572. Prompt analysis shows that the LLM suppresses temporal comparisons, uncertainty, and other non‑renderable report content while emphasizing visible radiographic findings. However, unconstrained reformulation reduces BioViL‑T alignment with source prompts from 0.695 to 0.609. We therefore introduce JustLLMGRPO, which applies standard Group Relative Policy Optimization (GRPO) only to the LLM prompt policy while keeping Sana frozen. Group‑relative radiology‑aware image feedback retains visual focus while preserving source‑prompt alignment. On CheXGenBench, JustLLMGRPO reduces RadDINO‑FID to 26.780, a 50.6% improvement over direct prompting, while maintaining alignment (0.696 versus 0.695). It also achieves state‑of‑the‑art distribution coverage and downstream classification utility. These results show that substantial performance can remain latent in how radiographic information is expressed to an adapted generator. Code is publicly available at https://github.com/pxcai/JustLLMGRPO.

Authors:Nacira Agram, Reda Hmioui, Jan Rems
Title: A cylindrical neural approximation theorem for conditional laws of McKean-Vlasov equations with common noise
Abstract:
We introduce conditional cylindrical neural networks for approximating functionals of conditional laws in McKean‑Vlasov equations with common noise. Fourier moments of the initial law and truncated signatures of the time augmented common noise are mapped by a mixture density network to a Gaussian mixture approximation of the conditional law. A cylindrical neural network then evaluates the target functional through analytic integrals against this predicted measure. Rough path well posedness and stability provide a conditional law map that is continuous in the initial distribution and the rough driver and agrees almost surely with the classical conditional law at the Itô Brownian lift. Combining this continuity with Fourier separation, signature uniqueness, Wasserstein density of Gaussian mixtures, and neural universal approximation, we prove an L^2 universal approximation theorem for continuous square integrable functionals. The numerical study implements the resulting two stage procedure on six examples, including non Gaussian initial laws, nonlinear drift, multiplicative common noise, and a two dimensional state. Independent particle references are used when no closed form law is available. The learned conditional law and functional approximations consistently improve on the empirical particle plug in, and additional experiments examine feature sensitivity, training from one terminal observation per common noise scenario, and Itô‑‑Stratonovich consistency.

Authors:Zakhar Mrykhin, Valentin Malykh
Title: Prompt Embedding Probes (PEP): Hallucination Detection in LLMs from Hidden States
Abstract:
Large language models (LLMs) can generate fluent and useful responses but remain prone to hallucinations. We introduce Prompt Embedding Probes (PEP), a white‑box method for answer‑level hallucination detection from the hidden states of a frozen LLM. PEP extends standard linear probes by augmenting the input with a small number of learnable prompt embeddings. We evaluate PEP on TriviaQA, GSM8K, and MedQA using Qwen3 models at multiple scales. PEP improves hidden‑state‑based detection over standard linear probes in the main in‑distribution setting. We further evaluate PEP for pre‑generation prediction, cross‑model transfer, and out‑of‑distribution generalization. PEP remains effective in the pre‑generation and cross‑model settings, whereas robust cross‑dataset transfer remains difficult. These results show that prompt‑based adaptation can strengthen hidden‑state probing while keeping the backbone frozen and adding only a small number of trainable parameters.

Authors:Jizhou Guo, Yitao Luo
Title: Reinhardt's Maximum-Perimeter Polygon Problem at n=16, 32, and 64: Computer-Assisted Proof Candidates
Abstract:
A convex polygon is called small if its diameter is at most one. Reinhardt proved the universal perimeter bound \mathrmperim(P) \leq U_n := 2n\sin(π/(2n)), and the bound is attained whenever n has a nontrivial odd divisor. The remaining power‑of‑two cases have resisted exact solution beyond n=8. This paper presents computer‑assisted proof candidates for the first three open cases, n=16,32,64. In each case, the candidate theorem asserts uniqueness of the maximizing congruence class. The proof architecture is common to all three cases: pass to the difference body P‑P; encode its reconstruction by a sign code; prove that every global maximizer is saturated, so all difference‑body vertices lie on the unit circle; localize every competitive configuration near the regular angle vector; exhaustively screen the sign codes using exact arithmetic; eliminate all nonwinning dihedral orbits; and prove uniqueness inside the winning code by strong convexity and a quantitative KKT argument. The exact certificates cover 2^15 normalized codes for n=16, 2^31 normalized codes for n=32, and all 2^64 half‑codes for n=64, leaving respectively 16, 96, and 896 survivors before orbit elimination. The accompanying source package contains the verifiers, recorded outputs, and separate computational cross‑checks. These results have not yet received independent human expert review and are therefore deliberately presented as proof candidates rather than literature‑established theorems.

Authors:Anran Zhang, Jiaqi Jiang, jiahui Jin, Yuhan Zhao
Title: Give the Long-tail More SPACE: Promoting Provider Fairness in Next POI Recommendation
Abstract:
Next point‑of‑interest (POI) recommendation predicts users' future destinations from historical mobility sequences and has become a key component of location‑based services. However, mainstream models often concentrate exposure on a small set of popular POIs, leaving long‑tail merchants systematically under‑exposed. While provider fairness has recently attracted increasing attention, directly applying existing provider‑fairness techniques to POI recommendation is problematic: (i) users face execution constraints; and (ii) POIs face resource supply constraints. To address this, we propose SPACE (Supply‑ and Physics‑Aware Conditional Embedding generation), a model‑agnostic framework that improves long‑tail POI exposure via virtual user generation under explicit feasibility and supply control. SPACE consists of three stages: (1) community inference to capture heterogeneous user execution constraints; (2) unbalanced optimal‑transport allocation to decide how many virtual users each tail POI should receive from which communities under POI‑specific supply budgets; and (3) constraint‑guided latent diffusion to generate POI‑conditional, community‑consistent virtual user embeddings. The generated user‑POI pairs can be seamlessly used to train existing recommenders without modifying their architectures. Extensive experiments on three real‑world datasets demonstrate that SPACE substantially improves provider fairness while maintaining and often improving recommendation accuracy across multiple backbone models. Our code is publicly available at https://github.com/Anniran1/SPACE‑main.

Authors:Ling Lin, Yang Bai, Congcong Zhu, Jiangming Shi, Meng Wang, Yang Long, Jingrun Chen, Ling Shao, Huazhu Fu
Title: Advantage-Guided Gate: Reshaping Open-Ended Reasoning for Vision-Based Spatial Intelligence
Abstract:
Multimodal large language models (MLLMs) have demonstrated significant potential in complex spatial scene understanding and reasoning tasks. However, their open‑ended reasoning process is prone to decision errors and error accumulation, leading to instability in answer quality. To address this, we propose an advantage‑guided gating framework that dynamically intervenes in and corrects deviations during the reasoning process. Specifically, we model step‑by‑step reasoning as a finite‑horizon decision process and introduce Monte Carlo value evaluation on the reasoning tree to provide intermediate supervision signals. The framework includes Step‑Advantage Gate and Trajectory‑Advantage Gate, which dynamically select high‑value reasoning steps and high‑quality complete reasoning trajectories, respectively. During training, we perform supervised learning for the gates using reasoning trees generated via multi‑branch sampling, and combine shared‑parameter initialization with task‑specific heads to achieve cross‑task robustness and diversity. During inference, the model greedily selects high‑value prefix reasoning steps while choosing the optimal reasoning head based on the problem type, thereby significantly improving the accuracy of the final answer. Furthermore, we constructed the Reasoning‑Tree‑160k dataset and performed two‑stage learning on it. Extensive experiments demonstrate that this advantage‑guided gating framework effectively enhances the performance of benchmark MLLMs in visual‑based spatial understanding and reasoning tasks. The code is open to the public for research: https://github.com/LingLin‑ll/Advantage‑Guided‑Gate.

Authors:Vasileios Tzouras, Paraskevas Pegios, Lazaros Nalpantidis
Title: AgriField-40K: Adapting Vision Models to Agriculture With Efficient Continual Pretraining
Abstract:
Field‑based agricultural computer vision is important for precision agriculture, yet it largely depends on expensive annotations and costly adaptation of large pretrained models. We introduce AgriField‑40K, a field‑centric dataset curated from 17 public resources and covering diverse crops, weeds, and field conditions. Building on this, we present AgriMAE, a parameter‑efficient continual pretraining baseline that adapts a masked autoencoder pretrained on natural images by training only lightweight adapters. We further explore semantic feature reconstruction as an alternative pretraining objective and evaluate transfer across multiple tasks. AgriMAE consistently improves downstream performance and can match or even outperform full fine‑tuning while using up to 9× fewer trainable parameters, showing that AgriField‑40K is a practical resource for continual pretraining in agricultural vision. Project page: https://dtu‑pas.github.io/agrifield40k/

Authors:Liangliang Zhao, Junying Wang, Danni Yang, Yifan Chang, Bin Fu, Yu Qiao, Bowen Zhou, Yihao Liu
Title: Distilling Physical Priors into Streaming World Models
Abstract:
Streaming world models predict future visual states online while maintaining physically coherent dynamics over long horizons. However, their rollouts often violate basic physical constraints. A common approach distills pretrained bidirectional DiTs into few‑step causal generators. However, this paradigm suffers from two fundamental limitations: generic bidirectional teachers acquire limited physical priors from visually oriented pretraining, and the limited priors suffer further loss during bidirectional‑to‑causal distillation. We present PhyS, a three‑stage framework for distilling physical priors into streaming world models. To acquire physical priors from real‑world interactions, we construct PhyS‑120K, a dataset of 120K real‑world physical‑interaction videos spanning rigid‑body dynamics, soft‑body deformation, fluid phenomena, and phase transitions. Each video is annotated with structured descriptions of object properties and causal state transitions. Physics‑aware supervised fine‑tuning injects the physical priors into a bidirectional 14B DiT teacher, which we then distill into a lightweight 1.3B causal DiT for few‑step autoregressive streaming generation. Finally, we use online reinforcement learning to incentivize the distilled model to generate physically plausible rollouts and further propose Temporal Credit Routing (TCR) to address temporal credit assignment. TCR evaluates physical consistency over overlapping temporal windows and routes the resulting group‑relative advantages to temporally aligned denoising actions. On PhysicsIQ, PhyS improves the Wan2.1‑14B teacher by 18.2% and the Self Forcing, Rolling Forcing, and Causal Forcing by 23.7%, 14.8%, and 31.4%, respectively. Results also improve the physics‑aware video benchmarks VideoPhy, VideoPhy2, and PhyGenBench. The dataset, code, and more sample videos are available on our Project Page.

Authors:Yize Wu, Ke Gao, Ling Li, Yanjun Wu
Title: EasyBalance: Cross-Layer Load Balancing in Distributed MoE Inference
Abstract:
Load Balancing has emerged as a critical problem in expert‑parallel distributed inference of Mixture‑of‑Experts (MoE) models. As routing distributions are typically skewed across experts, devices hosting lighter‑loaded experts must idle to wait for the heaviest during expert computing, leading to inefficiency. Existing load‑balancing approaches primarily rely on expert replication or migration within each layer, which introduce additional overhead and limit their flexibility and scalability. To address this problem, we propose EasyBalance, a cross‑layer load balancing strategy that requires no modifications to the expert‑device mapping, enabling instant adaptability and incurring essentially no additional overhead. Our key insights are that (1) experts of other layers can be viewed as naturally redundant for the current layer, and (2) cross‑layer MoE workloads can be jointly executed to mitigate their individual imbalance. Based on these observations, EasyBalance greedily schedules a subset of cross‑layer workloads to run at each MoE step and defers the remaining workloads for future balancing opportunities, effectively leveraging cross‑layer imbalance mitigation. Extensive experiments across models, tasks, and configurations demonstrate that EasyBalance consistently accelerates distributed MoE inference, reducing GPU idling by mostly over 40%. Code is available at https://github.com/yize‑wu/EasyInfra.

Authors:Baotong Tian, Cynthia Lu, Vincent K. M. Cheung, Ting-Kang Wang, Jonathan Churchill, Zhiyao Duan
Title: VIOLET: High-Fidelity Violin Synthesis with Techniques and Dynamics
Abstract:
Neural synthesis for musical instruments has the potential to revolutionize current practices that use concatenative synthesis and a sample library. However, most research focused on piano synthesis and expressive performance generation; little work has been done on continuously articulated instruments like the violin, let alone rendering them with playing techniques and dynamics. We present VIOLET, a latent‑diffusion framework for controllable violin synthesis, which uses a Diffusion Transformer (DiT) with rectified flow to synthesize high‑fidelity audio from MIDI notes, playing techniques, and continuous dynamics. To train VIOLET, in addition to using a few existing datasets, we curate a new dataset named CSV‑TD, which contains 39 h of 48 kHz synthetic audio and time‑aligned annotations of MIDI notes, note‑level techniques, and continuous dynamics curves. Objective and subjective evaluations show that VIOLET synthesizes violin performances with high technique adherence, accurate pitch and timing alignment, and good dynamics control. It outperforms the current state‑of‑the‑art neural violin synthesis system and approaches a top commercial virtual instrument in terms of technique clarity, naturalness, and dynamics following.

Authors:Amir Sabbaghziarani, Hanting Ye, Maria Gorlatova, Yi Ding
Title: FlexSplat: Flexible Feed-Forward 3D Gaussian Splatting without Point Cloud Correspondence
Abstract:
We present FlexSplat, a feed‑forward framework for novel view synthesis (NVS) from uncalibrated, object‑centric multi‑view image collections. A recent line of query‑based methods reconstructs a compact set of 3D Gaussians by treating them as transformer queries that are refined with multi‑view deformable attention; these methods, however, assume that camera poses are given. FlexSplat removes this assumption: a geometry transformer is trained jointly with the Gaussian decoder to predict per‑image camera parameters and depth, which in turn ground a depth‑guided Gaussian parameterization and a multi‑view deformable cross‑attention that aggregates evidence across all input views into a single, view‑consistent set of primitives. An uncertainty‑weighted depth‑consistency objective lets the jointly trained geometry adapt to the reconstruction task, while the cross‑view consensus formed during decoding absorbs the residual error of the estimated cameras and depth. The representation uses a compact Gaussian budget that is decoupled from the input resolution ‑ unlike pixel‑aligned methods, the primitive count does not grow with the image grid ‑ and is not dictated by the number of views. On ShapeNet‑SRN and Google Scanned Objects (GSO), FlexSplat matches or approaches posed state‑of‑the‑art reconstructors while requiring neither camera poses nor ground‑truth depth, and matches the best perceptual (LPIPS) quality among the compared methods on GSO. Our results indicate that a jointly trained geometry front‑end is sufficient to bring calibration‑free operation to query‑based Gaussian reconstruction while staying within 0.7 dB PSNR of posed methods and matching their perceptual quality.

Authors:Zhongheng Zhou, Yi Sun, Huiguo He, Yuyi Zhang, Peirong Zhang, Yulin Fang, Dezhi Peng, Minghui Liao, Lianwen Jin
Title: TongGuOCR: A Layout-Aware and Token-Augmented OCR MLLM for Chinese Historical Documents
Abstract:
Chinese historical documents preserve valuable cultural heritage, but many collections remain accessible only as scanned page images, preventing full‑text retrieval, collation, and computational analysis. Optical character recognition (OCR) can bridge this gap, but accurate transcription remains challenging because historical documents often contain complex layouts, rare characters, and nontrivial reading orders. We propose TongGuOCR, a layout‑aware and token‑augmented multimodal large language model (MLLM) for OCR of Chinese historical documents. First, a Layout‑Aware Preprocessing module constructs and refines locally coherent recognition blocks to preserve local context while reducing interference across regions. Second, a Token‑Augmented Recognition module augments the transcription target at two complementary levels: character‑level vocabulary expansion gives each rare glyph a direct one‑token representation and shortens its decoding path, while line‑to‑line transition modeling injects discrete spatial displacement tokens that guide the decoder along complex reading paths without requiring precise coordinates. Experiments on two Chinese historical document OCR benchmarks show that TongGuOCR outperforms representative traditional task‑specific OCR models, general‑purpose MLLMs, and OCR‑oriented MLLMs. On the more challenging M5HisDoc benchmark, TongGuOCR achieves 93.76 AR and reduces NED from 10.43 to 6.15 and RO‑ED from 7.53 to 3.49 relative to the best competing score for each metric. An online demo is available at https://jzzh2004.github.io/TongGuOCR.

Authors:Shilei Zeng, Xurui Li, Yaohan Tang, Yu Zhou
Title: DeCo: Zero-Shot Industrial Anomaly Generation through Decoupling and Recoupling
Abstract:
Industrial anomaly inspection is severely hindered by the scarcity of real anomalous data. Zero‑shot industrial anomaly generation addresses this by generating anomalies on specific products without requiring any of their real anomalous images. However, existing methods suffer from two critical limitations, i.e., inaccurate anomaly information acquisition and uncontrolled anomaly‑product fusion. To overcome these challenges, we propose DeCo, which decouples the anomaly structure from its source product, and explicitly recouples it with the normal textures of the target product. During anomaly information acquisition, Dual‑Routing Flow (DR‑Flow) binds the texture‑invariant anomaly structure to an abnormal token, while a parallel constraint, Product‑Invariant Flow (PI‑Flow), prevents the abnormal token from binding the source product. During anomaly‑product fusion, we propose a hybrid injection to recouple the acquired anomaly structure with the target product, and Product Compatibility Correction (PCC) to compensate for the incompatibility between the acquired anomaly structure and the product. Extensive experiments demonstrate that DeCo establishes a new state‑of‑the‑art. Training downstream detection models on our generated data yields massive pixel AP improvements of 5.1% on MVTec AD and 8.2% on VisA. Code is available at https://github.com/HUST‑SLOW/DeCo.

Authors:Ali Janati, Kaoutar El Maghraoui, Xinyi Luo, Wenyuan Shen, Owen Zou, Yankai Mao
Title: Router Sensitivity Under Lightweight Fine-Tuning Identifies Prunable Experts in Mixture-of-Experts Models
Abstract:
Mixture‑of‑Experts (MoE) models decouple total parameters from per‑token compute, but deployment still requires storing every expert. Recent theory shows that pruning experts with the smallest router‑norm changes during fine‑tuning can preserve accuracy, but assumes full fine‑tuning. We test whether lightweight adaptation can recover this signal. We briefly fine‑tune with a parameter‑efficient adapter, rank experts by the induced \ell_2 router change, and prune the least‑changed experts in one shot. On Mixtral‑8×7B‑Instruct (44.83% MMLU‑Pro), router‑only LoRA trains 0.002% of parameters and outperforms all‑module LoRA at matched rank with half the experts removed (27.54% vs. 24.42%); signal quality declines as adaptation spreads to attention and expert weights. Accuracy improves monotonically with LoRA rank, reaching 28.76%. IA3, which leaves router weights frozen, matches direct router adaptation, whereas unconstrained additive adapters degrade the signal. Router‑guided MMLU‑Pro accuracy decays quasi‑linearly rather than collapsing, remains nearly 1.8 times that of magnitude‑based or random pruning at maximal compression, and reduces memory by 49% and per‑token latency by 37%. At 25% compression, retention is competitive with methods using full activation statistics. The criterion also transfers to Qwen1.5‑MoE fine‑tuned for mathematics, retaining 49.7% mean accuracy over eleven benchmarks with half the experts removed while random pruning falls to single digits. Router sensitivity under lightweight fine‑tuning therefore makes provably motivated expert pruning practical at scale.

Authors:Kaigen Go, Yinlai Jiang, Hiroshi Yokoi, Shunta Togo
Title: A Mixed-Stiffness Anthropomimetic Fingertip Broadens the Operating Range for Coin Grasping
Abstract:
Robotic grasping of thin, flat objects such as coins on hard surfaces remains challenging because conventional methods require reorienting the object, accessing its underside, or adding a dedicated nail mechanism. We previously showed that a rigid nail arrests soft‑pad deformation and thereby forms a geometric constraint that improves precision grasping. Here we asked whether an additional constraint‑forming boundary, created within the pad by material choice rather than by anatomy, could extend the conditions under which that constraint holds. We fabricated anthropomimetic fingertips with Shore E10 silicone at the center and Shore A60 at the sides, and compared them with uniformly soft E10 fingertips. An automated apparatus performed an oblique rotational tip pinch in which the pad engaged the coin's lateral surface, lifting it from flush contact with no gap beneath it. Over variations in horizontal approach distances, vertical finger displacements, and index‑finger rotation, the mixed‑stiffness pair maintained high success rates across more tested settings than the uniform pair during both geometric‑constraint formation and the transition to a stable grasp. The nail‑free pair failed in all 36 conditions of Experiment 1‑1. However, the uniform pair performed better when coin position along the finger axis was varied, a condition‑dependent trade‑off. After tuning for coin size, both fingertip types grasped all six Japanese denominations. These results suggest that the operating range for thin‑object grasping depends not only on pad softness but also on where stiffness is placed within a nail‑supported pad, making boundary placement a candidate fingertip design variable.

Authors:Zihua Yang, Zhencheng Xie, Junyang Chen, Liang Xie, Yiqun Zhang, Mengke Li, Yang Lu
Title: GRACE: LLM-Grounded Semantic Metric Spaces for Scalable Mixed-Data Clustering
Abstract:
Clustering mixed tabular data requires a unified metric space to bridge the inherent heterogeneity between continuous numerical measurements and discrete categorical symbols. Traditionally, algorithms rely entirely on dataset‑internal statistics to estimate categorical relationships, which confines the learned metric to empirical co‑occurrences and ignores conceptually obvious yet statistically unobserved affinities. Although LLMs offer external world knowledge, applying their text‑centric reasoning to highly abstract tabular concepts presents significant challenges. Bridging this modality gap to construct a semantically complete metric typically requires embedding LLMs into iterative metric learning loops to dynamically optimize cross‑modality representations. This incurs intractable computational overhead, forcing a compromise between semantic enrichment and scalability. Therefore, we propose GRACE, an LLM‑grounded framework for scalable mixed‑data clustering. GRACE shifts semantic acquisition to the attribute‑value level via a multi‑perspective LLM querying strategy, mapping heterogeneous values into knowledge‑informed descriptions. Crucially, this one‑shot grounding extracts general‑purpose semantic representations that embed heterogeneous attributes into a unified space, decoupling expensive LLM invocation from iterative optimization. Furthermore, GRACE cross‑validates these external semantics against dataset‑internal statistical evidence to ensure alignment with the dataset‑specific cluster structure. Ultimately, GRACE matches the scalability of conventional statistics‑driven baselines while achieving superior clustering accuracy and conceptual interpretability over 11 competing methods. The source code is available at https://github.com/develop‑yang/GRACE‑GRACE‑A

Authors:Donghu Kim, Youngdo Lee, Hojoon Lee, Johan Obando-Ceron, Byungkun Lee, Aaron Courville, Pablo Samuel Castro, Jaegul Choo, Clare Lyle
Title: V-Simba: Unleashing the Architectural Potential of RL in Visual Continuous Control
Abstract:
Improving sample efficiency remains a core challenge in reinforcement learning (RL), especially in real‑world settings like robotics, where data collection is costly. This challenge is pronounced in visual RL, where high‑dimensional inputs often obscure learning signals. While prior work in visual RL has focused on algorithmic solutions, such as better dynamics models or exploration strategies, recent advances in state‑based RL show that architectural design alone can lead to significant gains in sample efficiency. This raises an important question: Can these architectural principles transfer to visual RL? In response, we introduce V‑Simba, a simple yet effective visual RL architecture inspired by the Simba architecture from state‑based RL. Built on top of Soft Actor‑Critic (SAC) with data augmentation, V‑Simba modifies the architecture by adding normalization layers to stabilize training and using pointwise convolutions to reduce computation. Despite its simplicity, V‑Simba matches or outperforms the state‑of‑the‑art methods across the DMC, Adroit, and Meta‑World benchmarks, while being more computationally efficient than DrQ‑v2. We make our code publicly available at https://github.com/DAVIAN‑Robotics/V‑Simba.

Authors:Shilei Zeng, Linxin Guan, Xurui Li, Yaohan Tang, Yu Zhou
Title: UniScale: Arbitrary-Scale Industrial Anomaly Generation
Abstract:
Industrial anomaly inspection faces a major challenge due to the lack of real‑world anomaly samples. While generative models are used to create anomaly data, existing methods still struggle when handling small‑scale anomalies. This failure occurs because extreme downsampling in diffusion models causes the information of small anomalies to be lost in the latent space. To address this, we introduce UniScale, a unified training and inference framework for high‑fidelity industrial anomaly generation across arbitrary scales. During training, we introduce an Error‑Suppressed Multi‑Scale Training (EMT) strategy, which enables the model to learn the rich location‑aware textures of anomalies, while suppressing upsampling‑induced interpolation errors in texture acquisition, ensuring the model is capable of learning small‑scale anomalies, while remaining effective for regular scale anomalies. For inference, we propose Generation‑then‑Fusion Denoising. It decouples anomaly generation from background integration, preventing small anomalies from being overwhelmed. Extensive experiments demonstrate that our method outperforms state‑of‑the‑art competitors in both anomaly generation quality and downstream detection performance. It achieves a relative IS(a) improvement of 45.86% (from 1.81 to 2.64) on VisA and 37.70% (from 1.22 to 1.68) on MVTec AD 2, while also improving the downstream pixel‑level IoU by 4.22% on VisA and AUROC by 6.55% on MVTec AD 2. Code is available at https://github.com/HUST‑SLOW/UniScale.

Authors:Debopriyo Banerjee, Kapil Rajesh Kavitha, Angana Borah, Xudong Han, Yuxia Wang, Parameswari Krishnamurthy, Utkarsh Agarwal, Atharva Kulkarni, Swaran Lata, Ayush Munot, Dhruv Sahnan, Aaryamonvikram Singh, Preslav Nakov, Monojit Choudhury
Title: SurakshaEval: An Indic Safety Benchmark for Multilingual LLMs
Abstract:
Existing safety evaluation datasets for large language models (LLMs) predominantly focus on English and Western contexts, often overlooking the linguistic diversity and culturally grounded safety risks present in other languages. To address this gap, we introduce SurakshaEval, a novel safety benchmark composed of human‑written prompts spanning real‑world scenarios, explicitly designed for ten major Indian languages ‑ Assamese, Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Punjabi, Tamil, and Telugu, along with English. SurakshaEval includes both generic prompts common across India and region‑ and language‑specific prompts that capture localized sociocultural sensitivities. We benchmark a broad range of state‑of‑the‑art LLMs on SurakshaEval, establish baseline safety performance, and identify recurring failure modes, including over‑refusal, missed detection of implicit bias, and insufficient contextual awareness in regionally sensitive settings. Our results show that even strong multilingual LLMs struggle to reliably meet nuanced safety requirements when operating in Indic languages, particularly in native scripts. These findings highlight the urgent need for safety evaluation frameworks that incorporate region‑specific data and structured assessment protocols, enabling the development and deployment of AI systems that operate securely, ethically, and in alignment with diverse societal values. Our code and data are available at https://github.com/debobanerjee/SurakshaEval. Warning: This paper contains text that may be offensive or unsafe.

Authors:Zhuangfan Huang, Chusheng Fang, Xiaosong Li, Yang Liua, Xiaoqi Cheng, Haishu Tan
Title: IRPol-Fuse: Energy-structure coordination for infrared polarization fusion under low visibility
Abstract:
Robust perception under low‑visibility conditions requires fused imagery that jointly preserves infrared thermal saliency and polarization‑derived structural details. However, existing infrared‑polarization image fusion (IPIF) methods often overemphasize dominant infrared responses, causing weak yet informative polarization textures in dark regions to be suppressed. To address this issue, we propose IRPol‑Fuse, an energy‑structure coordinated IPIF framework for challenging low‑visibility scenarios. The proposed framework contains three key modules: Polarization Attention Fusion for adaptive infrared‑polarization allocation, Infrared Highlight Injector for highlight‑guided infrared preservation, and Polarization Texture Injector for polarization texture restoration and fine‑detail recovery. We further construct LI‑PI, a dedicated infrared‑polarization evaluation dataset for low‑visibility and visually concealed scenes. Experiments on LI‑PI and the public LDDRS dataset demonstrate that IRPol‑Fuse achieves favorable performance in thermal target preservation, structural detail recovery, and visual naturalness. Region‑aware evaluation and downstream object detection further verify that the proposed energy‑structure coordination strategy effectively preserves both infrared target saliency and polarization‑derived structural information. Code is available at https://github.com/1hzf/IRPolar‑Fuse .

Authors:Ilia Azizi
Title: Conformal Calibration for Multi-Modal Regression with Missing Modalities
Abstract:
Prediction intervals for multi‑modal regression with tabular variables, text, images, or other input sources are difficult to calibrate when those sources disagree or one is missing. A single global quantile averages these regimes together instead of calibrating to the modality pattern observed at test time. We address this through a modality‑aware conformal calibration layer. The layer trains or reuses one predictor per modality, computes a disagreement score from their predictions, and uses that score in split conformal calibration under a strict split protocol. We use the score in two complementary ways. First, a continuous disagreement‑scaled method reallocates interval width across examples while preserving the usual marginal split‑conformal guarantee. Second, a Mondrian (stratified) method calibrates within groups defined by disagreement or modality availability fixed before calibration, giving group guarantees under joint exchangeability of the calibration and test examples. Across four multi‑modal datasets, the disagreement‑scaled layer matches or improves the marginal conformal baseline in 59 of 60 paired runs for interval continuous ranked probability score (CRPS) and in 52 of 60 for interval width, while keeping empirical coverage near the 95% target. In stress tests with missing modalities, mask‑matched recalibration recovers up to 19.5 percentage points of coverage in the hardest fixed‑mask regime. The result is a simple, model‑agnostic reliability layer for multi‑modal regression systems. A project page is available at https://unco3892.github.io/modality‑aware‑conformal.

Authors:Zhongpai Gao, Benjamin Planche, Meng Zheng, Anwesa Choudhuri, Chaoyi Zhou, Terrence Chen, Ziyan Wu
Title: XClipGS: Exact Half-Space Clipping for Medical Volume Gaussian Splatting
Abstract:
Gaussian‑splatting proxies enable interactive rendering of volumetric medical scans, but a clipping plane exposes anatomy not constrained by external‑view training and intersects primitives that conventional splatting can only keep or drop whole. We present XClipGS (eXact Clipping), which treats these as two separate problems: the render‑time clip operator and supervision of the hidden interior. Under the local affine model used by EWA splatting, the ray integral of a half‑space‑restricted Gaussian factorizes exactly into its ordinary 2D footprint and a conditional Gaussian CDF whose argument is affine in pixel coordinates. The resulting closed‑form per‑pixel operator introduces no learned clipping parameters or auxiliary network and remains differentiable with respect to the primitive and plane. We use multi‑distance reference views with varied clipping‑plane axes and offsets to supervise the interior through the same operator. We also introduce a paired clipped/unclipped cut‑face protocol with difference‑referenced cut error (CDE) and culled‑side leakage (Leak), because global image metrics dilute errors near the plane. On eight CT and MRI volumes with plane offsets not used for training, XClipGS attains the highest PSNR on every volume (33.56 versus 32.34 dB for ClipGS) while rendering at over 650 FPS, far above real time, versus 278 FPS. On voxel‑axis cut‑face views, it raises average band SSIM from 0.809 to 0.860 and leaks roughly 40 times less. Without retraining, it also achieves the best average across all four metrics on arbitrary‑normal planes; on a fixed interior, it matches RaRa's face fidelity with about 16 times less leakage. Project page: https://gaozhongpai.github.io/XClipGS/

Authors:Guilherme Vieira Neto, Marcos Eduardo Valle
Title: Ghost Features and Spooky Transfer Learning for Hypercomplex-Valued Neural Networks
Abstract:
Hypercomplex numbers extend the concept of complex numbers by introducing additional imaginary components. Besides increasing dimensionality, operations on the imaginary parts provide algebraic and geometrical properties that can be beneficial for solving machine learning problems. In this paper, we show how to create hypercomplex‑valued neural network layers where the real part corresponds to the output of a traditional real‑valued layer. The additional imaginary parts of these hypercomplex‑valued layers produce what we call ``ghost features,'' which contain enhanced information that is not present in the output of the real‑valued layer. Moreover, ghost features can be effectively integrated into a trained neural network through a process we refer to as ``spooky transfer learning.'' This approach allows us to harness the richness of ghost features, leading to more efficient neural networks. The source code and Jupyter Notebook are available at https://github.com/mevalle/v‑nets/.

Authors:Liam Chalcroft
Title: Tokenizer Generator Coupling in Medical Image Generation
Abstract:
Latent medical image generators usually treat the tokenizer as fixed preprocessing. We test whether this separation is valid in a controlled ChestMNIST study at 64x64, crossing discrete tokenizers, generator families, and sampler settings under a shared latent grid, with continuous‑latent reference cells. In this controlled setting, rankings depend jointly on the tokenizer, generator, and sampler: the best quantizer changes with the generator, and validation‑based sampler selection changes the apparent generator ranking. We retrain the vocabulary‑1024 interaction block at three seeds and the interaction survives (6 of 9 pairwise quantizer comparisons exceed three seed standard deviations), and we scope the wider single‑seed grid accordingly. Reconstruction PSNR alone is not a reliable selection criterion; we instead introduce a generator‑free statistic, neighbour‑conditional predictive gain, that separates the quantizer families by downstream generation quality (rank‑AUC 1.00) where reconstruction PSNR and marginal token entropy do not. On LFQ‑1024, retuning D3PM and SE‑D3PM (selected on a held‑out validation split) moves them from default FID‑192 0.44/0.41 to 0.09/0.10 at lower NFE, replicated across seeds; the continuous references were not given an equivalent sampler sweep. We report FID‑192 as an internal ranking metric; it ranks consistently with standard FID‑2048 (Spearman 0.80) and with a label‑free classifier two‑sample test (0.78). We interpret these results through a rate‑distortion‑modelability framing, where modelability is conditional on the generator, sampler, and inference budget. All experiments are at 64x64 on low‑resolution medical‑style images, unconditional, and evaluated with non‑clinical FID‑based metrics, and we scope every claim to that setting. Code: https://github.com/liamchalcroft/medtokenizers and https://github.com/liamchalcroft/medlatents.

Authors:Ziqiao Yu
Title: SpikeWorld: Fast-State Adaptation for Frozen Spiking World Models
Abstract:
A predictive model receives a self‑supervised signal whenever the consequence of an action is observed. Using that signal after deployment is difficult when dynamics and semantics share parameters: freezing prevents adaptation, whereas weight updates require optimizer state and may alter the learned representation. Here we introduce SpikeWorld, a 1.45M‑parameter sparse spiking model jointly trained for heterogeneous sensory prediction, semantics, image‑text binding and action‑conditioned dynamics. At deployment, all trained parameters are frozen. Delayed next‑state residuals update two external paths: cumulative fixed‑bank losses select the bounded action correction, while route‑specific residual matrices refine next‑state prediction. Neither path uses labels, teacher outputs, rewards, success signals or the true shift value. Joint optimization improves action next‑state MSE by 17.10% while also improving multimodal prediction, semantic accuracy and image‑text retrieval. On held‑out shear and attenuation streams, the combined external state improves aggregate prediction by 5.48% and 30.01%; its fixed‑bank action path improves tracking by 24.20% and 3.94%, respectively. In a six‑arm study comprising 450 new Meta‑World trajectories (75 per arm), SpikeWorld raises frozen‑policy reward by 7.90 (95% CI [2.48, 14.06]); the 13.33‑point success difference is descriptive (CI [0, 40]). For identical sensory inputs, model parameters and inherited semantic outputs remain bitwise unchanged. A 16‑byte RLS estimator obtains the highest non‑oracle reward on linear attenuation, showing that the contribution is not superior linear identification, but its integration with a frozen multimodal spiking checkpoint. Reference code is publicly available at https://github.com/Oooorca/SpikeWorld.

Authors:Quang Minh Dinh, Tuan Kiet Doan
Title: CosmosAlign: Adapting a World Foundation Model for Generative Traffic Video Forecasting
Abstract:
Generative traffic video forecasting aims to synthesize long‑horizon, temporally coherent future videos of traffic scenes from a short observation history and textual descriptions. In this paper, we present CosmosAlign, a generative traffic video forecasting framework built upon the pretrained Cosmos3‑Nano world foundation model. Our approach is motivated by the observation that successfully adapting large pretrained world models to downstream forecasting tasks depends primarily on distribution alignment rather than increased model capacity. To this end, we propose a two‑stage LoRA adaptation strategy that first aligns the conditioning‑mode distribution with the target forecasting task, and then aligns the training captions with the model's native structured prompting interface through an LLM‑based re‑captioning pipeline. During inference, we further improve prediction quality using a fully training‑free procedure consisting of consensus‑based medoid sample selection and motion‑adaptive blending of static scene regions. CosmosAlign achieves a final score of 76.49 on the AI City Challenge 2026 Track 5 benchmark, ranking first on the final leaderboard. Our code is publicly available at https://quangminhdinh.github.io/CosmosAlign/.

Authors:Yeeun Choi, Youngbeom Yoo, Joon-Young Lee, Hyolim Kang, Seon Joo Kim
Title: Keep It Simple: Multi-Key Episodic Memory Retrieval for Ultra-Long Video Understanding
Abstract:
When videos extend from hours to days, directly processing them end‑to‑end becomes impractical for current Multi‑modal Large Language Models (MLLMs). This ultra‑long setting necessitates a two‑stage paradigm: query‑agnostic memory construction followed by retrieval‑based inference. Prior work invests in complex memory construction to pre‑model high‑level relations in videos, despite not knowing the downstream query at build time. We instead prioritize high‑recall retrievability during memory building, and defer query‑specific, high‑level relation composition to inference time. To this end, we propose MERIT(Multi‑key Episodic Retrieval with Inference‑time Temporal expansion), a simple yet effective agentic framework for ultra‑long video understanding. First, we formulate an episodic multi‑key representation that enables precise retrieval of fine‑grained memories through a simple key‑matching mechanism. Second, we introduce a neighbor filtering mechanism to capture broader semantic context without the massive computational overhead of global memory construction. This is achieved by expanding the temporal scope exclusively around the retrieved segments at inference time. By leveraging simple key‑matching with this on‑demand temporal expansion, MERIT achieves state‑of‑the‑art performance across three long‑video benchmarks: EgoLifeQA, LVBench, and Video‑MME (Long).

Authors:Changzhi Liu, Yilun Liu, Sikuan Yan, Volker Tresp, Yunpu Ma
Title: Mendel Gödel Machine: Recursive Self-Improving Coding Agents via Comparative Evolution
Abstract:
Self‑improving coding agents that iteratively rewrite their own source code have demonstrated impressive performance on coding tasks. However, existing solutions generally derive self‑modification from a single failure trajectory at a time, overlooking rich comparative signals available in the agent's expanding archive of past attempts. According to Mendelian principles of controlled inheritance, we introduce Mendel Gödel Machine (MGM). In addition to the general single‑trajectory clonal mutation, MGM includes two new types of self‑modification that better utilizes evidences accumulated: the reaction‑norm mutation edits an agent based on its trajectories on multiple tasks simultaneously, and the cross‑lineage hybridization edits an agent using the trajectory of a reference agent from another lineage on the same task. Under an additive fitness landscape model, we prove theoretically and demonstrate via controlled surrogate simulation that the new strategies facilitate a faster and better convergence over single‑trajectory baselines. Experiments on SWE‑bench and Polyglot confirm MGM's consistent improvement in performance, efficiency, and generalizability.

Authors:Shuai Lyu, Yuning Gong, Ruiling Gao, Xiaoran Shang, Zhonghong Ou, Ping Zong, Yifan Zhu, Yuan Sun, Yang Qin, Peng Hu
Title: Multi-Branch Policy Optimization for Multimodal Large Language Models
Abstract:
Group‑based reinforcement learning methods for multimodal large language models typically rely on trajectory‑level credit assignment that applies a single advantage to all tokens in a response. However, multimodal reasoning involves substantially higher perceptual uncertainty than text‑only settings, where the model must repeatedly re‑examine visual information to verify intermediate interpretations, and different visual groundings can lead to divergent reasoning paths, making such uniform credit assignment particularly inadequate and causing relative advantages to progressively degenerate toward zero. To address these challenges, we propose Multi‑Branch Policy Optimization (MBPO), a tree‑based framework that constructs reasoning trees at vision‑language decision boundaries, enabling sibling branches to explore diverse visual hypotheses and assigning segment‑level credit through branch‑relative advantages. We further introduce a temporal replay buffer to reuse informative segments while controlling policy staleness. Experiments on several multimodal reasoning benchmarks show that MBPO outperforms representative baselines, improving both learning signal quality and optimization efficiency. The code is publicly available at https://github.com/ShuaiLyu0110/MBPO.

Authors:Felix Schaller
Title: Open-World Hierarchical Perception: Taxonomic Abstraction over Class-Agnostic Proposals for the Safe Handling of Out-of-Vocabulary Road Objects
Abstract:
A closed‑set detector for autonomous driving must assign every object one of a fixed set of labels. On an object outside that set (a horse‑drawn carriage, road debris, livestock on a rural road) it can only force a confident but wrong specific label or drop the object. Prior work in this series replaced the flat label set with a hierarchical taxonomy and a runtime abstraction rule, but evaluated it only on the boxes a closed detector already produces. This paper takes the layer open‑world: we place taxonomic abstraction on top of class‑agnostic region proposals so objects the closed detector never boxes can still be classified or flagged; we report a feasibility study of three open‑world signals (class‑agnostic segmentation, appearance‑based out‑of‑distribution scoring, monocular depth) that shows why no single 2D cue suffices and how they compose; and we run the evaluation the earlier papers could not, a ground‑truth leave‑classes‑out benchmark on real annotated objects. Holding out seven COCO classes and classifying their 235 ground‑truth crops, a flat closed head emits a confident wrong specific label 100% of the time (37% of them in the wrong super‑category, e.g. an animal named as a vehicle), whereas the hierarchical layer emits zero confident wrong specific labels and safely handles 94% of the objects (a correct super‑category, or an explicit UNKNOWN OBSTACLE). We are explicit that this is a safety result, not a specificity one: the correct super‑category is recovered only 26% of the time and the remaining 69% are conservatively flagged unknown. The contribution is an open‑world perception layer that never makes a confident categorical mistake on an out‑of‑vocabulary object, together with an honest account of its cost.

Authors:Seulchan Lee, Leesai Park, Minhyeong Kang, Sanghyun Kim
Title: Projection-Retraction MPPI: Exact Constraint-Manifold Control for Manipulators
Abstract:
Model Predictive Path Integral (MPPI) control is widely used in manipulation for its gradient‑free, parallel handling of non‑convex costs. Manipulation tasks, however, often impose constraints that hold throughout the motion: a closed kinematic chain that two grasping arms keep exactly, or joint limits and obstacle clearances that are never crossed. MPPI handles such constraints only through the cost, as soft penalties that hold approximately and fail under a strong task cost. To address this, we propose Projection‑Retraction MPPI (PR‑MPPI), which enforces the constraints inside the sampled dynamics. At every rollout step, the sampled velocity is projected to satisfy both constraint types: the equality restricts it to a subspace, and each inequality to a half‑space within that subspace, so inequality handling never breaks the equality. This projection, however, satisfies the constraints only to first order, and a finite step leaves a small drift off the equality. Therefore, we retract the returned command back onto the constraint to numerical tolerance and independent of task weighting. We validate PR‑MPPI on 14‑DoF dual‑arm systems. In simulation, the returned commands satisfy the closed‑chain equality to numerical tolerance through a joint‑limit stress test and randomized obstacle avoidance. On real hardware, the arms of a Unitree H1‑2 humanoid reactively avoid a moving obstacle. Code and experiment videos are available at https://rcilab.github.io/prmppi.

Authors:Jinlong Yang, Jinke Wu, Lizilin, Yao Zhou
Title: BRACE: Taming Sharp Irregularities via Barycentric Rational Forecasting for Fast Diffusion Transformers Inference
Abstract:
Diffusion Transformers (DiTs) have demonstrated exceptional performance in high‑fidelity image and video generation. To alleviate their massive computational overhead, temporal feature caching has been proposed to bypass redundant computations. However, existing cache‑then‑forecast methods driven by derivative‑based polynomials often cause severe quality degradation under high acceleration due to unstable long‑step predictions. To address this bottleneck, we propose Barycentric Rational Forecasting with Chebyshev Enhancement (BRACE). Motivated by the observation that DiT feature trajectories are globally smooth yet frequently exhibit sharp irregularities and local non‑smoothness, BRACE shifts the paradigm from derivative‑driven polynomial extrapolation to feature‑driven rational forecasting. Specifically, it maintains a local sliding window to cache sparse historical features and leverages adapted Chebyshev weights to formulate a barycentric rational function, directly aggregating these raw features to ensure numerical stability. Extensive experiments demonstrate that BRACE achieves state‑of‑the‑art quality‑efficiency trade‑offs across various DiT architectures with negligible computational overhead.

Authors:Xuan Yao, Yuze Zhu, Junyu Gao, Zongmeng Wang, Changsheng Xu
Title: SC$^{2}$-WM: A Self-Correcting World Model with Closed-Loop Feedback for Vision-and-Language Navigation in Continuous Environments
Abstract:
Vision‑and‑Language Navigation in Continuous Environments (VLN‑CE) requires agents to make fine‑grained navigation decisions under partial observability. However, most existing methods rely on open‑loop execution, lacking mechanisms to detect and correct internal state drift during inference. We propose SC^2‑WM, a self‑correcting world model framework that introduces internal feedback for closed‑loop decision making in VLN‑CE. Our method derives feedback from world‑model foresight to perform state‑level plan refinement before action execution. To handle challenging scenarios, we further introduce conditional world‑aware adaptation, which enables model‑level correction by selectively updating the world model at test time when feedback indicates model capacity insufficiency. Experiments on standard VLN‑CE benchmarks demonstrate improved navigation robustness and generalization. Our code is available at https://github.com/sunrise‑ikun/SC2_WM.

Authors:Xinmeng Yu, Jiaxin Gao, Jianguo Zhang, Dongmei Jiang, Ran Cheng
Title: AutoPSO: A Metaframework for Automated Particle Swarm Optimization
Abstract:
Particle swarm optimization (PSO) is a widely used metaheuristic, prized for its simplicity and small parameter set. Although decades of research have produced numerous PSO variants that improve performance by modifying key components (e.g., parameter schedules, swarm topologies, or updating rules), two fundamental challenges persist. First, most existing approaches are problem‑specific and hand‑crafted, leading to poor cross‑task generalization and forcing practitioners to navigate an impractically large design space, which also hinders systematic reuse of prior effective mechanisms. Second, mainstream implementations remain CPU‑bound, constraining scalability and substantially increasing computational cost in real‑world applications. To address these challenges, we propose AutoPSO, a highly automated metaframework for constructing customized PSO algorithms. AutoPSO formulates PSO‑based optimization as a bi‑level process: an outer search explores the joint space of effective PSO components, while an inner loop instantiates candidate variants to solve the target task and provide feedback. The outer search operates over a curated, open‑design component pool, supporting flexible replacement of the component set and the outer optimizer. Crucially, by leveraging EvoX for population tensorization and batched evaluations, AutoPSO can efficiently assess thousands of particles within practical time budgets. Comprehensive experiments on numerical benchmarks and neuroevolution robotic control tasks demonstrate that AutoPSO consistently discovers novel PSO variants that significantly outperform strong baselines. Ablation and scalability studies further highlight the contribution of individual algorithmic components and confirm that AutoPSO achieves increasing performance gains with larger swarm sizes. Code is available at https://github.com/EMI‑Group/autopso.

Authors:Cheng Ruoxi, Ma Haoxuan, Zhang Hongyi, Zhang Junming, Duan Ranjie, Xia Qiaolin, Wang Hao, Lu Yu, Shi Haibo, Ma Xingjun
Title: Search-G1: Grounded Search Agents via Representation-Based Intrinsic Rewards
Abstract:
Search‑augmented language agents should retrieve external information only when necessary and ground their answers in retrieved evidence. Existing external rewards provide either sparse outcome supervision or richer feedback from process annotations and LLM judges. Outcome rewards scale readily but cannot distinguish grounded retrieval from redundant search, whereas richer signals require costly annotation or inference during training. Internal rewards based on policy‑side signals such as entropy, likelihood, or information gain are graded and inexpensive to evaluate, yet mainly reflect model confidence rather than evidence grounding. We propose Search‑G1, a representation‑based intrinsic reward framework that measures the operational grounding of an agent's answers through two intervention‑calibrated readouts. A prompt‑state readout predicts closed‑book sufficiency, whose complement defines policy‑relative retrieval necessity; an answer‑commit readout estimates evidence reliance from answer‑stage sensitivity to evidence deletion. Together, they provide additional credit to correct searched trajectories when retrieval is estimated necessary and the answer is evidence‑sensitive, favor correct direct answers when closed‑book knowledge suffices, and penalize repeated search. After calibration, reward scoring requires neither process annotations nor LLM‑as‑judge inference during policy optimization. Because reinforcement learning changes policy representations, Search‑G1 periodically refits both readouts on trajectories from the latest checkpoint, allowing the reward to co‑evolve with the policy. Experiments across multiple search‑based question‑answering benchmarks and two model scales show that Search‑G1 improves the grounding‑‑search‑cost trade‑off, producing shorter response‑side trajectories at competitive task accuracy. Code is available at https://github.com/Rosy0912/Search‑G1.

Authors:Pengfei Zhou, Jiajun Song, Zhiwei Tang, Yixing Ma, Xiaopeng Peng, Donghui Si, Yuhang Xu, Huiqi Song, Yiyuan Miao, Yichen Qian, Weihua Chen, Wangbo Zhao, Bohan Zhuang, Jiasheng Tang, Yang You
Title: Unified Hallucination Fuzzing for Multimodal Large Language Models
Abstract:
Hallucination remains a persistent challenge for Multimodal Large Language Models (MLLMs), severely limiting their reliability in high‑stakes applications. Existing evaluations, predominantly based on static benchmarks, suffer from narrow taxonomical coverage and rapid performance saturation, failing to reflect model robustness in evolving real‑world scenarios. To bridge this gap, we present a systematic evaluation framework integrating a comprehensive benchmark with self‑evolving stress testing. First, we introduce UniHall, a fine‑grained dataset grounded in a unified taxonomy spanning Object, Instruction, and Knowledge dimensions. Second, to address benchmark saturation, we propose Self‑Adaptive Multimodal Fuzzing (SAMF), a self‑adaptive framework that employs evolutionary mutation strategies to explore the boundaries of model hallucinations. Crucially, to ensure reliable assessment of dynamic inputs, SAMF incorporates a structured metric suite driven by an ensemble of multi‑modal oracles. Our extensive experiments reveal that state‑of‑the‑art MLLMs exhibit significant performance degradation under fuzzing compared to conventional settings, exposing a dissociation between reasoning capabilities and factual grounding. Furthermore, we identify a helpfulness‑hallucination trade‑off, where reinforcement learning alignment inadvertently exacerbates sycophancy in instruction‑following tasks. The framework, code and benchmark are available at https://github.com/LanceZPF/EvalHall.

Authors:M. Waleed Kadous, Benjamin Olsen
Title: JaleesBench: Are AI Assistants Good Spiritual Company?
Abstract:
Large language models are already advisors to millions of people of faith who bring them real decisions. The pressing question for a person of faith is not what a model knows or professes but what its counsel does to the person who receives it. We introduce JaleesBench, which measures whether an AI agent is a righteous companion, judged by the residue an exchange leaves on the user, in the manner of the perfume‑seller and the blacksmith. It comprises 140 two‑turn scenarios drawn from a classical compilation organized by virtue (Riyad al‑Salihin), under six adversarial pressures and three framings, scored by two frontier judges against each scenario's own supporting texts. Across eight systems: (1) generic frontier models are only middling companions out of the box but a one‑page guide makes them genuinely good ones, on par with the domain‑tuned assistant: the frontier APIs climb from +0.28/+0.23 to a Guided +0.84‑0.87, so most of the expert's edge is companionship instruction that fits in a prompt; (2) every system caves under relational pressure, insistence and personal appeal; (3) the domain‑tuned assistant's advantage is overwhelmingly its retrieval‑and‑prompting layer, not its base model (+0.74 over the identical underlying model); and (4) it can be used to improve existing systems: guided by its diagnosis, a single steadfastness instruction lifts a deployed Islamic assistant from +0.48 to +0.84 (Faith unstated, after pressure), matching the best guided frontier systems while preserving first‑response quality. The construct is faith‑general; we instantiate it for Islam as the first of a planned cross‑tradition family. Code, scenario bank, and rubric are open source (github.com/iaser‑ai/jaleesbench), with an interactive results browser at s.iaser.ai/jb.

Authors:Yiqun Zhang, Yunfan Zhang, Mingjie Zhao, Sen Feng, Yiu-ming Cheung
Title: How to Ask the AI: A User Perspective Survey for Large Language Model Prompting
Abstract:
AI tools like ChatGPT and DeepSeek, powered by Large Language Models (LLMs), allow users to obtain instant and effective content responses simply by typing requests, such as ``plan a three‑day Vienna trip'', ``solve the attached mathematical problem'', ``draft an email to inquire review progress'', etc., which are also known as LLM prompts. Crafting clear and well‑structured prompts leads to more appropriate LLM feedback, which effectively bridges human‑LLM interaction. Although prompting appears accessible to non‑expert users, precisely organizing effective prompts is a highly systematic and skillful process, presenting potential challenges even for experienced users. This survey explores the principles, taxonomy, and organization of prompts from a user‑centered perspective. Differing from the existing surveys that primarily focus on technical principles and application scenarios of LLMs, this paper provides actionable guidelines for formulating effective LLM prompts across diverse real‑world tasks and specifically contributes by: 1) developing an intuitive evaluation strategy for prompt efficacy, 2) providing prompting workflow demonstrations on representative applications, and 3) maintaining a dynamically updated open‑source project to ensure the core takeaways remain up‑to‑date. These measures lower the threshold for users to correctly understand and craft prompts that align with evolving application scenarios. This work will be maintained as a living GitHub project \hrefhttps://github.com/Yunfan‑Zhang/TAI_Guideline‑Table\textcolorbluehere.

Authors:Zongchuang Zhao, Xin Zhou, Tianyang Xu, Zhengyang Sun, Kaixuan Zhou, Honglin Li, Dingkang Liang, Xiang Bai
Title: SimWAM: A Simple World Action Model for End-to-End Autonomous Driving
Abstract:
World‑Action Models (WAMs) improve end‑to‑end autonomous driving by transferring video dynamics priors to action prediction, but existing methods incur costly test‑time future imagination. We present SimWAM, a simple yet effective WAM that leverages future‑video prediction as a training‑time supervision signal. It co‑trains a pretrained video expert and a lightweight action expert with joint flow matching. An isolated attention mask keeps action prediction independent of future frames, allowing trajectory prediction without explicit future‑frame generation at inference. Since the two experts share no parameters and interact only through a unified attention interface, the video backbone could be replaced and the action expert scaled independently without modifying the learning objective or inference pipeline. We further apply reinforcement learning to optimize a compositional driving reward beyond trajectory imitation. Our SimWAM achieves 91.5 PDMS on NAVSIM, surpasses state‑of‑the‑art WAM‑based planners with substantially lower latency, and transfers zero‑shot to nuScenes. These results position SimWAM as a simple yet solid baseline that could readily benefit from advances in video generation for efficient autonomous driving. The code and model weights are available at https://github.com/H‑EmbodVis/SimWAM/.

Authors:Youjun Zhao, Alex Warren, Gary K. L. Tam, Rynson W. H. Lau
Title: MirrorWorld: Taming Video Diffusion Models for Mirror Reflection Generation
Abstract:
Recent advances in video diffusion models (VDMs) have enabled high‑fidelity video synthesis. However, generating mirror reflections remains challenging because the content within a mirror must remain consistent with the surrounding scene. Existing VDMs are not specifically designed to model scene‑to‑mirror relationships, which can lead to reflections with incorrect content or inconsistent spatial arrangements. We observe that mirror reflection generation involves two complementary challenges: determining what scene content should be reflected and how the reflected content should be spatially arranged within the mirror region. Motivated by this observation, we propose MirrorWorld, a reflection‑aware video inpainting framework that models scene‑to‑mirror relationships during generation. Specifically, we introduce Semantic Relation Distillation (SRD), which transfers relational information from a frozen visual foundation model to encourage semantic associations between visible scene content and mirror regions. We further propose Geometric Transformation Alignment (GTA), which learns a transformation that guides the spatial arrangement of reflected content. The two components play complementary roles, with SRD modeling what should be reflected and GTA modeling how it should be arranged. To facilitate research on this problem, we construct a benchmark for video mirror reflection generation by repurposing four existing video mirror datasets into a unified reflection reconstruction task. Experimental results show that MirrorWorld achieves improved reflection reconstruction quality over representative image‑based reflection generation methods and strong video inpainting baselines.

Authors:Hanke Xie, Haopeng Lin, Jiale Qian, Dake Guo, Yuepeng Jiang, Zhichao Wang, Wenxiao Cao, Jingbin Hu, Guobin Ma, Wenhao Li, Huakang Chen, Chengyou Wang, Ming Tao, Zhonghua Fu, Lei Xie, Xinsheng Wang
Title: SemBridge: Semantic Token Anchoring for Continuous-Latent Autoregressive Speech Generation
Abstract:
Continuous‑latent autoregressive speech generation has emerged as a promising alternative to discrete‑token modeling by avoiding quantization loss and preserving richer acoustic information. However, continuous acoustic targets do not ex‑ pose linguistic structure as explicit token‑level prediction tar‑ gets. Consequently, the autoregressive language model (LM) must acquire linguistic structure indirectly through acous‑ tic prediction, which can compromise the content fidelity of generated speech. We propose SemBridge, a training‑only semantic‑token anchoring framework for continuous‑latent autoregressive speech generation. SemBridge uses discrete se‑ mantic tokens to directly supervise autoregressive LM states and employs a Semantic‑Aligned Acoustic VAE to organize the continuous target space under the same semantic refer‑ ence. The semantic supervision is used only during train‑ ing, while inference remains entirely continuous. We evalu‑ ate SemBridge on zero‑shot text‑to‑speech (TTS) and score‑ conditioned singing voice synthesis (SVS). Across multi‑ ple benchmarks, SemBridge improves content accuracy, as measured by word and character error rates (WER/CER), while maintaining competitive speaker similarity and percep‑ tual quality. Experimental results demonstrate that explicit semantic‑token supervision for autoregressive state learning is an effective and general direction for continuous speech generation. Speech samples are available.1 The model code and checkpoints will be available at https://github.com/ASLP‑ lab/SemBridge

Authors:Ananya Sahu, Mohit Bansal, Elias Stengel-Eskin
Title: CreativeInstruct: Scalably Teaching LLMs to Balance Quality, Creativity, and Diversity
Abstract:
While post‑training improves the capabilities of large language models (LLMs), it generally lowers their output diversity and creativity, negatively impacting tasks that explicitly require creativity (e.g., story generation) as well as those that require it implicitly, e.g., reinforcement learning (RL). We instead propose CreativeInstruct, a scalable instruction‑tuning method that teaches LLMs to balance creative, base‑model‑like generations with the quality of post‑trained models, by learning to inject special [StartCreativity] spans that bias generation toward creativity. Furthermore, we introduce a structural diversity metric based on graph edit distance, which captures narrative level variation missed by purely lexical and semantic metrics. On narrative generation, CreativeInstruct matches or exceeds the diversity of both multi‑model baselines and distilled variants of their outputs, without sacrificing quality or requiring multiple models at inference time. These results are mirrored in our human evaluation, where we find that annotators rate CreativeInstruct generations as more creative than the post‑trained LLMs' generations in 70.3% of cases. We also show the benefits of creative models as a substrate for RL: GRPO applied to a CreativeInstruct checkpoint improves by ~4% on AMC and ~5% points on MATH over the same training applied to the post‑trained checkpoint.

Authors:Ali Janati, Kaoutar El Maghraoui, Andrei Kanavalau, Anass Belfatmi
Title: Post-Grokking Collapse at the Representation-Readout Interface in Muon-Trained Transformers
Abstract:
Under the standard split, Muon gets hidden matrices and AdamW embeddings/output head. Muon groks modular addition faster, but its solutions do not hold. All nine configurations on (a+b) \bmod 113 grok and later lose generalization. Across five seeds the selected AdamW reference falls below threshold on four, reaching 27.59%. Instability persists across two moduli, two widths, two training fractions, subtraction, and depth. The failure arises at the representation‑readout interface, identified only jointly up to an invertible map unselected by the loss. After solving the training set, the gradient falls to order 10^‑6 and the optimizers respond differently: step‑size elasticity is ‑0.03 for Muon versus +1.5 for AdamW, and the Muon group moves 8.0 times faster per parameter. From bit‑identical states, freezing either group prevents failure. Freezing embeddings/readout removes it in five runs over 451,400 post‑grokking steps and five paired seeds: unfrozen arms record 137‑321 sub‑threshold evaluations, frozen arms none. Removing Muon's normalization and orthogonalization is no substitute: it collapses representation from 326 effective conjugate pairs to 4, shows no recurrent collapse, and fails terminally. Fourier filtering separates circuit failure from masking. Across 43 checkpoints over five seeds and three regimes, the task‑aligned family reaches exactly 100% alone. In circuit failure it no longer solves the task; in masking it remains perfect while the full model reaches 45.85%, giving a positive margin on every example, including errors, but being outvoted by a near‑equal adversarial remainder. Rescaling it restores 99.9%; grokking is the same condition resolving upward. The task selects the family, swapping (k,k) for (k,‑k) under subtraction. Across an abrupt collapse, standard Fourier support is unchanged and the power‑distribution cosine remains 0.9899.

Authors:Zixuan Lan, Luzhe Sun, Matthew R. Walter, Jiawei Zhou
Title: SABRE: Scalable and Automated Benchmarking of VLMs under Stress
Abstract:
Vision‑language models (VLMs) are improving rapidly, but benchmark development lags behind, making weaknesses hard to identify. Building stress tests is costly: samples must satisfy controlled conditions, remain answerable, and challenge current models. We present SABRE, a scalable, automated pipeline that converts a Test Primer (a Markdown Task Design with Data Schema) into structured specifications, generated or edited images, and question‑answer pairs. Automated filtering removes candidates solved by a Filtering VLM, while human review verifies candidate validity and supports annotation correction and localized image repair. We instantiate SABRE‑Prior to test whether VLMs follow visual evidence instead of relying on world priors ‑‑ learned expectations about familiar objects and scenes. Its 600 images and 1,000 questions span Context (unexpected entities in familiar scenes), Texture (counterfactual materials), Attribute (noncanonical component counts), and Language Elicitation (answers suggested by language but unsupported by the image). Across six VLMs, macro‑average accuracy ranges from 17.8% to 31.3% (22.6% mean). A real‑image Attribute control is comparably difficult for the Filtering VLM. SABRE‑Counting and SABRE‑Spatial pilots show that the workflow supports other stress‑test settings. These results establish SABRE as a reusable framework for constructing and refreshing VLM stress tests rather than a single fixed benchmark.

Authors:Elena Dumitrescu, Gert Lek, Lydia Y. Chen, Jérémie Decouchant
Title: Diffusion LLMs as Targets and Adversaries: Mechanistic Safety Exploits
Abstract:
Diffusion Large Language Models (DLLMs) replace autoregressive next‑token prediction with iterative parallel denoising, yet their internal safety mechanisms remain poorly understood. In this work, we investigate DLLMs both as targets and as adversaries, exposing mechanistic vulnerabilities in diffusion‑based alignment. We first show that safety alignment in DLLMs remains sparse and transferable across architectures. DLLMs initialized from autoregressive predecessors inherit the same mechanistic safety footprint as their source models, enabling transfer attacks via direct safety neuron mapping and pruning. Self‑pruning increases attack success rates (ASR) from 2.6% to 73.8% on LLaDA and from 1.9% to 86.6% on Dream, while transfer pruning from Qwen2.5 increases ASR from 1.9% to 73.2% on Dream and from 7.0% to 86.3% on Fast‑dLLM. Building on these findings, we introduce SN‑Guided Diffusion, a fully offline black‑box jailbreak framework that steers the diffusion process away from safety‑triggering regions using a weighted safety neuron loss, which achieves near‑perfect prompt separability (AUROC = 1.0 for benign‑vs‑jailbreak discrimination). Across multiple open and proprietary targets, our method achieves a transfer ASR of up to 77.1% on Llama‑3‑8B‑Instruct, 86.9% on Qwen2.5‑7B‑Instruct, and 74.3% against Gemini‑2.5‑Flash‑Lite, while requiring only 20 generation episodes per prompt. Compared to prior jailbreaking frameworks, our method achieves competitive transferability with orders‑of‑magnitude lower generation cost. Our codebase is available at https://github.com/ellyoana/sn‑guided‑diffusion.

Authors:Rodrigo Ferreira Rodrigues, Karim Radouane, Jose G Moreno, Lynda Tamine
Title: GeoBenchLLM: A Comprehensive Benchmark for Evaluating LLMs on Geo-Related Tasks
Abstract:
In the context of geodata, existing Large Language Models have often been studied in a homogeneous setting, which has considerably limited insights into their generalization capabilities. In this paper, we present \benchName, a comprehensive benchmark for probing LLMs on geo‑related tasks. We leverage a careful selection of twelve publicly available datasets from diverse geo‑related tasks and domains, and evaluate a set of LLMs on geo‑spatial and temporal understanding using our benchmark. Our results show that reasoning and size have a strong impact on overall performance. GeoBenchLLM is publicly available at https://github.com/Rfr2003/GeoBenchLLM.

Authors:Sasan Mansouri, Daniel Saad, Mark Wahrenburg, Manu Weissel, Fabian Woebbeking
Title: FinRank: An Evidence-Grounded Benchmark for Financial Question Answering and Retrieval over SEC Filings
Abstract:
Financial question answering is typically evaluated by answer correctness, yet in SEC filings a plausible and even numerically correct answer can be grounded in the wrong evidence. Similar facts and disclosures recur across sections of a filing, across reporting periods of the same firm, and across comparable firms. FinRank targets this provenance‑sensitive retrieval problem by requiring systems to identify evidence for the intended entity, reporting period, and disclosure context. The benchmark contains 1185 manually authored question‑answer records over the 10‑K and 10‑Q filings of 22 companies. Each record includes a reference answer, gold supporting passages, and hand‑curated hard negatives drawn from confusable passages within filings, across reporting periods, and across comparable firms. FinRank evaluates passage retrieval, reranking, and hard‑negative discrimination as separately measured tasks. Baseline results demonstrate the difficulty of this setting: among the evaluated systems, even a 7B instruction‑tuned embedder reaches only 44.8% Recall@10 on the pooled evidence corpus; sub‑billion‑parameter encoders gain at most 3.5 points over BM25, a finance‑adapted embedder trails BM25 by 9.7 points, and pairwise accuracy falls by 13.0‑20.5 percentage points when random negatives are replaced with the curated hard negatives. FinRank provides an evidence‑first benchmark for developing financial question answering systems that are not only accurate but also grounded in the correct disclosure.

Authors:Karim Radouane, Jose G Moreno, Lynda Tamine
Title: Geo-Spatial Concept Probing of Large Language Models: Abstraction, Compositionality, and Grounding
Abstract:
Understanding concepts is fundamental to generalization. Despite their impressive performance on a wide range of tasks, Large Language Models (LLMs) still struggle with genuine concept understanding. Prior work has evaluated conceptual understanding in LLMs using natural‑language benchmarks or narrowly scoped synthetic tasks, but these settings often conflate multiple skills or lack precise control over the underlying concepts and their properties. To support controlled probing of concepts in LLMs, we design tests on their core properties: abstraction, compositionality, and groundness. We set up a concept‑centric benchmark, targeting spatial concepts such as direction, distance, topology, and their compositions, and use question answering tasks serving as a proxy. We conduct extensive experiments across multiple LLM architectures and training regimes to analyze how model scale and design impact conceptual understanding. The results reveal clear limitations in current LLMs and provide insights into the factors shaping their ability to acquire and compose structured concepts. Our findings shed light on how concept‑based LLMs can be redesigned for improved information access and knowledge management. The code will be available at https://github.com/rd20karim/concept‑probing.

Authors:Jia Wang, Jiaming Cai, Zunying Hu, Zhanjie Wu, Jinyuan Liu, Hua Cheng, Yun Peng
Title: H2AL: Hyperbolic Hierarchy-aware Aggregative Learning for Registration-based Few-shot Medical Image Segmentation
Abstract:
Registration‑based Few‑shot medical image segmentation (RFMIS) aims to generate pseudo‑labels for unlabeled images by warping a labeled image through registration. However, existing methods primarily perform pixel‑level optimization and inference in Euclidean space, treating anatomical structures as flat and disjoint. This neglect of inherent hierarchies degrades pseudo‑label quality and weakens the discrimination of ambiguous regions, limiting the segmentation performance. To overcome this challenge, we propose a Hyperbolic Hierarchy‑aware Aggregative Learning framework for RFMIS, termed H2AL, that enhances both deformation plausibility and anatomical discrimination for dual‑task learning. Specifically, we introduce a Hyperbolic Hierarchy‑aware Infusion (H2I) module, which leverages the hierarchical modeling capability of hyperbolic space to learn precise hierarchy‑aware representations via transformation‑guided supervised hyperbolic contrastive learning, and injects such hierarchical priors into Euclidean space through a gated infusion block while preserving semantic richness. Furthermore, we propose an end‑to‑end joint optimization algorithm by gradient aggregation, where the gradients from the registration and segmentation decoders, embedding semantic and hierarchical cues, are aggregated to update the shared encoder to promote collaborative learning across tasks. Extensive experiments on two anatomical regions, with five experimental settings, demonstrate the effectiveness and efficiency of our method in both registration and segmentation. The code is publicly available at https://github.com/JiamingCai469/H2AL.

Authors:Taha Shieenavaz, Shabnam Zareshahraki, Loris Nanni
Title: Aftab: A Comprehensive Benchmark of CNN Encoders and Advanced Value Functions in Parallelized Q-Networks
Abstract:
Recent advancements in deep reinforcement learning have increasingly favored simplified, highly parallelized paradigms. Notably, the Parallelized Q‑Network (PQN) algorithm enables off‑policy value learning without relying on experience replay buffers or target networks. However, the representational capacity and computational efficiency of visual encoders operating in these buffer‑free settings remain comparatively underexplored. In this work, we systematically investigate the architectural design space of Convolutional Neural Networks within PQN. We evaluate eight distinct CNN topologies while explicitly characterizing their parameter and computational requirements. We further study the effect of multiplicative representation learning and advanced value estimation by integrating the Hadamax encoding paradigm with categorical, ensemble, and dueling value heads. Extensive experiments on Atari‑57 show that our final composite architecture, Aftab, achieves an Interquartile Mean (IQM) Human‑Normalized Score of 6.592, compared with 2.715 for the standard PQN baseline, together with a 0.86 Probability of Improvement over PQN. We additionally evaluate Aftab on Procgen‑Hard to assess performance under procedurally varying visual environments. Aftab achieves a normalized learning‑curve Area Under the Curve (nAUC) of 0.541 compared with 0.216 for PQN. Overall, the results demonstrate that carefully designed encoder topology, multiplicative feature interactions, and advanced value‑estimation heads can substantially improve performance within a parallelized, replay‑free Q‑learning framework while preserving its memory‑efficient training paradigm. The complete Aftab framework, including model definitions, training configurations, reproducibility settings, and raw experimental logs, is open‑sourced at https://github.com/tahashieenavaz/aftab

Authors:Chen Shao, Yue Wang, Zhenyi Zhu, Zhanbo Huang, Tobias Käfer, Zonghan Wu, Danai Koutra
Title: When GNNs Fail: Quantifying and Overcoming Temporal Correlation Volatility in Time Series
Abstract:
Modeling multivariate time series by representing them as graphs, where individual series act as nodes and pairwise temporal corre‑ lations serve as edges, has gained significant traction. Recent advances in Graph Neural Networks (GNNs) have demonstrated strong perfor‑ mance by assuming a static graph topology and aggregating information from neighboring series. In this work, we investigate the representa‑ tional power of GNNs for forecasting under both static and dynamic settings (i.e., when pairwise correlations evolve drastically over time) and identify critical limitations in current architectures. To formalize this, we first propose Temporal Correlation Volatility (TCV), a model‑ agnostic metric designed to quantify the distributional evolution of these latent structures. We establish a clear connection between TCV and performance degradation, demonstrating that many popular models, including Transformers, generalize poorly in high‑TCV settings and are often outperformed by simple structure‑agnostic baselines. To address these limitations, we propose Graph Layer for Inference in Dynamic En‑ vironments (GLIDE), a novel GNN layer enhanced by two theoretically grounded design mechanisms: (D1) Path‑based Message Passing, which captures path‑based neighborhoods and (D2) Static and Dynamic Propagation Separation, which identifies optimal dynamics via local static approximation. These components significantly improve learning under dynamic topology while preserving robustness in static scenarios. Ex‑ tensive experiments on synthetic and real‑world benchmarks show that GLIDE improves average performance by up to 45.6% across static and dynamic settings, with the largest gain reaching 85.7%. The source code is available at https://github.com/ChenS676/GLIDE.

Authors:Chengyi Peng, Haoyu Yang, Meixing Shi, Yuxiang Cai, Yankai Jiang
Title: EliSeg: Verified Target Construction for Report-Grounded Abnormality Segmentation
Abstract:
Radiology reports describe clinical observations but do not specify executable segmentation targets. They may contain present, negated, prior,uncertain, or irrelevant findings, while multiple valid abnormalities may coexist. Existing segmentation methods largely bypass this ambiguity by receiving a target identity or spatial prompt before inference, which acts as a hidden target oracle. We study report‑grounded abnormality segmentation, where a model must determine target eligibility, cardinality, and finding‑to‑mask correspondence directly from an unfiltered report before delineating the corresponding regions. We propose EliSeg, an atcor‑‑verify‑‑revise framework that integrates target construction with mask generation. A grammar‑constrained Actor proposes target slots and masks, an independent text‑only Verifier reconstructs the eligible finding inventory, and Revision selectively re‑executes the shared Actor when their target structures disagree. EliSeg requires no predefined target identity, finding prompt, point, or bounding box. Experiments on MIMIC‑CXR‑ILS show that EliSeg consistently outperforms direct segmentation methods and extract‑then‑segment cascades across findings, while effectively suppressing masks for ineligible report mentions. Ablation studies confirm the complementary roles of verification and revision, and evaluation on CheXlocalize demonstrates effective transfer of the EliSeg to an external dataset.Code is available at https://github.com/Maybach‑dream/EliSeg.

Authors:Kendong Liu, Yuxin Yao, Junhui Hou
Title: CANIS: Generation-Assisted 3D Canonicalization via an Image-Semantic Bridge
Abstract:
Canonicalizing 3D object orientation is fundamental to 3D understanding and analysis. Existing approaches often rely on geometric cues, although 3D canonicalization ultimately requires a semantically meaningful orientation. To address this gap, we propose CANIS, a category‑agnostic, generation‑assisted framework that introduces the semantic orientation prior of a frozen image‑to‑3D generative model into 3D canonicalization, without canonicalization‑specific training or category‑specific templates. Specifically, CANIS first renders the input object from candidate viewpoints, selects an informative view, and generates a proxy in a canonical orientation. During generation, a sparse structural latent encoded from the input guides the proxy to preserve the geometry of an object. CANIS then uses the selected image as a semantic bridge between the input and the proxy. Image patches identify semantic regions on the proxy, and depth back‑projection locates the corresponding regions on the input. The resulting semantic anchors constrain geometric matching, from which we estimate the rigid transformation that canonicalizes the input. Experiments on synthetic benchmarks validate CANIS and its key components, while qualitative results on partial observations and OmniObject3D suggest its applicability to incomplete and real‑world scans. CANIS also improves downstream 3D classification, part segmentation, and dense correspondence under arbitrary rotations. Project page: https://kenkenzaii.github.io/Canis.

Authors:Eric Cullhed, Albin Thörn Cleland
Title: Stoicheia: Character-Level Masked Diffusion for Ancient Greek Textual Restoration, Parsing, and Metrical Scansion
Abstract:
We introduce Stoicheia, a 405M‑parameter character‑level masked‑diffusion encoder for Ancient Greek whose input factors into five aligned, independently maskable planes: letters, word and sentence boundaries, diacritics, capitalization, and punctuation. A single backbone can therefore restore lacunae, re‑segment, accentuate, and punctuate unspaced text without task‑specific retokenization. We pretrain it on an open, revision‑pinned corpus of 380M words and release eleven checkpoints: ten rotated, decontaminated folds, guaranteeing that for any given literary passage at least one released model has never seen its text, and one with no exposure to documentary texts. Three experiments ‑ reconstruction of damaged inscriptions and papyri, morphosyntactic tagging and dependency parsing, and macronization with metrical scansion ‑ each carry a matched random‑initialization control, isolating what character‑level diffusion pretraining contributes: 5.6 CER points on inscription reconstruction, 12.9 LAS on parsing, and 6.0 points of balanced accuracy on macronization. On Ithaca's own test split, with identical frozen samples and strict scoring, Stoicheia reduces character error relative to both prior state‑of‑the‑art systems, from 24.6 (Ithaca) and 23.5 (its 2025 Aeneas‑framework successor) to 15.5, and raises top‑1 accuracy from 63.0 and 64.0 to 74.5.

Authors:Idil Gözel
Title: Learning Suffers More Than the Policy Class Under Partial Observability: A Closed-Form Analysis
Abstract:
When a reinforcement learning agent cannot observe the full state, we usually blame its policies: it cannot see enough to represent a good one. We show that in a solvable case the bigger problem lies elsewhere. Even when a good policy is available and the agent's value function is expressive enough to describe it exactly, learning still ends up somewhere far worse. We study a partially observed linear‑quadratic problem in which a standard actor‑critic learner can be solved in closed form. At our default setting the best policy the agent can represent is already close to optimal, costing 10.4% more than the ideal controller that observes everything. Learning does not find it. The algorithm instead comes to rest at a policy that is 35% worse than the best one available to it, and we can say exactly where and why. The cause is a bias in what the critic learns rather than a limit on what the actor can express. Because the agent cannot attribute what it sees to the part of the state it cannot observe, the critic misreads that unexplained variation as sharp curvature in its own value estimates, and the actor follows that error away from the optimum. We derive closed‑form expressions for the resulting policy, for its cost, and for the one design choice that removes the problem, which is how far the learner looks ahead before trusting its own value estimates. Deep reinforcement learning experiments follow these predictions closely. Notably, giving the agent memory of past observations does not help, while changing how far it looks ahead does.

Authors:Vasanth Iyer
Title: Dual-Node NVIDIA DGX Spark over Tailscale: A Remote-Access Testbed for Distributed LLM Training and Cyber-Threat-Intelligence Fine-Tuning
Abstract:
Compact AI systems make local language‑model experimentation increasingly accessible, yet practical evidence for multi‑node training on desktop‑class accelerators remains limited. This report presents a proof‑of‑concept deployment of distributed NanoChat pretraining across two NVIDIA DGX Spark systems, each with a GB10 Grace Blackwell system‑on‑chip and 128 GB of unified memory, administered remotely over a Tailscale mesh VPN and connected for training by a dedicated 200 Gb/s QSFP56 direct fiber link. PyTorch torchrun, DDP, and NCCL were configured with one process per node, a depth‑20 NanoChat model, a local batch size of 32 per node, and a 2,048‑token context, giving a global batch of 131,072 tokens per step. The run sustained a step time of about 69.4 s (about 1,890 tokens/s), processing about 653 million tokens over four days. We document link configuration, container setup, interface binding, a step‑zero evaluation bug that triggered NCCL timeouts, checkpointing, and troubleshooting lessons, as a reproducibility reference for small labs. We also built a cybersecurity fine‑tuning dataset from 77 CISA advisories (338 training, 37 validation conversations) and ran a 17‑question held‑out evaluation comparing a baseline SFT checkpoint against a CTI‑augmented checkpoint with an Ollama‑hosted LLM judge. CTI‑specific categories improved while general‑knowledge categories regressed, for a small overall change from 2.06 to 2.29 on a 0‑10 scale. The same cluster supports a 400‑level AI course (CS 426) and a query engine for CompTIA Security+ POGIL activities in CBS 255, showing modest local infrastructure can serve both research and teaching. The study establishes feasibility rather than a scaling‑efficiency claim, since single‑node throughput used for comparison was estimated, not measured under matched conditions. Runbook and scripts are available (see Code Availability).

Authors:Jiaqian Wang, Yutao Qi, Wenjin Hou, Yuanxi Che, Muning Wen
Title: From Test-Time Scaling to Reusable Memory: Measuring Crystallization in Text-to-SQL
Abstract:
Test‑time scaling can correct difficult text‑to‑SQL queries, but the extra computation is normally discarded after each answer. Systems increasingly retain verified repair episodes, yet evaluations still report one end‑to‑end score. It cannot distinguish replay on recurring questions from help on unseen questions, or identify the responsible memory choice. We call measuring this future value the crystallization problem. Our controlled evaluation holds the single‑shot solver fixed and varies one memory choice at a time. We separately measure replay, cross‑question retention, and held‑out same‑database transfer. On BIRD, storing verified corrected queries improves held‑out first‑attempt accuracy by 4.34 percentage points. This gain captures 44.4% of the accuracy headroom provided by on‑demand repair on the same questions. Controlled interventions identify database‑specific content as the main operating ingredient. Reliable verification and broader retrieval coverage yield supported gains; richer formats and elaborate retrievers do not. Open‑source code, evaluation artifacts, and reproduction instructions are available at https://github.com/ai‑jiaqian/text‑to‑sql‑memory‑crystallization.

Authors:Minchao Jiang, Xiaoxuan Ma, Shunyu Jia, Haoru Wang, Zhang Liang, Wentao Zhu
Title: InstanceSplat: Instance-Aware Feed-Forward 3D Gaussian Splatting for Scene Understanding
Abstract:
Feed‑forward 3D Gaussian Splatting (3DGS) enables efficient and generalizable 3D reconstruction, but current feed‑forward 3DGS methods for scene understanding remain largely category‑oriented. In contrast, instance‑aware 3DGS methods typically rely on per‑scene optimization and often decouple reconstruction from instance and semantic learning, limiting reciprocal interactions among them. We present InstanceSplat, a unified feed‑forward 3DGS framework for generalizable 3D reconstruction and instance‑aware scene understanding from pose‑free multi‑view images. In a single forward pass, InstanceSplat constructs an instance‑aware Gaussian representation that jointly encodes appearance, geometry, instance identity, and language‑aligned semantics. Shared 3D Gaussians ground instance identities across views, producing renderable and cross‑view‑consistent instance features. To allow reconstruction and scene understanding to benefit from each other, we further design an instance‑centric learning strategy that connects reconstruction, instance learning, and semantic learning through shared instance structure. Specifically, instance cues guide reconstruction, language‑aligned semantics strengthen the discrimination of confusing same‑category instances, and instance regions aggregate semantic evidence into coherent object‑level predictions. Experiments on novel‑view synthesis, instance segmentation, and open‑vocabulary semantic understanding under varying input‑view settings and on an unseen dataset demonstrate state‑of‑the‑art performance, practical efficiency, and strong generalization.

Authors:Lumin Chen, Qingyao Tian, Jinpeng Li, Haoyu Jiang, Huai Liao, Xinyan Huang, Hongbin Liu, Dong Yi
Title: Geometry-Aware Camera Localization for Bronchoscopy
Abstract:
Camera localization in bronchoscopy remains a challenging problem due to stringent accuracy requirements, real‑time constraints, and limited training data. Compared to natural scenes, the confined anatomical structures demand millimeter‑level precision, while intraoperative guidance necessitates low‑latency inference. However, existing methods often fail to effectively exploit preoperative geometric priors, limiting their robustness and accuracy. To address these limitations, we propose a unified geometry‑aware bronchoscope localization framework (GABL) that effectively fuses preoperative structural priors with paired intraoperative video to estimate 6‑DoF camera poses. Specifically, to address visual ambiguity in complex airways, we propose a graph‑guided coarse‑to‑fine localization scheme that effectively leverages structural priors for precise pose estimation. Furthermore, to mitigate pose jitter and bridge the visual‑structural gap, we integrate a Transformer‑based tracking model with a novel RGB‑depth matching objective, jointly enforcing spatio‑temporal and geometric consistency. Extensive experiments demonstrate that our method yields remarkable reductions of 8.37% and 31.76% in translation and rotation errors over the prior state‑of‑the‑art, alongside 4 times inference speedup (33.6 FPS) for robust real‑time bronchoscope localization. Project website: https://paulili08.github.io/GABL/.

Authors:Bohao Tang, Zhen Qin, Yuqi Pan, Zheng Li, Pengfei Liu, Ya Zhang
Title: Modular TTT: Rethinking Test-Time Training as Composable Modules
Abstract:
Test‑time training (TTT) views sequence modeling as an online learning problem in which fast weights are updated by an internal learning rule. Despite the growing number of TTT variants, existing approaches typically hard‑code each variant separately, which makes it difficult to design new TTT methods and to isolate the role of each component. To address this, we propose Modular TTT, a framework that represents the inner learner as a directed acyclic graph and exposes the fast‑weight network, loss function, learning rate, weight decay, and normalization as explicit design dimensions. Modular TTT automatically composes primitive‑level train‑view forward, train‑view backward, and causal query‑view rules into the full graph‑level TTT computation, including the fast‑weight state transition. Using Modular TTT, we systematically ablate the components of TTT and find that small learning‑rate initialization, weight decay, and a single‑layer nonlinearity improve performance, while MSE and inner‑product losses perform similarly. Deeper fast‑weight networks and normalization tend to hurt performance because they induce excessively large activations, while residual connections and gating provide little measurable benefit. Guided by these findings, we train the best resulting variant as 410M‑ and 1.45B‑parameter models on 100B tokens, and observe training loss and benchmark performance comparable to Gated DeltaNet.

Authors:Francisco Javier Becerra Sanchez, Antonio Ken Iannillo, Radu State
Title: SoK: Cryptographic Key Recovery for Cryptoasset Custody and Financial Technologies
Abstract:
Cryptoasset systems often bind cryptographic key control to financial control: losing a wallet seed, custody share, hardware device, or smart‑account credential can remove spend authority, while compromised recovery can enable theft. Existing work treats recovery through separate vocabularies‑‑key backup, secret sharing, account recovery, credential re‑issuance, social recovery, and asset migration‑‑making mechanisms and tradeoffs difficult to compare. This paper presents a Systematization of Knowledge (SoK) on cryptographic key recovery for cryptoasset custody and financial technologies. Starting from a 118‑paper systematic‑review discovery corpus, we derive a 77‑paper synthesis corpus and code each retained system in a master matrix covering recovered objects, recovery semantics, mechanisms, enrollment and storage, authorization, trust placement, failure events, post‑recovery state, validation evidence, deployment status, privacy, usability, and limitations. The matrix supports an axis‑first taxonomy that separates secret‑restoring, hybrid, control‑restoring, forensic/extractive, and framework‑oriented recovery. Our central observation is that recovery is not a single operation: systems may reconstruct an original secret, regenerate a seed, restore a share, reissue a credential, migrate signing authority, restore account control, move assets, or extract forensic artifacts. We derive a generalized construction model, check it against production‑facing designs, and identify six findings: recovery semantics are heterogeneous; recovery shifts trust; liveness improvements create abuse paths; post‑recovery lifecycle management is uneven; protocol evidence outpaces user evidence; and recovery metadata remains underprotected. These gaps motivate a research agenda for recovery‑aware financial technologies.

Authors:Qiyanhui Lu, Han Wu, Rongjian Xu, Tingzhang Luo, Cheng Fan, Xinghao Chen, Minjing Dong, Jufeng Yang, Jianyuan Guo
Title: RoRA: Role-Oriented Regional Allocation for Visual Token Pruning in MLLMs
Abstract:
Multimodal large language models (MLLMs) encode images as long visual token sequences, making prefilling and KV‑cache storage expensive. Existing training‑free pruning methods select tokens by importance, diversity, or spatial coverage, but treat retained tokens as interchangeable and do not explicitly track which object‑related regions are already covered. We present RoRA, a training‑free framework that casts visual token pruning as role‑oriented regional evidence allocation. Given a fixed budget, RoRA partitions tokens into a protected semantic core, complementary context, and fine‑grained detail. It first calibrates text‑conditioned attention with a positional prior and a prompt‑calibrated object prior, then builds Attention‑Anchored Regions (AARs) from high‑confidence anchors as lightweight proxies for covered object support. Context is explored mainly outside AARs, while a small AAR‑guided budget restores local detail; pairwise similarity is used only for context‑stage redundancy filtering. Under matched budgets, RoRA consistently outperforms strong training‑free baselines across LLaVA and Qwen‑VL families, retaining most of the unpruned accuracy even at aggressive pruning ratios, e.g., 96.5% of full performance at 88.9% pruning on LLaVA‑1.5, and improving over D2Pruner by about 5% on Qwen3‑VL at 75‑90% pruning. At a 66.7% pruning ratio, RoRA requires only 0.7 ms for token selection and reduces end‑to‑end inference time by 24.6%, corresponding to a 1.33x speedup over unpruned inference on an NVIDIA H800.

Authors:Zehui Li, Zihao Sun, Jiawei Xu, Zheqi He, Xiaoqiang Zhang, Jing-Shu Zheng, Lu Liu, Dahui Gao, Xiuwan Chen
Title: LifelongCrossNav: Persistent 3D Semantic Memory for Cross-Floor Multi-Object Navigation
Abstract:
Object‑goal navigation has made substantial progress in semantic perception and exploration, yet persistent memory for multi‑object navigation and cross‑floor navigation are still commonly addressed separately. We present LifelongCrossNav, a framework for sequential multi‑object ObjectNav in unknown multi‑floor indoor environments. Within each episode, the agent receives an ordered sequence of object‑goal queries while continuously maintaining a shared sparse 3D semantic voxel memory. This memory incrementally accumulates geometric structure, traversability states, and vision‑language features, allowing subsequent object‑goal queries to retrieve previously acquired scene information without rebuilding the map. To support persistent search across floors, LifelongCrossNav combines support‑aware 3D traversability mapping, stair‑specific perception, and direction‑aware stair traversal. A unified navigation policy coordinates same‑floor frontier exploration, live and historical point‑of‑interest retrieval, stair navigation, and target‑object search and approach. We further introduce HM3D‑MFMON, a benchmark for sequential Multi‑Floor Multi‑Object Navigation built on HM3D scenes, including a dedicated subset in which completing the full sequence of object‑goal subtasks requires at least one floor transition. Experimental results show that LifelongCrossNav consistently outperforms a representative planar persistent semantic‑map baseline on HM3D‑MFMON, demonstrating that persistent 3D semantic memory and cross‑floor traversability modeling effectively support sequential multi‑object navigation in multi‑floor environments. Project page: https://flageval‑baai.github.io/LifelongCrossNavPage.

Authors:Zhiyuan Liu, Tinghong Ye, Chenghao Liu, Yizhuo Li, Songfang Huang
Title: MemOPD: On-Policy Distillation through Memory State Alignment for Long-Horizon Agents
Abstract:
Long‑horizon agents accumulate growing contexts during interaction, impairing performance and stability. Compact memory mitigates this problem by compressing and rewriting the history retained between model invocations. Learning what to retain typically relies on proximal policy optimization (PPO) with final task rewards, but sparse rewards provide little guidance for individual memory updates. This limitation motivates on‑policy distillation (OPD), which supplies dense teacher supervision on student rollouts. For such supervision to be valid, the teacher must evaluate each sampled action under the same state in which it was generated. However, the context rewriting performed during memory compression can break this alignment. When sampled responses are retained and re‑encoded for later invocations, flattening the interaction into a persistent history may cause the teacher to score the action under a state that the student never visited during rollout. The action therefore remains on‑policy by provenance, but not necessarily by state. We therefore propose Memory‑Aligned On‑Policy Distillation (MemOPD). MemOPD records the inputs and sampled outputs of each model invocation, restores its original token positions and causal visibility, and packs the reconstructed invocations for efficient teacher scoring. The teacher provides full‑vocabulary supervision at the sampled action positions, while PPO preserves the final task objective. Experiments verify state alignment across several context updates and show that it improves F1 by 7.0% over persistent‑history teacher scoring in a matched control. Overall, MemOPD‑3B improves F1 over PPO by up to 416.2%, while packing yields up to a 1.63x speedup in actor computation during training. The code for this work is publicly available at: https://github.com/TPssp/MemOPD.

Authors:Hanshu Yao, Janfeng Zhong, Niu Lian, Jinpeng Wang
Title: DocMemo: Dynamic Evidence Discovery via Probabilistic Memory-Guided Retrieval for Multi-Modal Document Understanding
Abstract:
Long‑document understanding requires locating sparse and heterogeneous evidence across hundreds of pages, yet existing systems remain limited by static retrieval and fragile cross‑round memory. Mainstream single‑round methods commit to a fixed top‑k page set at the outset and struggle to recover from early retrieval errors; recent iterative approaches allow multi‑round evidence acquisition, but they do not investigate the propagation mechanism of cross‑round states, making it difficult to track the dynamic changes in page relevance. To address these limitations, we propose DocMemo, a memory‑guided framework that formulates long‑document reasoning as dynamic evidence exploration. DocMemo maintains a tri‑level retrieval state consisting of Document Schema Memory, Page Belief Memory, and Question Episodic Memory, which respectively capture structural priors, dynamic relevance estimation, and query‑specific reasoning trajectories. During reasoning, DocMemo continuously refines cross‑round page selection through Bayesian page belief updating with Thompson sampling, spatial proximity propagation, and structure‑aware adaptive‑granularity evidence access, while supplementing page‑level evidence with fine‑grained visual regions. Experiments on 3 benchmarks show that DocMemo achieves state‑of‑the‑art performance and validate the efficacy of structured memory and dynamic page belief updating. Code is available at https://github.com/Harrygof/DocMemo.

Authors:Wei Xu, Zhu Wang, Yifan Guo, Changlong Cheng, Yin Zhang, Zhihui Ren, Bin Guo, Zhiwen yu
Title: XGait: A Multi-Modality Wireless Sensing Dataset for Indoor Human Tracking and Identification
Abstract:
Wireless sensing has emerged as a promising approach for tracking and identification using commodity Internet of Things devices. However, the features derived from a single wireless modality are often fragile to variations in environmental layouts and walking trajectories. Furthermore, most existing studies are based on datasets collected in specific scenarios with limited trajectory diversity and sensing modalities, preventing a robust evaluation of system generalization. \textcolorblueTo address this gap, we introduce XGait, a multi‑modality wireless sensing dataset that synchronously captures human walking using Wi‑Fi and acoustic transceivers across three indoor scenarios, with vision‑based measurements serving as ground truth. Specifically, XGait contains more than 22K walking samples from 27 participants, covering diverse directions and trajectories to support both indoor tracking and identity recognition. To bridge the heterogeneity of wireless sensing modalities, we propose a unified Doppler spectrogram representation that maps Wi‑Fi and acoustic signals into a shared time‑‑frequency space, along with a standardized benchmark pipeline for pre‑processing, temporal alignment, and feature construction, enabling reproducible evaluation and systematic cross‑modal analysis. Extensive evaluations demonstrate that Wi‑Fi and acoustic sensing exhibit complementary strengths, particularly under complex trajectories and challenging propagation conditions, thereby paving the way for novel research in the field of multi‑modality wireless sensing. The dataset and code are available at https://github.com/warrior‑087/XGait.

Authors:R. G. Bahumanya, Harshith V. M., Shreyank N. Gowda, Anala M. R
Title: Explanation Stability of Test-Time Adaptation in Computational Pathology: A Large-Scale Benchmark
Abstract:
Test‑time adaptation (TTA) has become a practical way to adapt deployed models to unlabeled target data, a setting that is especially relevant in computational pathology where staining, scanner, and cohort shifts are routine. While most TTA methods are evaluated by their effect on accuracy, clinical use also depends on whether the model's explanations remain reliable after adaptation. In this paper, we take a closer look at this largely unmeasured effect. We study explanation stability under TTA across two histopathology benchmarks, Camelyon17 and NCT CRC‑HE, using five architectures ranging from convolutional networks to vision transformers and a pathology foundation model, seventeen TTA methods, and four attribution families. Across 2,958 adaptation runs, we observe a clear and systematic pattern: TTA methods differ sharply in how much they move model explanations, with frozen‑backbone methods leaving attributions almost unchanged and continual methods such as CoTTA and RoTTA causing the largest drift. This effect is not uniform. Convolutional networks are substantially more sensitive than transformer and foundation‑model backbones, and explanation drift increases with adaptation strength while remaining largely insensitive to batch size. Surprisingly, explanation stability is only weakly coupled to adaptation quality. Some methods preserve explanations almost perfectly while degrading calibration or accuracy, producing silent failures that would be missed by accuracy‑only or explanation‑only evaluation. These findings show that explanation stability is a distinct reliability axis for TTA in computational pathology. We release the metric, protocol, and full benchmark to support future work on adaptation methods that are not only accurate, but also stable and clinically auditable. Code: https://github.com/bahumanyarg11/tta‑explanation‑stability‑pipeline

Authors:Jie Ren, Zhehao Jiang, Yinhong Yang, Haorui Jia, Han Jiang, Ben Li, Yao Yao, Cheng Lin, Qiu Shen, Zhenshan Bing, Xiao-Xiao Long, Xun Cao
Title: C2Dex: Contact-Consistent Reconstruction and Retargeting for Dexterous Manipulation from Monocular Video
Abstract:
High‑quality demonstrations for dexterous robot manipulation are costly and difficult to collect, whereas monocular human videos provide a scalable source of diverse manipulation behaviors. However, transferring such demonstrations to dexterous robots remains challenging: monocular hand‑object interaction (HOI) reconstruction often produces temporally unstable contacts and physically implausible interactions, while conventional retargeting methods struggle to preserve task‑relevant contacts and local interaction geometry across different hand embodiments. We present C2Dex, a video‑to‑dexterous‑manipulation framework built around a shared interaction representation: stable object‑side contacts recovered by aggregating noisy frame‑wise observations in the canonical object space. These stable contacts serve a dual role: as trajectory‑level constraints that guide reconstruction toward temporally coherent and physically plausible human HOI trajectories, and as explicit transfer targets for the dexterous hand, where Laplacian interaction optimization preserves the local hand‑object geometry across embodiments and residual reinforcement learning refines the trajectory in simulation. Experiments on DexYCB and TACO show that C2Dex achieves end‑to‑end trajectory success rates of 57.78% and 26.67%, respectively, substantially outperforming the strongest baselines (17.78% and 10.00%) under identical evaluation criteria. Real‑robot replay experiments further demonstrate physical feasibility across diverse contact‑rich manipulation tasks. Project page: https://k‑jie.github.io/C2Dex/

Authors:Albert Saiapin, Kim Batselier
Title: Tensor Network Kernel Machines: A JAX Framework for Machine Learning and Nonlinear System Identification
Abstract:
Developing nonlinear models that are both expressive and computationally efficient remains a challenge in machine learning and nonlinear system identification. Tensor network kernel machines (TNKM) address this challenge by combining nonlinear feature representations with compact low‑rank tensor‑network parameterizations. However, practical and extensible software frameworks for developing TNKM models remain limited. In this work, we introduce "tnkm", an open‑source Python library for constructing and training TNKM models using JAX. The library provides a unified interface for combining different feature maps, tensor‑network architectures, and optimization strategies, including alternating least squares and gradient‑based methods. We demonstrate the capabilities of "tnkm" on nonlinear benchmark problems, showing that the implemented models achieve competitive prediction accuracy while retaining compact parameterizations and efficient training. The proposed framework facilitates reproducible development and application of tensor‑network‑based learning methods.

Authors:Simon Scholz, Mersedeh Sadeghi
Title: CAS2UML: A Handwritten Sketch-to-PlantUML Dataset for Class and Activity Diagrams
Abstract:
Automated UML generation from sketches and images is gaining renewed attention with the rise of large language models and multimodal AI. However, reproducible evaluation remains difficult due to the lack of public datasets with executable groundtruth models. We present CAS2UML, a public dataset of 557 handdrawn UML diagrams, including 271 class diagrams and 286 activity diagrams, each paired with manually validated PlantUML code. We also provide a PlantUML‑based validation tool and reusable scripts for checking the syntactic correctness and renderability of generated UML artifacts, enabling reproducible benchmarking of sketch‑to‑UML approaches. The dataset, validation tool, processing scripts, documentation, and demonstration video are publicly available at: Dataset: https://huggingface.co/datasets/Seym0n /cas2uml_hand‑drawn_to_plantuml_dataset; Tool and Scripts: https://github.com/Seym0n/handwritten‑uml‑dataset; Video: https://www.youtube.com/watch?v=KQrYeGgT3hs.

Authors:Alisa Pesotskaia, Emin Zerman
Title: IceHorizon: A Dataset for Horizon Detection in Ice-Covered Maritime Environments and Comparative Evaluation of Detection Methods
Abstract:
Horizon detection in images of ice‑covered waters is a challenging problem for maritime navigation due to low contrast between water and sky, cluttered ice structures, and varying illumination conditions. This paper presents a comparative evaluation of six horizon detection algorithms, including four classical computer vision methods and two hybrid approaches combining deep learning with classical line detection. A new bespoke IceHorizon dataset consisting of 30 ship‑based and 8 drone‑based videos is used to evaluate detection accuracy, horizon coverage, and computational performance. The results show that hybrid methods achieve the highest accuracy and most reliable horizon estimates. In contrast, purely classical methods exhibit reduced robustness, particularly in visually ambiguous scenes. Performance on ship‑based imagery was consistently higher than on drone‑based imagery, indicating a strong dependency on acquisition characteristics. The created dataset and codes used in this study are made publicly available to support further research on this topic. The code is available at https://github.com/allythe/HorizonDetection. The dataset is available at https://doi.org/10.5281/zenodo.20411867

Authors:Jiankun Wang, Yisen Gao, Ziwei Zhang, Xingcheng Fu, Jiaxin Bai, Chen Gao
Title: Does More Retrieved Evidence Help Visual Retrieval-Augmented Generation with Diffusion Language Models?
Abstract:
Visual retrieval‑augmented generation (RAG) commonly expands the retrieved evidence set to improve answer‑page coverage, implicitly assuming that all available evidence should be passed to the generator. We show that this assumption does not hold for diffusion language models (DLMs): retrieving more pages increases answer‑page recall, whereas unconditionally passing all retrieved pages to the generator often reduces answer accuracy, primarily because of semantic conflict. A latent‑source analysis explains this mismatch through source‑coherence loss in parallel denoising, where position‑wise proposals can combine incompatible visual sources into unsupported answers. We further find that such interference is already visible in the first‑step answer‑block distribution, making it possible to assess evidence before decoding. To preserve retrieval coverage while limiting harmful visual exposure, we propose the Entropy‑Based Candidate Filter (ECF), a training‑free evidence‑admission framework. To reduce irrelevant content within individual candidates, ECF constructs multi‑granularity evidence units; to identify beneficial additional evidence, it uses blank‑controlled block confidence and retrieval rank to determine whether and which candidate should enter the final context. Across three multimodal DLMs and five visual QA benchmarks, ECF improves answer accuracy by 2.62 percentage points on average over the strongest fixed top‑k input and, with LLaDA2.0‑Uni, by 2.37 percentage points on average over the best competing training‑free result for each dataset. These results show that broader retrieval benefits visual DLM‑RAG through selective evidence admission rather than unconditional evidence expansion. Code is publicly available at https://github.com/wjkuser/ECF.

Authors:Yu Xue, Haoxuan Qu, Zhuoling Li, Hongbin Xu, Jianxiong Yin, Simon See, Hossein Rahmani, Jun Liu
Title: HRDiT: Training-Free High-Resolution Image Generation with Off-the-Shelf Diffusion Transformer Models
Abstract:
Training‑free text‑to‑high‑resolution image generation has recently attracted growing research attention. However, existing studies on this task primarily focus on adapting off‑the‑shelf U‑Net‑based diffusion models to high resolutions, with limited progress on adapting off‑the‑shelf Diffusion Transformer (DiT) models despite their strong text‑to‑image generation capabilities at limited resolutions. In this work, we find two key challenges particularly hindering the application of off‑the‑shelf DiT models for high‑resolution image synthesis in a training‑free manner, namely, spatial disorder and long generation time. To address these challenges, we propose a novel method tailored to adapt off‑the‑shelf DiT models for high‑resolution image synthesis. Extensive experiments show the efficacy of our method. Our code is available at: https://github.com/zylwithxy/HRDiT.

Authors:Jiaqi Zhang, Zheng Pang, Rongrong Gao, Qiyuan Zhang, Yang Yang
Title: Local Epistemic Uncertainty Guided Active Sampling for Plug-and-play Diffusive Image Restoration
Abstract:
Diffusion models have demonstrated remarkable effectiveness in image restoration tasks. However, when guiding image reconstruction, existing Diffusion Model‑based Image Restoration (DMIR) methods typically rely on fixed data constraints and uniform step sizes, thereby overlooking the dynamic nature of the generative process. Such rigid designs render the models vulnerable to spatially non‑uniform degradations, thus resulting in structural distortions and loss of fine details. Meanwhile, uniform step sizes introduce computational redundancy, whereas naïve step reduction strategies tend to accumulate approximation errors. To address these limitations, we propose a Local Epistemic Uncertainty Guided Active Sampling framework (LEADer). In the spatial domain, LEADer leverages pixel‑wise uncertainty to dynamically modulate the prior strength within the null space, which effectively balances detail preservation and artifact suppression. In the temporal domain, it quantifies sampling stability via the uncertainty trace to enable adaptive trajectory pruning, thereby accelerating convergence. Theoretical proofs demonstrate that our framework achieves strict data consistency, while the trajectory pruning strategy admits a deterministic error bound, thereby guaranteeing stable convergence under skip sampling. Notably, our plug‑and‑play method can be seamlessly integrated into various DMIR baselines. Extensive experiments show that LEADer improves the performance of multiple state‑of‑the‑art DMIR methods, while significantly reducing sampling time with negligible memory overhead. Code is available at https://github.com/JiaqiZhang‑Sengoku/LEADer.

Authors:Hongyu Luo, He Wang, Huihao Jing, Hong Ting Tsang, Yuxuan Liu, Wuganjing Song, Yauwai Yim, Chunyang Li, Yangqiu Song
Title: Can Language Models Imagine Without Seeing? Ekphrasis: Measuring Visual Creative Ideation in Text-Only LLMs
Abstract:
Current evaluations do not isolate whether text‑only language models can originate visual concepts before image generation. Fluent visual prose can hide visual‑plan failures: an answer may appear creative while repeating familiar visual clichés or failing to specify a renderable scene. We define Visual Creative Ideation (VCI) as the ability to produce textual visual plans that are useful, expressive, and population‑novel, and introduce Ekphrasis, a 400‑task benchmark spanning Abstraction, Combination, Transformation, and Adaptation. Ekphrasis scores anonymized pairwise comparisons with dimension‑specific checklists, aggregates preferences with Bradley‑Terry models, and uses Typed Idea Graphs to convert task‑specific population clichés into novelty references. Across 14 language models, VCI separates usefulness, expressiveness, and novelty rather than reducing to fluency: strong models achieve similar overall scores through different profiles, and useful plans can remain visually clichéd. A cross‑modal grounding study further shows that text‑level VCI ordering largely survives faithful rendering and blind image‑level preference judgment, supporting Ekphrasis as a measure of visual ideation beyond prose quality.

Authors:Alex Kwon
Title: Explicit, Not Longer: What Makes Epistemic Stance Survive Memory Compression
Abstract:
Agent memory systems compress what they store, and compression is built to drop qualifiers, so a claim's epistemic standing tends not to survive being written to memory. We ask what governs whether it does. Matched notes carry the identical claim and identical stance and differ only in where that stance sits; one model compresses both under the same budget among the same filler notes, and a blind reader that never sees the condition scores the result. Across 60 claims in seven registers, writing the stance as a labelled field rather than a bracketed aside raises retention by about 15 points on two models (37 claims to 2 on one, 30 to 8 on the other; permutation p=0.00005), and a pre‑registered replication on Haiku, its prediction and decision rule committed before the run, gives +15.6 points, 38 claims to 1. Ablating the format on both models gives the same net effect from different parts: labels help on both (+9.7 and +12.8) and length helps on neither, but wording the stance as a full sentence is the largest component on one model (+12.5) and worth nothing on the other (+0.6). Either model alone would have licensed a confident and different mechanism, so we claim only the intersection: make the stance explicit, not merely longer, and expect the best way of being explicit to depend on the model. A deterministic readout with no model reproduces the two‑cell direction and five of seven ablation contrasts, but not length or labels, which we therefore do not claim on one instrument. Fifty hand labels (kappa=0.75) agree on direction; we print their seven disagreements in full. We also report nine withdrawn claims, three of them former title claims of this paper.

Authors:Paul-Peter Arslan
Title: Does Splitting a Triage Decision Across Agents Hide Bias or Help Catch It? A Multi-Agent Simulation Study of LLM-Based Resource Allocation Under Audit Capacity Constraints
Abstract:
Prior benchmarking work has shown that a single large language model (LLM), forced to make life‑or‑death resource‑allocation decisions, exhibits measurable demographic bias. Real deployments, however, rarely use a single agent: they use pipelines, with review steps meant to catch exactly this kind of failure. We study what happens to bias when the same decision is distributed across a role‑differentiated multi‑agent pipeline (assessment, allocation, independent audit) instead of made and checked by one model alone. Using a synthetic disaster‑triage simulator with paired cases that are clinically identical except for one demographic attribute, we run 192 episodes (2,304 resolved case pairs) on GPT‑4o‑mini comparing a single‑agent control condition to a nine‑agent pipeline under three independently varied pressure dimensions. We find no measurable difference in how often biased outcomes occur between the two conditions (6.9% vs. 6.1%, p = 0.498). We do find a large and significant effect of audit capacity on whether bias is caught: 30.0% of biased outcomes go entirely undetected, rising to 43.8% when the auditor is overloaded and falling to 18.4% when it is not. Decomposing this effect shows it is driven almost entirely by coverage (whether a case is reviewed at all, which collapses from 100.0% to 65.6% under load, p < 0.001) rather than by degraded judgment on the cases that are reviewed (81.6% vs. 85.7%, p = 1.000, direction reversed). A follow‑up experiment shows that reordering the audit queue by estimated risk, rather than first‑come‑first‑served, recovers most of the lost coverage under the same capacity constraint (65.6% to 91.7%, p = 0.028). We discuss the implications for any system that adds independent oversight to an LLM agent pipeline under resource constraints, and report the study's limitations honestly: one model, modest sample sizes, and no adversarial replication.

Authors:Yingtao Ren, Ziyi Zhao, Yiwei Fu, Xiao Luo, Yu-Cheng Chang, Chin-Teng Lin
Title: When Context Bites: Detecting RAG Poisoning via Document-Level Attention Collapse
Abstract:
Retrieval‑augmented generation (RAG) is indispensable for enhancing large language models. However, RAGs are increasingly susceptible to poisoning attacks, in which adversarial documents are injected to manipulate generator outputs. Previous methods rely on output‑side signals such as perplexity and consistency checks to detect such attacks. Nevertheless, our analysis reveals that deliberate attacks often induce false confidence, where poisoned outputs exhibit even lower perplexity than benign ones, rendering uncertainty‑based detection ineffective. To address this challenge, we explore the internal dynamics of the generator and identify a distinctive signature termed Attention Collapse. Unlike the dispersed attention in benign generations, attacked generations exhibit a decrease in entropy as attention concentrates on poisoned documents. Building on these findings, we propose \textttD‑SCAN (Document‑level Signal Collapse Analysis), a lightweight detection framework that monitors attention dynamics to identify attacked generations. Extensive experiments on multiple attack benchmarks demonstrate the effectiveness of our method. Moreover, D‑SCAN can detect attacks even when they fail to alter the final answer. Code is available at https://github.com/yingtaoren/D‑Scan.git.

Authors:Rui Xu, Yang Yong, Shunzi Yang, Ruihao Gong, Chengtao Lv
Title: MaskFlow: Precise, Consistent and Seamless Regional Image Editing
Abstract:
Regional image editing has attracted considerable attention for its spatial controllability. Although instruction‑based and mask‑reference‑based editing methods can achieve strong semantic alignment, reliable regional control remains challenging, where an edit must be accurately localized and naturally integrated with the preserved context. We propose MaskFlow, a training framework for precise localization, consistent background preservation, and seamless boundary transitions. MaskFlow incorporates the mask into the probability path and flow‑matching objective, coordinating generation within the editable region with source preservation outside it. The proposed Soft‑Poisson de‑seaming module further refines the predicted vector field during both training and sampling to improve the smooth integration of the edited foreground with the preserved background. We also design a data synthesis pipeline to construct MEData, a mask‑based image editing dataset for training regional image editing models and facilitating further research. Experiments on natural scenes and infographic images demonstrate consistent improvements over competing methods in both quantitative and qualitative evaluations. Project page: https://reychiaro.github.io/MaskFlow

Authors:Oliver Lemke, Alexander Liniger, Abel Gawel, Marco Hutter
Title: Vernata: Self-Supervised Learning of LiDAR Point Representations
Abstract:
LiDAR serves as a primary sensing modality for robots operating in outdoor environments. However, the performance of deep learning models in this domain is severely limited by the scarcity of labeled data, a direct result of the high cost of 3D annotation. Self‑supervised learning addresses this scarcity by learning general‑purpose features from unlabeled data. In this work, we present a multi‑modal, multi‑teacher distillation framework for self‑supervised learning on outdoor LiDAR point clouds. Building upon the Sonata architecture, we introduce Vernata, consisting of three extensions: sparse view augmentation to improve robustness against varying point densities, a memory bank mechanism to stabilize resource‑constrained training, and cross‑modal distillation utilizing dense, high‑resolution 2D image features to enable fine‑grained semantic guidance. We evaluate our method on the GrandTour, TartanGround, and Waymo datasets, as well as data collected from our own robotic platforms. Our experiments demonstrate a significant performance improvement over Sonata baselines, yielding mIoU scores of 54.7 on TartanGround (+5.9 points, +12.1%) and 57.1 on Waymo (+7.3 points, +14.7%). Finally, we show that the self‑supervised approach maintains strong performance even in reduced‑modality settings (lacking color or normals), achieving competitive mIoU scores of 49.4 and 50.2 on the respective datasets.

Authors:Dilek Yargan, Jörg Waitelonis, Mahsa Vafaie, Harald Sack
Title: BZKO: An Ontology for the Card Index of German Post-War Compensation Records
Abstract:
The Central Federal Card Index (Bundeszentralkartei) of Germany is a key archival resource documenting compensation claims submitted by victims of National Socialist persecution and their relatives, within the German Wiedergutmachung process. To enable semantically enriched representation, integration, and reuse of this historically significant collection, we present the BZK Ontology (BZKO). We propose a two‑layer ontology for historical archival data that separates ontologically grounded domain semantics from interoperability‑oriented extension constructs. The approach combines BFO‑based realism with archival standards (RiC‑O, PROV‑O, PiCo), enabling provenance‑preserving semantic integration, while maintaining logical rigor, modularity, and reuse across digital humanities infrastructures. The proposed approach establishes a reusable semantic foundation for the integration of Wiedergutmachung archival materials into digital humanities infrastructures and lays the groundwork for future knowledge graph generation, ontology validation, and the incorporation of additional historical entities and uncertain temporal and spatial information. The ontology is available on https://github.com/ISE‑FIZKarlsruhe/bzko.

Authors:Jing Chen, Yang Sun, Li Zhang, Lin Xu, Jie Shi
Title: Long-Horizon Agent Trajectory Attribution: A Unified Benchmark and Fine-Grained Annotation Framework
Abstract:
Large language model (LLM) agents increasingly operate through long‑horizon trajectories involving user instructions, tool use, external observations, and memory. Existing benchmarks primarily evaluate behavioral outcomes but provide limited support for fine‑grained attribution analysis. We introduce trajectory attribution and develop a benchmark and annotation framework for this task. The benchmark organizes heterogeneous trajectories under a unified component schema and provides annotations of the primary attribution component, together with attack and execution chains where applicable. Instantiating the benchmark with trajectories from AgentDojo and the Stage and Canary settings of Agent3Sigma yields more than 1,300 annotated trajectories covering task‑aligned actions, unsafe actions, and safety refusals. The benchmark defines two evaluation tasks, primary attribution localization and attribution‑chain recovery, and provides reference baselines based on incremental trajectory contribution and component‑level leave‑one‑out perturbation. It captures diverse attribution settings, including local and long‑range attribution as well as structured attribution chains. Reference baseline results exhibit substantial performance differences across these settings, providing an initial characterization of the benchmark's attribution challenges. Beyond this initial instantiation, we release a reusable annotation skill that enables trajectories generated by new agent models to be standardized, annotated, and evaluated under the same framework. Project resources and future releases are available at https://github.com/chenjing‑2024/agent‑trajectory‑attribution.

Authors:Minseok Kang, Hyunwoo Kim, Chanyoung Kim, Minwoo Kim, Jaekoo Lee, Dahuin Jung
Title: Prune Once: Retraining-Free Task-Agnostic Pruning for Vision-Language Models
Abstract:
Vision‑language models (VLMs) have achieved remarkable generalization across diverse multimodal tasks through large‑scale pre‑training, yet their rapidly increasing computational and memory requirements pose significant challenges for deployment in constrained environments. Existing pruning strategies often depend on task‑specific criteria or LLM‑oriented importance measures, making them unsuitable for task‑agnostic pruning, where no task‑specific samples are available at pruning time and the pruned model remains broadly applicable. We introduce a retraining‑free VLM pruning framework called PORTA that derives a task‑ and modality‑agnostic importance formulation based on activation variation, estimated from generic calibration data, which reliably captures feature‑level representation utility across modalities. PORTA further incorporates an adaptive sparsity allocation mechanism that assigns layer‑wise pruning ratios based on output feature variability, avoiding the limitations of uniform sparsity and reducing performance degradation at high compression levels. Extensive experiments across VLM architectures, such as CLIP, BLIP, and Qwen2‑VL, demonstrate that PORTA achieves competitive downstream performance under high sparsity without requiring any retraining, supporting efficient VLM compression. Code is available at https://github.com/cau‑hai‑lab/PORTA.git.

Authors:Zhentao Tan, Ruijie Quan, Yi Yang
Title: From Points to Edges: Edge-Conditioned Spectral Operators for Physics-Sensitive PDE Learning
Abstract:
Neural operators have become a central tool for solving partial differential equations (PDEs), with spectral operators offering efficient global mixing across spatial locations. However, many PDEs contain physics‑sensitive local structures that are critical to the underlying physical behavior. For example, in Darcy flow, local material interfaces are often reflected by sharp changes in the permeability field and can strongly influence the solution. Existing spectral operators primarily adapt modal mixing based on center‑point representations, making them insufficiently responsive to such localized structural variations. We propose the Edge‑Conditioned Spectral Operator (ESO), a novel spectral operator framework that modulates global spectral mixing using local edge‑wise variations. By incorporating the Pairwise‑Variation Modal Mixer (PVMM) to inject local edge information into spectral mode selection, ESO preserves the global approximation capability of spectral neural operators while enabling the learned kernel to adapt to physics‑sensitive local structures. Furthermore, we introduce a task‑adaptive Physics‑Aware Reweighting (PAR) that emphasizes physically important regions, identified by taskspecific physical quantities. Across nine PDE benchmarks, ESO consistently achieves state‑of‑the‑art performance. Visual and region‑wise analyses further demonstrate that ESO reduces solution errors near coefficient jumps, high‑gradient flow structures, and other physically sensitive regions. The code is available at https://github.com/Tanpig‑X/ESO.

Authors:Xuechao Zou, Shun Zhang, Kai Li, Yi Zhou, Xinyu Sun, Yuhui Chen, Zhe Wu, Congyan Lang, Junliang Xing
Title: Multi-Agent Forensic Reasoning for Generalizable Deepfake Video Detection
Abstract:
The malicious use of generative artificial intelligence to create highly realistic deepfake videos raises serious ethical concerns and poses substantial challenges to AI safety. However, existing deepfake video benchmarks provide limited coverage of recent synthesis methods and generally lack reliable fine‑grained textual annotations. Meanwhile, conventional detectors and multimodal large language models (MLLMs), whether operating as a single model or relying on a single analytical perspective, often fail to capture subtle forgery artifacts, limiting their generalization to emerging AI‑generated methods. To address these limitations, we introduce FaceVid‑Forensics‑100K, a large‑scale deepfake video dataset comprising 100,000 videos and spanning 33 synthesis methods across face swapping, face reenactment, and entire‑face synthesis, including recent generators such as Seedance 2.0. The dataset provides fine‑grained textual annotations of visual observations and verdict‑consistent forensic explanations, automatically synthesized through a multi‑model aggregation and conflict‑resolution pipeline powered by advanced MLLMs. Building on this benchmark, we propose a multi‑agent forensic reasoning framework that employs four specialized domain‑expert agents to independently analyze forgery cues from four perspectives: texture, lighting, motion, and physics. A judge agent then reconciles their reports to produce a final prediction together with an explanation. Extensive evaluations on out‑of‑domain test sets show that, despite being composed entirely of small open‑source MLLMs, our framework outperforms all methods including closed‑source GPT and Gemini models and ranks first across all reported metrics on this benchmark. The project page is available at https://xavierjiezou.github.io/ARGUS/.

Authors:Nguyen-Truong Thinh, Yuxuan Du, Phongsakon Mark Konrad, Arpit Narechania
Title: Fact-Check Your Information (FYI): A Design Probe to Understand How People Actually Fact-Check Data-Driven Articles
Abstract:
Data‑driven journalism and policy reports frequently rely on statements grounded in statistical evidence, referred to as data claims. Verifying such a claim requires connecting it to the underlying structured dataset. However, existing systems typically isolate automated fact‑checking from manual data exploration, leaving it unclear how readers coordinate AI assistance with manual inspection of the evidence in practice. We present FYI, a browser extension that embeds fact‑checking in the reading environment, and use it as a design probe to study how people detect, verify, and determine the validity of data claims against the underlying dataset. FYI provides four complementary tools spanning the spectrum from full automation to manual data exploration. In an exploratory study (N=22), participants used FYI to fact‑check claims in a data‑driven article. We find that participants adopted three distinct workflow archetypes‑‑‑AI‑first with manual confirmation, manual‑first with AI supplement, and parallel co‑review‑‑‑with visualization serving as the primary mechanism for auditing AI conclusions. Trust in AI shifted dynamically, growing when multiple tools converged and eroding when AI outputs were inconsistent. These findings suggest that fact‑checking systems should treat AI as a starting point that human verification complements rather than a definitive authority, elevate visualization as a core verification capability, and support flexible, user‑driven workflows. We release FYI as open‑source software for further research at https://github.com/DataVisards/FYI.

Authors:Oseong Choi, Hoeinn Kim, Jihoon Lee, Byungsoo Kang, Taeyeong Jang
Title: Progressive Alignment of Recommender Foundation Model through Multi-Phase Post-Training
Abstract:
Foundation model(FM) for recommendation has shown strong ability to model long‑horizon sequential user behavior. In practice, a single pretrained foundation model is often adapted to diverse downstream serving surfaces through Supervised Fine‑Tuning(SFT). However, optimizing task‑specific objectives such as clicks or likes does not necessarily align the serving policy with the business metrics that determine recommendation quality. We propose a three‑phase progressive post‑training framework that explicitly separates downstream adaptation from business‑metric alignment. The adaptation stage is decomposed into Linear Probing(LP) and Full Fine‑Tuning(FFT): LP first stabilizes randomly initialized downstream heads within a frozen pretrained representation space, and FFT then jointly specializes the full model for the target task. On top of this stabilized policy, Reinforcement Fine‑Tuning(RFT) aligns the model with practical business objectives using a learned reward model. Rather than directly optimizing the serving policy on sparse business targets, we train the policy on dense implicit feedback and use business‑metric supervision only for reward modeling. Offline experiments show that the progressive LP‑FFT‑RFT framework outperforms single‑phase alternatives, and that reward‑based alignment yields a stronger serving policy than directly using the reward model itself for ranking. Large‑scale online A/B tests further show that the proposed framework improves production recommendation quality over a conventional non‑foundation baseline. A reference implementation is available at https://github.com/webtoon/rec‑fm‑progressive‑alignment

Authors:Hao Li, Yunzhi Zhuge, Wenning Hao, Pingping Zhang, Xiaoxiong Zhang, Dong Wang, Huchuan Lu
Title: AnyTrack: Unifying Visual Object Tracking with Any Modalities
Abstract:
Visual object tracking aims to continuously locate specific targets within sequential frames, evolving from single‑modal methods to multi‑modal ones. However, existing multi‑modal trackers are typically designed for fixed modality combinations, requiring separate models for different inputs. This leads to a poor adaptability to missing or imperfect modalities, and limited generalization. To address these issues, we propose a novel unified framework called AnyTrack for object tracking with any modalities. Specifically, we design a Modality‑aware Interaction Module (MIM) to facilitate dynamic interaction across diverse modalities. This module bridges modality discrepancies and aggregates temporal cues to maintain spatio‑temporal consistency during cross‑modal interaction. Furthermore, we introduce a Context Understanding Module (CUM) to establish spatial correspondence between visual features and target locations via global‑local prompts. This module employs target‑aware context modeling to enhance foreground‑background discrimination for precise localization. Finally, to support the training and evaluation under diverse modalities, we extend existing multi‑modal object tracking benchmarks by incorporating grayscale images, language descriptions, and audio clips. Extensive experiments with both complete and missing modality settings demonstrate that our AnyTrack achieves state‑of‑the‑art performance, validating its effectiveness and flexibility. The source code is available at https://github.com/IdolLab/AnyTrack.

Authors:Yuanfu Sun, Yuanhang Ren, Kang Li, Chuanhao Ji, Jiaxi Li, Jiajin Liu, Ninghao Liu, Qiaoyu Tan
Title: GraphVerse: A Comprehensive Visual Graph Reasoning Benchmark for Multimodal Large Language Models
Abstract:
Recent Multimodal Large Language Models (MLLMs) have achieved remarkable progress across diverse vision‑language tasks, creating an urgent need for more challenging benchmarks. Yet existing evaluations still provide limited insight into whether these models can truly reason over structured visual information. Visual Graph Reasoning (VGR) offers a compelling testbed for this challenge, requiring models to integrate perception, structural understanding, and multi‑step reasoning over graph‑based visual inputs. However, prior VGR benchmarks often reduce the task to visual perception followed by text‑based reasoning, restrict evaluation to single‑image settings, rely on answer‑only metrics, and underrepresent realistic graph‑centric scenarios. To bridge the gap, we introduce GraphVerse, a unified benchmark that jointly evaluates perception, visual reasoning, and text‑based graph reasoning in MLLMs under both single‑image and paired‑image settings. At its core is a suite of Graph‑centric Image Editing (GIE) strategies that modify graph images while preserving their semantics, turning them into active tests of visual reasoning. We further propose VGR‑Score, a process‑sensitive metric that evaluates reasoning quality beyond final‑answer accuracy. Extensive experiments reveal several key limitations of current MLLMs in VGR, while also validating the effectiveness of GIE strategies and the transferability of GraphVerse to broader multimodal reasoning capabilities. The code is available at https://github.com/sunyuanfu/GraphVerse.

Authors:Zibo Shao, Baochen Xiong, Chengdong Xu, Linhui Xiao, Kaichen Li, Haoran Gong, Yan Li, Yaguang Song, Xiaoshan Yang
Title: AgentPatch: Coarse-to-Fine Weak-Task Repair for Merging Agentic Multimodal Large Language Models
Abstract:
Agentic multimodal large language models (MLLMs) extend multimodal perception and reasoning with planning, tool use, and interaction in dynamic environments. Yet current models are specialized for particular tools or environments, complicating consolidation into a single generalist. We formulate Agentic MLLM Merging and identify two challenges: asymmetric capability preservation, whereby capabilities with different interaction complexity are retained unevenly, producing weak tasks after merging, and behavior‑critical forgetting, whereby losing decisive actions can derail long‑horizon execution. We propose AgentPatch, a training‑free coarse‑to‑fine repair framework. It selects a stable merged backbone, restores diluted weak‑task‑specific signals through Weak‑Task Unique Residual Recovery, and applies an Agent‑Guided Behavior‑Critical Patch that recovers decisive behaviors under explicit capability protection. AgentPatch produces a single static checkpoint without routing or ensembles. Experiments across six agentic and multimodal benchmarks show that AgentPatch improves diverse merged backbones, alleviates weak‑task degradation, and better balances weak‑task recovery with the preservation of complementary search and agentic visual processing capabilities. Code is available at https://github.com/ziboshao/AgentPatch.

Authors:Novendra Setyawan, Chi-Chia Sun, Mao-Hsiu Hsu, Wen-Kai Kuo, Jing-Ming Guo, Jun-Wei Hsieh
Title: CoDAT: Collaborative Dual-Attention Transformer with Low-Cost Temporal Modeling for Efficient Edge Action Recognition
Abstract:
Real‑time human action recognition on Internet‑of‑Things (IoT) edge devices requires models that capture rich spatio‑temporal cues within strict latency, memory, and power envelopes. Current 3D CNNs, video transformers, and shift‑based ViT deliver high accuracy but come at computational costs that preclude edge IoT deployment. This paper proposes CoDAT, a Collaborative Dual‑Attention Transformer that replaces conventional multi‑head attention with a lightweight dual‑branch module: Spatial Convolutional Attention (SCA) for local aggregation and Strided Single‑Head Attention (SSHA) for global context. SSHA jointly compresses the spatial resolution and channel dimensions of the query, key, and value tensors via stride‑based sparse projection, then fuses the resulting global and local features at a markedly reduced cost. To enable temporal communication across frames, a parameter‑free TShift module is embedded in each block. Extensive experiments on Jetson AGX Orin and Raspberry Pi 5 demonstrate that CoDAT achieves an energy‑accuracy balance in both image and action recognition. On ImageNet‑1K, CoDAT‑M runs 2x faster than EfficientViT384 and FastViT‑S12 at comparable accuracy, and CoDAT‑L matches ViT‑S with 3x fewer parameters at 2x higher throughput. On Kinetics‑400 and MA‑52, CoDAT achieves competitive Top‑1 accuracy against state‑of‑the‑art CNN, transformer, and hybrid baselines while running up to 2.9x faster than VSwin‑T, 2x faster than ViT‑Temporal‑Shift variants, and 5x faster than UniFormer‑B. On UCF‑101, CoDAT‑S384 matches TokShift and LAPS while being 6x faster and requiring up to 13x fewer FLOPs, establishing an efficiency‑accuracy balance for real‑time action recognition in edge IoT perception systems. Code is available at https://github.com/novendrastywn/CoDAT .

Authors:Haiping Liu, Qian Zhao, Lijing Lin, Jingyuan Sun, Hongpeng Zhou
Title: CellWorld: From Gene-Level Reconstruction to Latent Cell Prediction in Spatial Transcriptomics Foundation Models
Abstract:
This paper shows that latent‑space predictive pretraining can provide a scalable route to foundation models for spatial transcriptomics. Existing spatial transcriptomics foundation models primarily reconstruct masked gene identities or expression values, potentially encouraging the reproduction of assay‑specific technical variation and limiting representation transferability. To avoid directly reconstructing such variation, we shift the prediction target from observed gene measurements to latent cell representations and introduce CellWorld, which predicts the latent representations of masked cells from visible spatial context and a limited partial‑expression hint. We pretrain four CellWorld variants, spanning 5.74M to 94.56M trainable parameters, on a corpus of 46 million human cells. Our controlled scaling experiments show that performance improves with model capacity, particularly on spatial tasks, while spatial transfer depends more on sufficient optimization and broad biological source diversity than on cell count alone. Across four held‑out datasets, even CellWorld‑Small, with 5.74M trainable parameters, outperforms every baseline on all 11 linear‑probe benchmarks and all seven fine‑tuned spatial benchmarks. Most notably, a frozen CellWorld‑Large pretrained on only 5% of the corpus with broad biological source coverage outperforms every fully fine‑tuned baseline across all seven spatial benchmarks. Code is available at https://github.com/UoM‑HealthAI/CellWorld.

Authors:Linghao Kong, Inimai Subramanian, Micah Adler, Dan Alistarh, Dan Gutfreund, Nir Shavit
Title: The Sparsity Whisperer
Abstract:
Pruning reduces the inference cost of large language models, but existing criteria primarily preserve large activations or reconstruct layer outputs. We argue that this overlooks a key computation performed by particularly sparsity‑sensitive neurons in the MLP up and gate projections: separating similar inputs into dissimilar outputs. This suggests that effective pruning should preserve not only activations, but also the differences between outputs more broadly. We introduce a family of difference‑informed pruning methods built upon this principle. Wisp is a first‑order, update‑free method that scores weights using input‑difference norms, and Wisp+ refines this score neuronwise using the input pairs each neuron separates most strongly. Finally, Whisper is a second‑order method that uses a lightly regularized difference Hessian as its reconstruction objective. Across Llama 2 and 3.1 models from 7B to 405B parameters, our second‑order variant consistently improves over strong reconstruction‑based baselines, while our update‑free variants improve over activation‑aware baselines, especially in constrained settings. The improvements over Wanda and SparseGPT extend to structured sparsity, downstream evaluations, and other model families. Augmenting stronger techniques such as RIA and ALPS with our difference‑informed criteria yields further improvements, shifting the overall accuracy‑runtime frontier outward at negligible additional cost. These results suggest that preserving output differences is a broadly useful and composable signal for post‑training LLM sparsification.

Authors:Jinha Kim, Younghun Roh, Jaeyeon Kim
Title: Retrofitting Linear Attention into Diffusion Language Models
Abstract:
Diffusion language models (dLLMs) offer a promising alternative to autoregressive models by accelerating inference through parallel decoding. Recent dLLMs commonly use blockwise semi‑autoregressive decoding, generating blocks autoregressively while denoising tokens within each active block in parallel. However, despite KV caching, each denoising step still attends to all previous blocks, repeatedly incurring prefix‑attention cost. Motivated by this bottleneck, we ask whether dLLM inference can be further accelerated by linearizing attention over previous blocks. We introduce block‑hybrid attention, which retains exact softmax attention within the active denoising block while applying linear attention over previous blocks. We show that this hybrid attention can be retrofitted into a pretrained dLLM with minimal post‑training: LLaDA‑Hybrid replaces 6 of the 20 attention layers in LLaDA~2.1, a 16B open‑source dLLM, largely following LoLCAT (Zhang et al, 2024). The conversion takes only approximately 60 hours while preserving benchmark performance: 72.0% vs. 75.6% on HumanEval, 63.0% vs. 57.7% on MBPP+, and 86.7% vs. 88.3% on CMATH. With a Triton implementation, LLaDA‑Hybrid achieves up to 1.7× higher decoding throughput and supports more concurrent requests before exhausting memory, showing that pretrained dLLMs can be efficiently linearized for faster inference. Our code is available at: https://github.com/Diuven/LLaDA‑Hybrid.

Authors:Sreerekha Rajendran
Title: Pre-Inference Routing for Cost-Efficient Document Field Extraction
Abstract:
Most document‑extraction systems use a single model for all documents. This is simple but can be costly for easy cases and less effective for difficult ones. We examine whether we can predict a document's difficulty before extraction using inexpensive, document‑based signals, and use this to choose between a cheaper and a stronger extractor. We find that routing only helps if two conditions hold: the cheaper model fails often enough to make routing worthwhile, and those failures can be predicted from visible features such as image quality and layout. We turn these into a practical test and apply it to five genres. When both conditions are met, the calibrated router reduces cost by 31‑33% on receipts and 77% on degraded ad‑buy forms while keeping quality within 0.02 F1 of always choosing the large model. Routing does not help if either condition is missing, as with clean digital invoices or nutrition labels that are already easy to read. A small labeled pilot can predict whether routing will work, and in the two cases where we ran it first, the prediction was correct. A simple bag‑of‑words router works about as well as engineered features, showing that the main limit is the genre, not the router design; we use interpretable features to help explain which genres can be routed. The router must be retrained for each dataset and does not transfer across datasets, even within the same genre. These results hold for two model pairs with cost differences of 5x and 3x.

Authors:Oren Nelson
Title: Beyond Attention: Signed Integrated Gradients Attribution in a BiomeGPT-Style Microbiome Transformer
Abstract:
In a feature‑tokenized transformer (arXiv:2106.11959) such as BiomeGPT (doi:10.64898/2026.01.05.697599), each input token is built by fusing a fixed identity with a sample‑specific measurement: a fixed species and a variable abundance, T = S + A. To interpret downstream classification in such models, prior work inspects the attention weights of the special [CLS] token (arXiv:2106.11959, arXiv:1810.04805, BiomeGPT) to rank sample tokens by importance. These weights have two critical limitations: they are nonnegative, so they cannot separate disease‑supporting from health‑supporting evidence (arXiv:2201.12114), and they act after token fusion, obscuring how the input sources S and A each affect the output. To address this we use Integrated Gradients (arXiv:1703.01365), a signed, fusion‑aware attribution method, and propose a source‑derived baseline T' = S + A_0 for feature‑tokenized models such as BiomeGPT, which preserves species identity as a fixed biological coordinate while isolating the effect of abundance variation. Applied to a disease‑versus‑health decision margin, it yields polarity that explicitly separates pathogenic from protective microbial signals. We show that this gradient‑based approach uncovers species‑abundance directional relationships and sensitivity diagnostics entirely obscured by unsigned [CLS] attention weights. We further recommend second‑order Integrated Hessians (arXiv:2002.04138) to expose microbiome community interaction rules: how a perturbation in one member alters the model's sensitivity to another, and which other species drive ambiguous cases toward disease or health at a given abundance level. This provides a principled approach to explainability in BiomeGPT that generalizes to other smooth and differentiable feature‑tokenized transformers. Code is available at https://github.com/nohren/token‑source‑attribution

Authors:Zhuoxin Zhan, Akbar Rafiey, Avery Ma, Leila Pishdad, Layla El Asri
Title: StepJack: Benchmarking Computer-Use Agent Safety Against Multi-Step Indirect Prompt Injection
Abstract:
Computer‑use agents (CUAs) face a growing threat from indirect prompt injection, where adversarial instructions are planted in the environment such as web pages. In this paper, we introduce multi‑step indirect prompt injection, a new attack class against CUAs in which the adversarial goal is decomposed into multiple innocuous‑looking sub‑steps and distributed across a chain of pages referenced along the agent's navigation path. We develop a pipeline to automatically decompose an adversarial goal under the constraint that the execution of the decomposed sub‑steps must achieve the original goal while optimizing the innocuousness of each decomposed sub‑step. With this pipeline, we build StepJack, a CUA safety benchmark with 480 test examples. On this benchmark, we evaluate six state‑of‑the‑art CUAs and find that at a fixed decomposition depth, multi‑step attacks raise attack success rate (ASR) on three of six CUAs, by up to 31.2 points (e.g., GPT‑5.4‑mini: 41.7% at single‑step to 72.9% at three‑step); averaged over the five CUAs that can reliably follow the reference chain (all but EvoCUA‑32B), ASR rises from 31.3% at single‑step to 36.9% at three‑step. Dataset and code are available at https://github.com/BorealisAI/StepJack.

Authors:Masoumeh Sharafi, Muhammad Osama Zeeshan, Soufiane Belharbi, Alessandro Lameiras Koerich, Marco Pedersoli, Eric Granger
Title: Test-Time Adaptation with Online Personalized Energy-Based Cache for Fine-Grained Video Expression Recognition
Abstract:
Facial expression recognition (FER) in videos is challenging because models must identify subtle, temporally evolving affective states that vary across individuals. Although vision‑language models provide transferable visual‑semantic representations, models trained on subject‑independent data often degrade under subject‑specific distribution shifts at inference time. Existing test‑time adaptation (TTA) methods commonly update model parameters during inference, increasing computational cost and latency. Cache‑based methods avoid parameter updates, but they usually require enough target samples to form reliable class prototypes, which is difficult early in adaptation and for rarely observed classes. We introduce Energy‑Based Cache Personalization (EB‑CaP), a subject‑based online TTA method for video FER that generates class‑specific prototypes personalized to each target video. EB‑CaP uses a lightweight energy‑based model to sample prototypes from the current unlabeled video and populate a personalized cache online, without accumulating large amounts of target data or storing diverse source prototypes. Its energy function relies only on pretrained CLIP: similarities between the target video embedding and class text embeddings guide prototype sampling. In parallel, positive and negative caches store reliable and uncertain target embeddings. An adaptive entropy gate controls cache updates according to the evolving confidence distribution, while a diversity gate limits redundant samples. Final predictions combine cache‑derived scores with the current CLIP scores. Experiments on BioVid, StressID, and BAH show that EB‑CaP outperforms state‑of‑the‑art TTA methods while maintaining low computational and memory overhead. Code is available at https://github.com/MasoumehSharafi/EB‑CaP.

Authors:Iftach Shoham, Tali Dror, Oren Gal, Haim Permuter, Gilad Katz, Eliya Nachmani
Title: Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing
Abstract:
Speech recordings often contain missing, corrupted, or incorrect regions that must be reconstructed or modified without re‑synthesizing the entire utterance. Speech inpainting restores missing segments, whereas speech editing replaces spoken content according to an edited transcript. Both tasks require the generated speech to express the intended words while remaining consistent with the surrounding speaker identity, prosody, timing, and recording conditions. Discrete diffusion is particularly well suited to these tasks because it can iteratively refine masked tokens while jointly conditioning on both left and right acoustic context. We introduce SIEDD, a discrete diffusion framework for text‑guided speech inpainting and editing over hierarchical codec tokens. Its core architecture, HiCoDD, follows the RVQ generation order by representing previously generated codebooks as clean, committed acoustic context and applying diffusion only to the current refinement codebook. This separation enables leakage‑free joint training while matching sequential coarse‑to‑fine inference. The model further combines phoneme‑level conditioning, span‑localized classifier‑free guidance, and duration prediction to support both fixed‑duration inpainting and variable‑duration text edits. On the RealEdit benchmark, SIEDD achieves the best overall speech‑editing performance among the evaluated methods. It also outperforms the evaluated autoregressive baselines across all speech‑inpainting settings, on both single and multiple gaps. These results demonstrate that explicitly modeling the codec hierarchy substantially improves context‑preserving speech reconstruction and editing. See our full code at https://github.com/iftachShoham/SIEDD.

Authors:Pedro T. Barcelos, Otávio Parraga, Marcelo M. Mussi, Lucas M. Fraga, Lucas S. Kupssinskü, Rodrigo C. Barros
Title: Latent Fact-Checking: Detecting Misinformation through Activation Engineering
Abstract:
The proliferation of misinformation online has driven demand for scalable detection systems. While most existing approaches rely on surface‑level linguistic features or external knowledge retrieval, we examine truthfulness as a geometric property of a language model's representation space. We introduce a misinformation detection framework grounded in activation engineering, which leverages the latent geometry of transformer models. Our approach elicits a misinformation direction in the residual stream by contrasting activations from paired truthful and false statements, following the difference‑in‑means principle of Contrastive Activation Addition (CAA). At inference time, the last‑token activation of an unseen claim is projected onto this direction, and the projected representation is fed to an Multilayer Perceptron (MLP) for classification. The procedure requires no fine‑tuning of the backbone model, no external evidence retrieval, and no task‑specific supervision beyond the contrastive pairs used to estimate the direction. We evaluate the method across 11 models from the Gemma, Llama, and Qwen families, ranging from 270M to 12B parameters, on three fact‑checking benchmarks: AVeriTeC, LIAR, and FACTors. The falsehood direction is recoverable across model scales and architectural families, and last‑token projection matches or surpasses zero‑shot and few‑shot prompting baselines on LIAR and FACTors, with the largest gains observed for smaller models. Performance on AVeriTeC is more limited, which we attribute to its evidence‑grounded labeling scheme. These findings provide evidence that truthfulness is a structured, linearly separable concept in the latent space of pretrained language models, and point toward interpretability‑driven misinformation detection as a practical complement to retrieval‑based pipelines. The code is available on https://github.com/Malta‑Lab/LaFaCt.

Authors:Junxiong Zhou, Xuechen Li, Chonghao Qiu, Lang Qiao, Xiaowei Jia, Qi Yang, Chishan Zhang, Leikun Yin, Nanshan You, Vipin Kumar, David Mulla, Ce Yang, Zhenong Jin, Licheng Liu
Title: UAV3DCrop: Benchmarking 3D Reconstruction in Repeated Multi-Angle UAV Crop Surveys
Abstract:
Accurate 3D crop monitoring underpins data‑driven precision agriculture by enabling field‑scale analysis of plant structure, growth dynamics, and management response. Modern 3D reconstruction methods perform strongly on generic benchmarks, but rendered appearance may not translate into metrically and agronomically useful geometry in crop fields. We introduce UAV3DCrop, a public benchmark of repeated multi‑angle unmanned aerial vehicle (UAV) crop surveys. It contains 88,830 RGB images at 5280 × 3956 pixels, with a ground sampling distance of 3.6‑5.8 mm, from 91 scenes spanning corn, soybean, wheat, and oat. Track A evaluates seven scene‑optimized methods ‑‑ Neural Radiance Field (NeRF) and 3D Gaussian Splatting (3DGS) variants ‑‑ on held‑out views, photogrammetry‑referenced depth, and canopy‑height recovery. Track B tests four pretrained feed‑forward models on zero‑shot camera‑pose and geometry estimation. The scene‑optimized methods rank differently across the three targets: Splatfacto‑big leads appearance, whereas Scaffold‑GS leads depth and is statistically tied with Splatfacto for canopy height. Among feed‑forward models, MapAnything leads on seven of the eight metrics, while the remaining models vary more across crops and fail severely on absolute scale in a way that alignment conceals. Repeated acquisitions reveal further sensitivities that differ by output type and by model, associated with position within the acquisition sequence and with tie‑point multiplicity. Current 3D reconstruction methods are therefore not yet interchangeable for agronomic use: no single method wins on appearance, geometry, and canopy height at once, and only one of four feed‑forward models recovers usable metric scale. The dataset is publicly available at https://link‑dev.github.io/UAV3DCrop/

Authors:Xian Sun, Wei Chow, Yingshuo Wang, Junhao Liu, Wei Gao, Qing Wu, Lingdong Kong
Title: Learning When to Trust via Selective Context Preference Optimization
Abstract:
Language models increasingly condition their answers on external signals, and a single misleading one can turn a correct answer wrong. The obvious remedy, training models to resist such signals, hides a failure mode: a model that ignores all context looks robust yet is useless when the context is worth trusting. We recast the problem as selective trust and introduce MIST, a human‑annotated benchmark that renders each reasoning item under four matched conditions (clean, misleading, correct‑context, and irrelevant‑context), together with SC2W, a paired metric counting how often a misleading signal flips a clean‑correct answer to wrong. Across a comprehensive benchmark study, we observe that such a susceptibility is universal. We then propose SCOPE, which mines clean‑correct/misleading‑wrong failures and optimizes a standard Direct Preference Optimization (DPO) objective over matched preference pairs balanced equally across all four conditions, rather than over misleading items alone. Our approach substantially reduces SC2W on popular open‑sourced models while preserving accuracy when the added context is clean, correct, or irrelevant. With this work, we argue that models should be judged on selective trust, not on resistance alone.

Authors:Fanzhe Meng, Guoxin Chen, Jiale Zhao, Shuang Sun, Zhiyu Lin, Wayne Xin Zhao, Ruihua Song, Ji-Rong Wen, Kai Jia
Title: CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks
Abstract:
Training terminal agents requires executable and verifiable tasks that are not merely solvable, but appropriately challenging for learning. Executable validation establishes feasibility, yet does not reveal how a task behaves relative to a given solver setting. In this paper, we present CalibForge, an autonomous terminal‑task synthesis system that uses verified solver behavior to revise candidate tasks through adversarial solver calibration. Multi‑solver calibration targets disagreement within a heterogeneous solver pool, whereas contrastive solver calibration targets a designated strong‑pass/weak‑fail relation; both operationalize a solver‑relative learnable zone anchored in demonstrated solvability. Using CalibForge, we construct 5,431 calibrated terminal tasks. Our ablations show that both strategies yield more effective supervision than authoring and validation alone or ordinary single‑solver feedback. Models trained on the full collection achieve 32.58% and 47.57% on Terminal‑Bench 2.0. The largest improvements over the corresponding base model reach 24.71 percentage points on Terminal‑Bench 2.0, 27.68 points on SWE‑bench Pro, and 30.04 points on Doc2Repo. Together, these results support solver‑relative learnability as a practical target for constructing effective and transferable agent training data.

Authors:Xinye Wang, Junxiao Liu, Shujian Huang
Title: RP-OPSD: Reasoning-Pivot-Guided On-Policy Self-Distillation for Multilingual Reasoning Transfer
Abstract:
Multilingual reasoning transfer is crucial for extending reasoning capabilities of large language models (LLMs) beyond high‑resource languages. On‑policy self‑distillation (OPSD) and its variants have emerged as a promising paradigm, providing dense token‑level supervision on student‑generated rollouts, yet their objectives do not explicitly prioritize reasoning signals most critical to cross‑lingual transfer. We characterize that target‑language reasoning comprises the generation of both surface text and reasoning pivots, which are decisions that advance or redirect the reasoning process and shape subsequent inference. This motivates concentrating privileged distillation around such pivots. We therefore propose RP‑OPSD, Reasoning‑Pivot‑guided On‑Policy Self‑Distillation, using the distributional shift between matched teacher views with and without an English reference solution as an operational proxy to guide privileged distillation and reference anchoring. Experiments on mathematical reasoning benchmarks covering 17 languages and multiple difficulty levels show that our method outperforms strong multilingual reasoning baselines and OPSD variants. Further analysis reveals that RP‑OPSD concentrates privileged distillation on reasoning‑control and problem‑condistioned state‑update tokens, while downweighting it for tokens that mainly support surface realization. Our code is available at https://github.com/NJUNLP/RP‑OPSD.

Authors:Chenglong Wang, Ziming Zhu, Yifu Huo, Bei Li, Qiaozhi He, Yan Ding, Xiaoyang Hao, Yuxin Gao, Tianhua Zhou, Xiaojia Chang, Tongran Liu, Jingbo Zhu
Title: RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction
Abstract:
Recent advances in reward modeling show a paradigm shift from discriminative reward models to generative reward models. However, despite their strong capabilities in response ranking, generative reward models have not realized their potential in reinforcement learning (RL). Our analysis reveals that this limitation arises from a mismatch between the comparative nature of generative reward modeling and the scalar scoring paradigm adopted by existing RL algorithms. To bridge this gap, we propose a Ranking‑based Reward Construction (RRC) approach, which enables generative reward models to provide more effective RL learning signals by deriving rewards from relative preference rankings. RRC introduces two complementary strategies: self‑competitive ranking, which exploits comparisons among sampled responses, and anchor‑guided ranking, which enables scalable ranking‑based reward construction with a small set of reference responses. Experiments across open‑ended chat and reasoning benchmarks demonstrate that RRC substantially improves RL training with generative reward models, achieving consistent gains over existing reward construction approaches. Our code can be found at https://github.com/wangclnlp/RRC.

Authors:Zhiheng Wang, Bo Peng, Lai Wei, Chaochao Lu
Title: The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images
Abstract:
The "thinking‑with‑images" paradigm equips multimodal LLMs with active visual operations such as crop‑and‑zoom. However, models using these operations often achieve only marginal or negative gains over direct inference at substantially higher token cost. They may also repeatedly crop irrelevant regions and fail on questions that direct inference answers correctly. We ask whether the returned visual evidence causally affects the answer. To answer this question, we formulate visual tool‑use as a causal graph that separates observation‑mediated paths from action‑induced shortcuts. We then audit it through interventions at the three levels: policy (comparing tool‑use with direct inference), trajectory (corrupting all observations during rollout), and step (counterfactually replacing one individual observation under a fixed prefix). Our step‑level estimand, Visual Evidence Gain, isolates the contribution of each returned observation. Across six representative models and five fine‑grained perception benchmarks, we uncover policy miscalibration with two failure modes. In Calling Without Looking, returned observations have no causal effect on the answer. In Looking Without Planning, observations are informative but the call schedule is incoherent. A trajectory‑level diagnostic decomposes the policy‑level accuracy gain and shows that the gain is concentrated in a Calibrated minority. We term this discrepancy the illusion of visual tool‑use: despite aggregate accuracy gains, visual tool‑use is not causally effective across a broad range of rollouts. The code is available at https://github.com/OpenCausaLab/CauAudit.

Authors:Stefan Dziembowski, Grzegorz Fabiański, Daniele Micciancio, Rafał Stefański
Title: Game Hopping in Lean
Abstract:
We present HOPSCOTCH, a Lean 4 framework for mechanizing computationally sound, game‑based cryptographic proofs. Security definitions are expressed as indistinguishability between stateful probabilistic oracles, and proofs follow the standard game‑hopping paradigm. HOPSCOTCH uses a shallow embedding: oracles and reductions are ordinary Lean definitions, enabling direct integration with the full Lean ecosystem, including general mathematical theories from Mathlib, such as finite‑group theory. A game‑hopping proof in HOPSCOTCH is represented as an explicit formal object whose constructors correspond to the standard steps of a game‑hopping argument, making proofs easier to construct, automate, and inspect. We prove a general computational soundness theorem that interprets these proof objects by constructing reductions against the assumptions they use and deriving a concrete bound on the advantage of any distinguisher. Observational equivalence between oracles is established using a state‑abstraction methodology: a simple yet powerful approach that supports transformations such as adding or forgetting state and replacing eager sampling with lazy sampling. We illustrate the framework with formalized proofs of the IND‑CCA security of encrypt‑then‑MAC, the security of ElGamal encryption from DDH, the implication from one‑time secrecy to public‑key IND‑CPA security, and the GGM pseudorandom‑function construction. To the best of our knowledge, the last is the first mechanized proof of GGM for non‑constant depth.

Authors:ZhiYan Hou, Xinyu Tang, Hongyan An, Jianjin Zhang, Weizhen Wang, Yunyun Han, Gengsheng Li, Xiangzhao Hao, Haiyun Guo, Wenbin Hu, Jinqiao Wang, Yafeng Deng
Title: DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models
Abstract:
Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models using automatically verifiable outcome signals, but these signals are typically sparse and at the sequence‑level. On‑policy self‑distillation (OPSD) mitigates this sparsity by querying a privileged teacher at student‑visited prefixes and providing dense token‑level distributional supervision. Although this dense supervision alleviates signal sparsity, we find that standard OPSD still underexploits the temporal structure of the rollout. It assigns every local divergence the same coefficient, regardless of its position or the divergence sequence in which it occurs. In on‑policy autoregressive generation, the same divergence magnitude can follow different discrepancy histories, reflecting different evolutions of the mismatch between the teacher and student. Since the local scalar alone cannot distinguish these temporal contexts, standard OPSD cannot adapt its token‑level weights to the realized discrepancy sequence. To address this limitation, we propose Divergence‑Adaptive Supervision Horizons (DASH). DASH maps the gap between each local distillation signal and the sequence‑level mean to an adaptive propagation gate and then uses these gates to control backward multi‑step aggregation. By doing so, DASH adjusts token‑level supervision weights according to how local divergences evolve during generation. Experiments on three mathematical reasoning benchmarks across three model scales show that DASH improves over our matched vanilla OPSD reruns on every benchmark at all three scales. DASH reuses the teacher and student distributions that OPSD already computes, so the gains require no additional teacher or student forward pass. Code: https://github.com/DBtxy/DASH‑OPSD

Authors:Yuanjing Xu, Xinyan Liu, Weidong Chen, Zixuan Zou, Linhao Zhang, Zhuangzhe Meng, Antoni B. Chan, Weigang Zhang
Title: Depth-Guided Video Object Counting in Crowded Scenes
Abstract:
Our primary objective is to advance video object counting in crowded scenes, aiming to robustly count all instances of a target category based on given text or visual prompts. Existing methods rely on RGB information, limiting their discriminative ability in crowded and occluded conditions. To address this, we propose a Depth‑Guided Detector (DG‑Det) along with a general post‑processing pipeline. By integrating depth cues with multi‑scale RGB‑D cross‑attention and explicit occlusion prediction, our method enhances spatial understanding and achieves robust detection in crowded and occluded scenes. Furthermore, we introduce a unified de‑duplication framework to eliminate cross‑frame redundant counting. To facilitate future research, we also release a new RGB‑D Video Object Counting dataset featuring depth information and multiple object categories persequence. Extensive experiments demonstrate that our method achieves a 62.01% reduction in MAE compared to existing baselines, and also produces consistent improvements in RMSE. We provide the source code at https://github.com/streamer‑AP/DG‑Net and the dataset at https://huggingface.co/datasets/aerospace123/RGBD‑VideoCount.

Authors:Binze Wang, Jinyu Tian, Xingrun Wang, Xiaochen Yuan, Jianqing Li
Title: Reversible Unlearnable Examples: Towards the Copyright Protection in Deep Learning Era
Abstract:
Significant advancements in deep learning have been made possible by the utilization of large datasets, underscoring the critical importance of copyright protection. Adding meticulously designed perturbations to examples, making them unlearnable has become a crucial approach for safeguarding data copyright. Existing methods for creating unlearnable examples overlook the risk of data leakage, which can threaten data ownership. Thus, copyright protection in deep learning faces two main threats: illegal model training and malicious data leakage. We investigate that these two threats cannot be solved by straightforwardly combining existing availability attacks and watermarking techniques as their negative interaction effects. Therefore, in this paper, we propose a novel copyright protection mechanism for the aforementioned security concerns. Considering that the prevention of unauthorized model training requires powerful generalizability of unlearnable perturbations, we generate perturbations to induce the model to learn uncorrelated features of input images. It works by minimizing the mutual information of the input and output of the model. On the other hand, to eliminate the side impact of unlearnable perturbations on the watermark extraction, we design a dual extraction strategy by using two distinct watermark extractors. Extensive experiments on the image datasets ImageNet, CIFAR10, and Pets show that our proposed method could provide comprehensive copyright protection to images. The code is available at https://github.com/Yeah21/ReversibleUnlearnableExamples.

Authors:Nima Hatami, Karim Faez, Saeed Sharifian, Hamidreza Amindavar
Title: CFGPNet: Cross-Attention-Based Fused Gradient Programmed Network Framework for Multispectral Object Detection
Abstract:
RGB‑‑T object detection exploits the complementary strengths of visible and infrared imagery, supporting robust perception in low‑light, adverse‑weather, and complex multi‑scale environments. However, existing methods still suffer from insufficient cross‑modal interaction, unstable fusion from modality distribution gaps, and the high computational cost of heavy attention‑based architectures. To address these issues, CFGPNet is proposed, a Cross‑Attention‑Based Fused Gradient Programmed Network framework for multispectral object detection. CFGPNet uses an improved GELAN backbone with RepViT‑style re‑parameterized blocks to strengthen feature representation while preserving computational efficiency. A Cross Computation Efficient Attention (CrossCEA) module is introduced to enhance cross‑modal feature interaction and reduce redundant information transfer between visible and thermal branches. To generate compact and discriminative fused representations, an Attention Selection and Aggregation Fusion (ASAF) network combines dense feature aggregation with selective attention‑based emphasis. Moreover, a programmable‑gradient auxiliary branch is integrated into each CFGPNet variant to improve gradient delivery and optimization quality. Experiments on five public multispectral benchmarks, FLIR, M3FD, LLVIP, VEDAI, and MFAD, demonstrate that CFGPNet achieves strong and consistent performance across diverse scenes, object scales, and modality balances. In particular, the framework attains 80.7% mAP50 / 45.0% mAP50:95 on FLIR, 89.9% / 63.4% on M3FD, and 97.8% / 68.9% on LLVIP. It also reaches 83.3% / 56.9% on VEDAI and 83.4% / 61.8% on MFAD. These results show that CFGPNet is an effective, practical solution offering useful accuracy‑‑efficiency trade‑offs across three model scales. The code, data, and fine‑tuned models are available at https://github.com/NimaHatami99/CFGPNet.

Authors:Zishan Xu, Zhiyuan Yao, Yuxin Chen, Yifu Guo, Zhengxi Lu, Yuquan Lu, Jinyang Huang, Yan Xu, Yasheng Wang, Weinan Zhang, Xingshan Zeng, Weiwen Liu
Title: EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning
Abstract:
Training large language model agents for long‑horizon tool use typically relies on interactions with real or synthesized executable environments, whose construction and verification are costly, or on external simulators that are difficult to ground. We introduce EnvACE, an agentic reinforcement learning method that replaces external environment interaction during training with world rehearsal. The policy alternates between acting and rehearsal: it first generates a tool call, then plays the role of the environment to produce the response induced by that action, and conditions subsequent decisions on the rehearsed response. Both roles are jointly optimized end‑to‑end using task‑success rewards. Through world rehearsal, the policy internalizes the relationship between actions and their environment responses in its parameters, yielding an agent world model that directly supports decision making. Across BFCL‑v4, tau^2‑Bench, VitaBench, and FinMCP‑Bench, EnvACE achieves strong and transferable performance, outperforming environment‑scaling baselines in the overall evaluation. Controlled studies further show that world rehearsal consistently improves policy learning across model scales. At test time, the internalized world model enables private rehearsal before committed execution, yielding further gains under a moderate rehearsal budget without additional external interaction. Our findings establish world rehearsal as a new path toward scaling LLM agent training beyond the constraints of external environments. Our code is publicly available at https://github.com/Within‑yao/EnvACE.

Authors:Subin Jeon, Byungjun Kim, Hanbyul Joo
Title: HOPE: Hand-Object Pressure Estimation from Monocular Videos
Abstract:
Estimating physical pressure from vision is essential for understanding contact‑rich hand‑object interaction. However, prior vision‑based pressure estimation methods are largely limited to planar surfaces and single image input, making them difficult to apply to dynamic hand‑object interaction with diverse objects. We instead formulate pressure estimation as a hand‑centric video prediction problem with monocular video as input. This formulation predicts temporally evolving per‑vertex normal pressure and contact directly on the hand mesh, yielding a unified output space independent of object shape and sensor layout. Building on this formulation, we propose HOPE, a framework with two key components. First, we lift tactile‑glove pressure, planar‑sensor pressure, and distance‑based hand‑object contact annotations into a shared hand vertex space, allowing bare‑hand contact data to regularize pressure learning where metric labels are unavailable. Second, we introduce a vertex‑anchored video transformer that treats each vertex as a persistent token, aggregates visual features and hand pose over time, and uses a contact‑gated pressure head to enforce that pressure vanishes without contact. Experiments on OpenTouch, PressureVisionDB, and hand‑object contact benchmarks validate HOPE across object‑pressure, surface‑pressure, and contact‑supervised HOI settings. Despite using metric pressure supervision primarily from gloved‑hand videos, HOPE generalizes to bare‑hand egocentric and in‑the‑wild videos, producing joint contact and pressure predictions beyond the scope of contact‑only or planar‑pressure baselines.

Authors:Jiaxiao Wang, Dachun Kai, Huyue Zhu, Quanquan Hu, Zhenyang Xu, Xiaoyan Sun
Title: EvReflection: Event-Driven Micro-Dynamics for Reflection Removal
Abstract:
Despite remarkable progress in reflection removal, current methods primarily exploit static image priors from a single frame and still suffer from severe residual artifacts due to the inherent ambiguity between the reflection and transmission layers. In this paper, we propose leveraging event signals to break this ambiguity. By employing event cameras to capture micro‑dynamics, we reveal the differential motion between these two layers. We thereby present a novel event‑driven reflection removal network, EvReflection, that utilizes these dynamic cues for layer separation. Specifically, we design a Micro‑Dynamics Decoupler to disentangle layer‑specific motions from event streams as priors, which then guide a Parallax‑Attention Rectifier to cleanly remove artifacts from the RGB image. Furthermore, to address data scarcity, we develop a parallax‑aware simulation pipeline and construct the EVR^2 benchmark dataset, the first real‑world dataset for this task. Extensive experiments demonstrate that EvReflection achieves state‑of‑the‑art performance on both synthetic and real‑world benchmarks, surpassing the best competing method by more than 1.6 dB and 1.2 dB in PSNR, respectively. The code, dataset, and pre‑trained models are available at https://github.com/JiaxiaoWang/EvReflection.

Authors:Jia Xiong, Runkai Li, Chenxu Niu, Guangyuan Gao, Changwen Xing, Yifan Zhang, Xinlai Wan, Jieran Cui, Chen Bai, Yusheng Hua, Ying Wang, Ming Ling, Xi Wang, Tao Xie
Title: MicroEvo: Knowledge-Guided LLM Sampling for Efficient Microarchitecture Design Space Exploration
Abstract:
Microarchitecture design space exploration suffers from expansive search spaces and expensive PPA evaluation, leaving only a small simulation budget for design decision‑making. Existing methods perform blind search without considering microarchitectural dependencies and fail to learn from the iterative search effectively, leading to wasted evaluations and weak Pareto convergence. In this paper, we propose MicroEvo, a knowledge‑guided framework that couples off‑the‑shelf LLMs with Monte Carlo Tree Search (MCTS) for multi‑objective microarchitecture optimization. MicroEvo combines LLM‑driven evolutionary operators, a Pareto‑aware tree policy that balances Pareto contribution and diversity, an active knowledge accumulation mechanism that extracts and reuses optimization insights, and state‑aware directives that adapt the search behavior online. Experiments show that MicroEvo improves Pareto‑front quality by up to 36.2% over NSGA‑II and achieves 10.6x higher search efficiency, and also demonstrates strong scalability to a complex industrial‑scale core. The code repository is available at: https://github.com/GEAR‑SEU/MicroEvo‑ICCAD‑26.

Authors:Eoin Cummins, Zhongyi Huang, Alexandre D'Hooge, Zhuoro Mo, Yaolong Ju
Title: Audio-to-Score Transcription using Pre-trained Features, Data Augmentation, and the New SheetSage-A2S Dataset
Abstract:
Existing audio‑to‑score (A2S) systems primarily focus on classical music, and the application to popular music remains underexplored. This paper first presents the new SheetSage‑A2S Dataset, which includes 61 hours of audio with \textttkern score encodings for 9,468 clips originating from 6,066 unique songs, the first of its kind to facilitate A2S research for popular music. Additionally, we improve on existing A2S approaches by using data augmentation and MuQ, a pretrained feature‑extraction model for music audio, to enhance generalisation abilities and extract meaningful audio features. Results show that the proposed A2S model achieves 4.98% symbol error rate (SER) on the Quartets collection for classical music, which significantly outperforms the 15.3% SER from the existing state‑of‑the‑art \citealfaro‑contrerasTransformer2024. Additionally, our model achieves 20.92% SER on the SheetSage‑A2S dataset for popular music, serving as a strong benchmark for future research. The dataset, model, and code are made publicly available at: https://github.com/Multimodal‑Music‑Research‑Lab/SheetSage2Kern_model.

Authors:Zijie Wang, Chen Zhong, Wei He
Title: CogVis: Must Open-Vocabulary Change Detection Perceive the Scene Anew for Every Query?
Abstract:
Earth‑surface monitoring requires change detection models capable of recognizing arbitrary semantic categories. Open‑Vocabulary Change Detection (OVCD) addresses this need. However, existing methods often entangle temporal perception, semantic discrimination, and region verification, causing unstable results and redundant computation. Inspired by human visual change perception, we propose CogVis, a cognitive memory‑guided framework that reformulates OVCD as a perception‑memory‑verification paradigm. CogVis first employs a Scene Change Perceptron (SCP) to extract a reusable, category‑agnostic change prior from frozen bi‑temporal features, thereby decoupling temporal evidence from semantic category decisions. A Semantic Memory Calibrator (SMC) then compensates for category‑dependent score shifts by dynamically estimating an image‑query‑specific decision threshold. Finally, an Adaptive Region Filter (ARF) filters connected candidates using learned semantic, temporal, and structural reliability. Experiments on seven benchmarks spanning semantic change detection, binary change localization, and building‑damage assessment show that CogVis achieves state‑of‑the‑art performance across all evaluated datasets. By sharing scene‑level change perception, CogVis further avoids repeating category‑agnostic temporal perception across queries and improves inference throughput by 28.50%.

Authors:Hao Yu, Jiabo Zhan, Kang Liu, Linnan Zhao, Dongxu Yue, Rui Chen, Jinglin Wang, Chong Sun, Chen Li, Jing Lyu, Chun Yuan
Title: PaDoc: Layout-Grounded Parallel Decoding for Document Parsing
Abstract:
End‑to‑end document parsers provide a unified interface, but serialize page layouts and regional contents into one autoregressive sequence. This formulation forces independent regions onto a decoding path whose length grows with the total content, whereas crop‑based two‑stage parsers expose region‑level parallelism at the cost of repeated visual prefills and fragmented page context. To retain full‑page context while removing dependencies, we propose PaDoc, a layout‑grounded parser that treats the predicted layout as a branching structure over a shared page representation. Under a region‑sufficiency assumption, we derive a prefix‑conditioned factorization in which the layout stream and regional content branches advance concurrently, reducing the decoding depth to the longest layout‑content path. We realize this factorization within a single MLLM: packed variable‑length ancestor attention preserves the visibility under standard next‑token training, while masked parallel decoding creates branches that the evaluated vLLM backend serves as concurrent requests with cache‑resident shared‑prefix reuse. On OmniDocBench Full, PaDoc attains an Overall layout F1 of 91.1 and, among end‑to‑end parsers, a top‑tier Overall score of 94.24 together with the best Text Edit (0.038) and Formula CDM (95.59). On a 384‑page subset and one A800 GPU, it is the fastest end‑to‑end parser at five concurrency levels, improving valid‑page throughput by 67.4‑118% and reducing P95 latency by 39.2‑54.9% relative to a same‑backbone Sequential SFT baseline. Code is available at https://github.com/Longin‑Yu/Padoc

Authors:Poonam Poonam, Alexander Epple, Timo Ropinski
Title: Bar-JEPA: Extracting Values from Bar Chart with Joint-Embedding Predictive Architecture
Abstract:
Bar charts are commonly used in data visualization, and while they are easily understood by humans, it is non‑trivial to extract the underlying data computationally. For a machine‑learning‑based approach, training chart de‑rendering models usually requires labeled, real‑world data. Labeling data is a time consuming task, which is why annotated data is scarce. Models can learn more efficiently when provided with features of high semantic quality, which a joint‑embedding predictive architecture (JEPA) is designed to learn in a self‑supervised manner. We present a per‑bar, numerical value recovery pipeline for bar charts, where a JEPA encoder is used to produce semantically rich latent features. The decoder model consuming these features is simple and quick to train and outputs the coordinates of ticks and bars, which can be used to recover bar values. The effectiveness of self‑supervised finetuning and quality of the extracted features is evident when comparing our model to end‑to‑end supervised baselines. Code, datasets and checkpoints are available on \hrefhttps://github.com/dralois/Bar‑JEPAGitHub.

Authors:Jiming Su, Hantao Hua, Lujia Yin, Yiping Yao, Feng Zhu
Title: Hybrid-Adaptive Thread Tuning to Mitigate Simulation Execution Bottlenecks in High-Performance Reinforcement Learning Inference
Abstract:
In simulation‑in‑the‑loop decision‑making systems, reinforcement learning (RL) inference is often constrained by simulator‑side execution overhead, where workloads are highly dynamic and sensitive to runtime thread configurations. Existing multithreaded strategies struggle to match thread resources before or during execution, causing resource contention, scheduling overhead, and reduced throughput. Through empirical analysis, we identify the ratio of task execution time to scheduling time as the key factor determining the optimal thread count. Building on this insight, we propose AutoThread, a hybrid adaptive thread‑tuning method for mitigating simulation bottlenecks in RL inference. AutoThread employs a Physics‑Informed Neural Operator (PINO) as a thread‑count predictor and incorporates a finite‑source M/M/1 queueing model to constrain and guide prediction, enabling fast and accurate estimation under dynamic workloads. It further performs load‑aware online fine‑tuning to compensate for prediction errors and refine resource allocation. Experiments show that AutoThread improves average speedup by 18.4% over static strategies, achieves average throughput of 1.7x and 1.8x that of XGBoost and Reinforcer, respectively, and reduces execution time by up to 83.8% compared with state‑of‑the‑art methods. Our code and dataset are publicly available at https://github.com/suchenjm/AutoThread.

Authors:Jiale Han, Xiang Li, Jing Qian, Wenyuan Gu, Pin Gao, Ye Luo, Hongyuan Zha, Dacheng Tao, Benyou Wang, Lin William Cong
Title: From Economic Agents to Agentic Economies: A Systems Blueprint for Economic World Models
Abstract:
Economic World Models (EWMs) are generative economic models that simulate how economies evolve from within by modeling heterogeneous agents, their beliefs and actions, and the market and institutional mechanisms through which their interactions produce aggregate outcomes. This paper develops an implementation roadmap for building economic world models as generative engines in which heterogeneous agents act, interact, adapt, and co‑evolve with markets and institutions, thereby producing economic dynamics from the inside. We organize EWM systems into a six‑level capability ladder, from fixed rule‑based agent worlds to adaptive and LLM‑based agent worlds, self‑evolving agents, evolving institutional worlds, and sim‑to‑real economic twins aligned with real observations. A systematic literature survey across these levels reveals that existing work remains concentrated in lower‑level agent and simulation environments, while systems with self‑evolving agents, endogenous institutions, persistent empirical alignment, and validated economic mechanisms remain rare. By translating the EWM agenda into an implementation blueprint, this paper aims to accelerate the development of the next generation of economic simulation environments that can serve as high‑fidelity sandboxes for human decision‑makers and as training, planning, evaluation, and safety substrates for AI agents. We release a curated paper list and related resources to support future research.

Authors:Jiahao Huang, Zheng Lian, Jingyi Zhang, Zhide Chen, Xiaojiang Peng, Shaonan Wang
Title: OneEmo: A Unified Multimodal Reasoning Model for Emotion Perception, Understanding, and Interaction
Abstract:
Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in emotional intelligence. However, prevailing research predominantly focuses on task‑specific specialization, often neglecting inter‑task synergy and leaving latent reasoning potential underexplored. To bridge this gap, we introduce OneEmo, a unified affective generalist capable of mastering emotion perception, comprehension, and interaction. For this purpose, we first construct EmoWorld‑130K, a comprehensive dataset that distills specialized affective knowledge into explicit reasoning trajectories via a human‑in‑the‑loop workflow. Supervised fine‑tuning on this corpus reveals significant mutual benefits derived from multi‑task learning. Second, to fully unlock the latent reasoning potential, we propose Emo‑Chord, a novel reinforcement learning strategy that stabilizes optimization through unified multi‑task reward allocation. Extensive experiments demonstrate that OneEmo achieves state‑of‑the‑art performance against similarly sized baselines across most benchmarks. Notably, despite having significantly fewer parameters than commercial models, OneEmo delivers highly competitive results. This paper paves the way for more reliable and interpretable affective computing. The code is available at https://github.com/waHAHJIAHAO/OneEmo.

Authors:Guangyuan Wang, Li Hu, Dechao Meng, Zhongyi Zhang, Peng Zhang, Mingyang Huang, Ruoshi Zhang, Ke Sun, Zhe Zhang, Xingjun Wang, Gang Cheng, Bang Zhang
Title: Wan-Animate-2: Pushing the Application Boundaries of Character Animation
Abstract:
Character image animation remains a foundational yet challenging task in computer vision. Existing approaches can be broadly categorized into three paradigms: methods based on explicit motion representations suffer from extraction errors and identity drift; methods based on implicit motion features lose fine‑grained dynamics through compression; and in‑context learning approaches avoid intermediate representations but incur prohibitive computational costs. Furthermore, all current systems are designed for offline synthesis, unable to meet the real‑time requirements of interactive applications such as digital avatars and live‑streaming hosts. To address these limitations, we present Wan‑Animate‑2, an end‑to‑end character animation framework that directly consumes the driving video within a redesigned Diffusion Transformer. Our architecture achieves superior motion fidelity and identity preservation by eliminating intermediate motion extractors entirely. We further introduce text driven viewpoint control that decouples the output camera perspective from the driving video‑‑a capability rarely supported by prior character animation methods that rely on explicit motion representations. Beyond generation quality, we present Wan‑Animate‑2‑Lite, an efficient variant that reduces inference latency to real‑time thresholds through a three‑stage training paradigm: teacher forcing pretraining with error buffer mechanism, and Self‑Forcing distillation with chunk‑wise backpropagation. This enables streaming character animation for interactive applications, opening new deployment scenarios that were previously infeasible. Qualitative evaluations and user studies demonstrate that Wan‑Animate‑2 achieves high‑fidelity animation results across diverse characters and motion patterns. To foster further research and community development, we will release the Wan‑Animate‑2‑Base model weights to the public.

Authors:Zi-Han Wang, Zhengxi Lu, Zhiyuan Yao, Jinyang Wu, Jie Wu, Zhengzhou Cai, Yueqing Sun, Ziang Ye, Linji Hao, Qi Gu, Xunliang Cai, Yongliang Shen, Yujiu Yang
Title: AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning
Abstract:
Reinforcement learning (RL) with verifiable rewards constructs trajectory‑level advantage estimates, yet it often fails to credit the few pivotal decisions that determine outcomes in long‑horizon, multi‑turn agentic tasks. Recent work introduces privileged self‑distillation for credit assignment, providing denser supervision, but it remains unclear how such local signals should represent sequential credit. We propose AgentOPSD, a critic‑free, recursive method for turn‑level credit assignment in agentic reinforcement learning. AgentOPSD aggregates token‑level teacher‑student log‑probability gaps into turn‑level evidence and recursively updates a Bayesian belief state in log‑odds space. This yields a principled reweighting scheme that converts sparse outcome supervision into turn‑level credit signals and identifies pivotal turns through the marginal belief revision between consecutive states. The method is fully compatible with standard policy optimization and requires neither an additional critic nor extra rollouts. We evaluate AgentOPSD on ALFWorld, WebShop, and Search‑QA using Qwen2.5 models at two scales (3B and 7B). AgentOPSD outperforms GRPO and strong self‑distillation baselines, achieving 89.1% success on ALFWorld with Qwen2.5‑7B. Ablation studies attribute the gains to turn‑level aggregation and history‑dependent recursive belief updates.

Authors:Shuvam Swapnil Dash, Arpit Narechania
Title: PoseForge: Editable Pose Analytics for AI-Assisted Sports Coaching
Abstract:
Athletic coaching increasingly relies on video analysis, yet raw footage lacks tools to quantify motion or simulate valid technique corrections. Drawing on formative interviews with eleven cricket experts (coaches, performance analysts, captains, and players), we introduce PoseForge, a visual analytics system that extracts 3D skeletal poses from single‑camera sports videos for interactive movement analysis. In a cricket batting case study, PoseForge computes interpretable kinematic metrics such as feet gap and elbow angle, compares them against scientifically derived norms, and uses an AI coach to suggest targeted adjustments, presented visually and through natural‑language feedback (e.g., "increase feet gap by 10 cm"). Users can directly modify poses via mouse interaction or natural‑language instructions, with inverse kinematics maintaining anatomical plausibility and real‑time updates of metrics and comparisons. An evaluation with the same eleven cricket experts found PoseForge effective for diagnosing movement issues and exploring corrective alternatives, highlighting its applicability in low‑resource, academy, and grassroots coaching settings, while identifying opportunities for enhanced sport‑specific metrics and longitudinal tracking. PoseForge is available as open‑source software at https://github.com/DataVisards/PoseForge.

Authors:Attila Simkó
Title: MirrorNet: Can Medical Image Anonymization Really Protect Patient Identity?
Abstract:
Medical images are routinely de‑identified‑‑‑names, dates, and other metadata removed‑‑‑and then shared for research, teaching, and public benchmarks under the assumption that this renders them anonymous. Such de‑identification protects the metadata but not the pixels, and‑‑‑apart from scans that directly contain facial structures‑‑‑whether the image content itself identifies the patient has received little scrutiny. We investigate this question by learning a cycle‑consistent correspondence between a cross‑sectional medical image and a non‑medical, patient‑identifying image, using a pair of coupled, cycle‑consistent variational autoencoders. From a held‑out scan, the model recovers a recognisable likeness of the patient (identity‑region MAE = 0.163); conversely, it synthesises a scan from such an image. These results indicate that a de‑identified medical scan remains identifying‑‑‑it is, in effect, a photograph of the patient‑‑‑and that imaging data should be governed as biometric data rather than as anonymisable records. To support reproducibility, the code and trained models are shared at https://github.com/attilasimko/public‑repository.

Authors:Krzysztof Byrski, Rafał Tobiasz, Grzegorz Wilczyński, Mikołaj Zieliński, Dawid Baran, Dominik Belter, Jacek Tabor, Przemysław Spurek
Title: Floating Radiance Networks
Abstract:
Recent advances in neural scene representations enable photorealistic novel‑view synthesis, yet most methods remain tightly coupled to a single rendering paradigm, limiting their versatility and integration with conventional graphics workflows. We introduce Floating Radiance Networks (FlaRe), a neural scene representation combining explicit ray‑traceable geometry with continuous neural radiance functions. A scene is represented by floating planar generalized Gaussian primitives, each carrying a compact latent descriptor of a local radiance field. A lightweight decoder shared across the scene maps this descriptor, local surface coordinates, and viewing direction to color and opacity. This formulation preserves the expressiveness of neural fields while providing an explicitly addressable structure that can be efficiently queried and manipulated. Hardware‑accelerated primitive intersections enable interactive rendering and recursive ray‑tracing, including reflections, refractions, transparency, and shadows. The same representation further supports primitive‑level deformation, mesh extraction, and appearance stylization directly in its learned descriptor space. Experiments across standard reconstruction benchmarks demonstrate competitive rendering quality while using a compact set of primitives. Together, these results establish FlaRe as a versatile representation that brings high‑fidelity neural rendering, ray‑tracing, geometric manipulation, and appearance editing into a unified scene model. Source code is available online. Source code can be found at: https://github.com/KByrski/FlaRe

Authors:Paweł Batorski, Abtin Pourhadi, Akylgali Aitaza, Przemysław Spurek, Paul Swoboda
Title: MACRO: Markov Chain Routing of Transformer Layers
Abstract:
Standard Large Language Models (LLMs) execute layers sequentially. Dynamic layer routing, i.e. search for a different execution path through layers involving layer repetitions, skips and other moves, can improve performance. Existing routing approaches often require updating model weights, running expensive search loops per test instance, or demand ground‑truth labels during inference. In this work, we propose Markov Chain Routing of Transformer Layers (MACRO), a framework that learns task‑specific routes over LLM architectures without modifying underlying parameters. MACRO models layer routing as a context‑dependent Markov policy conditioned on layer indices, computation budget phases, directional displacements, and operator context, supporting skip, repeat, and residual hidden‑state addition operations. The Markov route distribution is updated via feedback on training data and decoded using a top‑k Viterbi algorithm to isolate high‑probability candidate programs. We evaluate MACRO across diverse reasoning and knowledge benchmarks on multiple open‑weight LLMs. MACRO achieves a +5.0% average accuracy improvement over the unrouted baselines, with largest gains on small models. We outperform the best dynamic routing approach Dr. LLM by +7.2%, while reducing route‑search time 9.4x (from 14.8 to 1.6 hours). Our code is publicly available at https://github.com/Batorskq/MACRO.

Authors:Fanzhe Wei, Li Liu, Ziyang Wang, Chenyu Wang
Title: Runtime Observability for Heterogeneous Attention Memory
Abstract:
Modern models no longer keep a plain KV cache: latent caches, learned sparse selectors and recurrent states each carry the model's memory in a different form, and each fails differently under compression. We give a runtime observability contract that covers all four memory classes with three operators, instantiate it on six model configurations across five architecture families, and compose the per‑stage bounds into an executable request‑level risk ledger. Contracts carry their error metric as a type ‑‑ composition is only defined when metrics match, and this check rejected our own first composed chain; the repaired chain crosses metrics through two proved bridges, and whatever no formal system can certify is measured instead, dropping the composed tier to empirical automatically: every claim is certified, partially certified, or empirical, composition inherits the weakest tier, and the tier is decided by the machine. Replayed over 12.4M entry reads and run under eight‑way concurrency with per‑request budgets and fail‑closed identity attribution, the ledger quantifies the honest trade‑off on today's witness and holds its risk budget with zero violations. A fused always‑on probe observes a declared one‑layer subset under CUDA graphs inside the serving noise floor. Applied to a served DeepSeek‑V4 stack with a packed compressed‑KV prototype, the same machinery localizes a silent corruption to a precise structural boundary ‑‑ exact in the eviction‑free, identity‑isolated regime, with every observed failure in an eviction or slot‑reuse regime ‑‑ through a machine‑adjudicated discrimination campaign whose calculus rejected two of our own confounded inferences along the way. All artifacts, guards, and the Lean development are released at https://github.com/metask‑ai/witprobe‑attention‑memory; every number in this paper regenerates from the shipped artifacts by one command.

Authors:Yongjie Qian, Ke Gao, Zhibin Zhang, Shaohui Peng, Ling Li
Title: RepoOMP: Repository-Aware Hotspot OpenMP Parallelization via Dependency-Aware Context Reduction
Abstract:
OpenMP parallelization of hotspots in mature repositories remains difficult because loop safety and optimization payoff often depend on non‑local evidence. Rule‑based tools under‑parallelize when legality is not locally provable, while agent‑based approaches become unstable when retrieval misses decisive dependencies or includes irrelevant code. We present RepoOMP, a hybrid framework that recovers parallelization‑relevant evidence before generation. RepoOMP builds a Multi‑granularity Attributes Performance graph (MAP), routes hotspots between deterministic rules and an LLM agent, and constructs a Structured Transformation Context (STC) that exposes dependency facts without flooding the model with unrelated repository text. We evaluate RepoOMP on 951 profiled hotspots from NPB, BOTS, FFmpeg, NCNN, and GROMACS. Under compilation, workload‑specific checks, and positive speedup, 372 hotspots are accepted, including 330 real‑world repository hotspots. RepoOMP achieves average speedups of 8.23× on NPB and 8.96× on BOTS. For the nine detailed real‑world kernels used in matched‑backbone and robustness analyses, RepoOMP reaches a cross‑backbone mean of 5.25×, improves speedup by 18‑‑28%, and reduces agent‑side token cost by 47‑‑68% relative to the unstructured Claude Code baseline. Across 330 accepted real‑world hotspots, median speedup is 2.25×. Overall, RepoOMP provides an evidence‑guided workflow for hotspot parallelization in repository settings. The open‑source repository is available at https://github.com/Qlalq/RepoOMP_Simplified.

Authors:Runrui Li, Lin Zhu, Hua Huang
Title: DTRNet: Dual Text-Radical Decoding for Handwritten Chinese Text Recognition with Faked Character Detection
Abstract:
In K‑12 educational scenarios, handwritten Chinese text recognition should not only transcribe student writing, but also detect faked characters. However, existing recognition models are usually confined to a predefined set of normal characters and therefore cannot explicitly identify faked characters. Existing detection methods exhibit complementary limitations: character‑level methods provide interpretable structural evidence but suffer from low efficiency, whereas line‑level methods are efficient but rely heavily on confidence scores, making them prone to missed detections and lacking explicit structural evidence. Thus, the key challenge is to preserve character‑structural evidence independent of contextual inference while maintaining line‑level efficiency. To this end, we propose DTRNet, a dual Text‑Radical decoding framework for line‑level faked character detection. DTRNet decouples context‑aware text recognition from character‑wise structural verification, where the text branch performs line‑level transcription and the radical branch predicts legal Ideographic Description Sequences (IDS) for lexicon‑based faked character judgment. We further introduce IDS‑Guided Confidence Adjustment (IGCA) to refine text predictions using structural evidence during inference. Experimental results demonstrate that DTRNet effectively detects faked characters while maintaining strong recognition performance and providing interpretable radical‑level evidence. Code, checkpoints, and the processed dataset are publicly available at https://github.com/BNU‑ERC‑ITEA/DTRNet.

Authors:Yaozi Zhong, Xingxing Yang, Shaohui Mei, Mingyang Ma
Title: Overcoming Attention Drift: Homogeneity-Heterogeneity Guided Feature Aggregation for Low-Light Remote Sensing Image Enhancement
Abstract:
Restoring high‑fidelity remote sensing imagery from extreme low‑light degradation is indispensable for reliable Earth observation and downstream machine vision. However, under severe noise and illumination corruption, existing methods suffer from attention drift, erroneously aggregating features across distinct physical boundaries and causing severe structural blurring and color distortion. To address this, we propose HALO, a dual‑prior‑driven enhancement framework that formulates enhancement as a guided feature aggregation problem driven by foundation model priors. Specifically, an illumination‑invariant semantic prior provides regional homogeneity as a positive bias for content‑consistent aggregation, while a pseudo‑3D topological prior provides boundary heterogeneity as a negative penalty to strictly prevent cross‑boundary confusion. To cooperatively incorporate these two priors, we propose a Homogeneity‑Heterogeneity Cooperative Attention Module (H2CAM) to resolve feature conflicts during cross‑modal prior fusion. Extensive experiments demonstrate that HALO achieves state‑of‑the‑art performance across 8 challenging synthetic and real‑world remote sensing benchmarks, significantly improving physical boundary sharpness and color fidelity while maximizing the preservation of discriminative features for downstream Earth observation tasks.

Authors:Hong Jiang, Junnan Zhu, Jingwang Huang, Xiao Sun, Yuming Yang, Jiang Zhong, Ruirui Chen, Jingman Shi, Hao Wu, Nayu Liu, Xinyi Jiang, Kaiwen Wei
Title: M$^3$R-Bench: A Unified Benchmark for Evidence-Grounded Multimodal Metaphor Understanding
Abstract:
Metaphor enables the understanding of abstract concepts through cross‑domain mappings while conveying affective attitudes. In multimodal scenarios, visual and textual information jointly construct Target‑‑Source mappings, requiring both conceptual understanding and cross‑modal reasoning. However, existing benchmarks mainly evaluate metaphor understanding through isolated subtasks and lack evidence‑grounded explanations, making it difficult to assess whether models establish mappings grounded in visual and textual cues.To address these limitations, we introduce M^3R‑Bench, a unified and evidence‑grounded benchmark containing 1,000 image‑‑text instances with human‑verified annotations. Guided by Conceptual Metaphor Theory and theories of nonliteral language understanding, M^3R‑Bench provides joint annotations for metaphor occurrence, Target‑‑Source mapping, sentiment, and stage‑wise explanations following ``evidence identification‑‑mapping establishment‑‑sentiment inference.''Evaluations on M^3R‑Bench reveal that existing models often overlook visual evidence, rely on superficial textual cues, and produce inaccurate Target‑‑Source mappings, exposing a cross‑modal evidence‑‑mapping mismatch. To address this mismatch, we propose M^3R‑Reasoner, which combines curriculum‑based reasoning supervision with task‑aware reinforcement learning to align model reasoning with metaphor interpretation. Experiments show that, with only an 8B‑parameter backbone, M^3R‑Reasoner outperforms larger proprietary MLLMs across four unified‑task metrics and improves Visual Evidence and Sentiment Justification scores over GPT‑5.5 by 28.45 and 30.11 points, respectively, while surpassing Claude‑Sonnet‑4.6 by 8.00 points in mean rubric score. The dataset and code are available at https://github.com/hongshi4/M3R‑Bench.

Authors:Haoyang Tong, Yu He, Fang Li, Lichen Ma, Jingling Fu, Dong Chen, Zhen Chen, Junshi Huang, Jie Cao
Title: Energy-Guided Flow Matching
Abstract:
Pixel‑space generative models bypass lossy latent compression, yet necessitate joint learning of global structure and fine‑grained details in a high‑dimensional space. Standard flow matching interpolates noise toward a fixed clean‑image endpoint, leaving the spectral evolution to be learned implicitly. In this paper, we introduce Energy‑Guided Flow Matching(EG‑FM) that explicitly models a coarse‑to‑fine generative trajectory by moving endpoint. Specifically, EG‑FM replaces the fixed endpoint with a heat‑kernel‑filtered endpoint that evolves smoothly from low‑frequency image to clean image.The fraction of high‑frequency signal in moving endpoint is released by an image‑specific energy‑guided scheduling, leading to the re‑targeting of velocity in flow matching.Our framework requires no adaptation of the backbone and training data, bringing negligible cost on the training and inference stages. In our experiment, EG‑FM consistently achieves lower FID on the ImageNet class‑conditional image generation task at 256 × 256 with fewer epochs, reaching an FID of 1.55 at 200 epochs and 1.45 at 600 epochs. We continue training the generation task on the setting of 512 × 512 resolution, yielding a FID of 1.58 after only 40 high‑resolution adaptation epochs.Furthermore, we transfer EG‑FM on text‑to‑image generation and achieve 0.85 on GenEval score and 83.9 on DPG‑Bench. Code is available at https://github.com/ysng123/EG‑FM.

Authors:Mattia Masiero, Ilya A. Petrov, Daniel Cremers, Gerard Pons-Moll, Riccardo Marin
Title: Ordered Diffusion for 3D Human Registration
Abstract:
3D human registration has historically been treated as a regression task, assuming a unique ground‑truth alignment exists between the template and an input point cloud. In reality, acquisition noise, occlusions, and unknown soft tissue dynamics introduce inherent ambiguity into human scans. Regression‑based methods consequently converge to an average prediction, often failing to represent a plausible geometry. In our work, we embrace such uncertainty by modeling the registration as a distribution of alignments. We propose ODin, which formulates registration as a 3D diffusion process that generates a point cloud aligned with the target geometry while preserving template semantics through consistent point ordering. To achieve this, ODin relies on global, local, and positional conditioning, guiding each point to its correct location. Our experiments demonstrate that such a generative formulation not only outperforms its regression‑based baseline, but also establishes a new state of the art, surpassing highly engineered methods while reducing the registration time by two‑thirds. Pre‑trained models and code are available at https://riccardomarin.github.io/odin/.

Authors:Vorch Team, Xiaoyu Chen, Yang Ding, Cong Han, Menglin Han, Yuxin Hong, Jiebo Hou, Zequn Jie, Xiang Li, Jing Liu, Qi Liu, Yulei Lu, Siyuan Luo, Lin Ma, Xin Ma, Yinlong Qian, Peng Shi, Fang Wan, Siqi Wang, Yaohui Wang, Yaole Wang, Yidi Wu, Siqian Yang, Mingyu Yin, Haoran Yu, Gang Yue, Lisai Zhang, Yuting Zhang
Title: Vorch-Omni: Multi-Task Orchestration of Sight and Sound
Abstract:
Recent advances in generative video modeling have enabled diverse generation, reference‑based synthesis, extension, and editing, but existing approaches often rely on fragmented task‑specific models. A general model must distinguish heterogeneous target, source, and reference signals to determine what to generate, preserve, or use as guidance, while reducing interference among tasks. Joint audio‑visual generation further increases this challenge by introducing diverse conditioning and output configurations across modalities. We present Vorch‑Omni, a unified multi‑task framework for audio‑visual synthesis based on an arbitrary‑condition‑to‑arbitrary‑output formulation. It flexibly treats video and audio signals as either conditioning inputs or generation targets. Token‑level conditioning masks and task identifiers distinguish targets, source content, and references, while position types separate temporal context from independent conditions. To capture semantic and structural information, Vorch‑Omni employs complementary visual conditioning pathways: a vision‑language model interprets sampled frames with text instructions, and a video VAE encodes conditions into latent tokens for direct guidance. We further build a distributed data pipeline to curate diverse temporally aligned audio‑visual clips, generate structured captions and metadata, and balance heterogeneous task distributions. Built on a single flow‑matching diffusion transformer without task‑specific architectural changes, Vorch‑Omni supports over 10 tasks, including text‑to‑video, text‑to‑audio‑video, image‑ and reference‑conditioned generation, temporal extension, audio‑driven generation, video transformation, and audio‑visual editing. This unified framework provides a scalable foundation for general‑purpose audio‑visual generation and manipulation.

Authors:Andrey Shutkin, Denis Parkhomenko, Ivan Kirillov, Kirill Chernyshev, Kirill Malakhov, Ilia Vasiliev, Ilia Trushkin, Valeriya Kobenko, David Chikovani, Alexander Ivanov, Azat Saginbaev, Egor Silvestrov, Ivan Mikheev, Konstantin Zakharov
Title: KVAE: Family of Tokenizers for Multimodal Generative Models
Abstract:
Latent diffusion modeling (LDM), a prominent paradigm, utilizes tokenizers to map input signal to compressed representation. This dependency positions tokenizer as an integral part of generation process itself, since it affects learning speed, quality of synthesized samples and lay foundation for later applications. This report presents series of KVAE tokenizers for audio, image and video, all designed for subsequent text‑conditioned generation: KVAE‑Audio, a continuous full‑band 48 kHz tokenizer with a 50 Hz latent of 64 channels; KVAE‑3D ‑‑ two causal video tokenizers for 4x16x16 and 4x8x8 compression; KVAE‑2D, an image model, compressing input by factor of 8 with 32 channels. We demonstrate that reconstruction (PSNR, LPIPS, PESQ, etc.) and generation results on objective (Frechet Distance, CLIP score, CLAP score, etc.) and subjective (side‑by‑side evaluation) metrics matches or surpasses frontier opensource tokenizers, such as VAEs from Wan‑2.2, HunyuanVideo‑1.5, FLUX.2, MovieGen, StableAudio and MMAudio. Considering difficulty of development, we share with community training details, model selection method and ablation on design choices. The code is publicly available at https://github.com/kandinskylab/kvae and https://github.com/kandinskylab/kvae‑audio.

Authors:Paweł Batorski, Przemysław Spurek, Paul Swoboda
Title: GROM: Gradient-Free Rapid One-Shot Machine Unlearning
Abstract:
Machine unlearning has become a critical capability for safely removing specific, sensitive knowledge from large language models (LLMs). Current state‑of‑the‑art approaches primarily rely on iterative, training‑time unlearning via fine‑tuning. However, even when utilizing parameter‑efficient dimensionality reduction techniques like LoRA, gradient‑based optimization remains computationally expensive and lacks explicit analytical formulations. It can also leave the targeted knowledge merely hidden rather than removed, to the point that simply quantizing the unlearned model restores much of what it was supposed to have erased. To resolve this, we propose a novel one‑shot unlearning approach, abandoning iterative optimization in favor of a direct, exact analytical solution. We frame the unlearning process as a ridge‑regularized least‑squares optimization problem, deriving a closed‑form additive update for targeted weight matrices. This update forces the selected layer to suppress unwanted content while strictly preserving its behavior on retained data. Computed from gradient‑free forward passes alone, with no backpropagation and no iteration to convergence, GROM applies the weight edit in mere seconds, which makes it orders of magnitude faster than traditional fine‑tuning. Extensive evaluations demonstrate that GROM achieves state‑of‑the‑art forgetting‑utility trade‑offs on TOFU‑5%, TOFU‑10%, MUSE‑Books, MUSE‑News and WMDP, significantly reducing computational overhead without sacrificing overall model performance. Because the update removes the targeted content from the weights instead of masking it, GROM also withstands the low‑bit quantization attack that recovers much of the content a gradient‑based baseline had appeared to forget. Our code is publicly available at https://github.com/Batorskq/GROM.

Authors:Bo Zhang, Wenxin Wang, Feng Chen, Zhihao Zhang, Zixuan Wang, Changsheng Li, Yinjie Lei
Title: Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding
Abstract:
Recent advancements in MLLM‑based long‑form video understanding have mitigated inference‑time computational cost and limited context lengths by selecting query‑relevant frames. However, existing approaches predominantly rely on external proxy scorers and rigid heuristic rules, inevitably suffering from misalignment with the target MLLM's intrinsic evidence and failing to accommodate the non‑uniform spatiotemporal information density. In this paper, we propose a fine‑grained dynamic visual selection framework named EviSelect, grounded in the target MLLM internal attention evidence. Our method efficiently probes visual evidence via sparse prefilling as a structured prior to guide distribution‑aware dynamic sampling. Specifically, we efficiently approximate attention maps of the target MLLM using highly compressed visual inputs and sparse attention, well‑aligned to the full counterpart. Conditioned on three complementary attention components derived from this prior, we design a lightweight selector that not only precisely locates query‑relevant timestamps but also adaptively adjusts the local sampling rate and spatial resolution. To enable evidence‑conditioned spatiotemporal sampling, we formulate the selector as a stochastic policy and optimize it via GRPO under a joint accuracy‑‑efficiency reward. By rewarding correct predictions under lower visual cost through group‑relative comparisons, our method encourages the policy to allocate computation dynamically according to the information density of each video. Across three long video understanding benchmarks, EviSelect achieves superior performance compared to existing methods while reducing selected visual tokens by about 50% and achieving a 3.9x end‑to‑end speedup.

Authors:Lisai Zhang, Yidi Wu, Qi Liu, Xin Ma, Yang Ding, Gang Yue, Siqian Yang, Jingyuan Chen, Lin Ma, Yaohui Wang
Title: Vorch-Director: Interactive World Story Model via Noise-Aware Error Rectification
Abstract:
Autoregressive continuation provides a natural path toward minute‑scale audio‑visual generation by repeatedly extending a short‑window generator conditioned on previously generated video and audio. However, models are trained on clean ground‑truth histories, while inference relies on their own generated histories, where accumulated errors cause identity drift, over‑smoothing, and audio‑visual desynchronization. Recent methods reduce this mismatch by reusing prediction residuals as synthetic corruption, but we observe that the effectiveness of residual correction critically depends on the flow‑matching noise level at which residuals are produced. We propose Vorch‑Director, a noise‑level‑aware residual correction strategy that associates each residual with its originating noise level and injects residuals from matched noise regimes during training. By aligning injected errors with the denoising process, Vorch‑Director produces more realistic autoregressive histories while retaining efficient teacher‑forcing training. Built on the audio‑visual LTX‑2 diffusion transformer, Vorch‑Director further introduces task embeddings to distinguish historical video, reference images, and target video, enabling unified conditioning for long‑horizon generation. Together with a clean conditioning sink and mixed‑task training, Vorch‑Director supports multi‑shot, multi‑subject, reference‑guided audio‑visual long‑video generation. We evaluate Vorch‑Director on ST‑Bench and introduce a new long‑horizon audio‑visual benchmark with metrics for quality drift and long‑range consistency. Extensive experiments demonstrate improved stability and audio‑visual fidelity over strong baselines.

Authors:Aohua Li, Jin Kuang, Yubing Lu, Pingping Liu
Title: HyTBE: Hyperbolic Target-Background Expert Model for Cross-Domain Infrared Small Target Detection
Abstract:
Infrared small target detection (IRSTD) has achieved substantial progress under domain‑consistent evaluation, yet detector performance often degrades markedly when generalizing to unseen infrared domains. Existing methods primarily improve detection by enhancing target responses and suppressing background interference. However, when trained on only a limited set of source domains, their learned decision rules are inevitably established from a restricted range of source‑domain target‑background relation patterns. We formulate this cross‑domain failure as target‑background relation shift: unseen domains may exhibit relation patterns that are not observed during training, thereby weakening the discriminative capability learned from the source domains. To address this problem, we propose HyTBE, a Hyperbolic Target‑Background Expert model that expands source‑domain relation patterns and adaptively adjusts visual representations using explicit relation cues. The Target‑Background Relation Intervention selectively perturbs either targets or backgrounds, broadening the observable relation patterns during training while maintaining valid supervision. Subsequently, the Hyperbolic Relation Modeling maps multi‑scale visual cues into a Poincaré ball and characterizes the target‑background relation of each feature token according to its relative distances to the target and background anchors. The Hyperbolic‑guided MoE Adapter further uses these hyperbolic relation representations to calibrate multi‑scale visual features and aggregate expert‑specific feature corrections for different relation patterns. Leave‑one‑domain‑out experiments on NUAA‑SIRST, NUDT‑SIRST, and IRSTD‑1K demonstrate that HyTBE achieves stronger cross‑domain generalization than competitive baselines.

Authors:Bryan Wong, Xun Xu, Huazhu Fu, Nancy F. Chen, Mun Yong Yi
Title: Beyond Relevance: Bayesian Evidence Acquisition for Agentic Whole-Slide Image Reasoning
Abstract:
Whole‑slide image (WSI) reasoning requires an agent to sequentially acquire visual evidence before answering a diagnostic question. Existing training‑free agentic frameworks formulate this process as iterative patch retrieval based on semantic relevance to the question. However, semantic relevance does not necessarily imply diagnostic informativeness in computational pathology, where competing diagnoses often exhibit similar and overlapping morphological patterns, making many patches semantically relevant yet diagnostically non‑discriminative. Consequently, relevance‑based retrieval may acquire redundant observations and leave diagnostic uncertainty unresolved. We propose BEACON, a plug‑and‑play agentic framework that reformulates WSI reasoning as a Bayesian evidence acquisition problem. BEACON maintains a probabilistic belief over competing diagnostic hypotheses and sequentially acquires patches by maximizing expected information gain (EIG) to reduce diagnostic uncertainty. An evidence controller then determines whether to answer, acquire additional evidence, or perform higher‑resolution inspection. Built entirely from off‑the‑shelf foundation models, BEACON requires no additional training or fine‑tuning. Extensive zero‑shot experiments across five WSI‑VQA benchmarks demonstrate that BEACON achieves the strongest overall performance among training‑free agentic frameworks while substantially improving evidence acquisition efficiency, establishing Bayesian evidence acquisition as a principled paradigm for uncertainty‑aware agentic WSI reasoning. The code is available at https://github.com/bryanwong17/BEACON

Authors:Mehrshad Saadatinia, Parsa Razmara, Ardalan Aryashad, Ali Abbasi, Seyedarmin Azizi
Title: CircuitSteer: Geometrically Aligned Multi-Layer Steering via Sparse Autoencoder Circuits
Abstract:
Controlling the behavior of large language models (LLMs) remains a critical challenge for AI alignment. Existing steering methods, such as Contrastive Activation Addition (CAA), typically rely on fixed single‑layer interventions derived from aggregate activation differences. These methods impose a single intervention across semantically diverse inputs and often fail to sustain consistent behavioral changes across layers, limiting the effectiveness of the steering. In this work, we introduce CircuitSteer, a novel framework that leverages Sparse Autoencoders (SAEs) to identify and manipulate coherent semantic circuits distributed across multiple layers. By constructing a feature flow circuit based on feature co‑activation and the geometric alignment of decoder directions, we isolate the specific multi‑layer subcircuits responsible for a target behavior. We then synthesize dense steering vectors from these sparse features and apply multi‑point interventions to guide the model's internal semantic trajectory. We evaluate CircuitSteer using contrastive examples across a diverse set of tasks, including toxicity, emotion‑intensity, sycophancy, and refusal, spanning two model families. Across all models and datasets, CircuitSteer is the only method to consistently produce fluency‑preserving interventions; competing methods either sacrifice text quality or lack coverage, failing entirely on complex behaviors like sycophancy and refusal. These results demonstrate that multi‑layer circuit steering, enabled by enforcing geometric alignment among selected features, yields strictly more robust and effective behavioral control than static single‑point interventions. Code is available at https://github.com/mehrshad‑sdtn/CircuitSteer.

Authors:Puyuan Zhang, Jianming Huang, Wenkai Ye, Wei Dong
Title: G$^2$ARD-GS: Geometry-Guided Anchor-Regularized Gaussian Splatting Distillation
Abstract:
Dense colored LiDAR maps provide accurate city‑scale geometry, but lifting them into 3D Gaussian Splatting (3DGS) retains millions of primitives, making the resulting models costly to store, transmit, render, and adapt. Aggressive primitive reduction alleviates this burden, but can remove the local surface support needed for stable novel‑view synthesis and downstream geometric use. We introduce G^2ARD‑GS, a geometry‑guided distillation method that converts a dense Gaussian prior instantiated either as a training‑free point‑cloud lift or a trained GS model into a compact, reusable representation. G^2ARD‑GS progressively consolidates the prior into surface‑aware representatives, then recovers appearance on the resulting fixed topology under construction‑time anchor constraints, with no primitives added or removed during recovery. Under limited supervision, geometry‑aware view selection allocates the available view budget. On MatrixCity, G^2ARD‑GS achieves the best PSNR, SSIM, and LPIPS across matched 5×‑‑30× compression budgets, outperforming PUP by 3.2‑‑6.8,dB in PSNR. When reused as frozen geometry, the compact model improves off‑trajectory appearance adaptation by 3.7‑‑4.9,dB over PUP 3D‑GS and preserves image‑to‑model registration accuracy on Cambridge KingsCollege at 30× compression. Project page: https://patrick1159.github.io/gardGS‑page/.

Authors:Mohammad Hosseini, Hamed Khatounabadi, Mohammad Fakharzadeh
Title: A Low-Power Wearable Respiratory Sensor for Non-Invasive Stress Monitoring
Abstract:
Respiration provides a continuously available window into physiological state and behavior. However, monitoring it outside controlled settings remains challenging because a wearable system must capture small body deformations while remaining comfortable, low power, and robust to changes in posture and motion. We present a compact non‑invasive respiratory sensing system based on a force‑sensitive resistor (FSR) embedded in an abdominal belt and integrated with a custom Bluetooth Low Energy acquisition board. The system combines a simple piezoresistive readout with a mechanical holder designed to transfer abdominal expansion to the sensor without analog amplification. We evaluate the complete sensing pipeline across multiple breathing patterns and body positions. In stationary settings, the recorded signals exhibit consistent amplitude changes and recurring peak‑to‑peak timing across breathing maneuvers; under light movement, these variations remain visible despite motion‑induced baseline shifts. We further design a five‑phase stress‑induction protocol and collect respiratory recordings from 12 participants. Using interpretable time‑domain features and standard classifiers, we examine whether the acquired signals distinguish relaxation from stress‑induction phases. In this preliminary experiment, the best‑performing model achieves 88.0% test accuracy, indicating that the extracted respiratory features distinguish stress‑induced phases from relaxation phases in this dataset. Overall, our results show that the proposed platform enables real‑time respiratory monitoring across diverse daily‑life scenarios and captures respiratory changes that distinguish stress‑induction from relaxation phases, supporting its potential for affective‑computing applications.

Authors:RA Team
Title: JoyAI-RA 0.5: Scaling Robot Manipulation Learning via Dual Action Alignment
Abstract:
Robot data is scarce, so generalist policies need to learn from heterogeneous sources, including human egocentric video, simulation, and real robots, which differ in supervision and embodiment, with action labels missing or mutually incompatible. Human egocentric data scale best but sit farthest from robot data, and naive pooling causes negative transfer rather than knowledge sharing. We propose JoyAI‑RA 0.5, a generalist Vision‑Language‑World‑Action (VLWA) framework that couples physical world‑dynamics priors with visual semantics and scales manipulation learning across such data via dual action alignment. Implicit action alignment infers latent actions from visual transitions, enabling action‑free human, simulation, and robot data to guide a latent‑action‑conditioned world model in learning physical dynamics. Explicit alignment grounds reliable human and robot trajectories in a unified physical action space through a canonical action representation and camera‑frame chunk‑relative end‑effector actions. An inner‑outer‑loop reinforcement stage then pairs efficient task adaptation with foundation‑policy improvement. On a real‑world AgiBot benchmark, JoyAI‑RA performs strongly on both seen tasks and unseen variations. The task score improves consistently as the volume of human egocentric pretraining data increases and shows no sign of plateauing at our largest scale. This suggests that abundant but weakly labeled human experience can be converted into a transferable training signal, making human video not merely a weak auxiliary source but a primary axis along which manipulation capability can be scaled. Project page can be found at https://joyai‑ra‑05.github.io/.

Authors:Jiaheng Chen, Jiaxing Li, Tinghe Zhang, Chaopeng Guo
Title: A Unified Framework for Trajectory Prediction with Explicit Planning and Reaction Decomposition
Abstract:
Trajectory prediction has shifted toward structured formulations with explicit social modeling. However, existing methods inadequately distinguish the functional roles of social influence in trajectory planning. Observing that agents typically form motion plans by anticipating others' future behaviors before making local reactive adjustments, we identify social interactions as playing staged roles, namely planning precedes reaction. We propose INTraJ, a unified framework that decomposes social influence into two stages: a planning stage constructs reference trajectories using future social information, and a reaction stage recovers local adjustments from the residual between full‑context prediction and the reference. INTraJ supports both multi‑target and single‑target paradigms. Extensive experiments on four standard benchmarks, including Argoverse 2, Argoverse 2‑ped, ETH/UCY, and SDD, demonstrate consistent improvements, particularly in FDE and long‑horizon consistency, with state‑of‑the‑art performance achieved in several settings. INTraJ reframes trajectory prediction as a planning‑driven two‑stage process, validating that staged social modeling is critical for stable predictions. The code is publicly available at https://github.com/11isnotavailable/INTraJ.

Authors:Guoan Xu, Zhengxue Wang, Yang Xiao, Ligeng Chen, Guangwei Gao, Dongchen Zhu
Title: URNet: A Unified Reparameterized Network for Efficient RGB-D Semantic Segmentation
Abstract:
Previous RGB‑D semantic segmentation methods commonly employ dual encoders to separately process RGB and depth inputs, followed by dedicated modules for cross‑modal feature fusion. However, such designs often inadequately capture depth representations and consequently limit effective cross‑modal interaction, while the additional encoder branch introduces redundant computation that hinders lightweight execution. To tackle these challenges, we propose URNet, a Unified Reparameterized RGB‑D Network that performs simultaneous multi‑modal feature extraction and cross‑modal fusion within a single encoder. Specifically, we adopt a reparameterization strategy to compact the network architecture and facilitate fast inference. Within each Reparameterized Block (RepBlock), a Linear Gated Attention (LGA) module is introduced to fully exploit complementary RGB and depth cues across different feature scales. Furthermore, considering that decoder design has been relatively underexplored in existing RGB‑D segmentation models, we develop a concise yet effective universal decoder, termed the Pyramid Merging Decoder (PMD). Extensive experiments on multiple RGB‑D segmentation benchmarks demonstrate that URNet achieves state‑of‑the‑art performance while maintaining high efficiency. Code will be available at https://github.com/Wild‑Stephen/URNet.

Authors:Hao Si, Zehua Chen, Qingquan Yang, Xiao Wang, Dengdi Sun, Wanli Lyu, Gaoting Chen, Guosheng Xu, Hang Su, Jin Tang, Jun Zhu
Title: SafeDivertor: Faithful Divertor Heat Flux Reconstruction from Macroscopic Plasma State Signals via Time-Frequency Prior Exploitation
Abstract:
Divertor heat‑flux analysis is essential for understanding plasma‑wall interactions and protecting plasma‑facing components in magnetic‑confinement fusion devices, while conventional infrared‑based inversion is usually performed after discharge and requires heat‑conduction modeling with device‑specific material properties, divertor geometry, and boundary conditions. Rather than accelerating this conventional infrared‑based inversion paradigm, we introduce a new online‑oriented signal‑based reconstruction paradigm that directly reconstructs time‑resolved radial heat‑flux profiles from multi‑source macroscopic plasma‑state signals available during discharge. To enable systematic study of this task, we construct DivMPS2HF, a multi‑source discharge dataset that provides the data foundation and benchmark for signal‑based divertor heat‑flux reconstruction. We further propose SafeDivertor, a task‑driven framework designed to address the key challenges of signal‑based heat‑flux reconstruction. It employs physical prior‑aware initialization to provide radial‑distribution guidance for target channels, input perturbation to reduce over‑reliance on specific heterogeneous signals, spectral‑aware reconstruction optimization to exploit time‑frequency priors and preserve transient dynamics, and progressive training to stabilize the optimization of these complementary objectives. Experiments on DivMPS2HF demonstrate that SafeDivertor achieves the best overall performance among the evaluated time‑series baselines across all five metrics, establishing a new performance benchmark for signal‑based divertor heat‑flux reconstruction. The source code will be released on https://github.com/Event‑AHU/OpenFusion

Authors:Menglin Han, Yang Ding, Yulei Lu, Haoran Yu, Xin Ma, Junyi Chen, Zhangkai Ni, Lin Ma, Yaohui Wang
Title: Vorch-Streamer: Extending Human Audio-Visual Generation to Real-Time Long-Form Streaming
Abstract:
Real‑time long‑form avatar audio‑‑video generation requires causal, continuous synthesis while maintaining audiovisual synchronization and visual consistency. Adapting a pretrained bidirectional model to this setting presents two key dilemmas. First, autoregressively reusing generated blocks as context creates exposure bias, causing errors and visual drift to accumulate over long rollouts. Second, a global speech utterance does not indicates a causal generator which portion should be spoken next when only limited local audio‑‑video context is available. We present Vorch‑Streamer, a post‑training framework that addresses these challenges and enables real‑time long‑form Text‑to‑Audio‑Video (T2AV) streaming. We construct a synthetic corpus of 80K avatar clips spanning 12‑‑21 seconds and first train a causal generator with mixed Teacher Forcing and Diffusion Forcing. We then apply long‑horizon Self Forcing with DMD distillation, exposing the model to its own rollout distribution while preserving the quality of the pretrained bidirectional teacher. To explicitly control speech progression, an external language model predicts discrete 25‑Hz speech‑planning tokens, whose continuous features condition the audio diffusion branch and align each causal block with the content it should speak. With bounded causal context and four‑step denoising, Vorch‑Streamer jointly generates audio and video from text at 27.12 FPS, exceeding the 24‑FPS real‑time playback rate while maintaining competitive audio‑‑lip synchronization and strong identity preservation over long‑form generation.

Authors:Yaole Wang, Xiaoyu Chen, Xin Ma, Yang Ding, Gang Yue, Jingjing Chen, Lin Ma, Yaohui Wang
Title: Vorch-IR: Long-Form Unified Multimodal Identity Replacement Video Generation
Abstract:
Video identity replacement seeks to transfer the identities of one or more subjects while preserving the motion, expressions, and temporal structure of a driving video. Existing methods largely target single‑person settings and often require task‑specific structural controls, such as masks or pose representations, limiting their flexibility in general multimodal editing systems. Progress on multi‑person replacement is further constrained by the scarcity of paired training data. We present Vorch‑IR, a unified framework that supports single‑ and dual‑person identity replacement, with optional background replacement, in a single model. Built on LTX2, Vorch‑IR jointly conditions on a driving video, indexed reference images, and a textual editing instruction. The reference images need not match the pose, layout, or spatial configuration of the driving video: their roles as subject or background references are specified through the instruction. Dense visual conditions are fused through self‑attention, while a vision‑language context establishes semantic correspondence through cross‑attention. We further develop an automatic data construction pipeline that synthesizes paired supervision for all four editing settings. Experiments using automatic metrics and pairwise human evaluation demonstrate strong identity preservation, motion fidelity, and temporal coherence across diverse scenarios. A temporal overlapping inference strategy additionally extends the short‑clip model to minute‑long generation without autoregressive continuation.

Authors:Chih-Chung Hsu
Title: RASP-QAOA: Resource-Aware Per-Instance Selection for Exact QAOA Simulation
Abstract:
Exact QAOA simulation spans several computational representations whose useful regions differ sharply across graph structure, circuit depth, precision, and available memory. Choosing only a backend name hides these differences: an executable choice also fixes the representation, adapter, precision mode, and memory policy. We introduce RASP‑QAOA, a per‑instance selector over ten such actions. It first removes actions that cannot implement the requested QAOA semantics or execution requirements, then orders the remaining actions using instance features; actions outside learned support are handled by analytical work estimates. On a content‑disjoint 60‑request H200 evaluation, RASP‑QAOA succeeds on all 31 requests for which at least one admissible action completes and validates. Within this set it reaches 27/31 top‑1 and 31/31 top‑2 selection, with 1.051 geometric‑mean regret. Its failure‑penalized PAR10 score is 0.0396 times that of development‑selected CUAOA (95% interval: 0.0085‑0.1644). A separate 30‑request crossover shows that graph structure changes 16 decisions and improves the paired penalized score, while a depth‑1 stump matches gradient boosting. The evidence supports resource‑aware representation selection at n <= 35, p <= 5, with gains driven by representation features rather than classifier complexity.

Authors:Keane Zhang, Varshini Chinta, Raj Sanjay Shah, Sashank Varma
Title: Human-Like Anaphor Resolution in Large Language Models
Abstract:
Anaphors are expressions that refer to other expressions, called antecedents. The process of connecting the two is called resolution. Cognitive science has identified multiple factors that affect the speed and success of anaphor resolution, including discourse structure, situation‑model properties, and semantic factors. Here, we investigate whether these factors also affect anaphor resolution in five Large Language Models (LLMs) with open weights: GPT‑2‑XL, Llama‑3.1‑8B, Pythia‑12B, Mistral‑7B, and Mistral‑24B. To model processing difficulty, we adopt the standard linking hypothesis that relates human reading times to model surprisal at the anaphor. As a second behavioral measure, we compare model accuracy to human accuracy on comprehension questions probing the antecedents of anaphors. The results show selective cognitive alignment: some LLMs exhibit human‑like sensitivity to discourse prominence and distance‑based factors in anaphor resolution, while showing weaker or absent sensitivity to semantic interference effects. These findings delimit the conditions under which LLMs approximate human anaphor resolution.

Authors:Mohamad Zamini, Diksha Shukla
Title: SCI-CLIP: Segment-Centric Inference with Reference Memory for Training-Free Open-Vocabulary Segmentation
Abstract:
Training‑free open‑vocabulary segmentation remains limited by a missing inference abstraction. Frozen vision‑language features are produced at patch level, yet dense prediction requires a unit that simultaneously governs feature interaction, spatial support, contextual recovery, and retrieval‑based correction. We present SCI‑CLIP, a segment‑centric inference framework built around the principle that the same region abstraction should organize all stages of dense open‑vocabulary prediction. SCI‑CLIP first induces a region‑consistent interaction graph over frozen visual tokens, then reconstructs dense features by propagating values over this graph, augmenting them with selective cross‑window support only where local evidence is insufficient. The same segment abstraction is subsequently used to construct and query an offline reference memory, aligning exemplar retrieval with the units on which prediction is made. SCI‑CLIP turns frozen CLIP‑style features into spatially coherent, context‑aware, and retrieval‑compatible dense predictions without any training. SCI‑CLIP consistently improves the structural quality of dense predictions, the robustness of contextual reasoning, and the alignment of exemplar‑based correction, yielding stronger open‑vocabulary segmentation across eight benchmarks. Project code is available at: https://github.com/mzamini92/SCICLIP.

Authors:Yanqi Wu, Runhe Lai, Xinhua Lu, Qichao Chen, Zhiping Zhou, Jia-Xin Zhuang, Weijiang Yu, Ruixuan Wang
Title: TruthLens: Object Hallucination Detection via Self-Evaluating Truthfulness Scores in LVLMs
Abstract:
Despite the remarkable progress of large vision language models (LVLMs), object hallucination remains a fundamental challenge that hinders their trustworthy deployment. A key finding motivates our work: real and hallucinated object tokens are clearly separable in hidden representations, yet this separability is largely lost at the language‑modeling (LM) head. We propose TruthLens, a self‑evaluation framework that teaches the LM head to expose a per‑object truthfulness signal without any auxiliary model or additional inference cost. Concretely, a rarely‑used special token is repurposed as a reference token. For each object‑token position, we extract the log‑probability assigned to this special token by the LM head, and define its difference from a predefined constant as the truthfulness score. The model is then fine‑tuned with an MSE objective that drives scores toward 1 for real objects and 0 for hallucinated ones, while a divergence constraint preserves the original generation capability. Despite being trained on only a limited set of object categories, TruthLens generalizes effectively to benchmarks with substantially larger label spaces. Extensive experiments across multiple LVLMs demonstrate state‑of‑the‑art performance; notably, on Qwen2.5‑VL‑7B, TruthLens outperforms the previous best method on MS‑COCO by over 17% in AUROC. Our code is available at https://github.com/wyqstan/TruthLens.

Authors:Dongchen Li, Jitao Liang, Wei Li
Title: ALTER: Modeling Longitudinal Changes via Regional Differencing for 3D CT Report Generation
Abstract:
Computed tomography (CT) is widely used for clinical diagnosis and longitudinal follow‑up, yet automatically generating accurate and complete radiology reports from three‑dimensional (3D) CT remains challenging. Existing methods improve fine‑grained correspondence between images and text by modeling anatomical regions, but remain centered on the current examination. Consequently, patient‑specific longitudinal changes within individual regions remain insufficiently modeled. Meanwhile, interval changes are often distributed across multiple anatomical regions, complicating a coherent assessment of the overall longitudinal state. We propose Anatomically Localized Temporal Evidence Representation (ALTER) to address these limitations. Global Prior Integration (GPI) incorporates the prior CT and report to establish historical context for the current examination. Regional Proxy Differencing (RPD) enables each current anatomical region to retrieve a historical proxy from a single shared encoding of the prior volume and to derive localized interval evidence. Interval Change Fusion (ICF) further combines current abnormality states with region‑distributed differences, converting their joint representation into change‑aware soft prompts that guide report generation. ALTER achieves state‑of‑the‑art results on most evaluation metrics across the RadGenome‑ChestCT validation and CTRG‑Chest‑548K test sets. Code and data preprocessing details are available at https://github.com/peytonkarlie/ALTER/tree/main.

Authors:Guanyu Wang, Zidi Zhang, Xu Chu
Title: FOCUS: Decoupling Expert Personas in LLMs to Enhance Domain Expert Capabilities
Abstract:
Large Language Models (LLMs) can exhibit diverse personas, and activating expert personas has been shown to improve domain expertise and task accuracy. However, existing persona control methods often suffer from cross‑domain coupling, which may lead to overly aggressive behavior in high‑caution domains such as healthcare, or excessive conservatism in risk‑sensitive domains such as financial trading. To address this issue, we propose FOCUS (\underlineFine‑tuning with \underlineOrthogonal \underlineControl for \underlineUncoupled persona\underlineS). FOCUS first automatically extracts expert persona vectors from LLMs, then applies orthogonal decomposition to decouple domain‑specific expert personas, and finally introduces an expert gating module to adaptively control persona activation according to task contexts. With a two‑stage training strategy and a gated selection regularizer, the model learns to activate appropriate personas for both single‑domain and cross‑domain tasks. Experiments on financial, legal, medical, and cross‑domain benchmarks show that FOCUS improves task accuracy and outperforms existing persona control methods. Our code is available at \hrefhttps://anonymous.4open.science/r/openpersona‑48F4this url.

Authors:Chih-Chung Hsu
Title: LC-Implicit-QAOA: Active-Workspace-Capped Exact Objective-and-Gradient Evaluation for Training over Bounded QUBO Light Cones
Abstract:
QAOA training repeatedly queries an objective and all shared gradients, making exact evaluation a feasibility bottleneck even when QUBO terms have bounded causal cones. Building on established causal‑cone restriction and adjoint differentiation, LC‑Implicit‑QAOA profiles cone structure and induced‑edge counts before local‑amplitude and named‑workspace allocation, then jointly selects equal‑size microbatches and checkpoint schedules under a named active‑evaluator workspace budget. "Implicit" means omitting both global state and global cost table, not implicit differentiation; infeasible requests are rejected before those allocations. An independently implemented complex128/float64 dense adjoint agrees with LC over 1,800 graph‑angle comparisons, with a worst relative gradient error of 1.56 x 10^‑13. LC completes all 104 target requests in a p=2 bounded‑cone grid; under a prespecified n <= 24 validation cap, the matched state‑plus‑cost reference is executed for 28 requests and deliberately not run on 76. Across 80 budgeted requests, measured allocated evaluator memory stays within budget, reaching at most 0.797 of it. On 3‑regular n=512, p=2, the adjoint reaches the same finite‑budget endpoint in 101 objective‑equivalent calls and 189 s, versus 909 calls and 1,565 s for central differences. LC targets fixed‑depth one‑ and two‑local diagonal QUBO costs with a transverse‑field mixer; it provides neither global states, sampling, nor a hardware‑independent fastest‑backend rule.

Authors:Deyi Zhu, Haoyu Fan, Yinan Zhu, Weichen Zhang, Shilin Ma, Xinlei Chen, Yansong Tang
Title: Uncertainty-Aware World Model for Aerial Image-Goal Navigation
Abstract:
Aerial image‑goal navigation requires an unmanned aerial vehicle (UAV) to reach a target location specified by a goal image. Existing world‑model‑based methods rank candidate trajectories using predicted futures, but typically rely on only one or a few point predictions, which is inadequate for large‑scale outdoor environments with substantial future‑state uncertainty. To address this limitation, we propose the Uncertainty‑Aware Navigation World Model (UA‑NWM), an efficient latent world model for aerial image‑goal navigation, which formulates trajectory scoring as conditional out‑of‑distribution detection. UA‑NWM represents plausible futures with an uncertainty subspace and decomposes the prediction‑‑goal discrepancy into uncertainty‑explainable and unexplainable components. Only the unexplainable residual is used for scoring, enabling robust selection without multiple future samples. Extensive experiments demonstrate that UA‑NWM consistently outperforms existing navigation world models while maintaining low inference latency. Real‑world UAV experiments further validate its practical applicability. Project page: https://duryi.github.io/UA‑NWM‑Project‑Page

Authors:Rishik Sathua, Haonan Chen, Katherine Driggs-Campbell
Title: ARGUS: Aligning Robot Scene Geometry Under Shifting Views with Large 3D Vision Models
Abstract:
Large‑scale visuomotor policies have demonstrated impressive performance across a wide range of robot manipulation tasks. However, despite this success, manipulation polices often entangle scene geometry with the corresponding viewpoint, learning where objects lie in an image rather than where it lies in the task space. This entanglement inherently limits the corresponding policy's ability to learn from viewpoint‑diverse datasets (ex. DROID, BridgeV2) and generalize beyond the viewpoints captured in their training data. In this work, we present ARGUS, an observation pre‑processing pipeline that uses large‑scale 3D vision models to align image observations from arbitrary camera viewpoints into a canonical viewpoint before passing it to downstream visuomotor policies. Experiments across training datasets with varying levels of viewpoint diversity, from fixed multi‑view camera configurations to highly varied camera placements, show that our method consistently outperforms prior approaches across both limited‑view and view‑diverse training regimes. In efficiency comparisons, ARGUS demonstrates an ability to learn from view‑diverse data, converging to high success rates 4‑6x faster than previous methods by leveraging a simplified observation space. Overall, our findings show that leveraging large‑scale 3D vision models reduces the learning burden on visuomotor policies, enabling more efficient learning from large‑scale, viewpoint‑diverse robot datasets.

Authors:Zhi Han, Chenxi Zeng, Liuhaichen Yang, Zihan Guo, Ming Zhou, Yang Li
Title: SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution
Abstract:
LLM agents increasingly execute long‑horizon tasks through tool use and environment interaction, shifting evaluation from final‑response scoring to verification of complete executions. For skill‑augmented agents, verification additionally requires the procedural knowledge encoded in task‑time skills, because this knowledge indicates what evidence to inspect and which failures are task‑critical. However, existing judge benchmarks often expose final responses or static trajectories, and rarely combine task‑time skills with directly inspectable artifacts and environments. We therefore introduce SkillTV‑Bench, a 681‑case benchmark of real agent trajectories from 50 tasks across eleven domains, designed to evaluate skill‑aware trajectory verification for both LLM‑as‑a‑Judge and Agent‑as‑a‑Judge methods. Additionally, we propose SkillTV‑Evolve, which externalizes verification knowledge as a reusable JudgeSkill that guides an agent judge to plan targeted inspections and issue evidence‑grounded verdicts. On a disjoint development pool, an automated evolution loop further refines the JudgeSkill using misjudged cases. On SkillTV‑Bench, the refined skill increases the same agent judge's accuracy by 14.8 percentage points. In offline rollout‑pool selection, it increases selected‑trajectory success from 22.9% with one rollout to 45.5% with ten rollouts. The code and data are available at https://github.com/HanZhi306/SkillTV‑Bench

Authors:Mohammad Asadi, Soheil Hor, Bardiya Akhbari, Jack W. O'Sullivan, Tahoura Nedaee, Layne C. Price, Raviteja Anantha, Euan Ashley, Ehsan Adeli
Title: Align-RAG: Alignment Is All You Need for TSFM In-Context Learning
Abstract:
Retrieval‑augmented forecasting promises to adapt frozen Time Series Foundation Models (TSFMs) to new domains without fine‑tuning, but recent methods typically rely on learned fusion modules, i.e., trained adapters that merge retrieved examples into the backbone's forecast, based on the assumption that frozen backbones cannot dynamically incorporate retrieved context on their own. We show this assumption is unnecessary. We introduce Align‑RAG, a training‑free method that applies a closed‑form per‑pair amplitude rescaling and integer‑lag phase shift to retrieved past‑future windows before they enter a frozen backbone's context. With no learned parameters, Align‑RAG outperforms the state‑of‑the‑art trained retrieval adapter on a frozen Chronos‑Bolt on all seven datasets of the standard benchmark (avg ‑3.75% MSE), showing that the gains previously attributed to learned fusion are recoverable without any training. Align‑RAG further improves zero‑shot MSE on four additional frozen TSFMs with various architectures by 2.5% to 13.7% per backbone with no per‑backbone tuning. To probe why alignment helps, we compare the frozen backbone's prediction shift under aligned demonstrations to the closed‑form ridge prediction shift on the same pairs. We find that aligned demonstrations induce prediction shifts that track a closed‑form ridge predictor on the same pairs, with a future‑shuffle control ruling out a futures‑averaging account. Together, these results indicate that frozen TSFMs already support dynamic in‑context use of retrievals, and that closed‑form alignment should be the default baseline for retrieval‑augmented forecasting before any fusion module is trained. Code available at: https://github.com/masadi‑99/align‑rag

Authors:Feier Wu, Wanke Xia, Xu He, Zilang Zhou, Si Chen, Dongxia Liu, Liyang Chen, Qimeng Wu, Zhengbo Zhang, Wenming Yang, Zhiyong Wu
Title: EffectLearner: World-Aware Object-Effect Reasoning for Real-World Video Object Removal
Abstract:
Video object removal must eliminate not only the target object but also its induced effects while maintaining high‑fidelity and spatiotemporally coherent restoration. Existing methods mainly learn object‑effect correspondences implicitly from predefined effect categories and fixed data distributions, limiting their generalization to complex real‑world scenes involving compositional effects, spatially detached or weakly correlated effects, long‑tail physical phenomena, and dynamically evolving interactions. We propose EffectLearner, a semantic‑reasoning‑enhanced framework that combines a VLM‑based Object‑Effect Reasoner with a DiT‑based Video Eraser. Guided by a structured effect‑analysis prompt, the Reasoner performs cross‑modal reasoning over a target‑highlighted video and extracts compact effect‑aware context, which guides the Video Eraser toward comprehensive object‑effect removal. Motion‑aware mask guidance and motion‑consistency supervision further improve removal coverage and spatiotemporal stability under object motion and evolving scene dynamics. To fully exploit the framework in challenging real‑world scenarios, we further construct EffectWorld, a paired video dataset specifically designed for complex object‑induced effects, and introduce a progressive training curriculum that combines common supervision with complex‑effect data. On the standard ROSE‑Bench, EffectLearner outperforms existing baselines on most metrics and achieves clear advantages on both EffectWorld‑Eval and the challenging EffectWorld‑Wild, demonstrating its ability to deliver high‑quality video object removal in complex real‑world scenes.

Authors:Yu Gu, Zhi Zheng, Yunpeng Ba, Xialiang Tong, Mingxuan Yuan, Zhenkun Wang
Title: Hyper-ES: Effective Evolution Strategies for LLM Reasoning via Descent Direction Merging
Abstract:
Evolution Strategy (ES) is a promising alternative to gradient‑based fine‑tuning for resource‑constrained Large Language Model (LLM) reasoning. However, directly applying ES to billion‑parameter LLMs is highly ineffective. In such high‑dimensional parameter spaces, most random perturbations are nearly orthogonal to useful update directions, leading to unstable optimization. We propose Hyper‑ES, a subspace‑based ES framework that avoids the weakness of ES in full‑parameter search while exploiting its strength in low‑dimensional optimization. Instead of asking ES to discover useful directions from random perturbations in the LLM parameter space, Hyper‑ES first performs a small number of inexpensive gradient‑based fine‑tuning runs to obtain descent directions. Although each direction may provide only a limited improvement on its own, their span forms a compact adaptation subspace that captures useful reasoning updates. Hyper‑ES then applies CMA‑ES to optimize layer‑wise DARE‑TIES merging coefficients within this subspace, allowing ES to search over combinations of meaningful descent directions rather than over arbitrary full‑model perturbations. We evaluate Hyper‑ES on three Qwen2.5‑Instruct and DeepSeek‑R1‑Distill backbones across six mathematical reasoning datasets. Results show that Hyper‑ES consistently outperforms GRPO‑LoRA by 1% while requiring 10% fewer space‑consuming gradient updates. Code at https://github.com/kuangrepi/Hyper‑ES.

Authors:Pratik Patil, Bhushan Bonde, Bhaskar Choubey
Title: A Quantum Circuit Framework for Protein Ensemble-Level Energetics
Abstract:
Proteins occupy heterogeneous free‑energy landscapes in which high‑entropy ensembles converge toward compact, low‑energy basins with multiple sub‑states. Molecular dynamics can access these landscapes at atomic resolution, but exhaustive sampling remains computationally demanding. Meanwhile, most quantum approaches target only single optimal structures, leaving full ensemble energetic heterogeneity unexplored. We introduce a residue‑level, gate‑based quantum circuit framework for coarse‑graining protein thermodynamics. Each amino acid is represented as a two‑state qubit (stabilised vs. excited solvation state) based on residue solvation energetics. A structure‑informed entanglement block then encodes covalent and non‑covalent contacts using parameterised controlled gates, embedding correlations across the residue‑interaction network. Sampling the circuit (~ 10^6 measurements) yields binary thermodynamic microstates used to compute protein energy distributions, residue‑level statistical couplings, energetic sensitivities, and information gains relative to total free energy. We showcase the framework on the benchmark Trp‑cage miniprotein 1L2Y (TC5b) and 9GDL, a disulfide‑stabilised Trp‑cage‑fortified exenatide chimera. For 1L2Y, the circuit reproduces a structured, folding‑funnel‑like energy distribution. Comparative analysis with 9GDL reveals shifts in global energy distributions and residue‑level stability profiles. Coupling and information‑theoretic analyses localise residues associated with ensemble reorganisation, while multi‑body couplings show the circuit resolves both direct and indirect statistical correlations. This framework expands quantum protein modelling beyond single‑structure optimisation toward ensemble‑level characterisation, capturing key features of rugged energy landscapes to guide protein design, mutation mapping, and allosteric pathway identification.

Authors:Ziyun Zeng, Zixuan Wang, Yongsheng Yu, Hang Hua, Jiebo Luo
Title: VideoArgus: Agentic Rubric-Grounded Unified Evaluation for Video Generation and Editing
Abstract:
Evaluating generated videos remains challenging because existing benchmarks rely on fixed evaluation content, cover only a subset of generation and editing settings, and provide limited evidence for their scores. We introduce VideoArgus, a unified rubric‑grounded framework covering five video generation and editing settings. For each input instance, VideoArgus generates an output‑blind, sample‑specific rubric once and reuses it to evaluate all corresponding candidate videos. The rubric defines concrete criteria, scoring rules, failure modes, and evidence plans, which guide criterion‑specific VLM QA and visual tools to produce evidence‑grounded criterion scores, rationales, and a diagnostic report. We further construct VideoArgus‑Bench, containing 1,026 curated input instances built from 653 high‑quality images and 416 high‑quality videos, with all benchmark rubrics pre‑generated, frozen, and released. On a separate 1,260‑video human‑alignment set, VideoArgus achieves higher within‑input Spearman and Kendall correlations with human judgments than the corresponding benchmark‑specific evaluators across all five tasks. Model rankings also remain largely consistent across different rubric‑generation and evaluation‑VLM backbones. All code and data are released. Visit our project page: https://zzzmyyzeng.github.io/VideoArgus

Authors:Zhaowei Han, Xiang Zhang, Bing Han, Kai Liu, Danqi Hu, Jie Liu
Title: KV-Skill: Forging Expertise in the Model's Native Language
Abstract:
Task knowledge is commonly stored either as text in the prompt or as an update to model weights. Text is modular but must be interpreted on every use, while weight adaptation makes the resulting capability difficult to load, remove, or share independently. We introduce KV‑Skill, a design space of external factorized operators that a frozen language model reads through a lightweight interface. KV‑Skill supports two complementary paths. Registration converts an authored text skill into a text‑derived operator and trains a shared per‑backbone interface. Reward learning develops a compact latent operator directly from task outcomes, with or without an authored skill. Neither path adds positions to the prompt. Across ten benchmarks and four backbones from three model families, converting text to a KV‑Skill consistently makes the same procedural knowledge more effective. On Qwen3.5‑4B LiveMath, registration reaches 77.2 accuracy, compared with 23.4 for the source text skill, 52.0 for SkillOpt, and 64.5 for SoftSkill. Under matched reward training and parameter budgets, KV‑Skill gives the best result in seven of eight matched settings against soft prefixes, prefix tuning, and LoRA. A post‑hoc rank analysis further shows that text‑derived operators retain nearly all of their benefit with one task‑aligned direction per injection layer, while matched random directions fail. Finally, one shared interface retains three independently loadable KV‑Skills without measurable forgetting. These results show that task knowledge can be acquired from text or experience, compressed into an external operator, and deployed separately from the backbone. Code is available at: https://github.com/shawnzhg/KV‑Skill

Authors:Harvey Mannering, Yilin Zhang, Ziao Liu, Zhiwu Huang, Jacqueline Matthew, Miguel Xochicale
Title: A Foundational EDM2-Based Generative Model for High-Resolution Synthetic Fetal Ultrasound Imaging from Open Datasets
Abstract:
Prenatal ultrasound imaging is key for assessing fetal health, but AI progress is limited by scarce, privacy‑restricted, and hard‑to‑annotate datasets. We propose a high‑resolution fetal ultrasound synthesis framework based on the EDM2 diffusion architecture, trained on multiple public datasets to generate 512x512 images across six anatomical classes. Our method achieved improved image quality with lower FID scores and enhanced downstream fetal plane classification, reaching 93.36% ensemble accuracy after fine‑tuning, surpassing real‑data‑only training. Clinical evaluation by an experienced fetal ultrasound specialist (10+ years) on 100 images yielded a mean realism score of 2.67/5, with real images rated higher than synthetic. Artefacts included smoothing, speckle irregularities, and anatomical inconsistencies. Code, data, models and other resources to reproduce this work are available at https://github.com/xfetus/fetal‑ultrasound‑edm2.

Authors:Yisu Zong, Jinjie Shi, Joshua Reiss
Title: Diff2Mix: Controllable Music Mixing via Diffusion Models and Differentiable Audio Effects
Abstract:
Automatic music mixing aims to combine multitrack recordings into a balanced and coherent musical piece. Because the content of different songs and the subjective preferences of mixing engineers jointly shape the final outcome, a practical system should deliver well‑balanced mixes while allowing for controllable stylistic variation. However, most existing methods treat automatic mixing and mixing style control as separate tasks, making it difficult for a single system to produce high‑quality mixes while remaining editable and style‑aware. To address this limitation, this paper presents Diff2Mix, a generative automatic mixing system based on diffusion models and a differentiable mixing console. This system offers two levels of optional user control: a reference audio enables overall production style control, and the differentiable mixing console provides explicit audio effects parameters for interpretability and fine‑grained optimization. We demonstrate our system's competitive performance through both objective and subjective evaluations in terms of mixing quality and control ability. We provide code and audio samples at our project page https://zys711.github.io/Diff2Mix .

Authors:Buzhao Liu, Xinhang Ma, Yevgeniy Vorobeychik
Title: Robust Context-Aware Detection of Malicious Instructions in Text
Abstract:
The remarkable instruction‑following ability of modern LLMs has enabled their practical use as the minds of agents that can autonomously complete increasingly complex tasks. Therein, however, also lies their vulnerability to attacks which embed malicious instructions in text, common variants of which are known as indirect prompt injection (IPI). A fundamental task in addressing this vulnerability is successful segmentation of a given text into benign and malicious sentences (if any). While a number of approaches for this task have been proposed, no detector combines query‑relative detection at the segment level, and none are hardened against adaptive evasion attacks realizable in agentic executions. We address the former limitation by developing an approach for malicious sentence classification that is both context‑ and query‑aware. Next, to harden the resulting classifier against evasion, we present two adversarial training methods. The first is directly adapted feature‑space adversarial training (AT) in which evasions are approximated using projected‑gradient‑based optimization in the embedding space. The second simulates realizable evasion attacks in the AT loop through LLM‑based paraphrasing. Crucially, we parametrize both AT variants to facilitate a smooth tradeoff between utility and attack robustness. In extensive experiments using indirect prompt injection benchmarks we show that the proposed approach outperforms state‑of‑the‑art IPI defense baselines under static attacks, while in the case of adaptive attacks, our AT variants provide significantly higher utility, lower attack success rate, and often both. Finally, we show that the best AT parameters can depend intimately on the particular application domain. Consequently, domain‑dependent tuning of malicious text detectors is likely necessary in practice. Our code is publicly available at https://github.com/tavia‑liu/CAD.

Authors:Vladan Stojnić, Ryan Ramos, Giorgos Kordopatis-Zilos, Noa Garcia, Giorgos Tolias
Title: Invisible Shortcuts: Why Vision Encoders Know Your Camera
Abstract:
Deep vision models exploit shortcuts, relying on cues that correlate with supervision signals. Prior work has focused on visible biases, such as object‑background or texture correlations. We identify a different source of shortcut learning: invisible metadata traces embedded at the pixel level, for metadata such as image processing and photo acquisition. We hypothesize that large‑scale semantic supervision, whether through categorical labels (ImageNet) or billion‑scale captions (LAION), naturally induces metadata‑semantics correlations during pretraining, leading models to convert low‑level signals into predictive features. By introducing controlled metadata‑semantics correlations, we show that stronger ones produce systematically higher sensitivity to metadata traces and larger performance degradation under metadata distribution shifts. We further explore mitigation strategies applied during and after pretraining that reduce sensitivity not only to targeted metadata but also to unseen ones, without sacrificing performance on downstream tasks. Metadata sensitivity also has a positive side: it partly explains the strong generated‑image detection ability of some encoders, while its mitigation can improve out‑of‑distribution generalization. Code: https://github.com/ryan‑caesar‑ramos/visual‑encoder‑traces

Authors:Michael Timothy Bennett
Title: Why the Third Axis Is Freedom
Abstract:
In generative training, a model produces an output and is penalised for its difference from an example. With one output per comparison, a model that produces one common answer can outperform a model retaining a broader repertoire. Explorative Modeling (XM) produces K outputs per comparison and updates on the closest, claiming exploration as a "third pretraining axis" associated with generative expressivity. Here I show the third axis is actually freedom, meaning the weakness of the constraint implied by a model's behaviour. Previous work showed freedom is a property of function rather than form. Parameters, architecture, minimum‑description‑length (MDL), and data can vary while the behavioural constraint remains unchanged. It was formally proved that weakest models are likeliest to generalise, and freedom selection beat MDL by 110‑500% in induction experiments. I prove average XM loss depends on the chance a candidate misses an acceptable region, with exploration raising miss probability to power K. For K>1, match probability rises with freedom. I then demonstrate empirically that XM optimises for freedom. In a Forward XM experiment, larger K increased or saturated measured freedom, and increased freedom at every tested value under context‑dependent targets. I trained XM candidate pools and compared validation selection with a freedom selector that read unlabelled parent contexts. Freedom won in 29 of 30 cases. Generative expressivity is a mode‑count proxy for freedom, that discards the extension structure that gives freedom its generalisation significance. XM is a means, freedom an end, and selecting for freedom improved XM under distribution shift.

Authors:Xi Xiao, Xingjian Li, Cheng Han, Tianyang Wang, Lin Zhao, Yunbei Zhang, Guosheng Hu, Runmin Jiang, Xi Li, Xiao Wang, Min Xu
Title: Adapting Vision Foundation Models with Cascaded Semantics
Abstract:
Prompt tuning, a leading parameter‑efficient adaptation paradigm in NLP, has recently been extended to computer vision. Visual prompt tuning (VPT) adapts pre‑trained vision transformers (ViTs) by updating a small set of additional prompt parameters. However, existing visual prompts are randomly initialized and do not exploit prior knowledge, such as instructions in NLP. We address this gap by injecting two complementary semantic priors into VPT. Fundamental image priors, including color, texture, and shape, are extracted with classical hand‑crafted operators and injected into the input space, while self‑attention maps provide instance‑aware semantics in the feature space. We further propose a cascaded scheme that integrates both priors throughout ViT adaptation. Experiments on 34 challenging image classification datasets demonstrate superior downstream adaptation while tuning only 0.74% of ViT parameters. Project page: https://xixiaouab.github.io/Cascaded‑Semantics/.

Authors:Ruiyu Wang, Yuzhang Xie, Xiao Hu, Carl Yang, Jiaying Lu
Title: BioMedJImpact: A Comprehensive Dataset and LLM Pipeline for AI Engagement and Scientific Impact Analysis of Biomedical Journals
Abstract:
Assessing journal impact is central to scholarly communication, yet existing resources rarely capture how collaboration and artificial intelligence (AI) research jointly shape venue prestige in biomedicine. We present BioMedJImpact, a large‑scale, biomedical‑oriented dataset built from 1.74 million PubMed Central articles across 2,744 journals. BioMedJImpact integrates bibliometric indicators, collaboration features, and an LLM‑derived AI engagement rate, defined as the proportion of AI‑related articles within each journal‑year. Specifically, AI engagement rate is extracted through a reproducible three‑stage LLM pipeline. We analyze how collaboration intensity and AI engagement rate jointly influence scientific impact across two temporal subsets (2016‑2019, 2020‑2023). Two main patterns emerge: journals with larger author teams tend to have higher citation impact, while AI engagement rate is positively associated with Impact Factor only in the 2019 subset. To validate the LLM pipeline for deriving the AI engagement rate, we conduct human evaluation, confirming substantial agreement in AI relevance detection and consistent subfield classification. Together, BioMedJImpact provides both a comprehensive dataset at the interface of biomedicine and AI and a validated framework for scalable, content‑aware scientometric analysis. Code and dataset are available at https://github.com/JonathanWry/BioMedJImpact.

Authors:Rui Yang, Michael Fu, Kla Tantithamthavorn, Chetan Arora, Joey Chua
Title: Towards a Risk Assessment of Malicious Skill Files in Coding Agents
Abstract:
Autonomous coding agents are increasingly embedded in enterprise software workflows with delegated authority over connected systems. Central to this architecture is the agent skills interface: folders of instructions and scripts that agents load dynamically to specialize their behavior. This interface also widens the attack surface, letting malicious shell commands hide within natural‑language skill files. We make three contributions. First, an adversarial skill‑synthesis method using six LLMs across four families to transform 471 real‑world shell commands into benign‑appearing skills, released as a benchmark of 2,826 skills mapped to 11 MITRE ATT&CK tactics. Second, a reproducible evaluation pipeline coupling run stratification, evidence anchoring, a refusal veto, and a deterministic declared‑intent override with a three‑judge LLM‑as‑a‑judge panel, validated against a blind human gold standard (Cohen's kappa = 0.85). Third, a large‑scale characterization of two enterprise‑grade agents across 5,629 completed runs. Gemini CLI is exploited in 95.5‑96.1% of runs and Qwen Code in 71.6‑74.0% (raw majority vote to declared‑intent‑corrected estimate, both within the human gold standard), nearly invariant to the generating model. Explicit safety recognition occurs in only 1.99% of runs. Enterprises must assess and mitigate skill‑interface risk before adopting coding agents. Our code and dataset are available at https://github.com/awsm‑research/AgentJailbreak

Authors:Junzhuo Liu, Weiwei Li, Jun Ling, Peng Wang
Title: When Privileged Guidance Misaligns: State-Matched Routing and Contextualized Self-Distillation for Multi-Turn Agents
Abstract:
Privileged on‑policy distillation provides dense supervision for multi‑turn agents by allowing a synchronized teacher to re‑score the student's response at every turn with access to training‑only references, such as successful trajectories. In interactive environments, however, the student's preceding actions continually change the execution state. As the student takes different actions or completes subgoals in a different order, its rollout may reach states not covered by the reference, making the reference an unreliable source of guidance for the state actually reached. Applying privileged distillation indiscriminately therefore creates state‑‑reference mismatch. This mismatch motivates a central objective: providing privileged reference guidance that remains compatible with the student's current execution state. We introduce State‑Matched Routing and Contextualized Self‑Distillation (SMRC‑SD), which explicitly determines when and how a privileged trajectory should guide an on‑policy student. At each turn, SMRC‑SD verifies whether the student's current execution state matches a supported state along the reference trajectory. Distillation is applied only at matched states, filtering out turns for which the reference lacks locally compatible guidance. For each matched state, SMRC‑SD further constructs state‑conditioned teacher context from the successful trajectory, grounding supervision in the state actually reached. Across ALFWorld and WebShop, SMRC‑SD consistently outperforms unconditional successful full‑path distillation. With Qwen3‑1.7B, it improves task success from 0.746 to 0.865 on ALFWorld and from 0.574 to 0.693 on WebShop. Controlled routing and context ablations support both selecting locally supported turns and constructing state‑compatible teacher context as contributors to these gains. Code is available at https://github.com/liujunzhuo/SMRC‑SD.

Authors:Jihoon Oh, Kento Kawaharazuka, Kei Okada
Title: VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances
Abstract:
Learning manipulation skills from human videos is promising for scalable robot learning. However, the embodiment mismatch between humans and robots makes this challenging. One promising solution is to learn object‑centric actionable affordances that are embodiment‑agnostic. In this work, we propose a framework that leverages egocentric human videos with state‑of‑the‑art 3D Structure‑from‑Motion and hand mesh reconstruction to extract actionable affordances such as visual, grasp, and trajectory affordances that explicitly encode where to interact, how to grasp, and how to move. We construct EgoAffordance, a large‑scale dataset comprising 204K episodes with 5.6M visual affordances and 11.6M grasp and trajectory affordances. Building on this, we introduce VLAff, a large vision‑language model‑based unified foundation model that learns cross‑modal correlations across all actionable affordances. Given a visual observation and instruction, VLAff generates visual affordance heatmaps, grasp poses, and trajectories, which are then converted into directly executable actions by utilizing 3D scene information. Through extensive experiments, we demonstrate that VLAff not only achieves state‑of‑the‑art performance on visual affordance prediction, but can also be effectively applied to real robot applications such as zero‑shot manipulation and affordance‑guided robot learning.

Authors:Sanghyeok Lee, Jihye Kang, Namhyuk Ahn
Title: StyleComposer: Training-Free Multi-Reference Style Composition
Abstract:
The style of a painting is not monolithic: color, texture, and structure may come from different sources. Existing reference‑guided methods transfer them as one style signal, leaving each attribute's source and strength outside the user's control. We ask where in a diffusion model one attribute can change while the others hold, and find that no single representation isolates all three. The proposed StyleComposer therefore routes each style attribute through the representation where it separates best and coordinates the routes over denoising time. Without training or inversion, it satisfies three references and the prompt jointly more closely than prior methods, and exposes one strength slider per attribute. Project page: https://lexxsh.github.io/StyleComposer

Authors:Adam Simson, Ankush Dutta, Quang Bui
Title: MS-MLB: An Open Machine Learning Benchmark for Blood-Based MS Classification
Abstract:
Multiple sclerosis (MS) is diagnosed through clinical assessment, magnetic resonance imaging, laboratory evidence when appropriate, and exclusion of better explanations. Blood RNA expression data may contain disease associated immune signal, but a blood RNA classifier cannot be treated as a replacement for clinical diagnosis. This paper presents MS‑MLB (Multiple Sclerosis Machine Learning Benchmark), a reproducible open benchmark for machine learning based MS research classification from whole blood RNA expression data. MS‑MLB uses the public GSE17048 cohort, converts it into an MS versus healthy control task, and evaluates multiple algorithms under a shared, leakage controlled pipeline that a researcher can rerun without reconfiguring the evaluation. The evaluation includes nested cross‑validation, an untouched stratified holdout set, bootstrap confidence intervals, ROC and precision recall analysis, calibration measurement, and an exploratory MS Research Score. In the final benchmark summary, Gradient Boosting ranked first by MS Research Score on the holdout set, with an MS Research Score of 93.83, AUC‑ROC of 0.989, sensitivity of 0.950, specificity of 0.778, F_1 score of 0.927, and Brier score of 0.050. Prior studies have applied machine learning to MS blood transcriptomic data, including PBMC stage classification and whole blood diagnostic signature modeling. The contribution here is different and narrower. To our knowledge, MS‑MLB is the first open benchmark focused on MS versus healthy control classification from GSE17048 whole blood RNA expression data with a documented external model submission pathway built into the framework. The score is intended for research comparison only and has not been clinically validated. The benchmark is accessible here: https://github.com/duckyquang/MS‑MLB.

Authors:Saman Rahbar
Title: The Ignition Index: Measuring Global Workspace Dynamics in Language Models
Abstract:
We introduce the Ignition Index (I), a validated scalar metric that operationalizes Global Workspace Theory's (GWT) all‑or‑none ignition prediction in transformer language models. The metric fits a four‑parameter sigmoid to per‑layer linear probe accuracy as a function of input signal strength, extracting steepness parameter beta‑hat: high values indicate abrupt, ignition‑like transitions; low values indicate graded build‑up. Across 11 models spanning five architecture families, shuffled‑label controls demonstrate 9.6‑fold selectivity for genuine linguistic structure over spurious probe capacity (p < 0.001, Mann‑Whitney U‑test). We find: (1) Feedforward transformers exceed SSMs by 89% in aggregate beta‑hat (p < 1e‑13, Cohen's d = 0.52), with Mamba exhibiting near‑linear profiles consistent with absent global broadcast. (2) Huginn‑3.5B exhibits 2.12‑fold higher ignition along its iteration axis than its depth axis, demonstrating that recurrent architectures manifest workspace‑like transitions along the recurrence dimension. (3) Pythia‑410M shows a PELT‑detected phase transition at training step 256 (+67%), preceding induction‑head formation. (4) Hypotheses linking ignition to model scale and signal strength were not confirmed, suggesting transformer architectures may saturate available ignition mechanisms. The Ignition Index provides the first validated quantitative bridge between GWT's dynamical predictions and mechanistic interpretability, with 9.6‑fold measurement selectivity and architecture‑level discriminability not previously characterized in the scaling literature. Code: https://github.com/saman‑rahbar/ignition‑index

Authors:Damien Sileo, Valentin Lacombe, Dimitri Kachler
Title: Reasoning Core: Designing Broad Procedural Data for Completion-Supervised Reasoning Training
Abstract:
Procedural generators produce useful verifiable reasoning problems at scale, but have received less attention as data for completion‑supervised fine‑tuning. We introduce Reasoning Core, a collection of 50 generators spanning mathematics, logic, planning, state tracking, formal languages, structured data, games, causality, and code, with semantic scorers, difficulty controls, and task evaluators. Under a matched completion‑supervised protocol, we compare Reasoning Core with Procedural Warmup, Reasoning Gym, and SynLogic across four base‑model settings and multiple training durations. In the primary 3B comparison, Reasoning Core achieves the highest mean scores on DROP, LogiQA, and ARC‑Challenge, exceeding both the baseline without procedural data and all three alternative procedural collections. Task‑level analyses show that semantic validity alone does not ensure training utility, highlighting compact targets and calibrated difficulty as important design factors. We ran audits combining model‑assisted review, human adjudication, and regression testing. Applied throughout Reasoning Core development and to the other collections, they reveal subtle mismatches among generation, rendering, targets, and scoring, a reminder that procedural generation alone does not guarantee correctness. The library, generated datasets, and audit material are publicly available.

Authors:Zisen Shao, Zihao Wei, Derong Jin, Ruohan Gao
Title: Objects as Audio-Visual Modal Sound Fields
Abstract:
While modern 3D reconstruction excels at modeling object geometry and appearance, it largely ignores the rich acoustic cues revealed through physical interaction. Object impact sounds convey material, stiffness, and structural properties that complement vision, yet existing impact sound modeling approaches either rely on expensive physics‑based simulation or require large datasets to generalize in a purely data‑driven manner. We introduce Audio‑Visual Modal Sound Field (AV‑MSF), a novel object‑level acoustic representation reconstructed from multi‑view images and only a few impact sound recordings. AV‑MSF builds on 3D Gaussian Splatting integrated with dense 3D visual feature to provide a strong geometry‑aware prior, and represents the impact sound field using compact, physically meaningful modal parameters, enabling robust few‑shot reconstruction. Experiments on two real‑world datasets show that AV‑MSF achieves state‑of‑the‑art impact sound rendering, outperforming both physics‑based and data‑driven baselines. Furthermore, we demonstrate downstream applications enabled by our representation, including contact localization and object sound editing.

Authors:Yinghui He, Ling Yang, Jiarui Liu, Yongjin Yang, Lechen Zhang, Yingcheng Wu, Zhenfei Yin, Mengdi Wang, Sanjeev Arora
Title: Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning
Abstract:
Long‑horizon reasoning in recent LLMs demands that the model switch between distinct skills inside a reasoning chain, such as first doing a math derivation, then using the result to plan a schedule. We call such problems cross‑skill long‑horizon tasks: multi‑step tasks whose steps require different reasoning skills and depend on earlier outputs. Existing benchmarks often evaluate individual skills, lacking a principled way to measure how well a model switches between skills. We address this gap from both the evaluation and training sides. We introduce Skill Entropy, a measure of the difficulty of switching from one skill to another. We then propose Skill^2‑Bench, a benchmark of cross‑skill long‑horizon tasks built over 558 skills across 9 verifiable and open‑ended domains. Each task is assigned a task‑level skill‑entropy score and grouped into three difficulty levels. Evaluating 8 frontier and 4 open‑source models on Skill^2‑Bench reveals a skill‑switching gap: accuracy decreases on higher‑entropy tasks. We then turn skill entropy from a benchmark scale into a training signal. We propose Skill‑Entropy RL, an RL framework where the model predicts not only the answer at each step but also the skill used to produce it. The reward combines step‑level correctness with a skill‑entropy reward that measures the alignment between the model‑predicted skill sequence and the gold skill sequence. On Qwen3‑4B‑Instruct and Qwen3‑1.7B, Skill‑Entropy RL improves the Skill^2‑Bench score from 34.4% to 68.4% and from 14.6% to 40.1%, respectively, outperforming competitive baselines. The same pipeline can be applied to off‑the‑shelf training data such as OpenR1‑Math, indicating that skill entropy is a reusable training signal. Code available at: https://github.com/Gen‑Verse/Skill‑Entropy‑RL

Authors:Yue Zhang, Yingzhao Jian, Yunqiu Xu, Xiaoxiao Sun, Hehe Fan
Title: SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding
Abstract:
Understanding 3D scenes is fundamental to embodied intelligence, requiring joint reasoning over heterogeneous information from multiple modalities, including visual and geometric cues. However, the relevance of these modalities often varies across queries. Existing Multimodal Large Language Models (MLLMs) typically rely on fixed modality combinations, overlooking query‑dependent modality needs. Such a rigid design can introduce semantic noise from irrelevant modalities while underutilizing more informative ones, leading to wasted computation and diluted reasoning. To address these challenges, this paper proposes SmartMage, a unified MLLM that dynamically orchestrates heterogeneous modalities for semantic‑aware 3D scene understanding. Specifically, SmartMage incorporates: (1) a Semantic‑guided Modality Adaptive RouTng (SMART) module that selects task‑relevant modalities using semantic priors, text‑modality alignment, and modality quality; and (2) a Modality‑Aware Gating Expert (MAGE) module that leverages modality priors to guide expert activation, fostering adaptive specialization in multimodal reasoning. Empirically, SmartMage achieves state‑of‑the‑art performance across five 3D scene understanding benchmarks, and attains competitive results on RGB‑only video understanding benchmarks. In our diagnostic benchmark ScanFacet, tasks are divided into fine‑grained semantic categories, enabling analysis of modality combinations preferred by each semantic type. The observed modality‑semantic patterns provide further evidence of SmartMage's effectiveness. Project page: https://yuecheong.github.io/SmartMage/.

Authors:Devender Singh
Title: The Loss Does Not See the Basis, but Adam Does
Abstract:
Gradient descent on a factored model W = UV^\top is implicitly biased toward low‑rank solutions, while Adam, starting from the same small initialization, is not. We trace the difference to the gauge symmetry of the loss, its invariance under (U, V) \mapsto (UQ, VQ). Gradient flow's low‑rank mechanism is available to an optimizer only if that optimizer is gauge‑equivariant, a condition necessary for the transfer but not sufficient for low‑rank recovery. Gradient descent, momentum, "shared‑scalar" Adam, Muon, and Shampoo satisfy it. Adam, RMSProp, and the other coordinate‑wise methods do not. A structure theorem characterizes the memoryless equivariant rules as exactly the Gram‑determined left preconditioners, and a transfer theorem carries gradient flow's pathwise properties to common‑scalar flows. We then sort nine update rules on underdetermined matrix sensing by recovery error against the planted ground truth. A one‑parameter family from coordinate‑wise to shared‑scalar preconditioning restores the bias monotonically, isolating anisotropy as the cause. A "spectral schedule" reconciles two opposing reports about Muon: equal‑rate updates recover exactly low‑rank targets but lose their edge as the spectral tail grows. In transformers, Adam separates two gauge‑equivalent initializations at the first step, where the equivariant optimizers stay at float precision, and ends with the per‑head invariants W_Q^\top W_K 56% apart in relative Frobenius distance, a gap no per‑head rotation can close. On two hyperspectral datasets at matched training loss, gradient descent cuts held‑out error by 43‑44% at the lowest sampling density, and at lower effective rank. Basis choice is therefore not a tuning detail but a decision about which interpolant the optimizer selects.

Authors:Aniri, Jinhe Bi, Peng Liao, Zengjie Jin, Volker Tresp, Fei Shen, Yunpu Ma, Tat-Seng Chua
Title: OPD-V: Visual On-Policy Self-Distillation with Modality Balance
Abstract:
On‑Policy Self‑Distillation (OPSD) has become a standard post‑training approach for improving visual reasoning in multimodal large language models (MLLMs). Existing methods draw privileged information from diverse input sources to guide self‑distillation. Yet these designs overlook Modality Imbalance, a challenge inherent to MLLM reasoning. When textual information dominates generation, the model cannot fully integrate its multimodal input. Consequently, carefully designed privileged information remains underused, limiting the effectiveness of OPSD. To examine this limitation, we construct a Positive Teacher with the Zoom‑In Image and a Negative Teacher with the Mask Image, which exhibit different degrees of Modality Imbalance. Changes in their reasoning correctness and token logits reveal that Modality Balance can itself serve as privileged information. Motivated by this finding, we introduce OPD‑V, a visual OPSD paradigm that instantiates such information through the Positive Teacher and Negative Teacher. Positive Modality‑Balance Logits Margins define a Modality‑Balance Trust Region that selects the on‑policy tokens used for self‑distillation. Experiments across 6 benchmarks, 4 MLLM backbones, and 5 post‑training methods show that OPD‑V consistently improves reasoning performance while reducing training cost.

Authors:Yanting Wang, Chenlong Yin, Runpeng Geng, Jinyuan Jia
Title: Agent Against Agent: An Agentic System for Automatic Prompt Injection Red Teaming
Abstract:
Prompt injection poses significant security risks to LLM agents. Efficient and effective red‑teaming is therefore critical, both for evaluating these risks and for collecting training data to improve defenses. Existing state‑of‑the‑art prompt injection red‑teaming methods primarily rely on reinforcement learning (RL), producing attacker models that often generalize poorly to new target LLMs. In this work, we develop PIMiner, an agentic system for prompt injection red‑teaming. During training, PIMiner is trained on a sequence of (dataset, target model) pairs and builds a strategy library from scratch. At test time, the learned strategy library can be directly transferred to a previously unseen target LLM without additional training. PIMiner requires only a small number of queries to a target agent (e.g., 10) per test sample. Experimental results demonstrate that PIMiner achieves strong performance. On IPIArena, it attains a 76.2% ASR against Gemini‑2.5‑Pro, 61.9% ASR against GPT‑5.1, and 42.9% ASR against Claude‑Sonnet‑4.5. On AgentDojo, it achieves an 86.7% ASR against Gemini‑2.5‑Pro, 53.3% ASR against GPT‑5.1, and 40.0% ASR against Claude‑Sonnet‑4.5.

Authors:Réemi Andrieu, Damien Sileo
Title: Same Formulas, Different Semantics: Do Language Models Follow Modal Logic Specifications?
Abstract:
Reasoning about necessity and possibility depends on assumptions about accessibility between worlds and about which objects exist at each one. The same inference may therefore hold under one modal system and fail under another. Evaluating language models on such problems requires testing whether their judgments follow the stated semantics rather than a familiar logic. We construct paired modal problems with identical premises and conjecture but different frame or domain conditions; automated reasoning verifies opposite labels. A balanced core prevents the semantic condition alone from revealing the answer. On this core, four of five recent models perform below the condition‑only baseline under direct prompting. Yet enabling reasoning mode raises DeepSeek V4 Flash from 4.4% to 88.1% on unchanged prompts. Following stipulated modal semantics thus depends strongly on inference mode as well as model identity. When frame conditions are omitted, models often agree but fit different familiar logics best. We release the formulas, oracle artifacts, countermodels, and responses.

Authors:Peer Saleth, Segun T. Aroyehun, Fabio Carrella, Christoph M. Abels, Stephan Lewandowsky, David Garcia
Title: German parties shifted towards intuition-based rhetoric after the far right's parliamentary breakthrough
Abstract:
The spread of misinformation is widely perceived as a threat to democratic deliberation, yet how political elites' rhetorical commitments to truth shift alongside the rise of populist actors remains poorly understood. Analysing 4.5 million tweets and 59,170 parliamentary speeches by German political elites between 2015 and 2025, we measure evidence‑based and intuition‑based rhetoric using a validated distributed dictionary representation. Across both arenas, intuition‑based language has become more prominent, and right‑leaning actors consistently exhibit the lowest Evidence Minus Intuition (EMI) scores. The parliamentary entry of the extreme‑right Alternative for Germany (AfD) in 2017 coincides with sharp downward shifts in EMI across the broader chamber, while a more gradual decline is observed on Twitter. These findings document an association between far‑right visibility and a changing approach to truth in elite discourse in a multiparty European democracy.

Authors:Liangyang Ouyang, Ruicong Liu, Xuangeng Chu, Kaipeng Zhang, Yoichi Sato
Title: HelloWorld: Enabling Socially Interactive Characters in Video World Models
Abstract:
Despite the remarkable recent progress of video world models, social interaction between users and the characters within these worlds remains unsupported. To fill this gap, we present HelloWorld, a video world model that enables social interaction with in‑world characters. With a single button press, users can prompt the on‑screen character to respond toward the camera, e.g., turning to the viewer, waving, nodding, or speaking a short greeting. To make these interactions natural, we propose a self‑distillation pipeline that finetunes the video generation model on data synthesized by itself. Each synthesized clip contains both social interactions and camera motion, allowing the model to learn camera‑pose conditioning without degrading interaction quality. At inference, we further introduce a training‑free module that determines when the interaction occurs. Upon a button press, it modulates the cross‑attention masks of the DiT so that the interaction‑related text prompt attends only to the frames within the press window, temporally localizing the character's response. We further build HelloWorldBench, a 400‑sample benchmark with three social interaction metrics alongside three conventional metrics, for evaluation. Experiments demonstrate that HelloWorld surpasses a variety of baselines in interaction quality, while maintaining state‑of‑the‑art picture aesthetics and camera‑pose following. Project page: https://github.com/AlayaLab/HelloWorld

Authors:Narges Rashvand, Ghazal Alinezhad Noghre, Shanle Yao, Gabriel Maldonado, Hamed Tabkhi
Title: VQ-VAD: Vector-quantized Motion Representation Learning for Human-centric Video Anomaly Detection
Abstract:
Video Anomaly Detection (VAD) is inherently challenging due to the scarcity of anomalies and the large visual variability in surveillance footage, including changes in lighting, viewpoint, and human appearance. To mitigate visual noise and address privacy concerns, recent work has shifted to pose‑based VAD, which focuses on motion dynamics rather than raw video data. However, existing pose‑based approaches model human behavior in continuous latent spaces, limiting their ability to learn compact motion patterns necessary for robust behavior analysis. We address this by proposing Vector‑Quantized Video Anomaly Detection (VQ‑VAD), a novel human‑centric anomaly detection framework that learns discrete motion representations. VQ‑VAD adapts Vector‑Quantized GAN (VQ‑GAN), originally developed for image generation, to operate on keypoint sequences and construct a motion codebook of normal behavior. Trained exclusively on normal motion sequences, VQ‑VAD detects anomalies by identifying high reconstruction errors when an observed motion sequence cannot be mapped to the learned codebook. We conduct extensive experiments across three complementary evaluation settings, including in‑domain, cross‑domain, and cross‑dataset generalization, on four anomaly detection benchmarks. VQ‑VAD achieves strong in‑domain accuracy (81.83% on HR‑SHT [15]), effective cross‑domain transfer from CMU Panoptic [14] (76.69% on HR‑SHT [15] without retraining), and competitive cross‑dataset robustness. The code base for this work is available at https://github.com/TeCSAR‑UNCC/VQ‑VAD.

Authors:Yuxuan Huang, Xingyu Zeng, Tianhang Zheng, Chaochao Lu
Title: Gradient Immunity: Null-Space Resistance to Malicious Fine-Tuning
Abstract:
Released aligned large language models remain vulnerable to malicious downstream finetuning. Existing defenses are largely designed for the fine‑tuning‑as‑a‑service (FTaaS) paradigm or rely on downstream users to follow additional safety procedures, and therefore do not directly address the setting we study: a provider controlled partially protected open‑weight (PPOW) release setting in which most weights remain trainable while a small safety‑critical component is preserved at release. We propose a Unidirectional Safety Gate (USG), instantiated as a Null Space Cubic Layer together with an Inverse Adapter inserted after the final Transformer layer. During downstream fine‑tuning, the cubic layer suppresses or blocks gradients from harmful samples whose hidden states fall in a calibrated protected region, while the Inverse Adapter restores the base model's forward behavior. In practice, we calibrate a threshold using defender‑held harmful data, allowing protection to generalize to nearby in‑distribution harmful samples. Across six evaluated model‑dataset settings, USG keeps post‑finetuning attack success rate close to the pre‑release level under a fixed release threshold, while maintaining high safe‑pass rates on easier settings and exhibiting a clearer safety‑utility trade‑off on unsafe samples from BeaverTails. These results suggest that release‑time representation‑space blocking can raise the cost of malicious downstream adaptation without requiring downstream cooperation. The code is available at https://github.com/OpenCausaLab/Gradient‑Immunity.

Authors:Peiyan Li, Yuze Zhu, Yixiang Chen, Qisen Ma, Yuan Xu, Jiabing Yang, He Guan, Yan Huang, Hongtao Wu, Xiao Ma, Tao Kong, Liang Wang, Tieniu Tan
Title: BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation
Abstract:
Leveraging pre‑trained vision‑language models (VLMs) to construct vision‑language‑action (VLA) models has emerged as a promising paradigm for 3D robot manipulation. However, existing 3D VLA methods remain data‑hungry, exhibit limited generalization under distribution shifts, and lack explicit memory of past observations. These limitations hinder their application to data‑scarce, open‑world, and memory‑dependent manipulation scenarios. Our previous work, BridgeVLA, improves data efficiency and generalization by preserving the input‑‑output alignment of a pre‑trained VLM during 3D action learning: raw point clouds are projected into multi‑view images, and intermediate heatmaps are predicted before generating robot actions. In this work, we develop BridgeVLA++ by equipping BridgeVLA with a unified spatio‑temporal memory architecture that models persistent spatial context and temporal interaction history. The resulting memory‑augmented framework can reason over observation histories while preserving BridgeVLA's data efficiency and generalization capabilities. Extensive experiments show that our framework achieves strong performance on spatial manipulation tasks while exhibiting robust generalization. BridgeVLA++ further achieves state‑of‑the‑art performance on two challenging memory‑dependent manipulation benchmarks without sacrificing the data efficiency and generalization of the original BridgeVLA. In addition, BridgeVLA++ performs effectively in bimanual manipulation settings and is validated on an additional real‑world robotic platform, demonstrating its scalability across tasks, environments, and robotic platforms. These results establish BridgeVLA++ as a unified 3D vision‑language‑action framework that simultaneously supports data‑efficient learning, robust generalization, and effective memory‑aware robot manipulation. Project website: https://bridgevla‑plus.github.io/.

Authors:Junlin Han, Shengbang Tong, David Fan, Minghao Chen, Philip Torr, Filippos Kokkinos, Mike Lewis
Title: Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes
Abstract:
Vision offers a critical axis for advancing foundation models, driving a shift towards natively unified multimodal pretraining. Despite this momentum, the design space and the fundamental mechanisms of how modalities interact during unified training remain underexplored. We provide empirical clarity through a systematic exploration of multimodal pretraining. Our controlled experiments on both synthetic and large‑scale real‑world datasets yield four key insights into the physics of multimodal pretraining: (i) Knowledge Flow: We disentangle how language, visual understanding, and visual generation transfer knowledge across modalities, revealing distinct patterns of influence and asymmetry; (ii) Synergy vs. Competition: We show that data "complexity" largely determines whether modalities are synergistic, identify architectural choices that promote synergy: such as shared attention and normalization with modality‑specific feed‑forward layers, and find that these behaviors generalize across different visual tokenizer designs; (iii) Early Unification: Unifying modalities from the very early stages and training them jointly is shown to be more effective than late alignment or sequential training. This process uncovers a vision laziness phenomenon, where delayed integration leads models to rely on language priors; (iv) Recipes: We derive efficient pretraining recipes that achieve strong generative performance using only 5% of the compute budget. These core findings are subsequently validated at scale by training multiple 13.5B MoE models on 2T tokens. We hope this study provides a principled foundation for understanding and scaling multimodal pretraining.

Authors:Shanglin Yuan, Weiheng Zhao, Xin Shi, Haoyi Jiang, Xianda Guo, Liu Liu, Wenyu Liu, Wei Sui, Xinggang Wang
Title: DreamWAM: Beyond RGB Future Prediction for World Action Models
Abstract:
World Action Models (WAMs) learn action‑relevant representations by predicting how the observed world will evolve. Most existing WAMs define this future in RGB space, where task‑relevant state transitions are entangled with nuisance variations in texture, illumination, background, and viewpoint. We argue that WAMs should explicitly predict action‑relevant future state rather than relying on RGB prediction alone. We introduce DreamWAM, which reformulates future prediction as structured world modeling beyond RGB, representing future states through complementary views of appearance, motion, geometry, and semantics. During training, DreamWAM combines joint latent denoising of RGB and motion with lightweight gated residual branches for geometry and semantics. Shared attention between VideoDiT and ActionDiT allows the action branch to learn from these future‑state predictions, while all beyond‑RGB supervision branches are disabled at inference and deployment remains RGB‑only. Across both no‑rollout and joint video‑action inference, DreamWAM consistently improves the matched RGB‑only baselines on LIBERO, from 97.30% to 98.40% and from 98.00% to 98.90%, respectively. The gains become larger under unseen LIBERO‑Plus perturbations, from 51.36% to 63.44% and from 69.16% to 75.47%. The same robustness extends to real‑world manipulation, where DreamWAM attains an average success rate of 74.4% across unseen changes in lighting, background, and object layout, compared with 55.6% for Fast‑WAM‑Joint. These results show that robust world‑action learning depends not only on predicting the future, but on representing it in a form that matters for action. The code and models are publicly released at https://github.com/hustvl/DreamWAM.

Authors:Sihan Hu, Lyuhan Huang, Youjin Deng, Kun Chen
Title: SciCode-Verified: How Benchmark Defects Underestimated the Scientific-Coding Ability of Language Models
Abstract:
SciCode is the standard measure of the scientific‑coding ability of language models: research‑level problems that demand both frontier scientific theory and its implementation as working numerical code. It is a component of the Artificial Analysis Intelligence Index and a standing evaluation in government and national‑laboratory suites. Yet its scores have recently plateaued: the strongest 2026 models cluster tightly around 60% subproblem accuracy, and a successor model ties its predecessor. We trace this stagnation to defects in the benchmark itself. A per‑problem, domain‑expert audit of all 65 test problems uncovers 263 defects; 192 of them, spread across 91% of the main problems, cause correct, instruction‑following solutions to be wrongly rejected‑‑‑through non‑reproducible gold answers, over‑tight tolerances, or self‑contradictory specifications. Critically, 78% of these score‑suppressing defects require specialized physics or mathematics knowledge to detect, not mere clerical proofreading. We corrected every confirmable defect to produce SciCode‑Verified. The corrections add only the specifications a well‑posed problem requires, repair grading, and tighten the tests that were too lenient; every change is recorded with its justification and independently re‑checked by a second domain expert. We re‑evaluate twelve frontier model snapshots on the corrected benchmark and find a substantial recovery: subproblem accuracy rises from 45‑‑60% to 84‑‑98%, and main‑problem accuracy from 9‑‑27% to 69‑‑92%. State‑of‑the‑art models are far more proficient in scientific coding than SciCode has suggested‑‑‑the bottleneck was not model capability, but the quality of the evaluation instrument. We release SciCode‑Verified with its complete audit trail as the corrected public standard.

Authors:Bohai Gu, Yueyang Yuan, Taiyi Wu, Dazhao Du, Jian Liu, Xiaoyi Pang, Jie Zhang, Xiaocheng Lu, Haobin Zhong, Xiaotong Zhao, Alan Zhao, Song Guo
Title: WorldCycle: Self-Verifiable Reinforcement Learning for Long-Horizon Video World Models
Abstract:
Interactive video world models are essential for long‑horizon planning and exploration, yet they suffer from compounding errors. Post‑training methods such as reinforcement learning (RL) can improve these models, but they hit a verification bottleneck: for arbitrary action sequences, no ground‑truth future state exists to measure long‑term drift. Our key insight is that reversible action cycles make this verification possible: a sequence composed with its inverse must analytically return to the initial state, yielding annotation‑free supervision on long‑horizon correctness. Building on this, we introduce WorldCycle, a self‑verifiable RL framework that constructs closed action cycles and their repeated executions from ordinary action sequences, and optimizes two complementary rewards: a spatial closure reward enforcing symmetry between mirrored forward and reverse segments, and a temporal consistency reward aligning states across repeated cycle executions. These rewards force the model to learn actions as consistent state operators rather than memorized temporal patterns, and extend naturally to out‑of‑distribution composite action cycles that the base model handles poorly. We further release CycleBench, a diagnostic benchmark for state‑returning ability under complex action structures. WorldCycle reduces state returning drift by up to 44% and lifts composite‑action accuracy nearly 4x over the base model, providing a vital foundation for physically grounded world models.

Authors:Xu Guo, Zhengxuan Wei, Xinghui Li, Hanzhuo Huang, Xinyu Liu, Xiangyang Luo, Min Wei, Yiran Zhu, Qiulin Wang, Yulong Xu, Xintao Wang, Pengfei Wan, Qi Fan, Xiangwang Hou
Title: ContextMaster: Interactive Multi-Shot Video Creation via Fixed-Budget Sparse Context Routing
Abstract:
Recent video models increasingly support generation, reference conditioning, and editing within a single model, yet typically expose them as separate operations over fixed inputs. Practical creation unfolds across multiple shots, requiring one model to generate from text, follow a reference, or edit source footage while maintaining shared history. We formalize this setting as interactive multi‑shot video creation (IMVC) and introduce ContextMaster, a unified model with a role‑aware context representation for these operations. An interactive model must retain access to an expanding history without allowing the context read cost at each denoising step to grow. ContextMaster combines reusable clean context states with fixed budget sparse context routing and uses ConstraintSink to keep task constraints visible. To address the dual challenges of sparse context access and inference with few denoising steps, we propose a two‑stage privileged context distillation framework, which transfers full context behavior from a dense teacher through consistency distillation and then refines deployment rollouts with distribution matching. Experiments on the three primitive tasks demonstrate improved task fulfillment and consistency across shots over specialized baselines. User studies further validate flexibly composed workflows, while the model reaches 16 FPS on a single GPU.

Authors:Xuehang Guo, Pengyuan Li, Tom Hope, Tirthankar Ghosal, Manling Li, Qingyun Wang
Title: Consistency-Driven Co-Evolution for Self-Supervised Cross-Representation Learning
Abstract:
As chart images, tabular data, and visualization code play increasingly important roles across diverse domains, cross‑representation understanding across these modalities poses fundamental challenges for AI systems: the relationships across representations are inherently one‑to‑many, supervision is ambiguous and costly, and model optimization lacks a principled signal that is both direction‑adaptive and representation‑generalizable beyond task‑specific objectives. We introduce CoCoEvolve to improve consistency across chart, table, and code representations. Instead of treating cross‑representation mapping as a one‑to‑many problem, we define explicit one‑to‑one correspondences and optimize models using agreement between representations, without additional annotations. During training, CoCoEvolve@Train performs co‑evolution across the chart‑table‑code cycle, while CoCoEvolve@Test applies the same consistency objective at inference time for test‑time co‑optimization. We also present CoCoEvolve@Eval, an evaluation suite covering all six cross‑representation tasks. Across four benchmarks, CoCoEvolve improves performance in both training‑time and test‑time settings. Our project page: https://xhguo7.github.io/CoCoEvolve/.

Authors:Chengyang He, Tanishq Duhan, Gadiel Sznaier Camps, Fangyuan Wang, Yuhong Cao, Jiankai Sun, Ge Sun, Mac Schwager, Guillaume Sartoretti
Title: PRIMAL3: Pathfinding via Reinforcement and Imitation Multi-Agent Learning - Leveraging LaCAM3
Abstract:
We present PRIMAL3, an ultra‑large‑scale learning‑based framework for multi‑agent pathfinding (MAPF) that integrates reinforcement learning, topology‑aware communication, LaCAM3‑guided training, and PIBT‑based action refinement. PRIMAL3 targets failures at topologically critical states, where agents must coordinate decisively around bottlenecks, dead ends, and persistent conflicts. Each agent is represented using features derived from cut vertices, dead‑end regions, shortest‑path distances, and blocking estimates. Two complementary graphs capture agent interactions: a same‑direction following graph propagates multihop context along compatible paths, while a different‑direction conflict graph differentiates agents competing for shared space through masked attention and relative features. During training, we propose to let policy entropy identify uncertain agents, for which LaCAM3 provides confidence‑triggered action interventions and label‑smoothed imitation targets. During execution, a priority‑aware PIBT module refines the proposed joint actions using persistent, learned, and distance‑aware priorities together with policy‑aware fallback preferences while maintaining collision‑free execution. The resulting framework combines learned exploration with structured expert guidance without requiring LaCAM3 at inference. Experiments demonstrate that PRIMAL3 substantially outperforms state‑of‑the‑art learning‑based baselines and scales to ultra‑large instances with up to city‑level 100,000 agents. Real‑world experiments further demonstrate the feasibility of deploying PRIMAL3 on physical robotic systems and ablation studies validate the individual contributions the components we proposed. Project page: https://marmotlab.github.io/PRIMAL3/

Authors:Fanfu Xue, En Yu, Bohang Liu, Hongjun Wang, Yang Yang, Xindi Wang, Jiande Sun
Title: Deliberate Before You Fly: Vision-Guided Spatial Deliberation for UAV See-and-Reach Navigation
Abstract:
UAV see‑and‑reach navigation requires an aerial agent to approach a language‑specified target visible in its initial view and stop reliably near it. Existing methods typically map vision‑language representations directly to action outputs without explicitly modeling intermediate fine‑grained spatial decisions. This direct mapping causes semantic‑control misalignment, leading to inconsistent maneuvers and unreliable termination. To address this issue, we propose DBFly, a vision‑language waypoint prediction framework that introduces explicit vision‑guided spatial deliberation before waypoint generation. Specifically, DBFly introduces a spatial maneuver decision chain that progressively performs target‑direction anchoring, spatial diagnosis, and maneuver decision, enabling high‑level maneuver intent to explicitly guide continuous waypoint generation. DBFly further constructs an implicit flight corridor by transforming the initial target‑direction prior into a persistent geometric reference and deriving an online corridor state from the UAV's current position, thereby providing soft geometric guidance for spatial diagnosis and maneuver correction. In addition, DBFly develops a terminal‑convergence‑aware stopping strategy that characterizes terminal states through both target proximity and short‑horizon motion convergence, enabling more reliable stopping near the target. Extensive experiments across seen, unseen‑object, and unseen‑scene test sets demonstrate that DBFly improves the success rate over the SOTA baseline by an average of 25.07 percentage points. The project homepage is available at https://xuefanfu.github.io/DBFly‑Page.

Authors:Haotian Yang, Zhile Yang, Kin-Man Lam, Patrick Le Callet, Xin Sun
Title: Global Attention-Fused Image Cropping with Attention-Guided and Global-Aligned Crop Evaluator
Abstract:
Image cropping aims to improve image aesthetics by preserving important content within an appropriately composed region. However, most existing methods focus primarily on salient regions and therefore have limited sensitivity to the global relationships among the main image components. To address this limitation, we propose Global Attention‑Fused Image Cropping (GAFIC), which consists of an Attention‑Guided Feature Fusion (AGFF) and a Global‑Aligned Crop Evaluator (GACE). AGFF aggregates the importance of local regions to construct a global representation that captures both image structure and local details. GACE aligns candidate crop features with this global representation, enabling crop evaluation to remain sensitive to boundary changes. We further combine three ranking losses across multiple scales to obtain accurate and stable crop scores. Extensive experiments on the GAIC and CPC datasets demonstrate that GAFIC outperforms existing image‑cropping methods, particularly in terms of accuracy and stability. Unlike pixel‑level retargeting methods such as seam carving, inpainting, and diffusion‑based synthesis, GAFIC does not synthesize or modify the retained pixels; instead, it selects an aesthetically preferred crop from the source image, making it suitable for scenarios where pixel integrity and efficient batch processing are important. The source code is available at https://github.com/AIVRC/GAFIC.git.

Authors:Yi Yang, Cong Qin, Xiaodan Liu, Chishui Chen, Qing Dong, Yan Zhang, Cao Liu, Zhao Yang, Lu Pan, Jiaye Lin, Yi Feng
Title: Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation
Abstract:
Large language model agents are commonly trained through reinforcement learning with sparse trajectory‑level rewards, which offer limited guidance on how strongly individual tokens should be updated. On‑Policy Self‑Distillation (OPSD) addresses this by re‑scoring generated tokens under a privileged replay view to obtain dense, token‑level supervision. However, we identify a confounding issue: the resulting support may reflect both the privileged information contained in the replay view and score shifts induced by the replay scaffold, making it difficult to attribute the support specifically to that information. This issue is especially pronounced when future environment observations serve as privileged information, since replaying them requires reconstructing an extended scaffold that itself perturbs token scores. To resolve this confounding, we propose Observation‑Calibrated Self‑Distillation (OCSD), which contrasts two structurally matched replay views, Full and Observation‑Ablated, differing only in whether the actual future observation is present, to derive an observation residual that discounts score changes shared by the replay scaffold. OCSD then applies this residual to modulate token‑level GRPO updates at high‑uncertainty steps, while preserving the trajectory‑level update direction. Experiments on ALFWorld, WebShop, and Search‑QA across three Qwen3 model scales show that OCSD consistently outperforms strong baselines. Diagnostic analyses further confirm that the calibrated residual aligns better with local environment feedback. Our code is publicly available at https://github.com/yiy1x/OCSD.

Authors:Yuexi Yang, Alyssa Wu, Ji Luo, Richeng Xuan, Zhichao Hu, Yuhong Liu, Zhen Qin
Title: RepoProbe: Benchmarking Architecture-Aware Repository Comprehension with Checklists
Abstract:
The integration of Large Language Models (LLMs) into software engineering has shifted the focus from function‑level generation to repository‑scale assistance. However, existing benchmarks largely rely on bug reports from GitHub Issues, which often allow models to bypass genuine understanding via pattern matching on error logs. This misalignment under‑measures Edit Bias, which refers to premature generation, where models prematurely propose code modifications instead of understanding the existing repository architecture. Furthermore, current LLM‑as‑a‑Judge scalar scoring suffers from high variance and low interpretability. This work introduces RepoProbe, a novel benchmark for evaluating repository‑level code understanding through open‑ended Q&A using GitHub Discussions, which focuses on open‑ended architectural inquiries rather than defect reporting. To ensure rigorous evaluation, we propose a Checklist‑Based Verification Protocol that decomposes answers into atomic, verifiable facts, thereby replacing subjective ratings with objective verification. Our evaluation of state‑of‑the‑art (SOTA) LLMs reveals a persistent gap between high clarity and evidencegrounded technical correctness. It also quantitatively confirms the prevalence of edit bias, in which models prioritize code generation instead of architectural analysis. Finally, we demonstrate that our verification protocol significantly improves evaluation reliability compared to traditional evaluations with scalar scoring.

Authors:Shijun Ding, Chen Qian, Weiwei Shang, Junlin Xiong
Title: From Transparent Labware Segmentation to Collision Avoidance: A Real-Time Edge-Aware Perception Pipeline
Abstract:
This paper presents an edge‑aware instance segmentation framework that enables real‑time robotic collision avoidance with transparent laboratory glassware using purely visual perception. Transparent vessels defy conventional segmentation due to refraction, specular reflection, and the absence of stable interior texture, yet their boundary contours remain comparatively reliable visual cues. Exploiting this observation, we augment a one‑stage real‑time instance segmentation backbone with a lightweight edge‑detection branch, edge‑guided attention fusion, and a parameter‑free SimAM module, and further construct LabGlass‑IS, a 3485‑image, 21‑category instance segmentation dataset of real laboratory glassware. The enhanced model achieves the highest Boundary F‑score of 97.80 among compared methods, outperforming the YOLO‑prompted FastSAM framework by 18.93 BF points. Furthermore, it maintains an inference speed of 7.1ms per frame and requires only 2.85% of the parameters of the closest accuracy competitor. Multi‑view triangulation of mask centroids further provides 3D positions for conservative bounding‑volume collision constraints. Real‑robot trials achieve a 93.3% collision avoidance success rate, indicating the feasibility of the proposed perception‑to‑action pipeline for robot collision avoidance among fragile transparent objects. Our code is available at https://github.com/havishamy/TransYOLO_3D. Our video is available at https://havishamy.github.io/paper‑videos/.

Authors:Agatha Duzan, Asa Cooper Stickland
Title: Chain-of-Thought Monitoring Can Be Unreliable in Implicit-Influence Settings
Abstract:
Chain‑of‑thought (CoT) monitoring is increasingly treated as an important safety layer for frontier reasoning models. Most monitorability evaluations study explicit‑influence settings: setups where the prompt directly incentivizes the model to hide something, e.g., by instructing it to perform a hidden side‑task. A complementary axis for CoT‑monitor evaluations is implicit‑influence settings, where the prompt contains no instruction to hide, but the model's behavior is still shaped by features of the task or context, e.g. an irrelevant detail about a candidate that biases a hiring rating. We introduce the first benchmark that directly compares CoT monitorability under the two regimes. We test how model behavior changes in the presence of a nudge to choose a particular option. The nudge is delivered either as a casual aside (implicit), or as a direct instruction to act on the nudge and to conceal having done so (explicit). The benchmark spans four task formats (binary choice, numeric rating, multiple‑choice QA, open‑ended coding) and seven frontier extended‑thinking models. Under explicit influence, a CoT monitor detects 60‑94% of behavior shifts: even models instructed to conceal it leak the instruction into their CoT. Under implicit influence, the same factors still shift behavior, but detection falls by 41‑46 percentage points in two of our four settings. Realistic system‑prompt additions (of the kind a developer might deploy to reduce off‑topic bias) lower implicit detection further, to as low as 5%, while preserving the behavioral influence itself. These results suggest that monitorability estimates obtained in explicit‑influence settings may over‑estimate monitorability, and that monitorability can be further decreased by well‑intentioned deployment choices. Our benchmark and code are available at https://github.com/agatha‑duzan/implicit‑vs‑explicit‑influence

Authors:Mahshad Rastegarmoghaddam, Davoud Nikkhouy, Shima Samadzadeh
Title: Toward Integrating Adaptive Experience Replay and Online Uncertainty Estimation in Safe Actor-Critic Optimal Control
Abstract:
Safe actor‑critic control often treats barrier filtering, uncertainty estimation, and experience replay as separate modules, even though each changes the data used for learning and control. We develop an integrated architecture in which the uncertainty estimate updates the obstacle geometry used by a control barrier function, filter interventions and estimation residuals determine replay priority, and the critic learns from the executed rather than nominal action. We instantiate the architecture on a two‑dimensional robot‑navigation task with corrupted obstacle measurements and compare six component‑matched configurations under common training budgets, random seeds, sensor streams, exploration, and disturbances. Evaluation includes a moderate post‑training test, an eleven‑level perception‑noise sweep, and an exploratory extreme‑stress test at multiplier 6.0. In the extreme test, the integrated configuration recorded no contacts and reached the goal in all five evaluation seeds. Its mean cost was 7.63\pm0.44 and its obstacle‑belief root‑mean‑square error was 3.52\pm0.55 cm. The uncertainty‑estimation ablation also recorded no contacts but reached the goal in four of five seeds, with mean cost 8.96\pm2.08 and belief error 11.08\pm1.23 cm. A finite‑training bound clarifies replay exposure, and a robust barrier condition states the required estimation‑error and feasibility assumptions. The results support coupling estimation, safety filtering, and replay on this benchmark; broader safety and convergence claims require further study.

Authors:Alireza Javanmardi, Vippin Kumar Jeetmal, Christen Millerdurai, Alain Pagani, Didier Stricker
Title: Multi-View Face and Gesture Animation with Dynamic Gaussians
Abstract:
Creating photorealistic 3D human avatars with realistic upper‑body motion remains challenging. Existing approaches either focus on the head and overlook hand gestures, or reconstruct the full body but fail to preserve fine‑grained facial fidelity and hand pose accuracy. As a result, current methods struggle to capture the subtle dynamics of facial expressions and hand gestures that are crucial for natural human communication. While methods based on full‑body parametric models enable avatar reconstruction from monocular or multi‑view inputs, they often lack accurate facial animation and detailed hand articulation. To address these limitations, we propose MVFGA, a novel multi‑view‑consistent pipeline for generating realistic upper‑body avatars. Our approach models the face and hands separately and fuses them with a parametric upper‑body mesh model, enabling the capture of fine‑grained facial expressions and hand poses for accurate upper‑body avatar reconstruction. We then splat 3D Gaussians onto the obtained mesh, enabling high‑quality rendering of dynamic avatars from novel viewpoints. Furthermore, we introduce MVFGA‑MoCap, a multi‑view upper‑body motion capture dataset featuring controlled facial expression sequences, diverse hand gestures, and free‑form communication. Experiments show that MVFGA generates visually realistic avatars with high‑fidelity facial expressions and hand motions, outperforming baselines for upper‑body avatar animation. Project page: https://dfki‑av.github.io/MVFGA/

Authors:Ajeet Kumar Yadav, Sankaran Balasubramaniam, Aritra Chatterjee, Vinod Aduru, Yogesh Simmhan, Pandarasamy Arjunan
Title: A 6G Integrated Sensing and Communication Framework for Railway Intrusion Detection and Collision Prediction
Abstract:
Integrated Sensing and Communication (ISAC) combines sensing and communication to efficiently utilize wireless resources and is emerging as a key paradigm for next‑generation wireless networks. By leveraging the wide bandwidth, high frequencies, and massive antenna arrays of 5G‑Advanced and 6G systems, ISAC enables physical‑layer sensing using Channel State Information (CSI). The 3rd Generation Partnership Project (3GPP) Release 19 identifies 32 potential ISAC use cases, with particular emphasis on detecting and tracking moving objects. In this work, we address the Sensing for Railway Intrusion Detection use case, where intruders, including wildlife, entering a railway track can pose serious collision risks. We generated 22,695 CSI matrices with corresponding ground truth using a 3D‑rendered railway environment and the Sionna radio simulator. We developed a machine learning model combining a three‑dimensional Convolutional Neural Network (3D CNN) and Bidirectional Long Short‑Term Memory (BiLSTM) network to detect intruders in the track danger zone and estimate their real‑time position relative to the train, velocity, and time to collision. On synthetic CSI data, the model achieves 99.57% intruder‑detection accuracy on a balanced test set and a combined Mean Absolute Error (MAE) of 0.4240 for position, velocity, and time‑to‑collision prediction. These results demonstrate the potential of CSI‑based ISAC sensing with machine learning for reliable railway intrusion detection. The complete codebase for CSI generation, preprocessing, and model development is publicly available at https://github.com/EdgeIntelligenceLab/6g‑isac‑railway‑intrusion‑detection.

Authors:Haiyang Zhou, Wangbo Yu, Chaoran Feng, Xunyu Zhou, Yonghong Tian, Li Yuan
Title: UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models
Abstract:
The abundance of casually captured monocular videos and images on social media provides a valuable source for immersive content creation, where generating novel views from such sparse observations can greatly enhance user experiences. However, producing photorealistic and geometrically consistent views with precise camera control remains challenging when input coverage is extremely limited. Reconstruction‑based approaches such as NeRF and 3D Gaussian Splatting (3DGS) deteriorate severely under sparse inputs and fail to explicitly handle occlusions. Generative methods ease data requirements but still struggle with large‑baseline view synthesis due to inaccurate or implicit geometric guidance. To overcome these limitations, we introduce UniWorld‑View, a unified framework for controllable large‑baseline novel view synthesis from monocular inputs. UniWorld‑View integrates explicit 3D guidance with generative diffusion modeling to enable precise camera control and geometrically consistent view generation. The geometric guidance is obtained through an occlusion‑aware point cloud rendering strategy that resolves visibility ambiguities and provides accurate priors for diffusion‑based synthesis. By coupling this rendering strategy with powerful video diffusion backbones, UniWorld‑View achieves high‑fidelity novel view generation even under extreme camera motions and wide‑baseline changes, and can further provide multi‑view videos for downstream dynamic 3DGS reconstruction. Experiments on the WorldScore benchmark and zero‑shot NVS benchmarks demonstrate the effectiveness of UniWorld‑View in controllability, geometric consistency, and visual fidelity.

Authors:Yueqiang Zhang, Liang Deng, Yi Zhang, Baoqiong Wang, Wenjun Chen, Shuixin Pan, Yulan Guo, Qifeng Yu
Title: Differential 6-DOF Pose Estimation with Provable First-Order Immunity to Camera Calibration Errors
Abstract:
Accurate six‑degree‑of‑freedom (6‑DOF) motion estimation is essential for robotic manipulation, autonomous systems, and structural displacement monitoring. Conventional 3D‑2D methods estimate absolute camera poses independently at each time and recover platform motion through camera‑to‑platform extrinsics, making them sensitive to extrinsic calibration errors, especially for micromotion. We present a differential pose estimation method that directly recovers platform motion from inter‑frame image displacements and known 3D control points. By differencing perspective projection equations, using a depth‑invariance approximation, and modeling motion on SE(3), the method avoids independent absolute‑pose estimation and supports both monocular and multi‑camera systems. We prove that translational extrinsic errors cancel exactly, while rotational errors induce a bounded perturbation determined by calibration error, motion magnitude, and observation geometry. We also derive generic observability conditions, a Cramer‑Rao lower bound, and a bias‑eliminated consistent estimator, and characterize the validity limits of the approximations. Extensive synthetic and real‑world experiments establish a new state of the art for 6‑DOF platform micromotion estimation, outperforming representative PnP and generalized‑PnP methods in accuracy, calibration robustness, and computational efficiency. With five control points and 0.5‑pixel image noise, the monocular solver obtains a combined pitch‑yaw rotation RMSE of 10.09 arcsec, a translation RMSE of 3.70 mm, and a runtime of 0.34 ms. The binocular solver achieves a rotation RMSE of 10.58 arcsec, a translation RMSE of 3.91 mm, and a runtime of 0.27 ms. Code will be released upon publication at https://github.com/zyoungszu/pami2026.

Authors:Mohammadsaeed Haghi, Mahdi Salmani, Nima Kelidari
Title: Differentiating Through Dual Prices: End-to-End Policy Learning Under Capacity Constraints
Abstract:
Many social services assign scarce resources, such as housing assistance or hospital interventions, to people who arrive one at a time: each arrival must receive a decision immediately, and the long‑run usage of every resource must stay within its capacity. We study how to learn such an assignment policy from logged observational data. The standard pipeline is decision‑blind: fit one outcome model per arm by regression, price each capacitated resource from the fitted models, and assign each arrival the arm whose predicted outcome minus price is largest. We instead train the outcome models end‑to‑end, differentiating an off‑policy estimate of the deployed policy's value through the dual prices themselves. We study two formulations: an exact nonconvex one, and a convex relaxation whose optimum always satisfies the capacity constraints in expectation and which is suboptimal by at most a term linear in the smoothing temperature and logarithmic in the number of arms. Every method is evaluated in a queueing simulation with resources replenished at their capacity rates. Across six datasets, the two end‑to‑end variants take the top slots on a deployment‑adjusted value index at every delay cost, including zero; when capacities are binding, decision‑blind baselines frequently violate them and incur much longer queueing delays. On the largest dataset, a hospital cohort of seventy thousand patients, end‑to‑end training also achieves significantly higher policy value, a margin that survives a capacity‑matched neural baseline. Flexible decision‑blind regression remains the stronger pure predictor where ground truth is measurable; end‑to‑end training is best suited to settings where resources are genuinely scarce and feasibility matters.

Authors:Zhe Shan, Ziming Yang, Lei Zhou, Wenwen Zhang, Cong Lin, Xia Xie
Title: CSGen: A Multi-Domain Curvilinear Structure Generation Model via Hierarchical Multimodal Diffusion
Abstract:
Curvilinear structure analysis is an important and fundamental task in multimedia. However, the controllable generation of images with precise curvilinear structure objects remains an open challenge. To address this, we propose CSGen, a hierarchical multimodal diffusion model that synthesizes high‑fidelity images precisely aligned with multiple control conditions. The CSGen is built upon three key innovations: 1) We construct a multi‑domain and multimodal dataset, including over 24K samples from 5 domains and 7 different types of annotations, to train the unified generation model. 2) We propose a novel hierarchical progressive control strategy that decouples topology clues from visual context by a phased signal injection, mitigating semantic drift while ensuring the topological integrity of sparse structures. 3) We design a sparsity‑aware loss re‑weighting mechanism to address the extreme sparsity of curvilinear structures, significantly enhancing the attention on thin and fragile structures during optimization. Extensive experiments demonstrate that CSGen generates images with superior structure accuracy and visual realism, significantly improving downstream segmentation performance while maintaining robustness across diverse prompts. Our results confirm CSGen as a scalable, data‑centric paradigm for the analysis of complex curvilinear structures in diverse multimedia applications. Code and dataset are available at https://github.com/ShanZard/CSGen.

Authors:Haotian Yang, Zhile Yang, Huiyu Zhou, Xin Sun
Title: DAC-Pose: Dual-Agent Collaborative Framework for Pose-Guided Human Generation
Abstract:
AI agents have emerged as a powerful new paradigm in generative image synthesis, enabling systems to perform complex semantic reasoning rather than passive pixel‑level mapping. In pose‑guided human generation, conventional methods inevitably produce severe visual artifacts under drastic viewpoint shifts, fundamentally because they lack the cognitive capacity to logically deduce unseen regions and model complex spatial deformations. To bridge this gap, we propose DAC‑Pose, a novel agent‑driven multimodal framework that reformulates single‑view human generation as a collaborative dual‑agent system. DAC‑Pose integrates two complementary components, namely, the Prior Semantic Reasoning (PSR) agent and the Discrepancy‑Aware Visual Encoding (DAVE) agent. Functioning as a cognitive engine, PSR utilizes collaborative reasoning to deduce the fine‑grained attributes of unseen regions. Concurrently, acting as a specialized visual perception agent, DAVE quantifies and encodes viewpoint‑induced spatial misalignments, continuously feeding robust spatial constraints back into the generative process. This autonomous feedback loop between semantic deduction and visual perception ensures high‑fidelity detail synthesis. Extensive experiments on the DeepFashion and Market‑1501 benchmarks validate the superiority of our agent‑driven paradigm. Notably, DAC‑Pose excels in preserving texture alignment and identity consistency under drastic viewpoint changes. The code is available at https://github.com/AIVRC/DAC‑Pose.

Authors:Jiuhe Qu, Yingping Liang, Ying Fu
Title: HiSC: Hierarchical Spatial Clustering Token Compression for Efficient 3D Scene Understanding
Abstract:
3D vision‑language models (3D VLMs) enable spatial reasoning over multi‑view scenes but suffer from substantial token redundancy due to duplicated observations and large uninformative regions, leading to high computational cost. Although visual token compression has shown promise in accelerating 2D VLMs, it fails to capture the structured nature of 3D scenes and leads to incomplete spatial coverage and loss of fine‑grained details. In this paper, we propose HiSC, a training‑free framework for hierarchical spatial clustering token compression in 3D VLMs. HiSC lifts token compression from token‑level selection to cluster‑level processing by organizing tokens into spatially grounded clusters using joint geometric and semantic cues. Specifically, we first introduce a spatial graph‑based merging (SGraM) strategy that models cross‑view redundancy as spatial connectivity and consolidates physically consistent regions, effectively merging extremely similar redundant tokens prior to LLM inference. We then propose a spatial clustering‑based pruning (SCluP) paradigm within LLM inference, which performs hierarchical compression across clusters and within clusters, preserving object instance completeness while retaining fine‑grained details for important regions. Extensive experiments on diverse 3D reasoning benchmarks show validate the effectiveness of HiSC, particularly under high visual token pruning ratios. Besides, HiSC achieves over 90% token reduction with minimal performance degradation. Code is accessible at https://github.com/elecreak/HiSC.

Authors:Fang Li, Shihao Zou, Weixin Si, Yang Gao, Shuai Li, Aimin Hao
Title: TRCoRSurg: Temporal-Relational Co-Reasoning for Surgical Video Triplet Recognition
Abstract:
Understanding complex surgical scenes requires recognizing multiple interdependent entities, such as instruments, actions, and targets, while maintaining their relational consistency across time. Existing surgical triplet recognition methods struggle to jointly model intra‑frame label dependencies and inter‑frame temporal semantics in a unified manner. To address these limitations, we propose a unified framework that integrates spatial, relational, and temporal cues for robust surgical triplet recognition. Specifically, class‑specific spatial priors are first extracted through a multi‑scale encoder. These priors are then refined by a Label Correlation Modeling module with multi‑scale class activation map‑guided relational extraction (MS‑CAMRE), enabling the model to capture both static co‑occurrence patterns and dynamic contextual dependencies among triplet components. Furthermore, a Bidirectional Temporal‑Relational Fusion Attention (BTRFA) module harmonizes temporal and relational representations to achieve coherent temporal reasoning. We also introduce a new evaluation metric, the Triplet Consistency Error Rate (TCER), which quantitatively measures the model's ability to preserve causal and semantic consistency across triplets. Extensive experiments on the CholecT45 and ProstaTD datasets show that our method achieves state‑of‑the‑art performance, improving AP_IVT by 5.1 percent and 7.8 percent, respectively. Moreover, according to TCER, our approach achieves relative reductions of more than 36 percent and 25 percent on the two datasets, respectively, demonstrating the effectiveness of our framework in temporal‑relational co‑reasoning.

Authors:Benlei Cui, Ruize Wang, Junjie Li, Jinhao Chen, Longtao Huang, Yinghao Chen, Yuwen Zhai, Jingqun Tang, Ruijian Jia, Weiwei Wu, Pengfei Sun, Haiwen Hong
Title: MetaVideoAgent: Automated Video-Agent Evolution for Long-Form Video Understanding
Abstract:
Long‑form video understanding requires locating sparse, question‑relevant evidence in long, multimodal videos. Real‑world video distributions differ in modality‑specific information density, content structure, and evidence patterns, causing fixed video‑agent designs to incur redundant processing or fail when mismatched. Extending automated agent evolution from text to video is challenging because full long‑video execution makes candidate validation expensive, failures propagate across coupled evidence‑processing stages, and complex preprocessing, perception tools, and localization strategies make code‑level updates difficult to implement reliably. We introduce MetaVideoAgent, a framework that automatically evolves a video agent for a target distribution. It profiles information density and evidence requirements from sparsely sampled frames and associated queries to guide initial design, then compresses localized failures into independently executable minimal validation tasks. It constructs evidence‑grounded Gold Paths, audits Student trajectories, aggregates recurring failures across samples, and attributes them to responsible modules. A modular agent representation constrains each update to the primary responsible module and its necessary dependencies. We further introduce VA‑EvoBench, covering eight video distributions with separate evolution and held‑out splits. With four evolution iterations per distribution, MetaVideoAgent improves every initial agent and raises macro‑average accuracy from 38.44% to 51.47%, at an average evolution cost of 3.54M tokens per distribution. The evolved agents outperform the strongest prior fixed‑design video agent by 6.39 percentage points while using the fewest tokens and video frames per question among the compared video agents. We will release all code and data to support reproducible research.

Authors:Fang Li, Yang Gao, Shihao Zou, Weixin Si, Hongyu Wu, Qing Xia, Shuai Li, Aimin Hao
Title: VoxStruct3D: Structure-Leading Flow Matching for Voxel-Space 3D MRI Synthesis
Abstract:
High‑fidelity 3D MRI synthesis requires both globally coherent anatomy and fine‑grained voxel‑level detail. Although latent diffusion makes volumetric generation tractable, its image autoencoder introduces a reconstruction bottleneck that can limit the fine detail recoverable in the final volume. We present VoxStruct3D, a voxel‑space flow‑matching framework that directly models full‑resolution MRI volumes using a clean‑data prediction objective. Its Volumetric Voxel Generator (VVG) combines factorized 3D patch embedding with overlapping upsampling, time‑modulated residual refinement, and skip fusion, enabling neighboring tokens to jointly reconstruct shared voxel regions and suppress patch‑boundary artifacts. To complement direct voxel‑space modeling with an explicit anatomical prior, we further introduce a Structure‑First, Image‑Follows (SFIF) strategy. A frozen pretrained 3D medical encoder and a StructVAE extract compact structure tokens that preserve dominant anatomy, while a structure‑leading schedule keeps their trajectory ahead of the image trajectory. Patch‑Aligned RoPE spatially aligns the unequal token grids, and asymmetric attention enforces one‑way guidance from structure to image. Experiments on pathological and healthy T1‑weighted brain MRI datasets show that VoxStruct3D achieves the strongest overall performance across feature‑distribution alignment, sample diversity, and perceptual quality, producing anatomically coherent and visually realistic volumes.

Authors:Chen Zhong, Xiao An, Zijie Wang, Jiepan Li, Guangyi Yang, Wei He
Title: DIVE: Dynamic Iterative Visual Evidence Construction for Efficient Vision-Language Models
Abstract:
Visual inputs in vision‑language models (VLMs) are often encoded into substantially longer token sequences than text, making visual tokens a major bottleneck for efficient inference. Abundant recent methods address this bottleneck by scoring token importance and pruning low‑scoring tokens in a single pass. However, one‑shot scoring is insufficient because a token's prompt‑relevant usefulness depends on the evidence already retained. Motivated by this insight, we introduce DIVE (Dynamic Iterative Visual Evidence Construction), a training‑free framework that recasts visual‑token pruning as dynamic evidence construction. DIVE repeatedly selects the remaining token with the highest residual‑conditioned score, updates the visual and prompt residuals to discount the evidence already explained, and re‑evaluates the remaining tokens. This select‑update‑re‑evaluate process builds a retained set of complementary, prompt‑relevant evidence. Experiments across eight image‑understanding benchmarks show that DIVE consistently preserves performance across token budgets. With an 88.9% reduction in visual tokens, DIVE retains 98.2% of the uncompressed model's average performance. Code is available at https://github.com/Zhong‑Chenchen/DIVE.git.

Authors:Hyeonyu Kim, Sehwan Lim, Youngwon Choi, Taeyoun Kwon, Jaejin Kim
Title: Not All Redundant Tokens Are Alike: Analyzing Visual Token Pruning through Token Roles
Abstract:
Vision‑language models (VLMs) process an image as a sequence of visual tokens, which creates a substantial computational bottleneck during inference. Recent visual token pruning methods address this issue by removing seemingly redundant tokens, yet it remains unclear how these pruning decisions relate to the functional roles of visual tokens. In this work, we analyze visual token pruning through the lens of token roles identified by EmbedLens. We first show that representative pruning methods exhibit distinct token‑role biases, but these biases do not directly correlate with downstream performance. To better understand this behavior, we refine the token‑role assignment procedure and evaluate role‑protected pruning variants. Our results show that preserving non‑alive tokens can sometimes maintain or improve performance, suggesting that tokens with weak direct semantic alignment may still affect model behavior under pruning. Our code is publicly available at https://github.com/jaykim9870/Not_All_Redundant_Tokens_Are_Alike.

Authors:Suemin Jeon, Kaiyuan Tang, Chaoli Wang, Won-Ki Jeong
Title: Super-Gaussian: Interactive Scene Editing for 3D Gaussian Splatting and NLI-Based Volume Visualization in Virtual Reality
Abstract:
Despite the promise of virtual reality (VR) for intuitive spatial interaction, volume visualization (VolVis) in VR remains constrained by high rendering costs and motion discomfort. Recent advances have shown that representing volumetric scenes with 3D Gaussian splatting enables high‑performance rendering, making this representation well‑suited for VR. However, existing Gaussian‑based scene editing workflows remain limited by slow offline segmentation and fatigue‑inducing manual selection. To address these challenges, we present Super‑Gaussian, a novel VolVis framework that enhances scene editing and interaction in VR through intuitive 3D Gaussian selection and natural language interaction (NLI). Our approach groups Gaussian primitives into higher‑level units via feature‑aware clustering, enabling efficient selection of complex volumetric regions, such as tumors in medical images or filaments in cosmological data, without point‑by‑point interaction. Building on this, we introduce a hierarchical select‑and‑refine workflow that combines random‑walk‑based region propagation, cluster selection, and point refinement, allowing users to progressively specify regions of interest with reduced effort. We further support on‑the‑fly text labeling of selected regions using NLI, allowing users to semantically query, interpret, and manipulate content within a visualization‑perception‑action loop. By integrating multimodal interaction, including speech, visual feedback, and spatial manipulation in VR, our framework supports intuitive exploration, editing, and scientific analysis of volumetric data. We demonstrate the effectiveness of Super‑Gaussian through four case studies, quantitative selection benchmarks against existing Gaussian‑based techniques, and system‑level evaluations. Implementation details and experiments can be found on the project page: https://smin0136.github.io/super‑gaussian‑project/

Authors:Seunghyun Ji
Title: When does training on downscaled images yield the same gradients?
Abstract:
Diffusion transformers deliver strong image generation, but their training cost grows superlinearly with resolution. Recent work justifies training or sampling at reduced resolution on a spectral premise: at high noise, a downscaled latent preserves almost the full surviving signal. Whether a downscaled step also preserves the native training gradient signal, however, has remained unresolved. We reduce how that signal changes under downscaling to two terms: a noise‑dependent term governed by the downscale ratio, which decays at high noise as the spectral premise predicts, and a σ‑independent floor governed by the target grid's absolute token count, carried by the compute graph itself and removed by no noise level. The measured (route, σ) map corroborates the account and uncovers structure the spectral picture cannot express: on the 1024‑>768 route, a window (0.65 < σ< 0.95), predicted by no spectral criterion at any tolerance, where the downscaled gradient stays within a small margin of the native one. Training LoRA adapters with downscaled steps restricted to the routes and noise windows the map validates reduces training time by 14.6% at a fixed step budget while remaining near‑native in weight space. Code is available at https://github.com/sorryhyun/anima_lora.

Authors:Daijing Shi, Hongxiao Zhao, Yihan Fu, Zhan Chen, Jiayi Li, Yihang Zhu, Anjunyi Fan, Yaoyu Tao, Yuchao Yang, Bonan Yan
Title: MCHA: A Memory-Centric Hierarchical Architecture for Parallel-Sequential Computing
Abstract:
Emerging workloads, such as Multi‑Agent Reinforcement Learning (MARL), large‑scale neuromorphic computing, and probabilistic graphical models, intrinsically exhibit parallel‑sequential computing patterns. While these tasks demand massive parallelism to achieve high throughput, they are severely bottlenecked by irregular data access patterns centralized to main memory. Consequently, conventional architectures face fundamental limitations when executing these workloads, primarily manifesting as global buffer saturation and memory‑bound bottlenecks. To address these challenges, we propose the Memory‑Centric Hierarchical Architecture (MCHA), a reconfigurable hardware solution tailored for parallel‑sequential execution. MCHA leverages a hierarchical communication strategy that facilitates distributed, inter‑core data routing, thereby significantly reducing the bandwidth burden on the global memory. Complementing the hardware, MCHA introduces a novel parallel‑sequential programming model that utilizes event‑driven conditional triggers to effectively hide data transmission latency within the execution pipeline. We benchmark MCHA against a diverse suite of parallel‑sequential tasks, including MARL, motor variable control, and Markov random fields. Validated through our open‑source, cycle‑accurate simulator, MCHA demonstrates performance speedups ranging from 153.06× to 2456.96× over NVIDIA A100 GPUs on MARL workloads, while maintaining robust programming flexibility across other application domains. Furthermore, the architecture successfully reduces main memory access from 96% to 5.44%. When synthesized in a 28 nm process, the MCHA implementation occupies an area footprint of 2.92mm^2 and consumes 115.36 mW of power at 200 MHz. MCHA is open‑sourced at https://github.com/carabdis/MCHA.

Authors:Junbin Yuan, Muqing Cao, Yunwoo Lee, Brady Moon, Sebastian Scherer
Title: SCOPE: Field-of-View-Aware Path Planning in Unknown 3D Environments via Safety-Volume Certification
Abstract:
Safe navigation with a body‑mounted limited‑field‑of‑view sensor requires the complete robot‑inflated volume of an intended motion to be observed and verified free before execution. We formulate this requirement as online safety‑volume certification in an unknown voxel map and construct a certified graph whose vertices correspond exactly to positions with fully known‑free safety volumes. Based on this representation, we propose SCOPE (Safety Certification through Observation Planning and Execution), a planning framework that decouples optimistic goal‑directed guidance from certified execution. SCOPE converts the first uncertified point along an optimistic route into an explicit observation obligation, resolves it through target‑centric viewpoint search, and recursively clears intermediate obligations when useful viewpoints are not yet certified‑reachable. A certified preview mechanism and an observation‑aware trajectory optimization backend enable smooth execution. We prove conditional complete planning: under ideal monotone sensing and exhaustive finite‑domain graph search, SCOPE reaches the goal whenever a finite feasible sequence of certified sensing actions exists within its planning primitives. Across 60 randomized tasks in three unknown 3D environments, SCOPE reaches every goal while maintaining near‑zero entry into non‑certified inflated space. Preview reduces mean mission time by 27%, and real‑robot demonstrations in two representative scenarios validate the complete system.

Authors:Daohai Yu, Zhanpeng Zeng, Keyu Chen, Wenhao Li, Zhifeng Shen, Luxi Lin, Ruizhi Qiao, Xing Sun, Rongrong Ji
Title: Training-Free Hashing-Based Attention via Binary Principal Components
Abstract:
Long‑context large language models (LLMs) are increasingly deployed in real‑world applications, yet self‑attention remains a major efficiency bottleneck ‑‑ especially during decoding ‑‑ due to the necessity of repeatedly processing ever‑growing key‑value (KV) caches. Existing sparse attention reduce computation by attending to fewer KV pairs, but often suffer from substantial accuracy degradation, require additional training, or rely on expensive hashing. In this work, we present BinaryPC, a training‑free, data‑aware hashing‑based sparse attention for long‑context LLMs. BinaryPC constructs compact binary hash codes and corresponding hash function by computing binary principal components of data. Unlike Locality‑Sensitive Hashing (LSH) with data‑independent random projections or learned non‑linear hashing methods, BinaryPC constructs binary codes that explicitly preserve the structural information of data without requiring gradient‑based training. Comprehensive experiments across multiple model families and long‑context benchmarks show that BinaryPC preserves accuracy relative to full attention while achieving superior performance among sparse and hashing‑based baselines. On modern GPUs, BinaryPC improves end‑to‑end decoding throughput by 3.56× over the FlashAttention kernel. Our code is available at https://github.com/yudaohai666/BPC.

Authors:Zijian Zhuang, Yixiong Zou, Yuhua Li, Ruixuan Li
Title: Free-Lunch Augmentation by Revisiting Diffusion-Based Data Generation for Cross-Domain Few-Shot Object Detection
Abstract:
Cross‑Domain Few‑Shot Object Detection (CDFSOD) aims to transfer knowledge from data‑rich upstream generic domains to downstream expert domains using scarce training data, where the significant domain gap and data scarcity make it an unsolved challenge. To address this problem, we revisit a natural yet underexplored approach in CDFSOD: data augmentation, by directly synthesizing data through diffusion models to supplement limited training samples. However, due to large domain gaps, we find that current diffusion methods cannot produce good results, leading to performance even lower than using the original images. To address these limitations, we divide the domain gaps into visual gaps and semantic gaps for separate analysis. For the visual gap, we find that the diffusion model cannot distinguish noise from useful information on expert domains, which can be mitigated by adding weakened noise. For the semantic gap, we find that the background semantics shows much smaller gaps between domains than foreground semantics, and we can bridge this gap by background inpainting. Based on the above analysis, we propose a method (Selective Inpainting with Tailored Noise, SITN) to dynamically take different strategies for downstream data synthesis based on their different gaps from the general domain, including a Generation Module for adding tailored noise and a Selection Module to dynamically select the inpainting regions. Extensive experiments on 6 datasets of CDFSOD and 4 datasets of cross‑domain few‑shot segmentation (CDFSS) validate that we can synthesize helpful data, achieving new state‑of‑the‑art performance. Our codes is available at https://github.com/zzzzj311‑droid/Free‑Lunch‑SITN

Authors:Soroush Omidvartehrani, Mohammadamin Habibollah, Mohammadreza Daviran, Davood Rafiei
Title: EdgeLM: Edge Demonstrations for Language Models' Table Understanding
Abstract:
Large language models (LLMs) perform table‑centric prediction through in‑context learning, making demonstration selection critical to performance. Existing retrieval methods prioritize similarity to the query, but similar demonstrations often reinforce the model's likely prediction rather than reveal the distinctions needed for difficult decisions. We propose EdgeLM, a retrieval framework that instead selects edge evidence, demonstrations that are both relevant to the query and informative about the decision boundary. EdgeLM retrieves two complementary forms of edge evidence by selecting data edges, nearby examples with different ground‑truth labels, and model edges, similar examples previously misclassified by the deployed model. EdgeLM requires neither model retraining nor task‑specific engineering. Across five data wrangling tasks, fifteen datasets, and five open‑weight and proprietary LLMs, EdgeLM consistently achieves the best or near‑best performance in every setting, while ablations show that the two forms of edge evidence provide complementary benefits. Our code and datasets are publicly available at https://github.com/soroushomidvar/EdgeLM.

Authors:Lei Peng, Shuai Lv, Wei Hu
Title: ReGround: Restoring Visual Grounding in Multi-Step Reasoning through Self-Diagnosis and Visual Re-Examination
Abstract:
Vision‑Language Models (VLMs) often lose visual grounding during multi‑step reasoning: as reasoning chains grow longer, later inference steps rely increasingly on language priors rather than image evidence. We identify a consistent benchmark‑level signature associated with this degradation: across 2,510 re‑examined samples from four benchmarks, attention entropy over image tokens typically decreases during Round 1 and rises again after image re‑injection. However, we find that effective visual re‑examination requires two complementary ingredients: image re‑injection and targeted self‑diagnosis. Without targeted diagnosis, re‑examination can even hurt performance, whereas accurate self‑diagnosis yields substantial gains ‑‑ a swing of several points on key benchmarks, indicating that diagnostic quality is a key factor in whether re‑examination helps or hurts in our setting. We present ReGround, a two‑stage framework that teaches VLMs to self‑diagnose grounding failures and selectively re‑examine visual evidence, without architectural modifications or external tools. Through capability bootstrapping, a stronger variant from the same model family provides diagnostic scaffolding only during data construction, while the policy model learns to diagnose autonomously at inference time and retains most of the assisted gains. Experiments on eight benchmarks across two VLM backbones demonstrate consistent gains, especially on visually intensive multi‑step reasoning tasks, while incurring only modest inference overhead relative to tool‑augmented baselines. Project page: https://sespoir.github.io/reground‑page/ . Code: https://github.com/sespoir/ReGround .

Authors:Tinghe Zhang, Jian Xu, Jiaheng Chen, Jiaxing Li, Yucheng Xiao, Qiang Wang
Title: NodeJEPA: Structure-Conditioned Latent Prediction for Node-Level Graph Self-Supervised Learning
Abstract:
Self‑supervised learning on graphs is largely shaped by contrastive methods that depend on carefully designed augmentations, and by generative methods that reconstruct node attributes in the input space. Both paradigms can entangle representations with low‑level input statistics rather than with relational structure. Joint‑embedding predictive architectures (JEPA) instead learn by predicting latent targets rather than reconstructing inputs. Recent work has explored this idea for graph‑level representation learning, but how to design JEPA‑style objectives for node‑level tasks, and which structural signals the predictor should condition on, remains less clear. We present NodeJEPA, a joint‑embedding predictive architecture for node‑level graph self‑supervised learning. NodeJEPA masks structure‑aware k‑hop ego‑subgraphs and trains a context encoder to predict the latent representations of the masked nodes. These targets come from an EMA‑updated target encoder with stop‑gradient. A structure‑conditioned predictor integrates spectral and centrality descriptors through cross‑attention. Variance, covariance, and Laplacian spectral regularizers help stabilize the embedding geometry, and an optional curriculum gradually increases masking difficulty during training. Because prediction occurs in latent space, NodeJEPA does not rely on input reconstruction or hand‑crafted graph augmentations. We evaluate NodeJEPA on standard node classification benchmarks under linear probing and fine‑tuning protocols, and conduct ablations on masking, prediction, and regularization design choices. Our study offers a practical recipe for node‑level JEPA‑style latent prediction on graphs, and clarifies when structural conditioning helps representation learning. Code, configurations, and evaluation scripts are publicly available at https://github.com/OliverZ‑dot/Node‑Jepa.

Authors:Scott H. Hawley
Title: Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation
Abstract:
Collaborative music agents need internal representations rich enough to support both understanding and generation, yet flexible enough for a workflow where the human retains agency. We present a hierarchical self‑supervised ``world model'' for symbolic music: a 2.55M‑parameter Swin V2 encoder trained on MIDI piano‑roll images with JEPA‑style objectives (pitch‑ and time‑shift equivariance, masked embedding prediction, and a distributional regularizer), using no labels and no music‑theory vocabulary. Probing the frozen embeddings shows that the level at which a musical property becomes decodable tracks its musical time scale: phrase boundaries are read off the coarsest levels, note density and harmonic detail off the finest. Temporal and phrase structure emerge from the self‑supervised objectives alone, while harmonic content must be asked for; a small chord‑supervision head raises joint chord recovery from .18 to .54, and key detection, which is never supervised, from .16 to .70. Following the Representation AutoEncoder paradigm, a conditional flow‑matching model stands in for a trained decoder, flowing in pixel space from PCA‑reduced conditioning: it reproduces a target window at pixel F1 0.996, and the same per‑level conditioning dropout that controls how far variations stray also enables graphical prompting for masked inpainting with no inpainting‑specific sampler. The pipeline runs on CPU producing a suggestion in 2.8 s, or 0.6 s on Apple MPS, which we demonstrate in a live interactive demo. In concert with an LLM‑based brain, these capabilities supply the core of a collaborative music creation agent in service of, rather than in place of, human agency.

Authors:Yinghao Tang, Tan Zhenwei, Yiyao Wang, Wanli Gu, Xiaolu Zhang, Jun Zhou, Wei Chen
Title: FinReportBench: Measuring and Improving Institution-Grade Financial Report Generation
Abstract:
Large language models can produce fluent financial analysis, but fluency alone does not establish whether a report is suitable for institutional delivery. We introduce FinReportBench, an expert‑grounded benchmark for measuring and improving institution‑grade financial report generation. Expert review reveals recurring gaps in report identity, institutional components, source discipline, and visual delivery. We derive a 35‑item rubric through expert partial orders, multimodal evidence, and audits of decision boundaries, covering deliverability, report identity, and institutional completeness. Starting from 10,000 balanced Chinese and English financial‑research source records, we curate 244 bilingual tasks across three research objects and two input tiers. Each task separates the public query, reconstructed research trajectory, and hidden source packet. Three independent judge families reproduce the expert partial order at near‑ceiling rates, showing that bounded, observable criteria support reliable evaluation. Across nine model families, basic deliverability is nearly saturated, while report identity and institutional completeness remain the primary bottlenecks. The largest cross‑model gaps concern generation‑trace control, information density, and data discipline rather than basic report framing. We then use benchmark‑guided skill distillation to turn recurrent failures into reusable generation and self‑review constraints. Across five model families, the evolved skill improves mean G1 by 33.85 points and mean G2 by 13.83 points over paired no‑skill runs while preserving G0 for every pair. Code and benchmark artifacts are available at https://github.com/MisterBrookT/finreportbench.

Authors:Seyed Roozbeh Razavi Rohani, Khashayar Khajavi, Wesley Chung, Mandana Samiei, Mo Chen
Title: NeuMoSync: End-to-End Neuromodulatory Control for Plasticity and Adaptability in Continual Learning
Abstract:
Continual learning (CL) requires models to learn tasks sequentially, yet deep neural networks often suffer from plasticity loss and poor knowledge transfer, which can impede their long‑term adaptability. Drawing high‑level inspiration from global neuromodulatory mechanisms in the brain, we introduce Neuromodulation and Synchronization (NeuMoSync), a novel architecture that integrates dynamic, neuron‑specific modulation into deep neural networks to enhance their adaptability and plasticity. NeuMoSync extends standard neural network architectures with learnable feature vectors for each neuron that track network‑wide historical context and with a module operating at a higher level of abstraction. This module synthesizes neuron‑specific signals, conditioned on both current inputs and the network's evolving state, to adaptively regulate activation dynamics and synaptic plasticity. Evaluated on diverse CL benchmarks, including memorization (Random Label CIFAR‑10 and Random Label MNIST), concept drift (Shuffle CIFAR‑10 and Shuffle Mini‑ImageNet), class‑incremental learning (Class Split ImageNet and Class Split CIFAR‑100), and domain‑incremental learning (Permuted MNIST), NeuMoSync demonstrates strong performance in retaining plasticity and achieves improvements in both forward and backward adaptation compared with existing methods. Ablation studies validate the necessity of each component, while analysis of the learned modulatory signals reveals interpretable coordination patterns across tasks. Our work underscores the potential of integrating global coordination mechanisms into deep learning systems to advance robust, adaptive continual learning. The code is publicly available at https://github.com/RoozbehRazavi/NeuMoSync.

Authors:Al Zadid Sultan Bin Habib, Md Younus Ahamed, Prashnna Gyawali, Gianfranco Doretto, Donald A. Adjeroh
Title: iStructTab: Structured Feature Sequencing for Multimodal Learning of Image and Tabular Data
Abstract:
Multimodal learning of images and tabular data is often impaired by ineffective representations, resulting in redundancy, dispersion, and generalization problems. To tackle this challenge, we introduce Graph‑Enhanced Descriptor Sequencing (GEDS), a structured feature sequencing algorithm grounded in principles from the Column Permutation Problem (CPP). GEDS refines statistical descriptors of the features through similarity graph‑based computations, systematically determining an effective feature sequencing. We incorporate GEDS within an order‑aware efficient transformer framework, utilizing order‑aware memory tokens that explicitly adhere to the derived feature sequencing via a dedicated loss function. Experimental results across multimodal benchmarks demonstrate that iStructTab effectively minimizes feature dispersion, improving predictive performance and robustness, and highlighting the significance of structured feature sequencing in multimodal learning.

Authors:Mengyu Xu, Qiaoxin Yang, Zhihan Liu, Ruiyao Xu, Zachary Liu, Kezhen Chen, Chongyang Gao
Title: Equitable System-Prompt Selection via Constrained Mixed-Strategy GroupDRO
Abstract:
Large language models are increasingly used for information seeking, yet semantically equivalent questions phrased in different ways can receive answers of considerably different quality. System prompts are widely employed to steer response behavior, but they are typically optimized for average‑case quality, so some question phrasings may still receive incomplete or low‑quality answers. To address this, we formulate a constrained mixed‑strategy GroupDRO framework for system‑prompt selection. Instead of optimizing the system‑prompt text, the framework assigns weights to system prompts in an existing pool to minimize the worst‑case information‑quality loss across evaluation metrics and groups, while constraining the mean loss to stay close to that of average‑based selection. Because pool generation and selection are decoupled, the method applies to any system‑prompt pool and can leverage an ensemble of complementary system prompts rather than a single one. Across five LLMs on two bilingual medical and consumer‑finance benchmarks, the constrained method reduces the Overall Mean, Worst 25% Mean, and Worst by 13.1%, 13.2%, and 13.7% on average relative to no mitigation while keeping overall quality close to Average selection. Its multi‑prompt weights reveal complementarity across metric‑group pairs. Code and data are available at https://github.com/Rainxu09/equitable‑system‑prompt‑selection.

Authors:Jingzhi Gong, Jie M. Zhang, Gunel Jahangirova, Dong Huang, Mohammad Reza Mousavi, Mark Harman
Title: COMPAS: Difficulty-Aware Joint Search for Optimizing Code Generation
Abstract:
Code generation systems make each LLM call with a model, a prompt, and decoding settings. However, existing optimization methods usually tune only part of these choices or use one fixed configuration for all tasks: global optimizers search one configuration for all tasks, routers choose only a model, and prompt optimizers keep the model and decoding settings fixed. This leaves their joint, group‑specific interactions unclear. We therefore examine how these choices interact and observe that prompts and decoding settings interact, tuning effects vary by model, and the best configuration varies by task difficulty. Guided by these observations, we introduce COMPAS (Code‑generation Optimization over Models, Prompts, And Decoding Settings), a difficulty‑aware method that learns group‑specific quality‑cost fronts through low‑cost model selection and joint prompt‑decoding search, then routes each test task to its matching front online without further search. Under a matched search budget on LiveCodeBench, COMPAS improves pass@1 from 45.9% for the best baseline to 52.8% while reducing cost from 36.57 to 4.92. This also transfers to repository‑level code generation on SWE‑bench, resolving 76.0% of tasks versus 70.0% for the best baseline. Code and the reproducibility artifact are available at https://github.com/gjz78910/COMPAS.

Authors:Mike Vegeto
Title: Right Reset: Chunking by Prefix Removal
Abstract:
Removing the left context from a causal language model reveals a useful kind of boundary: an edge where the model processes the same right‑hand tokens with little change. We turn this observation into prefix‑removal probing and introduce Right Reset (RR), which measures preservation of the right‑hand hidden‑state trajectory. A dynamic program converts RR edge scores into variable‑length chunks. On flattened text formed by concatenating topically similar records after deleting their separators and layout, RR recovers 47.7% of the original records as clean units, versus 25.9% for a BGE embedding boundary, the strongest tested conventional baseline without task‑specific model training. The gain persists after rendering and OCR. Passive scores from the same Qwen3‑4B layer and direct prompting of a same‑scale instruction model perform substantially worse on flattened records. Across six language models, RR‑selected cuts also undergo consistently less local output disruption than unselected candidate edges. An observed‑token likelihood‑ratio readout is competitive in some architectures, indicating that the central contribution is the intervention: context dependence itself can provide a boundary signal when surface structure is weak.

Authors:Mukhtiar Ali, Harsh Dubey, Sugam Mishra, Chulwoo Pack
Title: CLIP-CC-Bench: Evaluating Paragraph-Level Video Descriptions in Video-Language Models
Abstract:
Benchmarking video‑language models has largely focused on short clips and single‑sentence metrics, leaving open whether current systems can generate accurate long‑form, paragraph‑level descriptions. We introduce CLIP‑CC‑Bench, an evaluation suite for long‑form video description built from 5 hours of movie content segmented into 90‑second clips, each paired with an expert‑written paragraph‑style reference. The evaluation suite employs an ensemble of five state‑of‑the‑art LLM‑based embedding models to increase reliability and mitigate single‑model bias, and applies two complementary methodologies: (i) coarse‑grained semantic matching and (ii) fine‑grained semantic matching to compare model‑generated descriptions against CLIP‑CC‑Bench references. Using this framework, we evaluate 17 state‑of‑the‑art video‑language models and report both their Borda‑aggregated rankings and their average scores on CLIP‑CC‑Bench. We further quantify the protocol's internal reliability through inter‑judge agreement and bootstrap ranking stability. We release standardized evaluation scripts, model outputs, and aggregation tools at https://github.com/Multimodal‑Intelligence‑Lab/CLIP‑CC‑Bench to support reproducibility. CLIP‑CC‑Bench provides a practical evaluation framework for long‑form video description, filling a gap left by existing short‑clip and QA‑only benchmarks.

Authors:Atri Vivek Sharma, Brian Formento, Alessio Lomuscio
Title: Eliciting Intrinsic Hallucinations in LLMs via Semantically Equivalent Adversarial Attacks
Abstract:
Large language models (LLMs) are often used in conjunction with external knowledge sources to improve their factual accuracy and decrease hallucinations, through methods such as Retrieval‑Augmented Generation (RAG). However, these systems remain susceptible to intrinsic hallucinations, where the model generates unfaithful or fabricated information that is not supported by the retrieved evidence. We propose a novel framework to assess model robustness against this phenomenon by stress‑testing using natural, semantically equivalent variations of a user query found via adversarial optimization methods. We apply our framework, which enforces strict semantic equivalence constraints and an intrinsic hallucination objective, to a range of adversarial attack techniques across white‑box, gray‑box, and black‑box adversarial settings. Evaluating these attacks on 5 open‑source and 5 closed‑source generator models across 3 datasets, we demonstrate that even state‑of‑the‑art models are highly susceptible to meaning‑preserving perturbations, which significantly degrade contextual faithfulness (by up to 50% for GPT‑5‑mini). Our findings indicate that faithful use of in‑context evidence remains fragile even in state‑of‑the‑art LLMs, motivating architectures and training objectives that enforce robust grounding independent of surface query form. Code is available at: https://github.com/atriviveksharma/intrinsic_hall

Authors:St John Grimbly, Nicolas Kuske, Evert A. Boonstra, Bruce A. Bassett, Charel van Hoof, Rowan Hodson, Benjamin Rosman, Ryan Smith, Mark Solms, Jonathan P. Shock
Title: Interoceptive Attention as Dynamic Homeostatic Prioritization in a Foraging Agent
Abstract:
Biological systems must regulate competing needs under limited perceptual bandwidth, where sharpening one estimate costs the capacity to sharpen the others. Any fixed‑budget system therefore has to decide where to allocate its perceptual precision. We study this in a foraging agent that must keep several bodily needs satisfied to survive, modelled with active inference. At each step it reads its own body‑state beliefs, identifies the most‑needed channel, and reallocates a fixed budget of interoceptive precision toward it, so that the same precision‑shaped likelihood feeds both belief update and planning. In AffectWorld, a four‑channel foraging gridworld, this selective allocation more than doubles learning‑phase survival at matched budget against a uniform‑precision agent (0.414 vs 0.199 across 11 layouts, n=32 seeds each, paired cluster‑bootstrap p \leq 10^‑4). Two further results sharpen the mechanism. The benefit runs through planning as well as perception, since denying the shaped likelihood to the planner alone removes about half of it. It is also need‑aligned, since aiming precision at the least‑needed channel does worse than spreading it evenly. The attended channel additionally learns its own dynamics about twice as fast, and stays ahead even at matched observation count, a behavioural trace of the same precision routing, visible in learning speed, not survival.

Authors:Arthur Freitas Ramos, Ruy J. G. B. de Queiroz, Anjolina Grisi de Oliveira, Tiago M. L. de Veras
Title: Topological Semantics for Scoped Computational Paths
Abstract:
Computational paths record equality as explicit finite traces of primitive steps. We give a topological semantics for a scoped rewrite presentation whose steps have continuous geometric realizations and whose named rewrites carry endpoint‑fixed homotopies. For every presentation we construct a quotient arrow space with a canonical final‑domain groupoid structure: multiplication is continuous on the quotient of explicitly composable representatives. We prove an exact four‑way criterion for this final composable topology to agree with the ordinary pullback topology, together with a compact‑Hausdorff sufficient condition. Thus the unconditional construction exposes, rather than hides, the product‑quotient issue in ordinary topological groupoids. The realization map to geometric homotopy classes is a continuous groupoid morphism and is faithful exactly under a separate geometric‑completeness condition. In the universal presentation, a continuous section identifies the coherent‑path quotient homeomorphically with the usual quotient‑topologized fundamental groupoid. We then give finite‑generator circle and genuine torus examples, with winding‑based normal forms and classifications by Z and Z^2. A Lean 4.24.0 development checks the theorem package; the mathematical presentation is independent of the implementation.

Authors:Xin Lu, Zihao Fan, Mingchen Zhong, Jie Huang, Xueyang Fu, Zheng-Jun Zha
Title: OmniVR: Joint Video-Audio Conditional Generation for Restoring Degraded Historical Films
Abstract:
Historical films suffer from co‑occurring visual and audio degradations‑‑‑blur, noise, flicker, hiss, clipping, and dropout‑‑‑yet existing methods restore each modality independently, leaving quality gaps and cross‑modal inconsistency. We present OmniVR, the first joint audio‑video generative restoration model. Built upon a 22B‑parameter audio‑video generation backbone, OmniVR formulates restoration as conditional generation within a unified multimodal DiT: the low‑quality video and audio are encoded as latent conditions, combined with a fixed restoration prompt, and jointly denoised to recover visual structure, temporal motion, and acoustic detail under one coordinated objective. Three key designs enable this adaptation: (1) a joint audio‑video degradation pipeline that simulates real old‑film characteristics from Internet‑collected data; (2) an architecture‑preserving text‑to‑audio‑video (T2AV) to audio‑video‑to‑audio‑video (AV2AV) transition with prompt annealing that maximally retains the generative prior; and (3) first‑frame image‑to‑video (I2V) anchoring with loss reweighting and waveform supervision for long‑video extrapolation and audio fidelity. We also propose OmniVRBench, the first benchmark that evaluates audio‑video restoration across visual quality, audio quality, temporal consistency, and audio‑visual synchrony on 200 real historical clips. OmniVR surpasses all prior methods on all six visual metrics, achieves the best audio quality, and produces natural colorization‑‑‑the first method to jointly address all three aspects. Code and weights will be publicly released. Project Page: https://xin1u.github.io/OminiVR_PAGE/

Authors:Yilong Dai, Yiming Sun, Yiheng Chen, Shengyu Chen, Peyman Givi, Xiaowei Jia, Runlong Yu
Title: TIDE: A Physically Diverse 3D Turbulence Benchmark Dataset for Advancing Scientific Machine Learning
Abstract:
Turbulence is a central testbed for machine learning on physical dynamics because its governing laws are known exactly. However, most existing studies remain in 2D, while 3D turbulence has fundamentally different physics and is far more costly to simulate. Existing 3D resources also typically provide only one realization per configuration, making it difficult to distinguish learning the dynamics from fitting the statistics of a single flow. In this paper, we introduce TIDE (Turbulent Incompressible DNS Ensembles), a 256^3 DNS corpus and benchmark for 3D incompressible turbulence, with 15 configurations on eight controlled axes, independent ensembles, pressure fields, and equation‑level verification. The benchmark includes five tasks, standardized learned baselines, controlled generalization splits, and physical‑fidelity metrics alongside pointwise error. Across the main forecasting configurations, current learned models barely outperform persistence and still make about twice the error of a spectral solver given the true equations. Moreover, lower pointwise error can coincide with severely distorted small‑scale dynamics, showing that accuracy alone does not ensure physical fidelity. Generalization results further show that most regime shifts reflect limited training coverage, whereas forced‑to‑decay transfer exposes a missing conditioning variable: operators trained under forcing continue to predict driven evolution when the external drive is removed. Closing these accuracy, fidelity, and conditioning gaps is the central open problem made measurable by TIDE.

Authors:Nie Lin, Takehiko Ohkawa, Sijin Chen, Ruoshi Wen, Zhuohang Li, Liqun Huang, Zhengming Zhu, Yiming Bao, Yunfei Li, Minjie Cai, Xiao Ma, Wei Xu, Yoichi Sato
Title: SiMDex: Mining Similar Egocentric Videos for Cross-Embodiment Dexterous Manipulation
Abstract:
Recent years have witnessed an explosive trend of scaling ego‑centric human videos for robot manipulation, yet it remains unclear which data actually benefits dexterous manipulation. We present SiMDex, a similarity‑based data mining framework that casts human data selection for VLA post‑training in dexterous manipulation as a recommendation problem. For each robot demonstration, SiMDex employs a three‑layer recall‑ranking‑re‑ranking pipeline to extract task‑relevant subsets from a pool of ~32M egocentric human samples, operating in a morphology‑agnostic action space that requires no changes to VLA architecture or training. Against a strong baseline trained with an equal amount of randomly sampled human data, SiMDex uses only ~1.49M mined samples (<5% of the pool) yet improves the overall success rate from 47.7% to 61.1%, showing that selective curation outperforms indiscriminate data mixing.

Authors:Luxshan Thavarasa, Sivasuthan Sukumar
Title: Test, then Route: How Language Models Execute In-Context Conditional Rules Across Models and Languages
Abstract:
When a language model follows an in‑context conditional rule such as "if P(x) then A else B," does it assemble a runtime circuit with one module that tests the predicate and another that routes the answer? We probe this with activation patching under a four‑donor design whose two swapped‑rule donors make the condition and the answer word disagree, so each layer reveals which of the two it carries. Across three open models from two families and six languages sharing one fixed item bank, a mid‑stack residual band carries the predicate's truth value: patching it reroutes the answer with predicate‑outcome flip near 1.0 and mapping flip near 0.0, meeting a strict pre‑specified isolation criterion in 17 of 18 cells, and the same localization holds across five predicate families. The router shows the opposite profile. A learned subspace flips A and B near‑perfectly within the trained pair yet transfers to a new pair at approximately 0 in every model, while in Gemma‑3‑4B (the only model probed cross‑lingually) it transfers at approximately 0.98 to the same pair in other languages. Under every probe we ran, the router direction is token‑bound and non‑transferable (largely answer‑readout in Gemma, pair‑specific in Qwen) rather than an abstract routing module. Test is modular; under these probes, route is not.

Authors:Palash R. Roy, Banani Roy, Kevin A. Schneider, Chanchal K. Roy
Title: MergeSE: Post-Hoc Model Merging for Software Engineering Tasks Without Retraining
Abstract:
Fine‑tuned code models often behave as domain specialists and can degrade sharply under distribution shift: in our clone‑detection setting, a model trained on same‑language clones drops 71% F1 on cross‑language clones, while multi‑task training falls to 0.151 F1 on unseen AI‑generated clones. Our companion study shows that post‑hoc model merging can address this fragmentation, achieving 93% of multi‑task performance without training data while generalizing 4× better to unseen clone types. However, no practical tool exists that lets SE researchers diagnose checkpoint compatibility, merge specialists, validate results on SE benchmarks, and export models for deployment. We present MergeSE, an open‑source CLI and web tool for training‑free model merging of HuggingFace encoder checkpoints. While motivated by OOD generalization in clone detection, MergeSE supports SE classification workflows more broadly through a built‑in registry of nine task types, including vulnerability detection, defect prediction, and code‑smell detection. MergeSE provides five operations: tasks, inspect, merge, evaluate, and export. It supports five merging algorithms, including TIES, DARE‑TIES, Wudi, PCB, and averaging; detects cross‑task classification‑head mismatches; produces seedable deterministic outputs; and includes bundled benchmark samples for smoke‑test reproduction. A full merge of two 124M‑parameter checkpoints completes in under 5 seconds on CPU. End‑to‑end validation confirms that MergeSE‑produced checkpoints match reference implementations and recover cross‑domain performance from domain‑specific specialists. The tool is available online at https://mergese.usask.ca, and the development repository is at https://github.com/srlabUsask/MergeSE.

Authors:Qile Wang, Ali Salloum, Carolina Coimbra Vieira, Benjamin E. Bagozzi, Mikko Kivelä, Kenneth E. Barner, Matthew Louis Mauriello
Title: Echoes in the Sky: Computational Thematic Analysis of Online Public Discourse on Bluesky Across Trump's Reelection
Abstract:
As political disruption intensifies online discourse, Bluesky has become an important platform for political discussion and public reaction. In this study, we examine large‑scale discourse on Bluesky related to U.S. policy developments associated with the Trump administration. Using the historical retrieval API, we collected all available posts matching Trump and related keywords from 2019 to 2026, yielding 38.5 million posts. We leverage a large language model (LLM)‑assisted clustering pipeline, combined with human validation, to identify 14 interpretable thematic domains in English‑language posts and 19 thematic categories across 258 executive orders (EOs) signed between January 20, 2025, and May 1, 2026. Our findings identify several dominant themes in Bluesky discourse, including executive governance, political identity, and national security, as well as recurring themes in EOs, including executive task forces, border enforcement, and foreign policy. We also find substantial variation in the persistence and volatility of issue attention, accompanied by an increasing proportion of negative sentiment over time. The dataset and resources are publicly available at https://github.com/Sensify‑Lab/Echoes‑in‑the‑Sky

Authors:Matteo Teodori
Title: NEBULA: A Language - Independent Specification for Opaque Rotating Refresh Tokens
Abstract:
Refresh tokens are among the most sensitive credentials in modern authentication systems: long‑lived, bearer‑style, and sufficient to mint access tokens for days or weeks. RFC 9700, the current Best Current Practice for OAuth 2.0 security, mandates that refresh tokens issued to public clients be rotated on every use with replay (reuse) detection, or be sender‑constrained. But the BCP specifies policy, not mechanism: it prescribes no wire format, no storage schema, no ordering of verification steps, no concurrency contract, and no semantics for edge cases such as lost‑response retries or key rotation. Implementations may therefore diverge in precisely the corner cases that determine security outcomes. We present NEBULA, a precise, language‑independent specification of the RFC 9700 refresh‑token model, together with ten conformant reference implementations (TypeScript, Python, Go, Rust, Java, PHP, C#, Ruby, Elixir, Dart). NEBULA tokens are opaque ‑‑ a 128‑bit public selector and a 256‑bit secret verifier, both CSPRNG output, carrying no claims and no signature ‑‑ so token validity is a property of server‑side state rather than of cryptographic verification. Its conformance methodology publishes the behavioural suite as data rather than as prose: 38 scenarios in one machine‑readable file that every implementation executes through a thin per‑language runner, so that drift by transcription is structurally excluded. We describe the specification ‑‑ including a compare‑and‑set rotation contract that closes a reproducible bypass of reuse detection under concurrent refresh ‑‑ analyse its security properties including its post‑quantum posture, and report on cross‑language conformance as a method for multi‑implementation security specifications. The specification, implementations, and conformance artefacts are open source under the Apache License 2.0.

Authors:Boyao Wang, Zhihan Lei
Title: SpecDrop: Parameter-Free Category-Conditioned Routing for Modular Specialization
Abstract:
Mixture‑of‑experts (MoE) networks pursue specialization through learned routers, gates, and load‑balancing losses, yet at matched total‑parameter budgets learned routers can underperform equal‑weight No‑Routing baselines. Is the bottleneck the routing algorithm, or the alignment between training‑signal granularity and the target categories? We probe the question with SpecDrop, a fixed parameter‑free routing scheme: each of K branches receives weight p_a for its assigned category and a small leakage p_i > 0 otherwise, merged through a category‑independent fixed denominator, with no learned routing parameters and no auxiliary losses; the category label is required at inference. On vision tasks where each image has one superclass label (CIFAR‑100 on ResNet‑110; ImageNet‑1K on ViT‑S/16), SpecDrop reaches 79.23% on CIFAR‑100 and 79.89% on ImageNet‑1K, exceeding parameter‑matched baselines that do not use the label (+4.75 over dense on CIFAR‑100; +6.53 over the No‑Routing+SE control on ImageNet‑1K). These gains quantify what category supervision buys when deployed through routing ‑‑ not an advantage over label‑aware deployments of the baselines: given the same label, masking a dense model's outputs is stronger for accuracy alone (85.2 / 83.7). SpecDrop's contribution is converting the label into trained‑in modular structure: 58%/100% branch‑category alignment, and masking gains of 0.00 (CIFAR) / +1.06 (ImageNet) ‑‑ the output‑space restriction is largely internalized during training. On fuzzy partitions, where training units span multiple categories (SlimPajama‑6B language modeling with a 30M Transformer; SuperNI instruction tuning over Llama‑3.2‑1B with LoRA), the routing mechanism reduces to the matched No‑Routing controls within seed noise, the null our thesis predicts. Granularity alignment, not algorithm choice, localizes when routing helps. Code: https://github.com/Beryex/SpecDrop

Authors:Yannick Stade, Robert Wille
Title: Guiding Compiler Optimizations for Neutral Atom Quantum Computers Through Visualizations
Abstract:
The scale of Neutral Atom (NA) quantum computers requires automated compilation tools. Designing the required heuristic methods demands a deep understanding of complex hardware trade‑offs, for which visualizations can provide crucial insights. This work introduces NAViz, the first publicly available app to visualize quantum computations on NA devices in real‑time. A case study demonstrates how NAViz was instrumental in identifying and resolving inefficiencies in an existing compilation strategy, leading to a new, more performant one. The tool is available as part of the Munich Quantum Toolkit (MQT) at https://github.com/munich‑quantum‑toolkit/naviz.

Authors:Hadi Mohammadi, Tina Shahedi, Robert A. Bagheri, Mehdi Dastani, Masoume M. Raeissi
Title: Learning Sexism Detection Using Multi-Agent Perspectivist Preference Optimization
Abstract:
When people label text for sexism, they often disagree, and not because some of them are wrong: they genuinely perceive sexism differently. Most NLP systems discard this disagreement by collapsing it into a majority vote. We propose the Multi‑Agent Perspectivist Preference Optimization (MAP‑PO) framework to keep these different perspectives. On the EXIST 2024 dataset of labeled English and Spanish tweets, we first cluster annotators by their labeling behavior rather than their demographic attributes. We then fine‑tune one Large Language Model agent per cluster to reproduce that cluster's annotation behavior, and coordinate the agents with preference optimization that combines individual and team‑level rewards. We evaluate MAP‑PO in four settings defined by two languages and two backbone language models, asking whether each agent reproduces the annotations of its own cluster and whether the agents together reproduce the majority label. Two findings hold in all four settings. First, without fine‑tuning the agents behave almost identically, so cluster‑specific training is necessary. Second, we show that training each agent only on the labels of its own cluster pushes the agents far beyond the clusters they should represent, while adding a shared team‑level training signal consistently keeps each agent calibrated to its cluster.

Authors:Myung-Hwan Jeon, Sankalp Yamsani, Joohyung Kim
Title: Kitchen Robotic Manipulation utilizing Foundation Models
Abstract:
Deploying robots in everyday human environments requires perception systems that are both robust and adaptable to diverse, dynamic conditions. In this work, we present a modular perception pipeline for household manipulation tasks, with a focus on dishware handling in kitchen environments. The pipeline integrates open‑vocabulary object detection, multi‑view segmentation, instance‑aware 3D reconstruction, and a 2D‑3D feature fusion strategy for 6D pose estimation and grasp planning. Its modular design enables systematic substitution of multiple visual and geometric foundation models, allowing us to identify the best‑performing configuration through extensive evaluation on a custom kitchen dataset. The best‑performing configuration (LLMDet + SAMv2 + DINOv2 + GeoTransformer) achieves an ADI of 89.12% on the 20‑scene kitchen benchmark with cluttered and occluded conditions. Furthermore, real‑world demonstrations confirm that the best configuration can be deployed on physical robots without environment‑specific retraining, successfully executing tasks such as sink‑to‑dishwasher transfer and cup stacking. It validates the adaptability and scalability of the pipeline and highlights its potential as a practical framework for household robotic systems. Our code and supplementary materials are available at https://raivlab.github.io/FM_kitchen .

Authors:Yang Yang, Qinyu Zhao, Mouxiang Chen, Xiaohui Li, Lixin Gu, Wenhai Wang, Hongjie Zhang, Wenwei Zhang
Title: ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs
Abstract:
Existing scaling strategies for Multimodal Large Language Models (MLLMs) typically expand either model parameters or sequential inference computation, incurring substantial memory or latency overhead. More importantly, most existing methods fail to alter the rigid, fixed computation allocation between the Vision Transformer and the Large Language Model components, limiting task‑specific optimization. To address this, we introduce the Parallel Vision‑Language (ParVL) scaling framework for MLLMs, which scales parallel computation by reusing the existing ViT and LLM backbone parameters across multiple vision and language branches. This framework raises a central question: given a fixed backbone parameter budget, how should additional shared‑backbone computation be allocated between the vision and language modalities? We instantiate each parallel computational stream with branch‑specific prefix parameters over a shared backbone, and train the entire model end‑to‑end via full‑parameter supervised fine‑tuning on roughly 13B tokens. We systematically study the computation‑allocation trade‑off between the ViT encoder and LLM decoder. ParVL improves overall multimodal performance over same‑recipe single‑branch baselines, and the best evaluated vision‑‑language allocation varies across tasks. Code is available at https://github.com/YangYangGirl/ParVL.

Authors:Zhenran Wang, Zhonghan Bian, Jinsong Li, Zhangyang Qi
Title: SocietyBench: Forecasting Counterfactual Social-World Evolution
Abstract:
Large language models (LLMs), and the agents built on top of them, are now benchmarked heavily on whether they can finish a task ‑‑ fix a bug, drive a browser, operate a GUI. A complementary social ability, namely how well a model understands and forecasts the way real social events unfold, has barely been measured. We introduce SocietyBench, an end‑to‑end benchmark that takes a one‑line event topic, collects Web news and social‑media posts across five platforms, distills them into a date‑indexed timeline that keeps factual events and a public‑opinion layer separate, and then turns every cutoff date on that timeline into an audited bank of forecasting questions. Questions are scored on two orthogonal 100‑point axes: probability calibration and temporal accuracy. Before any model sees a timeline, a three‑phase procedure replaces every named entity and shifts every date by a per‑event constant, turning a real arc into a counterfactual social world ‑‑ structurally identical to what happened, but stripped of the surface labels a model could match against pre‑training memory. On five heterogeneous events and 125 prediction points in Chinese and English editions, the strongest of six frontier LLMs reaches only 75.0 out of 100, against a trivial anchor of 50. The two axes come apart: a model can be calibration‑strong but time‑weak, or the reverse. Three agent frameworks built on a shared base model fail to improve on that base, and two model‑free heuristics trail every LLM. Per‑event gaps reach 21.4 points on a single axis, which is our main argument for evaluating on several events rather than one. All anonymized timelines, question banks, ground truth, and scoring code are released.

Authors:Zhenran Wang, Zhonghan Bian, Jinsong Li, Zhangyang Qi
Title: WorldCup Arena: Prospective, Leakage-Free Evaluation of Frontier LLMs on a Live Tournament
Abstract:
Benchmarks that measure the forecasting ability of large language models are almost always retrospective: the event has happened, the answer is somewhere on the Web, and the evaluation must defend itself against memorisation. We report the opposite design. Over the 39 days of the 2026 FIFA World Cup, six frontier LLMs ‑‑ all with extended thinking and native server‑side web search ‑‑ were asked before every kickoff, one match at a time, to fill in a seven‑market prediction card for all 104 matches, plus 12 group winners and a pre‑tournament outright pool; no answer existed when the question was asked, so the evaluation is leakage‑free by construction rather than by filtering, and the frozen archive holds 4,494 scored predictions. What the tournament establishes is a set of behaviours the six systems share. On match outcome they average 63.9%, level with backing the bookmaker's favourite ‑‑ which is in fact what they usually do. They agree with one another far more often than they are right, so a majority vote adds nothing. They under‑commit to draws and to goals, and crowd their scoreline picks onto a single prototypical result. Accuracy tracks how lopsided a fixture is rather than how much is known about it: it collapses in the closest ties, where the dossiers are richest, while questions about the tournament as a whole are answered well. On this task the current generation of frontier systems is not sharply differentiated: the standings hold up at the top and the bottom across the run and churn in the middle, and the margins stay narrow throughout. The briefing dossiers, fixtures and official results are released as a benchmark, together with the scoring code.

Authors:Changle Qu, Sunhao Dai, Hengyi Cai, Yuqi Zhou, Xinran Chen, Simon, Jun Xu
Title: TurnSight: Turn-Level Hindsight Self-Distillation for Tool-Integrated Reasoning
Abstract:
Tool‑Integrated Reasoning (TIR) enables LLMs to solve complex tasks through iterative tool interactions. However, existing reinforcement learning methods often rely on trajectory‑level supervision, limiting fine‑grained credit assignment in long‑horizon TIR scenarios. On‑policy self‑distillation offers denser signals through teacher branches with privileged context, but existing approaches typically derive such context from ground‑truth answers or retrieved skills, which may not reflect the states actually visited by the agent. Moreover, token‑level supervision fails to capture the turn‑level structure of tool interactions. To address this, we propose TurnSight, a turn‑level hindsight self‑distillation framework that derives supervision directly from execution‑conditioned hindsight. It then constructs multiple hindsight views with different lookahead horizons and selects reliable supervision through cross‑horizon directional agreement. Finally, the selected hindsight signal is normalized across sibling rollouts and used to adaptively modulate RL advantages while preserving their original optimization direction. Extensive experiments on three benchmarks demonstrate the effectiveness of TurnSight. Our codes are available at https://github.com/quchangle1/TurnSight.

Authors:Shuhan Xue, Zixin Ding, Yichen Shen, Yinjie Wang, Zhenfei Yin, Yingcheng Wu, Yuxin Chen, Mengdi Wang, Ling Yang
Title: PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents
Abstract:
Recursive self‑improvement requires agents to turn accumulated experience into better future behavior. Personal AI agents offer a concrete setting for studying this capability because they retain preferences, task histories, tool routines, and learned skills across sessions. Yet whether retained experience actually improves them over time has not been systematically tested. We introduce PAST‑Bench, a benchmark designed to isolate this question. Each agent runs through ordered sequences of fresh‑session tasks under matched conditions that turn retained experience on and off. It spans 26 scenarios and 204 episodes across memory, procedural reuse, information gathering, and update. We report both later‑task gains and whether those gains follow the intended save, retrieve, and update pathway. Across seven base models and four agent frameworks, improvement is real but uneven across capabilities. Agents with the same headline gain can differ markedly in whether that gain is supported by evidence of the intended pathway. Guided by these findings, we develop Hermes+, which extends Hermes with five targeted interventions across stages of the agent loop. Hermes+ raises the average gain from retained experience and provides clearer pathway evidence, with its strongest improvement on tasks requiring outdated state to be replaced, although the effect remains capability‑ and model‑dependent. Together, PAST‑Bench and Hermes+ provide an evaluation and diagnostic foundation for studying how persistent agents can progress from retaining experience to systematically improving through it. Code: https://github.com/Gen‑Verse/PAST‑Bench

Authors:Junhao Chen, Mingjin Chen, Jingjia Mao, Lin Chen, Saining Zhang, Minglin Chen, Ruocheng Wu, Liaoyuan Fan, Wenyi Li, Mingju Gao, Henghaofan Zhang, Zhihao Li, Hao Zhao, Yufei Wang, Ruqi Huang
Title: Agogic: Performance-Timed Music Tokens for LLM-Native Text-to-Symbolic-Music Generation
Abstract:
Text‑to‑music language models begin with a choice usually made by default: how to tokenize music. Normally entangled with backbone, data, and recipe, its effect has never been measured in isolation. We fix pretrained Qwen3.5 (0.8B‑27B), data, budget, and decoding, and swap only the representation across seven tokenizations, anchoring texture metrics to each representation's model‑free ceiling. The ordering is clean and surprising: representation, not model size, is the binding variable for distributional fidelity. Scaling the backbone 34x barely moves Frechet Music Distance (FMD), whereas switching representation halves it. PMT, a performance‑resolution stream we release (10 ms timing, per‑note velocity, multi‑track texture; 609 symbols), reaches FMD 159 at 0.8B against 272‑286 for beat grids (1.7‑1.8x lower, up to 2.8x elsewhere; non‑overlapping bootstrap CIs), so a 0.8B performance‑resolution model beats a 27B beat grid. It reappears on a 26M from‑scratch backbone and a second performance‑resolution tokenizer: a property of the class, not one lucky vocabulary. Nor is it a finer‑lattice artifact: snapping PMT's onsets to the beat grids' resolution still leaves it 67‑129 FMD ahead of both (n=500). The effect is distributional; whether it is audible is a separate question, left open by our probe, with a human study pre‑registered. Native caption adherence is weak but separable: a lightweight decode‑time constraint doubles instrument‑F1 (.28 to .60) and Correct‑Key (.16 to .35) at no distributional cost. We release the harness, 25+ checkpoints, two corpora (86.6k aligned across caption/MIDI/ABC/audio; 6.25M captioned, the largest for music), and an imprinting diagnostic: published text‑to‑MIDI systems reproduce their training distribution near‑invariant to the caption (72% vs. 71% chord‑time on disjoint domains). The field's next representation claim can now be measured, not asserted.

Authors:Zhen Fang, Yu Zeng, Wenxuan Huang, Yiming Zhao, Shiting Huang, Tianfei Ren, Qi Lu, Qingnan Ren, Qisheng Su, Lionel Z. Wang, Qingyu Yin, Shuang Chen, Zehui Chen, Lin Chen, Zhenfei Yin, Yao Hu, Shaohui Lin, Wanli Ouyang, Shaosheng Cao, Feng Zhao
Title: Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent
Abstract:
We introduce Video‑DeepResearch (Video‑DR), extending multimodal agents from static images to continuous video streams, a setting that demands dense spatiotemporal grounding coupled with open‑web exploration. Preliminary evaluations reveal two critical bottlenecks in current models: (1) modality bias, where agents bypass visual tools in favor of textual search, and (2) parametric knowledge leakage, where models rely on internal memory rather than genuine tool‑augmented execution. To address these challenges, we propose Video‑DR, featuring a decoupled perception‑exploration pipeline with stage‑wise tool unlocking that compels exhaustive cross‑frame visual grounding prior to web retrieval. Our framework adopts a two‑stage training recipe: supervised fine‑tuning followed by Group Relative Policy Optimization (GRPO), enabling autonomous exploration that breaks the imitation‑learning ceiling. Furthermore, we curate Video‑DR‑Bench, a human‑AI collaborative benchmark comprising 200 complex, multi‑hop VQA instances. Empirical results demonstrate that our Video‑DeepResearch‑35B‑A3B establishes a new state‑of‑the‑art of 64.0% average accuracy, surpassing proprietary Claude‑4.5‑Sonnet (59.0%) by 5.0 points and significantly outperforming GPT‑5 (52.5%) and Gemini 2.5 Pro (57.5%). The 30B‑A3B variant achieves 59.3%, competitive with Claude‑4.5‑Sonnet and demonstrating the effectiveness of our training paradigm even at compact scale. Code: https://github.com/Osilly/Vision‑DeepResearch.

Authors:Yicheng Xiao, Wenxun Dai, Xinran Qin, Lin Song, Maoquan Zhang, Hang Xu, Yukang Chen, Yitong Li, Guohui Zhang, Yuan Zhang, Xuying Zhang, Tommy Zhang, Jianlong Yuan, Peihao Li, Shuai Lu, Siming Fu, Chuyang Zhao, Xin Han, Jie Huang, Wenbo Li, Guoqing Ma, Wei Huang, Xiaojuan Qi, Haoyang Huang, Nan Duan
Title: JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion
Abstract:
Real‑time video editing requires low‑latency causal generation with bounded computational resources while preserving source fidelity and long‑term temporal consistency. We present JoyAI‑Video‑Edit, a 16B‑parameter autoregressive diffusion framework for real‑time, open‑ended video editing without access to future frames or a predefined video duration. Our method combines chunk‑wise autoregressive adaptation, Source‑Anchored Distribution Matching Distillation (SA‑DMD), and Long‑Horizon Autoregressive Distillation to reduce train‑‑inference mismatch, preserve source fidelity during two‑step generation, and mitigate accumulated temporal drift. Extensive automatic and human evaluations show that JoyAI‑Video‑Edit substantially outperforms existing streaming editors and remains competitive with strong offline systems on both short and long videos. The complete system achieves end‑to‑end 720p video editing at approximately 30 FPS on a single Nvidia B200 GPU. Code is available at https://github.com/jd‑opensource/JoyAI‑Video‑Edit.

Authors:Jinhe Bi, Chennan Zhou, Zengjie Jin, Aniri, Shuo Lu, Wenke Huang, Hu Cao, Xun Xiao, Zhihong Zhu, Volker Tresp, Fei Shen, Yunpu Ma, Tat-Seng Chua
Title: ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning
Abstract:
On‑policy training has emerged as a powerful post‑training paradigm for improving the reasoning capabilities of large language models, and is often enhanced by golden trajectories from stronger expert models. However, when the expert fails on harder problems, existing trajectory‑guided methods lose their main source of supervision, and these failed trajectories are typically discarded as negative samples. We argue that such failures, which we call Golden Negative Trajectories, can still provide valuable reasoning signals when treated not as demonstrations to imitate, but as flawed trajectories to reflect upon. We identify a Reflection Advantage: for hard problems, reflecting on a flawed trajectory can be easier and more effective than solving the problem directly from scratch. Motivated by this, we propose ReflectRL, a lightweight plug‑and‑play framework that learns from Golden Negative Trajectories during on‑policy training. ReflectRL first uses these trajectories to elicit Reflective Reasoning, then applies Reflective‑to‑Direct Policy Transition to transfer the acquired reasoning behavior back to Direct Reasoning. Experiments across 9 benchmarks, 4 LLM backbones, and 4 on‑policy training methods show that ReflectRL consistently improves reasoning performance with minimal overhead.

Authors:Yuanshen Guan, Zipeng Feng, Chengru Song, Zhiwei Xiong, Peiqin Sun
Title: Latent Reward Registers for Diffusion Preference Alignment
Abstract:
Aligning diffusion models with human preferences usually relies on a sparse terminal reward evaluated on the final generated samples, presenting a severe temporal credit‑assignment challenge across the multi‑step denoising process. We propose Latent Reward Registers, a mechanism that estimates terminal preference directly from intermediate noisy latents by prepending learnable, position‑free register tokens to the input sequence of a frozen Diffusion Transformer (DiT). This independent readout mechanism extracts latent reward evidence without altering the generator's hidden states or velocity field. The resulting dense, differentiable reward signal throughout the full denoising process facilitates two alignment strategies. For training, Reward‑Gradient On‑Policy Distillation (RG‑OPD) distills reward‑guided updates along on‑policy trajectories, bypassing the computationally expensive rollouts of standard policy gradients. For inference, Reward‑Guided Sampling (RGS) steers trajectories via magnitude‑matched reward gradients without parameter updates. Empirically, at high noise levels (u = 0.8), the registers reach the highest pairwise accuracy among the evaluated latent reward models. Furthermore, RG‑OPD outperforms online reinforcement learning baselines while reducing GPU hours by up to 33x, and RGS establishes a new state‑of‑the‑art among training‑free methods, strictly enhancing both alignment and perceptual metrics. Code and weights are available at https://github.com/Guanys‑dar/latent‑reward‑register

Authors:Mateusz Smendowski, Kamil Faber, Piotr Nawrocki, Nathalie Japkowicz, Roberto Corizzo
Title: PRISM: Powerful Time Series to Image (TS2I) Representations for Multivariate Anomaly Detection
Abstract:
Time series anomaly detection (TSAD) underpins applications in predictive maintenance, finance, and cloud computing, however performance remains sensitive to representation choices, especially in multivariate settings. While transforming time series into images has shown success in forecasting and classification, it remains unclear how multivariate, high‑dimensional series should be mapped to multi‑channel images and whether vision backbones can match time‑domain baselines in TSAD. We introduce PRISM, a plug‑and‑play meta‑workflow enabling systematic construction and evaluation of image‑based representations for multivariate TSAD. Our evaluation spanning over 7,000 experiments shows that well‑designed PRISM configurations are competitive with 24 time‑domain baselines, achieving the best VUS‑PR on 10 of 14 datasets, with an average improvement of 41% over the best competing method on those datasets. Further, we identify channelization ‑ how the channel dimension of multi‑channel images is constructed ‑ as a critical and previously understudied design dimension, and introduce MSM, a novel statistics‑based scheme achieving 11‑27% gains over PCA‑based alternatives. Finally, ImageNet‑pretrained encoders transfer effectively to TSAD, with frozen encoders retaining 92% of fine‑tuned performance while training 1.8 times faster. Our code is available at: https://github.com/Smendowski/PRISM.

Authors:Lu Gan, Hanyu Yan, Chaofeng Chen, Junqi Hu, Dan Zeng
Title: GeoMAR: Unleashing Geometrically Aligned Features for Masked Autoregressive Blind Face Restoration
Abstract:
Codebook‑based blind face restoration (BFR) often suffers from ambiguous conditioning features and a fragile prediction mechanism under severe degradation. To address these challenges, we propose GeoMAR, a framework designed to unleash geometrically aligned features with masked autoregressive (MAR) refinement for robust face restoration. For feature conditioning, we introduce a dual‑input extraction pipeline to extract component‑based geometric descriptions with explicit, spatially faithful anchors. These textual priors are integrated with low‑quality (LQ) features via an Aligned Geometric Priors Injector, which employs a KV‑Q exchange strategy to generate geometrically aligned features. For prediction mechanism, we reformulate the one‑step mapping into a multi‑step MAR process. This coarse‑to‑fine generation progressively refines complex facial regions based on increasingly reliable context. Experiments on one synthetic and three real‑world benchmarks demonstrate that GeoMAR achieves highly competitive perceptual quality and coherent visual structures compared with existing methods. The code is available at https://github.com/BRL‑SYSU/GeoMAR.git.

Authors:Ke Li, Jiayu Chen, Maoliang Li, Zihao Zheng, Hailong Zou, Hengyi Zhang, Xuanzhe Liu, Xiang Chen
Title: When and Where to Look: Adaptive Visual Evidence Scheduling for Efficient Long Video Understanding
Abstract:
Efficient long‑video understanding requires vision‑‑language models (VLMs) to reason over a small number of frames selected as sparse visual evidence. Existing relevance‑based methods rely on static one‑shot selection with fixed frame budgets and candidate pools, while agent‑based schedulers achieve adaptivity through costly multi‑round reasoning and interactive search. We propose EcoFrame, a training‑free framework for low‑overhead query‑adaptive visual evidence scheduling. EcoFrame leverages the VLM's inference feedback to determine when to increase the frame budget and where to search for additional candidate evidence. Specifically, entropy‑gated budget scheduling uses output uncertainty to stop early when the current evidence is sufficient or progressively expand the frame budget otherwise. Meanwhile, attention‑guided candidate proposal converts frame‑level attention into a temporal prior, enabling dense local search in informative regions while preserving global coverage when attention is diffuse. Experiments on Video‑MME, LongVideoBench, and MLVU demonstrate that EcoFrame achieves a better accuracy‑‑efficiency trade‑off across multiple VLM backbones. On Qwen2.5‑VL, EcoFrame achieves an average accuracy of 64.4, surpassing BOLT at 63.5, while providing a 1.85× speedup over AKS and BOLT. Compared with the agent‑based A.I.R., EcoFrame maintains comparable accuracy with up to a 13.5× inference speedup. Code will be available at https://github.com/AK‑DREAM/EcoFrame.

Authors:Alberto Acedo
Title: Omega-S: A Functional Resilience Index for LLM Fine-Tuning
Abstract:
Fine‑tuning a large language model on new data degrades what it previously learned. We present Omega‑S, a drop‑in penalty computed from the weight matrix alone: it needs no previous‑task data, no Fisher matrix and no stored copy of the old weights. It is three lines in an existing training loop and adds under 4% to the cost of a step. Retention. On Llama‑3‑8B with LoRA, fine‑tuned from code to prose and measured by HumanEval over ten seeds, Omega‑S retains more of the original capability than no regularisation on 9 of 10 seeds (0.173 ‑> 0.238 absolute pass@1; sign test one‑sided p=0.011, Wilcoxon p=0.006), as a retention ratio, 62.9% ‑> 84.1%. It also beats tuned weight decay on 10 of 10 seeds (p=0.002) and tuned EWC on 8 of 10 (p=0.014), every arm re‑measured in the same session. Mechanism, measured rather than asserted. Omega‑S is topological by construction, its objective built from Tr(A^3), but we measured which of its four factors actually moves and three do not: their elasticity with respect to the weights is at or below 1e‑4, against 9e‑3 for the degree‑variance term. As implemented, the composite reduces to a penalty on the variance of node degrees, which means row magnitude in square modules and directional alignment in non‑square ones. We report this because a method whose name promises one thing and whose gradient does another should say so. We also enumerate the open design choices, including a contrast‑preserving construction that does what it was designed to do and makes retention worse on all ten seeds. Repeating an identical configuration, same seed and same hardware, gives a standard deviation of 0.104 in retention ratio. We have not found this quantified for low‑rank fine‑tuning of language models, and it bounds every seed‑paired comparison in this literature, ours included. Code, per‑seed results and the full record of negative results are available.

Authors:Shuoqin Zhang, Tongtong Cheng, Xiru Gao, Jinzhuo Peng, Bin Zheng, Jiahao Tu, Ke Wang, Jia Pan, Zhe Hu, Kai Liu
Title: EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning
Abstract:
Human‑in‑the‑loop reinforcement learning (HIL‑RL) enables robots to learn contact‑rich manipulation from limited real‑world interaction, but deployment exposes three coupled limitations: static visual reward models fail under scene changes; independently sampled actions cause temporally inconsistent motion; and vision‑based policies remain sensitive to appearance shifts. We present EvoHIL, a unified framework that adapts the reward model, action generator, and visual do main within a staged human‑in‑the‑loop learning process. First, self‑evolving reward (SER) adapts the success classifier from human‑confirmed positives and provisional weak negatives. Second, Action Flow Stabilization (AFS) generates temporally coherent action chunks through flow matching, grounding policy updates in executed action prefixes and demonstrated behavior. Third, retention‑aware offline fine‑tuning replays relit interaction data while anchoring the AFS actor‑critic to prior behavior, adapting the visual domain without additional robot interaction. Across six manipulation tasks on Franka FR3 and SO‑101 arms under a controlled lighting shift, EvoHIL improves task success, agreement with human‑confirmation labels, motion smoothness, and completion time relative to human‑in‑the‑loop and imitation baselines.Project page: https://anonymous4366.github.io/EvoHIL/

Authors:Jiaming Chen, Yisen Gao, Yanping Li, Zifan Liu, Yumeng Zhang, Jun Zhang
Title: MAFIA: Query-Only Memory Attacks via Probing and Factual Injection against Audited LLM Agents
Abstract:
Memory‑augmented LLM agents rely on rich context for long‑horizon reasoning and acting, yet their memory modules expose a persistent attack surface for malicious records, making the study of memory poisoning threats imperative. However, existing query‑only attacks often fail to remain effective in two realistic and prevalent settings: large‑scale benign memory pools and active input auditing. Consequently, current approaches fall short when facing the dual challenges of high retrieval competitiveness and rigorous semantic checks. To overcome these limitations, we propose MAFIA, a query‑only Memory Attack framework via probing and Factual Injection against Audit, tailored to this extended threat model. Specifically, MAFIA introduces: (1) a placement strategy that ensures retrieval‑competitive injection via memory probing, budget allocation, and scheduling; and (2) a payload design that bypasses audits using compact factual cloaks, preserving malicious effects while maintaining high semantic similarity. Extensive evaluations reveal that MAFIA achieves up to a 90.7% attack success rate while suppressing audit detection from a peak of 83.3% to at most 7.4%, exposing critical vulnerabilities across agentic memory systems. Code will be made publicly available at https://github.com/JiamingChen1234/MAFIA.

Authors:Amir Mohammad Ezzati, Kiyan Rezaee, Bardiya Kariminia, Mohamad Amin Yousefi, Asal Mohammadjafari Mamaqani, Behrad Samimi, Mohammad Hossein Rohban
Title: UHP Detection: LVLMs have their Unique Hallucination Pattern in the Consistency Space
Abstract:
Large vision‑‑language models (LVLMs) demonstrate strong multimodal reasoning capabilities but remain prone to hallucination, where model predictions are not grounded in visual evidence. Existing black‑box hallucination detection methods estimate uncertainty through a single consistency metric, implicitly assuming that model uncertainty can be adequately characterized by a single measure. However, hallucinations exhibit diverse manifestations of uncertainty across different behavioral probes, making a single measure insufficient to characterize their underlying behavior. We propose \emphUnique Hallucination Pattern (UHP) Detection, a fully black‑box framework that models hallucination as a structured uncertainty pattern defined by two axes: perturbation modality (image vs.\ text) and logical polarity (a statement vs.\ its negation). Their intersection produces four complementary consistency groups that capture distinct manifestations of model uncertainty, from which both within‑group and between‑group features are extracted to train a lightweight classifier. Through comprehensive experiments on AMBER and PhD across three LVLMs, UHP Detection consistently outperforms prior black‑box and white‑box baselines, with improvements of up to +18.72% AUC‑ROC and +20.07% AUC‑PR over the strongest black‑box methods. Extensive ablation studies demonstrate that each consistency group contributes complementary information and that their combination forms a structured hallucination pattern. Furthermore, cross‑dataset evaluation shows that this learned pattern generalizes across benchmarks, indicating that hallucination behavior reflects a model‑specific consistency pattern. Code is publicly available at https://github.com/amirezzati/uhpdet.

Authors:Bakbergen Ryskulov, Iker García-Ferrero, David Montero, David Jansen, Ali Hashemi, Jezabel R. Garcia, Antonio Tiene, Román Orús
Title: Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss
Abstract:
Small language models are often the only option for deployment under tight latency, cost, and on‑premises constraints, but they are rarely trained from scratch: a compressed model is usually recovered through knowledge distillation (KD). This recovery step largely decides the final quality, yet it is expensive. We present a practitioner's study of how to make distillation training efficient, organised around two systems contributions. First, we show that offline KD (caching the teacher's top‑K logits once and training the student against the cache) matches online distillation at near‑identical training loss while removing the teacher from memory, running about 29% faster per iteration, and reaching up to 41% higher throughput on a single H200 GPU. Second, we introduce a \emphfused, chunked KL loss that never materialises the full vocabulary‑sized logit tensor, making peak memory linear in the sequence length. This removes the memory spike that otherwise caps context length and lets us train at four times the context (32,768 tokens) on a single GPU. A separate output‑head‑only toy benchmark isolates the loss kernel and confirms its memory and iteration‑rate scaling from 4K to 256K tokens. Together these make large‑scale healing and hundreds of ablations affordable. We also report supporting ablations on loss design and sequence packing. We release our chunked‑loss implementation: https://github.com/CompactifAI/Full‑Chunked‑KL‑Loss.

Authors:Leijun Zhou, Zhihao Liu, Xiang Qu, Chenxu Liu, Yifei Liu, Yanke Yu, Jingzhe Xu, Xuejun Wu, Buyue Qian, Xi Chen, Yaowei Zheng, Junhao Hu
Title: GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks
Abstract:
Agent self‑evolution updates an agent's persistent state from prior experience and reuses it to solve related tasks more effectively. Evaluating self‑evolution is difficult: existing benchmarks provide limited coverage of economically valuable task domains, do not always design training and test tasks such that test‑time gains can be attributed to training experience, and remain vulnerable to data contamination. We present GDPevo, an evolution‑native benchmark grounded in GDP‑related enterprise workflows, together with the fully automated data pipeline that generates it. Its core mechanism, rule hybridization, decomposes each enterprise workflow into atomic business rules, distributes subsets of these rules across training tasks, and recombines them in held‑out test tasks so that test‑time gains are attributable. GDPevo spans CRM, ERP, finance, healthcare, legal, and data‑centric workflows. Its V1 release contains 120 tasks in 12 groups, with five training and five held‑out test tasks per group. Full automation enables the pipeline to expand the suite to 240 tasks in 24 groups (V2) within two days, providing a practical response to contamination. Using GDPevo, we evaluate four agents, each comprising a harness and a model, under four supervision types. Self‑evolution consistently improves held‑out accuracy by up to 16.44 percentage points. But the best evolved agents remain far below the fully informed oracle ceiling of 91.6%, indicating that the self‑evolution ability of current agents remains far from fully realized. We publicly release the pipeline, benchmark, and full evaluation results at https://github.com/Prism‑Shadow/GDPevo.

Authors:Andrea Protopapa, Davide Buoso, Francesca Pistilli, Georgia Chalvatzaki, Giuseppe Averta
Title: GORDON: Graph-based Object-centric Rewards for Decomposition of Long-Horizon Manipulation
Abstract:
Learning long‑horizon manipulation skills with reinforcement learning remains challenging due to the complexity of reward design, the limited guidance of sparse rewards, and the high cost of manual subtask annotation. Visual demonstrations can provide supervision for reward learning, but rewards learned from raw pixels can be brittle and sensitive to visual variation, background appearance, and robot motion. In this work, we propose GORDON, a graph‑based object‑centric reward learning framework that learns dense rewards from action‑free video demonstrations. Each visual scene is represented as a graph of detected objects and spatial relations, and a graph neural network is trained in a self‑supervised manner to embed these graphs into a task‑aligned latent space. To align the representation with semantic task progress, we introduce an activity‑aware weighted pooling mechanism that emphasizes task‑relevant objects while masking robot‑dominated motion. The dense reward is then computed as distances in the learned latent space of the current state to demonstrated goal configurations, providing a measure of task progress. In long‑horizon tasks, the temporal profile of this reward reveals stage‑wise object‑state transitions, enabling automatic subtask discovery without manual segmentation. The discovered segments are then used to train subtask‑specific rewards and specialized policies that are composed sequentially. Experiments on seven manipulation tasks on MAGICAL and ManiSkill3 benchmarks show that our object‑centric reward improves reinforcement learning in short‑horizon settings and enables successful policy learning in complex long‑horizon tasks through automatic decomposition, achieving an average success rate of 74.4% across the long‑horizon tasks (on average approximately +35 p.p. vs. best learned baseline and approximately +25 p.p. vs. oracle).

Authors:Sebastián Andrés Cajas Ordóñez, Agastya Munnangi, Aldo Marzullo, Felipe Ocampo Osorio, Quang Bui, Mohammad Shahin, Armaan Grewal, Emmanuel Paul Kwesiga, Anqi Peter Li, Josephine Nanyonjo, Aaditya Panchal, Arshnoor Bhutani, Nikhil Jaiswal, Milit S. Patel, Maximin Lange, Leo Anthony Celi
Title: Agents Catching Agents: Shortcut Cascades and Benchmark Gaming in Clinical Multi-Agent Systems
Abstract:
Clinical decision support is moving toward committees of language‑model agents deliberating on a shared workspace. We ask whether such committees can be gamed by shortcuts, cues a benchmark rewards but a clinician would ignore. Across seven cohorts on six public datasets spanning text (MedQA‑USMLE, MedMCQA, MIMIC‑CXR reports), imaging (NIH ChestX‑ray14, MIMIC‑CXR‑JPG, CheXpert) and tabular ICU records (SUPPORT2), Gemini committees resist these cues in isolation (flip 5‑16%), yet a socially plausible shortcut spreads: when two peers assert the same wrong answer, the holdout under test adopts it in 38% of cases, as does a false "pre‑screen" system flag, on both capability tiers. Of three oversight agents, a gate cannot separate adoption from honest agreement (false‑positive rate 100%); a same‑lineage judge reading only the transcript flags adoption on text (precision 100%, recall 93%) but collapses onto the gate in imaging; a referee that privately re‑queries the holdout transfers to imaging (77‑88% precision, 13‑21% false‑positive rate). Tripling a cue's visual salience does not move contagion, whereas a second peer voice raises it by half again. Gaming a hidden rubric is near‑silent: only 1/10 text and 1/134 imaging drifters name the rubric they moved toward. What games a committee is social plausibility, and only a referee independent of self‑report catches it. Code: https://github.com/criticaldata/benchmaxxing

Authors:Chao Peng, Ruida Hu, Ajitha Rajan, Tegawendé F Bissyandé, Jacques Klein, Cuiyun Gao
Title: Can LLMs Test Terminal User Interfaces?
Abstract:
Terminal User Interfaces (TUIs) combine the stateful, screen‑oriented behaviour of GUIs with terminal deployment and are now common in developer tools. Yet they lack a dedicated testing methodology. We survey 197 real‑world TUI applications: only 12% of test code exercises the interface, and 45% of those tests never send input, checking a static frame instead. We turn these applications into a headless benchmark spanning ratatui/Rust, bubbletea/Go, textual/Python, and ink/TypeScript, packaging each as an instrumented Docker image. We record line and widget coverage where reliable, rendered terminal states, and crashes. Under equal wall‑clock budgets, we compare four frontier LLMs with random exploration. No model dominates. Random is a strong time‑budgeted baseline, but its crash advantage comes from higher throughput: per interaction, LLM guidance is more efficient and uniquely reaches input‑gated faults. Automatically deriving launch inputs yields the largest practical gain, enabling applications that otherwise never start. Line coverage poorly predicts crash discovery, weakening it as a proxy for test effectiveness. Automated TUI testing is feasible but far from solved, and honest baselines matter more than model choice. We release the coverage tool tuicov at https://github.com/tui‑testing/tuicov and the testing framework tuibot at https://github.com/tui‑testing/tuibot.

Authors:Longji He, Jeto Xu
Title: SAT-Edge-Agent: Hardware-in-the-Loop Edge-Agent Orchestration for Onboard Satellite Intelligence
Abstract:
Onboard satellite intelligence requires a task layer that translates mission intent into local tool calls, exposes execution state, and returns machine‑consumable artifacts under communication and power constraints. We present SAT‑Edge‑Agent, a hardware‑in‑the‑loop (HIL) edge‑agent system deployed on a commercial off‑the‑shelf ARM‑based heterogeneous edge system‑on‑chip. A browser workspace and FastAPI agent coordinate a local OpenAI‑compatible language service with a project‑internal YOLO‑style oriented‑object‑detection endpoint that returns FAIR1M metadata‑backed structured results. Two fixed FAIR1M workloads, one single‑image and one serial two‑image request, were repeated 20 times each and completed 20/20 attempts. Mean Full‑Agent latency was 29.353 s and 60.937 s, with empirical P95 values of 31.166 s and 66.882 s. Mean detector time was 861.386 ms and 1510.920 ms, only 2.93% and 2.48% of the corresponding Full‑Agent means. Profiling indicates that most visible latency occurs outside detector execution. Mean CPU utilization was 20.761% and 20.482%. A 200‑ms NPU‑load field averaged 100% for both workloads, but it represents a shared‑accelerator software field rather than detector‑only occupancy or calibrated utilization. The public evidence package provides sanitized request‑level records, redacted JSON, normalized SSE examples, and scripts reproducing the reported statistics. These results establish a reproducible HIL boundary for observable satellite edge‑agent orchestration, but do not establish detector accuracy, a new geolocation method, calibrated energy efficiency, or flight readiness.

Authors:Chenyi Wang, Xinkai Wang, Bokai Lin, Jialin Tian, Fucheng Zhang, Cewu Lu, Lixin Yang
Title: Track4Action: Distilling World-Centric 3D Tracker into Vision-Language-Action Policies
Abstract:
Action labels tell a vision‑language‑action (VLA) policy which robot commands to imitate, but not how those commands change the 3D world. The aligned demonstration clip contains this missing supervision because its K frame transitions record the geometry, motion, visibility, and camera change produced during the corresponding K actions. We introduce Track4Action, a framework that distills this realized transition from a frozen world‑centric 3D tracker into a current‑observation VLA policy. During training, Track4World encodes the clip V_t:t+K into a pooled tracker feature. Learnable track queries infer this feature from current VLA hidden states, match it in a shared space, and condition a flow‑matching action head through a feature‑wise gate. The tracker feature only defines the alignment target, so neither the clip nor the tracker is used at deployment. Track4Action reaches 82.3% on zero‑shot LIBERO‑Plus, improving the alignment‑free variant by 7.6 points and LaMP by 3.0 points. It obtains 80.44% and 81.48% on the clean and randomized RoboTwin 2.0 splits, and 67.5% average success across four physical bimanual tasks, 25.0 points above the alignment‑free variant. The gains across simulation and physical tasks support action‑aligned 3D tracker features as privileged supervision for tracker‑free VLA deployment. Our project page is available at https://wing0night.github.io/track4action‑project‑page.

Authors:Ruirui Zhang, Zhengkai Zhao, Pan Gao
Title: MultiCompose: Multi-Concept Personalized Composition with Per-Subject Attribute Binding
Abstract:
Text‑to‑image diffusion models enable personalization of specific visual concepts from a small number of reference images. However, generating a single image that contains multiple personalized subjects, each bound to user‑specified attributes such as clothing, accessories, and held objects, remains largely unaddressed. Without explicit spatial constraints, concurrently activated concept checkpoints produce overlapping cross‑attention responses, causing per‑subject identity degradation and attribute misalignment. Moreover, no established benchmark jointly evaluates these two failure modes in the personalized multi‑subject setting. We present MultiCompose, a composition framework that decouples per‑concept personalization from multi‑subject inference. A semantic preservation regularization maintains attribute binding capacity during fine‑tuning, while a two‑phase inference procedure automatically establishes subject layout and composes per‑concept predictions through spatially exclusive masks. We further introduce MSP‑Bench, a benchmark that jointly evaluates identity fidelity (ID), attribute binding accuracy (BIND), and attribute misalignment (MIS) through a dual‑pathway protocol. Experiments show that MultiCompose outperforms existing methods on both conventional metrics and MSP‑Bench, confirming the benchmark's ability to reveal failure modes that conventional metrics overlook. Code is available at https://github.com/I2‑Multimedia‑Lab/MultiCompose

Authors:Yongli Xiang, Zhifang Zhang, Bojun Yang, Ziming Hong, Lei Feng, Miao Xu, Tongliang Liu
Title: When Agents Learn to Be You: Benchmarking Privacy Leakage, Impersonation Risk, and Defenses in Persona Skills
Abstract:
Persona skills distill personal interaction histories into portable and executable artifacts for downstream agents. While enabling flexible personalization, this process concentrates fragmented personal signals, amplifies their impact through reuse, and challenges defenses designed for individual records or retrieval‑based memory. To systematically investigate the safety of the persona‑skill pipeline, we introduce AntiSkillBench, an end‑to‑end benchmark for evaluating risks and defenses across the persona‑skill pipeline. It comprises: (i) a dataset of 7,500 persona‑grounded dialogue traces, constructed from 50 behaviorally rich profiles spanning diverse task scenarios; (ii) an evaluation suite that measures skill‑level privacy leakage and agent‑level attribute disclosure and behavioral impersonation across three skill‑distillation strategies; and (iii) a defense evaluation covering four configurations across online and post‑hoc interventions, including active risk suppression and passive provenance protection. Experiments across three frontier agents show that persona‑skill risks persist across agent backbones and distillation protocols, extending from explicit attributes to communication styles and personality traits. Existing defenses exhibit limited and distillation‑dependent effectiveness, failing to generalize across risk and distillation strategies. These results highlight AntiSkillBench as a challenging benchmark for developing privacy‑preserving and authenticity‑aware persona skills.

Authors:Yiyao Wang, Zhen Wen, Yinghao Tang, Yixiao Fu, Lin Yuan, Xiaolau Zhang, Jun Zhou, Wei Chen
Title: LiveEvalBench: Toward Open-World Evaluation for Web Generation
Abstract:
Large language models are increasingly capable of synthesizing executable frontend projects, yet existing benchmarks still treat web generation as a static evaluation problem. We argue that frontend artifacts demand a different paradigm: they are interactive rather than static, admit diverse yet equally valid implementations, and evolve faster than rigid pipelines can accommodate. To address these gaps, we present LiveEvalBench, an automated framework that reformulates web‑generation evaluation as an agentic, adaptive, and extensible process. LiveEvalBench instantiates evaluation as a collaborative review workflow, in which a Build Engineer, a Code Engineer, and a UI Tester collectively gather evidence across the full lifecycle of a frontend project, from deployment and code inspection to browser‑based interaction. To handle implementation diversity, an adaptive protocol couples shared rubrics for cross‑model comparability with implementation‑grounded criteria tailored to each artifact. The framework further supports incremental integration of new evaluator roles and assessment dimensions without pipeline redesign. Experiments across diverse real‑world web‑generation scenarios show that LiveEvalBench aligns closely with human expert judgment and provides fine‑grained insights into frontier models' web generation capabilities. Code is available at https://github.com/wyysteelhead/LiveEvalBench

Authors:Chenghua Wang, Daliang Xu, Dongqi Cai, Duojin Sun, Hao Zhang, Haoze Qian, Huaiyuan Zhang, Jinshuo Cui, Junbo Cui, Kezhao Zhao, Longxi Gao, Mengwei Xu, Rongjie Yi, Tam Sikyuen, Tianyue Zhang, Weikai Xie, Xuanzhe Liu, Yingying Qin, Yiwen Lu, Yuan Yao, Yuezhi Zu, Yunhan Guo, Yuxin Zheng, Ziqi Guo
Title: PhyAI: Real-Time Physical AI at the Edge, Scalable Rollouts in the Cloud
Abstract:
Physical AI policies require inference throughout their lifecycle, including model evaluation, cloud reinforcement learning rollout, edge GPU serving, and onboard deployment. Although these settings share the same checkpoint and action semantics, they often rely on separate inference programs. To unify them, we build PhyAI, a Physical AI inference engine with a single runtime that keeps architecture‑specific conditioning, solver, cache, and output logic in model adapters while sharing graph execution, kernels, memory management, and parallel services. The same codebase runs vision‑language‑action (VLA) models and world‑action models (WAMs) on single or multiple GPUs across onboard, edge, and cloud deployments. We used the adapter interface to add MiniCPM‑Robot on the day of its release. PhyAI achieves 1.40x‑4.65x speedups over the official implementations of pi0, pi0.5, GR00T N1.7, and MiniCPM‑Robot. On Cosmos3‑Nano‑Policy‑DROID it reduces latency from 2.46 to 1.18 s on eight H20 GPUs (CFG=2, TP=4), a 2.08x speedup. Specialized runtimes remain faster in several configurations, so our goal is one runtime with competitive latency rather than the fastest result in every case. Detailed profiles reveal why different models need different execution policies: on a Hopper‑series GPU at batch size one, the pi0.5 action expert accounts for 8.8% of FLOPs but 57.2% of latency; at batch size 32 its share drops to 13.5% and throughput reaches about 100 samples/s. Cosmos3 remains generation‑dominated and gains only 14.3% throughput as batch size increases from 1 to 16. We further introduce the control‑time Roofline, which distinguishes inference‑bound from environment‑bound control; the measured pi0.5 points on four LIBERO suites are environment‑bound while Cosmos3 stays inference‑bound. Code and benchmarks: https://github.com/mingti‑org/phyai.

Authors:Zhijing Hu, Changjun Fan, Yufan Deng, Zhiguang Cao
Title: AutoSND: From Execution Evidence to Structural Policies for Automated Network Dismantling Heuristic Discovery
Abstract:
Network dismantling is fundamental to analyzing the robustness and vulnerability of complex systems, yet practical heuristics must balance effectiveness and computational efficiency, and are usually designed manually by researchers. Existing large language model based automatic heuristic design methods can generate and screen candidates, yet they have difficulty further transforming candidate quality or failure states during execution into structural‑level guid‑ ance for subsequent generation. We propose AutoSND, a three stage tree search framework for complete network dismantling pro‑ grams. Stage I broadly explores from simple heuristics and archives execution evidence. Stage II compiles candidate records into struc‑ tural policies concerning local signals, neighborhood access, and state update ranges. Stage III continues tree search conditioned on these policies and obtains the final quality prioritized and speed prioritized candidates, AutoSND‑Q/S. Experiments on 12 real world networks and 3 large real world networks show that AutoSND achieves better search performance and stability and discovers more competitive and structurally interpretable network disman‑ tling programs. The final candidates form an interpretable structure that uses residual degree as the backbone, adjusts node order with bounded local signals, and restricts the state update range. Code is available at https://github.com/MirrorNew/AutoSND.

Authors:Feixiang Liu, Likun Wang, Qiang Qiu, Hui Xu, Huawei Shen, Xueqi Cheng
Title: SEER: A Self-Grounded Evidence Interface for Controlled Spatial Relation Classification
Abstract:
Spatial relation questions require a model to identify the queried subject and object before comparing their layout. Yet a VLM can recognize both entities and still answer from the wrong instance or an ambiguous global view. We ask whether making query‑specific evidence explicit can mitigate this failure and propose SEER (Self‑grounded Evidence for Entity‑Relation Reasoning), a training‑free inference‑time evidence interface for frozen VLMs. SEER hides candidate relations during pair localization, constructs a query‑specific view with explicit subject/object roles, and retains the full image and sparse box geometry as complementary evidence. For relation‑choice protocols with exact inverse support, an optional refinement swaps the entity roles and changes the forward decision only when exactly one visual state obeys the corresponding inverse relation. On an image‑disjoint GQA‑Train900 test frozen before model scoring, SEER pools to +3.94 [2.17,5.72] over Full; the gain remains positive under label‑independent grounding‑order counterbalancing and on the 535 rows whose entity names are unique. The unchanged protocol yields +4.35 to +11.79 on all 2,434 filtered EmbSpatial pair‑relation questions across three models. Matched controls separate local refocus from role‑explicit conditioning. These results establish query‑specific evidence construction as the principal intervention, with reciprocal consistency as a smaller protocol‑specific refinement.

Authors:Vladimir Beskorovainyi
Title: A machine-readable catalogue of the Tsiolkovsky papers (fond 555, Archive of the Russian Academy of Sciences), and a way to measure how well its handwriting can be read
Abstract:
The personal archive of Konstantin Tsiolkovsky (1857‑1935) is held as fond 555 of the Archive of the Russian Academy of Sciences. The archive scanned the fond and published the images, but with no queryable catalogue, no full‑text search and no dataset: the holdings can only be browsed one page at a time. This paper describes a machine‑readable catalogue of all 2,019 files and 51,008 scans, a dating for 1,969 files taken from the archive's own descriptions, a page‑level classification of every scan into handwriting and typescript, and a growing corpus of machine transcriptions (currently 322 files, 5,454 scans). It also reports a way to measure handwritten‑text‑recognition accuracy in an archive with no ground truth. Archives of the typewriter era often preserve one text twice, as manuscript and as a typed copy; transcribing both and comparing isolates the reading error, since source and pipeline are identical and only page difficulty differs. Across 294 such pairs from 27 files, two readings of a handwritten page agree on a median 37% of words. On two files that also have a published edition the estimate can be checked against ground truth: it is unbiased to within a percentage point and ranks pages as the truth does (rank correlation 0.92 where the edition is a faithful witness). This bounds use: two variants of one work here share 19% of words, below the rate at which two readings of a single page agree, so the redactions cannot be collated word by word at this quality. That negative result is reported as such, and the constraint is built into the tool.

Authors:Qin Lei, Hao Wu
Title: Predictive Enhancement Calibration for Latent Breast MRI Virtual Contrast Enhancement
Abstract:
Virtual contrast enhancement (VCE) synthesizes enhanced breast MR images from pre‑contrast acquisitions. Modern latent generators offer strong image priors, but their bounded natural‑image autoencoders conflict with the non‑canonical intensity scale of MRI. We show that the upper bound can alter radiomic fidelity before generation, while scaling source and target independently creates a coordinate inconsistency. We propose Predictive Enhancement Calibration (PEC), which represents each pair in a shared, case‑adaptive coordinate during training and predicts its unavailable upper endpoint from the pre‑contrast image at inference. We integrate PEC with a pretrained FLUX latent flow transformer via parameter‑efficient reference conditioning. Target round trips first isolate representation loss before generation; near‑matched conditional models then compare PEC with fixed‑wide and separate coordinates under comparable training budgets and backbone settings. On the fixed internal MAMA100 development cohort, PEC improves all eight point estimates in this source‑only VCE setting, with paired evidence strongest for MSE and LPIPS.\noindentCode: https://github.com/tanlei0/pec‑breast‑mri‑vce

Authors:Kejian Zhu, Zhuoran Jin, Shangqing Tu, Hongbang Yuan, Yushi Bai, Kang Liu, Juanzi Li, Jun Zhao
Title: SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs
Abstract:
Supervised Fine‑Tuning (SFT) and Reinforcement Learning (RL) exhibit fundamentally different behaviors in enhancing multi‑task reasoning for large language models (LLMs). Our preliminary experiments revealed a phenomenon: SFT suffers from severe task conflicts under multi‑stage training, whereas RL enables stable coexistence across diverse tasks. Empirically, we trace this to the parameter level, observing that RL induces sparse and approximately orthogonal updates across tasks. We provide a theoretical explanation for this mechanism by analyzing multi‑task gradient interference. Our results reveal a distinction: interference in SFT is norm‑limited, scaling with the absolute gradient magnitude, whereas interference in RL is variance‑limited, bounded by the gradient variance induced by advantage normalization and on‑policy optimization. This small variance bound yields near‑orthogonal optimization directions across tasks. Leveraging this insight, we propose Parallel‑RL, a paradigm that decouples multi‑task training, significantly improving efficiency and flexibility.

Authors:Kejian Zhu, Zhuoran Jin, Dongqi Huang, Hongbang Yuan, Yupu Hao, Kang Liu, Jun Zhao
Title: Beyond Simply Environment Scaling: Designing Effective Environment Distributions for Multimodal Agent Learning
Abstract:
Recent works train agents by constructing large‑scale multimodal environment pools. However, we find that simply increasing the number of multimodal environments does not always benefit. We further analyze the limitations in current multimodal environment distributions through a series of experiments. Based on these findings, we study how to build more effective training environment distributions from two dimensions: diversity and difficulty structure. For diversity, we propose Ability‑aware Environment Selection (AES) to obtain diverse environment sets. For difficulty structure, we propose Hierarchical Difficulty Curriculum (HDC), which organizes curriculum learning through two difficulty levels: harness weakening and state‑scale progression. Experiments show that AES and HDC effectively improve multimodal agent training.

Authors:Hui Liu, Chen Jia, Fan Shi, Xu Cheng, Mianzhao Wang, Shengyong Chen
Title: Compass: Degradation-Simulated Reciprocal Learning with Lightweight Needle RWKV for Multimodal Crack Segmentation under Missing Modalities
Abstract:
In multimodal crack segmentation for industrial facilities, the key challenge is preventing missing modalities from degrading pixel‑level performance while maintaining low computational cost. Existing methods struggle to address semantic degradation caused by missing modalities. We propose Compass, a lightweight network for robust crack segmentation under arbitrary missing modalities. Compass comprises Degradation Simulation Distillation (DSD), Needle Block, and Evidential Topology‑Preserving Fusion (ETPF). DSD constructs a degradation simulation stream that mimics more severe missing conditions and performs reciprocal distillation with the original stream, decoupling complete perception from degradation adaptation. Within DSD, Feature‑Aware Prototype Transmitter (FAPT) performs modality agnostic prototype‑guided feature completion to maintain semantic integrity under incomplete modality conditions. As a lightweight backbone, Needle injects crack‑direction cues into WKV modulation and combines connectivity‑aware gating with anisotropic context probing for structure‑aware modeling. ETPF fuses multimodal features via Dempster‑Shafer evidential combination with uncertainty‑gated decoding, preserving crack topology while suppressing unreliable features. Experiments on three datasets demonstrate state‑of‑the‑art (SOTA) performance under diverse missing modality scenarios. Even with 90% depth modality missing on CrackDepth, Compass achieves F1 of 0.8216 and mIoU of 0.8434 with only 2.58M parameters. The code is available at https://github.com/Karl1109/Compass.

Authors:Zonghao Ying, Xiangfan Wu, Huiyu Wu, Xing Zheng, Huangsheng Cheng, Xiaorong Shi, Jing Guo
Title: SkillJack: Persistent Skill Backdoors in Self-Evolving Agents
Abstract:
Self‑evolving agents increasingly convert interaction histories into reusable skills that persist beyond individual tasks. While prior work studies memory and retrieval poisoning, such attacks only affect agents when poisoned records are retrieved as context. We uncover a new and more fundamental risk: poisoned experiences can be transformed by the agent itself into durable behavioral artifacts. We present SkillJack, the first attack that exploits the experience‑to‑skill pipeline of self‑evolving agents. Instead of directly manipulating runtime context, SkillJack hijacks the agent's own learning process to implant malicious behaviors into its reusable skill repertoire. We identify three key properties of this transformation: \emphsanitization whitewashing, where malicious intent is obscured during skill extraction; \emphcross‑layer promotion, where transient experiences become persistent capabilities; and \emphpersistence isolation, where the attack survives removal of its original source records. We evaluate SkillJack on two representative systems, SkillX and Anything2Skill, using a shared dataset of 150 trajectories across four policy‑risk categories. Results show that skill extraction substantially reduces attack detectability: in SkillX, safety detection drops from 98.5% for poisoned trajectories to 11.4% for extracted skills, while Anything2Skill shows a similar effect. Meanwhile, the implanted skills remain effective, achieving attack success rates of 56.2% and 89.2% on the two systems, respectively. Furthermore, 80.0% of skill‑mediated attacks persist after deleting the original poisoned records, and some skills unintentionally activate on benign queries. Our findings reveal skill evolution as a new attack surface and motivate provenance‑aware skill lifecycle protection. Our code is available at https://github.com/Tencent/AI‑Infra‑Guard/research/skilljack.

Authors:Nizhang Li, Zonghao Ying, Xiangfan Wu, Zonglei Jing, Xixun Lin, Hao Zhang, Wenxin Zhang, Jiaye Lin, Quanchen Zou, Xiangzheng Zhang
Title: SkillSentry: Adaptive Honey Worlds for Dynamic Safety Testing of Agent Skills
Abstract:
External skills extend the capabilities of large language model agents, but also introduce an execution‑time attack surface: a skill that appears benign under inspection may reveal harmful behavior only after particular environmental states, resources, or interaction histories are encountered. Existing scanners primarily rely on static analysis, predefined rules, or one‑shot semantic judgments, making such conditional behavior difficult to elicit and attribute. We present SkillSentry, a dynamic safety‑testing framework based on adaptive honey worlds. SkillSentry infers the intended capability boundary of a skill, constructs an LLM‑simulated environment with controlled decoy resources, and adaptively generates tasks to explore its behavioral states. It then compares skill‑enabled trajectories with matched no‑skill executions, grounding suspicious behaviors in source code and verified execution traces before making a final decision. We evaluate SkillSentry against seven scanner configurations. SkillSentry achieves 99.50% Recall and 96.26% average F1 on standard benchmarks. Under semantics‑preserving evasion, it reaches 92.95% average F1, compared with 80.07% for the strongest baselines. Our code is available at https://github.com/nizhangli062‑jpg/SkillSentry‑Adaptive‑Honey‑Worlds‑for‑Dynamic‑Safety‑Testing‑of‑Agent‑Skills.

Authors:Weichen Xu, Zhenhua Liu, Lin Luo, Yaobo Liang, Chengtang Yao, Qingyu Mei, Jian Cao, Xixin Cao, Xing Zhang, Jiaolong Yang, Baining Guo
Title: Continue or Replan? Bernoulli-Continuation Policy Learning for Adaptive Horizon Execution
Abstract:
Existing chunk‑based Vision‑Language‑Action (VLA) models execute a fixed number of actions (i.e., execution horizon) before replanning, turning replanning into a task‑agnostic periodic schedule that is independent of task progress. As a result, when no replanning boundary falls before a critical manipulation stage, it is executed from a stale chunk rather than a freshly replanned one. To address this limitation, we propose Bernoulli‑Continuation Policy (BCP), a lightweight, plug‑and‑play framework for adaptive horizon execution that keeps the base VLA frozen. Given a fixed‑length action chunk, its continuation head decomposes execution‑horizon selection into a sequence of continue‑or‑replan decisions, which imposes an ordinal, prefix‑sharing inductive bias over candidate horizons rather than treating them as independent classes. Since the optimal horizon for each chunk is not observable, we train this head with reinforcement learning from trajectory‑level outcomes and introduce a Replanning‑Efficiency Reward that jointly rewards task success and efficient VLA usage, discouraging the policy from collapsing to unnecessarily short horizons. On RoboTwin 2.0 with LingBot‑VLA as the base policy, BCP improves the average success rate by +11.08% on 13 low‑success tasks and from 89.88% to 93.94% (+4.06%) across all 50 tasks. Although trained only under the Clean setting, BCP generalizes to the Randomized setting, raising the average success rate by +4.06%. It also transfers to a different base policy π_0.5, achieving a better result on LIBERO (+1.7%) and, notably, on the harder LIBERO‑PRO (+6.8%). On a real robot, BCP lifts success from 74% to 92% and from 44% to 84% on two manipulation tasks. Meanwhile, its negligible overhead, combined with higher success, makes BCP's overall runtime even lower than the fixed‑horizon baselines.

Authors:Álvaro Sánchez-Paniagua Ríos, Juan P. Llerena, Alberto Lastra, Nuria Torrado, Edmundo J. Huertas
Title: Beyond the Gegenbauer Paradigm: q-Orthogonal Kernels for Machine Learning
Abstract:
The performance of Support Vector Machines (SVMs) critically depends on the kernel function choice, which enables implicit mapping of data into high‑dimensional feature spaces. While classical kernels like Radial Basis Function (RBF) remain popular, orthogonal polynomial kernels offer mathematically interpretable alternatives that can incorporate structured prior knowledge. This work extends the orthogonal polynomial kernel paradigm by introducing a novel family based on discrete q‑Hermite I polynomials, a class of q‑orthogonal polynomials that generalize classical Hermite polynomials through a deformation parameter q. We formally define the q‑Hermite kernel and establish its validity under Mercer's theorem. The kernel's inherent boundedness properties naturally prevent annihilation and explosion effects without requiring explicit scaling mechanisms. Extensive experiments across 20 benchmark datasets demonstrate that the proposed kernel achieves competitive performance compared to both classical kernels and other orthogonal polynomial kernels, while offering advantages in numerical stability and computational simplicity. Our results confirm that q‑orthogonal polynomials constitute a promising direction for kernel design, bridging mathematical elegance with practical machine learning applications, that provides conceptual and algorithmic resources that may be further extended to emerging quantum computing paradigms. To facilitate full reproducibility, we provide the complete implementation and experimental pipeline in an open‑access GitHub repository at https://github.com/Kokechacho/SVMs‑QSVMs.

Authors:Zhe Cao, Miaowen Wen, Fangjiong Chen
Title: When Correct Solutions Repeat: Rarity-Aware Credit Redistribution for GRPO
Abstract:
Reinforcement learning with verifiable rewards (RLVR) com‑ monly optimizes each correct completion as an independent learning signal. In GRPO, this completion‑level uniformity creates structure‑level skew: recurring correct solution forms accumulate positive coefficient mass in proportion to how often they are sampled, while rare forms receive limited credit. We formalize this behavior as multiplicity‑induced structure‑level credit concentration and introduce a partition‑ conditioned rule that redistributes positive advantages accord‑ ing to cluster rarity. Cue‑GRPO instantiates this rule with‑ out auxiliary‑model inference by using deterministic Strategy Cues to construct rollout‑local partitions of verified‑correct traces. Across Qwen2.5‑Math‑7B and Llama‑3.1‑8B‑Instruct, Cue‑GRPO improves AIME repeated‑sampling performance, with the largest gains at high sampling budgets. Credit Re‑ distribution (CR) under Judge Partitions (JP) further indi‑ cates that the proposed redistribution mechanism can oper‑ ate with judge‑derived partitions. Cue‑GRPO adds only 6% wall‑clock training overhead over GRPO. These results sup‑ port structure‑level credit redistribution as a practical design axis for RLVR, with Strategy Cues providing a low‑overhead implementation for competition mathematics. Code is avail‑ able at https://github.com/CzZ12/When‑Correct‑Solutions‑ Repeat‑Rarity‑Aware‑Credit‑Redistribution‑for‑GRPO.

Authors:Haoqian Kang, Liupeng Li, Kuofeng Gao, Jinpeng Wang, Zhenyu Lu, Bin Chen, Ke Chen, Yaowei Wang
Title: Balancing Efficiency and Efficacy: Training-Free Attention-Guided Switching Between Explicit and Latent Thoughts for MLLMs
Abstract:
Reasoning in Multimodal Large Language Models (MLLMs) requires both fine‑grained visual perception and rigorous logical deduction. Explicit text‑based Chain‑of‑Thought (CoT) is computationally expensive and prone to visual hallucinations, while existing latent reasoning methods typically require costly training. Furthermore, directly adapting training‑free LLM reasoning mechanisms to the multimodal setting yields unstable performance. We identify that this failure stems from their reliance on token‑level entropy, which fundamentally conflates perceptual ambiguity (e.g., unclear visual details) with logical uncertainty (e.g., complex reasoning steps). To overcome this bottleneck, we present a novel training‑free inference strategy for MLLMs that explicitly decouples perception and reasoning. We propose a novel metric, the vision‑to‑text attention ratio, to dynamically gauge the model's cognitive focus. Guided by this metric, our proposed framework, Attention‑Guided Switching (AGS), adaptively triggers latent reasoning for perceptual tokens to preserve high‑fidelity visual information in the continuous space, while enforcing explicit text generation for logical tokens to maintain structural anchoring. Extensive experiments demonstrate that our method achieves state‑of‑the‑art performance, significantly improving both accuracy and inference efficiency by reducing autoregressive steps and latency. Code is released at https://github.com/swordAndSnow/MM26‑AGS.

Authors:Yuannuo Feng, Zegang Peng, Yuxin Xie, Yubing Ye, Yizhe Chen, Wenshuai Yao, Wenyong Zhou, Wang Kang
Title: Approximate Speculative Decoding
Abstract:
Speculative decoding accelerates autoregressive generation by verifying a draft block with a target model in parallel. Under standard greedy verification, decoding stops at the first draft token that differs from the target argmax, discarding the remaining target‑scored suffix. Although accepting such a mismatch changes the decoding trajectory, it can make a contiguous suffix reusable when its tokens remain target‑greedy under the realized prefix. In this paper, we introduce Approximate Speculative Decoding (ASD), a training‑free verifier that replaces binary first‑mismatch truncation with budgeted longest‑prefix selection. ASD accepts selected mismatches subject to a local target‑logit regret gate, a per‑block exception cap, and a persistent request‑level regret budget, then reuses the contiguous target‑greedy suffix without additional approximate decisions or target‑model forward passes. ASD requires neither a new draft model nor fine‑tuning, and exactly reduces to standard greedy verification when the budget is zero. Experiments show that ASD improves fixed‑workload throughput by 3.05%‑‑15.26% over matched strict verification and averages a 7.78% gain across seven Qwen3‑14B + DSpark‑14B tasks. On DeepSeek‑V4‑Flash (284B) with DSpark it also raises verifier‑side acceptance by roughly 10%‑‑16% on GSM8K and MATH‑500 in an FP4‑to‑FP8 compatibility setting. The source code is publicly available at: https://github.com/Kissmetothemoon/ASD

Authors:Dimitrios I. Zaridis, Traianos Tsiokris, Vasileios C. Pezoulas, Daphni Plati, Eugenia Mylona, Eleni Georga, Nikos Tsiknakis, Antonis Sakellarios, Dimitrios I. Fotiadis
Title: OliveGemma: A 3 Billion Visual Language Model for Recognising the Mediterranean & European Diet
Abstract:
Image based dietary assessment offers a scalable alternative to self reported food diaries, yet fine‑grained food recognition remains challenging due to high intra‑class variability and visually similar dishes. This study presents OliveGemma, a vision language model for recognising and reasoning about Mediterranean and European cuisine. Built on the open‑weight PaliGemma‑2‑3B architecture, OliveGemma is fine‑tuned with LoRA on a unified corpus of 17,340 images from three European research project datasets (MedGR, ODIN, and VIPPSTAR), reconciled into a vocabulary of 216 composed dish categories and paired with 102,642 instruction style question‑answer items covering dish recognition, likely and visible ingredients, class boundary discrimination, visual evidence and overall visual food understanding. Under a 3‑fold cross‑validation scheme, OliveGemma achieves a top‑1 accuracy of 92.96% +/‑ 0.91%, exceeding the strongest CNN baseline (DenseNet‑121) by 7.31% and outperforming zero‑shot frontier models with exact instructions and bounded classes including Gemini Flash 3 and 3.5, GPT‑5.4 Mini, and Claude Haiku 4.6 by 8%, 46%, and 64% respectively. Furthermore, OliveGemma demonstrates competitive performance on Top‑3 and Top‑5 accuracy, being second best across CNNs and frontier models, surpassed only by DenseNet‑121. In addition, OliveGemma achieves 90.79% +/‑ 1.3% Exact‑Set on the likely ingredients of the food categories. These results demonstrate that PEFT adaptation of a small VLM can surpass substantially larger proprietary models on specialised food recognition. The model is publicly available at https://huggingface.co/JamesZar/OliveGemma‑3B and the experiments and results can be found at https://github.com/tsiokris/OliveGemma.

Authors:Mingtian Zhang
Title: MMLongBench-Doc-V2: A Corrected-Annotation, Semantics-Aware Revision of MMLongBench-Doc
Abstract:
MMLongBench‑Doc is a long‑document QA benchmark of 1,082 questions over 135 PDFs. Two properties of it push measured scores away from the quantity they are meant to capture: the reference metric compares extracted answers, so 1,358,000 loses to 1358000; and a non‑trivial share of ground‑truth annotations are wrong, ambiguous, or incomplete ‑‑‑ concentrated, because of how they were found, in exactly the questions capable systems answer correctly. MMLongBench‑Doc‑V2 corrects 106 annotations, each published with the page and arithmetic that settle it, and replaces the string metric with a pinned LLM judge asked whether a response means the reference. Ten questions whose document ships under the wrong filename are removed rather than counted wrong, along with one duplicated question, leaving 1,071 questions over 134 documents. The most reusable contribution is a decision procedure for when an empty set key may be widened and when widening would destroy a deliberate negative sample; applied to all 208 rows, it widened 14. V2 scores are not comparable with published V1 numbers. The corrected corpus, the per‑entry correction record and the evaluation harness are available at https://github.com/VectifyAI/MMLongBench‑Doc‑V2.

Authors:Hao Zhou, Haichuan Hu, Ye Shang, Quanjun Zhang
Title: Self-Evolving Coding Agents
Abstract:
Large language models are increasingly embedded in software engineering workflows as coding agents that can inspect repositories, invoke tools, execute tests, debug failures, and generate patches. Yet most existing agents remain largely static after deployment, even though software development is a dynamic, feedback‑rich process in which repositories evolve, dependencies change, tests fail, and repair attempts leave reusable experience. This tension has motivated a growing body of work on self‑evolving coding agents, where the agent improves its future behavior by updating its framework, memory, skills, tools, models, or collaboration structures from prior coding interactions. In this survey, we provide a systematic synthesis of this emerging area. We first define self‑evolving coding agents and distinguish them from conventional coding agents and general self‑evolving agents. We then develop an object‑centered taxonomy that characterizes what evolves in these systems, and complement it with two orthogonal perspectives: when evolution occurs and what software‑specific evidence drives it. Across the literature, we find that executable feedback, repository‑level context, and coding trajectories give software engineering a distinctive role as a natural domain for agent self‑evolution, but also introduce new challenges in feedback reliability, benchmark overfitting, safety, maintainability, cost, and generalization. By organizing existing work around these dimensions, this survey aims to clarify the conceptual boundaries of self‑evolving coding agents and provide a foundation for designing more adaptive, reliable, and software‑aware agentic systems. The papers we collect can be found at https://github.com/zhouhao1024/Awesome‑Self‑Evolving‑Coding‑Agents.

Authors:Nicolas Zumarraga, Lorenzo Steno, Ning Wang, Max Rosenblattl, Thomas Kaar, Maxwell A. Xu, Kevin O'Sullivan, Markus Kreft, Elgar Fleisch, Paul Schmiedmayer, Patrick Langer, Robert Jakob
Title: TimeRLM: Recursive Language Models Enable Precise Anomaly Localization in Long-Context Time-Series
Abstract:
Precise anomaly localization over long‑context time series is a crucial task in monitoring applications across clinical care, industrial operations, financial services, and logistics, where brief evidence may hide inside long spans of high‑frequency data. Time‑Series Language Models (TSLMs) are able to ingest time series data and verbalize findings on anomalies in natural language; however, recent benchmarks report a decrease in retrieval performance at long contexts, mirroring failure modes in text, vision, and audio. In the text domain, Recursive Language Models (RLMs) can recover much of this lost performance by keeping context external to the large language model (LLM), allowing the model to query it through code. We present TimeRLM, an RLM formulation for time‑series that sequentially manipulates the signal using code and vision capabilities. We further introduce AnomalyXL, a synthetic long‑context anomaly localization benchmark with programmatically injected anomalies that require precise retrieval. We implement five different task categories and two variants: AnomalyXL‑MCQ and AnomalyXL‑Localize. TimeRLM outperforms every evaluated TSLM and single‑pass baseline on four of the five AnomalyXL‑Localize tasks, reaching 0.682 IoU on localization and 0.745 on classify‑with‑evidence, versus at most 0.329 and 0.072 across all baselines. We post‑train TimeRLM using reinforcement learning. The resulting model further improves performance and requires approximately one‑third as many agent interaction turns as its untrained base model to produce a final answer. On unseen real‑world ECG, sleep and software observability recordings, the post‑trained TimeRLM retains or improves performance, surpassing TSLMs despite being trained exclusively on synthetic data. Our findings suggest recursive interaction with time‑series is an effective approach for long‑horizon retrieval.

Authors:Wei Wei, Yinyuan Zhao, Ruixuan Yu
Title: Residual Flow Matching with Dynamic Cross-Interaction for 3D Multi-Person Motion Prediction
Abstract:
3D multi‑person motion prediction requires modeling both individual kinematics and inter‑person interactions. While Flow Matching is effective for multi‑hypothesis generation to improve prediction accuracy, directly predicting skeletal sequences from pure noise often compromises structural consistency and introduces unreliable cross‑agent interactions during early noise‑dominated integration steps. To address this, we propose a Prior‑Guided Residual Flow Matching framework. First, a Deterministic Coarse Prior (DCP) establishes a kinematic anchor, formulating the generative process as a conditional flow over motion residuals to simplify the generative objective and preserve structural stability. Second, a Dynamic Cross‑Interaction (DCI) mechanism temporally synchronizes inter‑agent message‑passing with the integration progress, ensuring the extraction of reliable social contexts and improving multi‑person motion fidelity. Finally, a decoupled joint‑motion architecture with bidirectional fusion effectively preserves fine‑grained kinematic coherence. Extensive experiments demonstrate that our approach achieves state‑of‑the‑art prediction accuracy across multiple datasets. Code is available at https://github.com/Wei‑Wei‑a/Residual‑Flow‑Matching‑with‑Dynamic‑Cross‑Interaction‑for‑3D‑Multi‑Person‑Motion‑Prediction.

Authors:Alex Kwon
Title: FACTWASH: Catching AI Rewrites That Wash Hearsay into Fact
Abstract:
AI systems rewrite information constantly: conversations become stored memories, documents become answers. The rewrite can keep a claim while washing away what made it checkable, who said it, how sure they were, when it held. We call that failure factwashing, and release factwash, an open‑source write‑time gate that catches it deterministically, with named flags and evidence rather than an LLM judge. Building it answers a practical question: when does a cheap check suffice, and when do you need a model? What decides is whether the property has a bounded surface‑cue inventory. Explicit negation cues are close to enumerable, so a word list finishes and transfers, reaching 0.91 F1 on untuned text. Hedging and attribution have open‑ended realizations, so vocabulary plateaus near half recall, and a one‑question LLM witness recovers +17 and +15 points of cue‑detection recall at equal precision. Deployed, that witness may only lower a verdict, so it buys precision rather than coverage. We measure cue detection on 105,596 independently annotated sentences. A blind‑labelled corpus of memory writes then locates the failure: 55% of bad writes in conversational hearsay, 7% in business email (p < 0.001), so the first deployment question is not which detector to use but whether the failure occurs at all. On unmodified mem0 2.0.7, the gate flags 5 of 8 hedged‑hearsay writes.

Authors:Shanghao Liu, Renze Chen, Size Zheng, Yuanqiang Liu, Yun, Liang, Hailong Yang
Title: SPADE: An Input-Adaptive Sparse Attention Engine for Fast Video Diffusion Models Inference
Abstract:
Video diffusion transformers (vDiTs) generate high quality but pay quadratic self‑attention cost, making inference prohibitive at video‑token scales. The challenge is input‑adaptive sparsity: selecting critical Q/K/V tokens with negligible overhead and executing them for end‑to‑end gains. We present SPADE, a training‑free sparse‑attention engine of three parts: (i) vDiT‑SSR, a specification defining 3D blocking candidates and formalizing dynamic masks via Summarizer/Estimator expressions; (ii) runtime scheme generation using SICS and a head‑wise policy; and (iii) an executor with low‑overhead index search, flash block‑sparse attention, and kernel grouping. Across Hunyuan‑Video and Wan 2.1/2.2 for text‑to‑video and image‑to‑video generation, SPADE raises sparsity and speed while preserving quality, accelerating attention by 2.26x‑3.40x and end‑to‑end inference by 1.49x‑1.80x. Our code is open‑sourced at https://github.com/6somehow/DAC‑SPADE.

Authors:Siqi Fan, Minghao Li, Xiaoqian Ma, Wenhui Tan, Xiusheng Huang, Juntong Wu, Liujie Zhang, Shuo Shang, Weihang Chen
Title: Screenshots or Tools? Eliciting Tool Use and Managing Multimodal Context in Hybrid GUI-MCP Computer-Use Agents
Abstract:
Hybrid computer‑use agents can act through screenshots or call text tools. We find that having a tool available does not settle which way the effect goes. Under one identical GUI‑MCP harness on the OSWorld‑MCP benchmark (309 tasks), the same MCP tools improve a reasoning model by +4.0pp and degrade a non‑reasoning model by ‑5.9pp (5 runs each, both beyond 2 SE). What separates the two is tool‑decision behavior. The non‑reasoning policy ignores, misnames, or falsely terminates around tools. The reasoning model avoids these failures, yet still calls a tool on only 55/309 tasks, 23.9% of the tool‑reachable ones. We call this shortfall the adoption gap. Both levels of the problem share one cause: the model already has a cheaper route and is never trained to take it. Multi‑turn RL probes that cause. At the action level, a dense tool bonus raises spreadsheet adoption 0.03 ‑> 0.33 and carries into greedy decoding, but held‑out accuracy does not follow. Behavior is steerable; competence is not. The bottleneck lies in tool‑call semantics. At the context level, a successful tool call often makes the next screenshot redundant. Dropping it and halving image history cuts input tokens by about a third, at a small accuracy cost. Retraining under the same observation rule removes that cost. The compressed agent then reaches 37.8% against 33.0% for the uncompressed operating point, at 53% of the input cost, and closes the rich‑lean gap on a pre‑registered degraded subset to zero. Tools help when the model chooses and integrates them, and current hybrid agents leave many such choices unused. Code and checkpoints: https://github.com/redai‑infra/hybrid‑routing‑agent

Authors:Gustav Hanning, Shaohui Liu, Rémi Pautrat, Marc Pollefeys, Kalle Åström, Viktor Larsson
Title: PolyLayout: Multi-room Manhattan Layout Estimation
Abstract:
Estimating room layouts from multi‑view imagery is a core task for indoor scene understanding. Existing methods are typically limited either by poor generalization to new datasets or restrictive geometric assumptions of the room shape or camera configuration. Most also estimate rooms independently, failing to exploit shared building structure such as dominant directions, ground plane or ceiling height. We propose PolyLayout, a multi‑room layout estimation method that parameterizes room layouts as Manhattan 3D polygons and optimizes them jointly across multiple rooms. The optimization objective is predicted by a neural network on top of robust pre‑trained visual features and trained end‑to‑end with supervision only on output room layouts. At the same time, camera projection and polygon updates remain explicit and model‑based. This separation between learned scoring and geometry improves generalization to new datasets and camera parameters. During optimization, PolyLayout adaptively refines the polygon topology through iterative wall split and merge operations while jointly utilizing structural cues across rooms. We introduce two new multi‑view multi‑room layout benchmarks by providing layout annotations to existing datasets, and experiments show that PolyLayout outperforms prior approaches, both in terms of accuracy and robustness. Project page: https://ghanning.github.io/PolyLayout

Authors:Zihan Wang, Tong Liu, Zhiwei Wang, Tao Huang, Wentao Jiang, Sihan Ma, Shanshan Ye, Xiaohui Yang, Jing Zhang
Title: LocAnyMed: Vision-Language Grounding for Multimodal Medical Images
Abstract:
Medical visual grounding connects free‑form clinical queries to spatial evidence in medical images and is an important component of interpretable medical artificial intelligence. However, general‑purpose grounding models are predominantly trained on natural images, while existing medical localization resources remain fragmented across imaging modalities, datasets, and task formulations. To address this gap, we construct LocAnyMed‑200K, a multimodal medical visual grounding dataset containing approximately 200K image‑query‑answer examples across computed tomography, optical medical imaging, ultrasound, and X‑ray. We harmonize heterogeneous detection and localization resources into a unified free‑form instruction format that supports one or multiple bounding boxes, point coordinates, and no‑target outputs for negative queries. Full‑parameter fine‑tuning of LocateAnything‑3B on LocAnyMed‑200K improves F1@IoU 0.50 from 10.64 to 85.59 on a held‑out evaluation split, demonstrating that large‑scale domain‑specific supervision can equip a general grounding model with effective medical localization capabilities. Beyond spatial coordinates, a clinically interpretable grounding system should also communicate the evidence supporting its prediction. We therefore derive LocAnyMed‑CoT‑20K, a rationale‑augmented subset that connects anatomical context, visual observations, and spatial conclusions through structured reasoning and further improves cross‑source generalization through fine‑tuning. Together, these resources provide a unified foundation for studying both localization accuracy and rationale quality across heterogeneous medical imaging modalities. The code is publicly available at https://github.com/MiliLab/LocAnyMed.

Authors:Zhiyuan Zhu, Xinling Meng, Junxuan Yu, Jiongquan Chen, Qiongying Ni, Tuhang Shao, Yuhao Huang, Luping Zhou, Ruiyang Huang, Yuxue Wang, Rongliang Zhang, Xue Wang, Tianhong Tang, Likun Wang, Junbo Chen, Yong Jiang, Yongping Lu, Xin Yang
Title: Recurrent Contrastive Learning for Imbalanced Medical Image Classification
Abstract:
Medical image classification often suffers from class imbalance due to the inherent disparities in disease incidence. Existing approaches, such as class resampling and loss reweighting, mainly improve learning within the observed feature distribution, but do not explicitly enlarge the latent support region of tail classes. As a result, tail‑class representations remain overly compact and are easily encroached upon by head classes, leading to biased decision boundaries. In this work, we propose Recurrent Contrastive Learning (RCL) for imbalanced medical image classification. RCL progressively expands the support region of tail classes by recurrently reusing historical feature states across training phases. Specifically, we adopt DINOv3 with LoRA adapters as the backbone to provide robust feature embeddings. We then devise a Temporal Memory Queue (TMQ) to preserve corpus‑level features across training phases and provide diversified global references for contrastive learning. Based on TMQ, we construct Temporal Anchors (TARs) to form an anchor field around tail classes. This field enlarges the support region of tail classes, suppresses head‑class encroachment, and improves inter‑class separation. Extensive experiments on three imbalanced medical datasets demonstrate that RCL achieves consistent improvements over strong baselines. The code is available at https://github.com/dndins/RCL.

Authors:Mohsen Arjmandi
Title: Distractor-Aware Truncation: Disentangling Context-Length Effects from Signal Loss in Long-Context LLM Benchmarks
Abstract:
A standard claim in the literature on retrieval‑augmented and memory‑augmented language models is that shorter context is better when the relevant information is preserved. We test this claim by running every sample of two long‑context benchmarks ‑‑ BABILong and GraphWalks (BFS) ‑‑ at four context‑retention fractions (100%, 75%, 50%, 25%) under two truncation protocols. The first is the naive protocol implicitly used in much prior work: drop content from the middle of the prompt. The second is distractor‑aware: identify the task‑relevant content for each sample and drop only the rest. We evaluate three sizes of the Claude family (Haiku 4.5, Sonnet 4.6, Opus 4.7) and, to test cross‑provider generality, GPT‑5.5 from a different provider; we apply the same protocol to two further benchmarks (MRCR v2, Oolong). Under naive truncation, score collapses monotonically (paired Wilcoxon, Holm‑corrected p_adj < 0.05 in all eight BABILong and GraphWalks cells). Under the distractor‑aware protocol ‑‑ which preserves the signal by construction ‑‑ performance is preserved or improves: the two smaller Claude models show statistically significant gains on BABILong, while the larger models (Opus 4.7 and GPT‑5.5) sit at their full‑context ceiling. The naive collapse and its distractor‑aware recovery replicate on GPT‑5.5, ruling out a single‑provider artifact. The mechanism is direct: under the naive protocol the answer‑bearing content survives in fewer than 1% of samples at 25% retention; under the distractor‑aware protocol it is preserved by construction. The naive protocol is therefore not a measurement of context‑window effects; it is a measurement of how often middle‑removal happens to spare the answer. We conclude that future studies of context‑length effects must specify how they distinguish signal from distractor, or they are at best ambiguous between two opposite hypotheses.

Authors:Yuke Xing, Jiarui Wang, William Gordon, Zhu Li, Guangtao Zhai, Yiling Xu
Title: 3DGSI-Assessor: A Large-Scale Dataset and An LMM-based Method for 3D Gaussian Splatting Image Quality Assessment
Abstract:
3D Gaussian Splatting (3DGS) has become a dominant representation for real‑time novel view synthesis (NVS), yet its storage footprint makes compression indispensable for practical deployment. 3DGS training and compression introduce representation‑specific distortions such as floating artifacts and surface scattering, which conventional image quality assessment (IQA) metrics fail to capture. Moreover, the independent compression of geometric and color attributes may lead to decoupled dimension‑specific distortions that must be diagnosed separately, yet existing metrics report only a single overall score. To address these gaps, we present 3DGS‑IEval‑15K+, a large‑scale, multi‑dimensional IQA dataset for compressed 3DGS, comprising 15,200 images from 10 diverse scenes, produced by 6 representative 3DGS algorithms at systematically designed compression levels and rendered from 20 strategically selected viewpoints spanning both training views and challenging novel views, annotated with 45,600 mean opinion scores (MOSs) across overall, geometry, and color quality. Based on 3DGS‑IEval‑15K+, we propose 3DGSI‑Assessor, an all‑in‑one 3DGS IQA framework that integrates global semantic and dimension‑specific local features within a large multimodal model (LMM), predicting all three dimensions in a single forward pass. 3DGSI‑Assessor achieves state‑of‑the‑art performance on 3DGS‑IEval‑15K+, and exhibits competitive generalization on other NVS benchmarks. Dataset and code will be released at https://github.com/YukeXing/3DGSI‑Assessor.

Authors:Anjun Hu, Hanting Xie, Saranya Govindan, Jas Kandola, Kurt Cutajar
Title: Attacking and Defending Multi-Agent Collaborative Filtering Systems Through Connectivity
Abstract:
Multi‑agent collaborative filtering (CF) systems coordinate autonomous LLM‑powered user and item agents through natural‑language interaction to refine preferences and generate recommendations. These systems inherit vulnerabilities from both their data‑driven nature and their multi‑agent interactions, which manifest in distinct ways. Understanding how connectivity modulates vulnerability in these systems could facilitate the development of more robust recommendation pipelines. In this work, we adapt attacks and defenses from the general multi‑agent systems (MAS) literature to the agent‑based CF setting, evaluating them under systematically varied connectivity in the AgentCF framework, where CF connectivity is characterized along two axes: (i) candidate count (the number of item candidates per turn per user, measuring user‑side interaction density) and (ii) catalog concentration (the degree of item catalog overlap across users). Our contributions include: (1) Adaptation: we reproduce MAS‑inspired attacks and defenses in the agentic CF domain, confirming partial transferability of original observations. (2) Characterization: we characterize how the two aspects of connectivity shape attack and defense outcomes, revealing role asymmetries between user and item agents, non‑monotonic temporal dynamics in attack efficacy, and divergent patterns across dissemination and extraction attack goals. Additionally, as an exploratory extension, we assess the applicability of epidemic‑inspired static metrics in ranking CF configurations by expected attack outcome, potentially enabling cost‑efficient robustness assessment. Implementation is available at https://github.com/anjunhu/ConnACF

Authors:Zichuan Fu, Shirong Wang, Wenlin Zhang, Guojing Li, Yimin Deng, Jingtong Gao, Junjia Qi, Hanyu Yan, Yefeng Zheng, Xiaopeng Li, Wanyu Wang, Xian Wu, Xiangyu Zhao
Title: GUI-Lens: Coarse-to-Fine Cropping for GUI Grounding with General-Purpose VLMs
Abstract:
GUI grounding maps natural‑language instructions to click locations and is essential for reliable GUI agents. The task remains difficult on high‑resolution, densely populated interfaces because a vision‑language model (VLM) may recognize a requested control without locating it precisely enough for interaction. Most existing methods provide various forms of localization assistance, but still rely on a direct click prediction, allowing visual ambiguity or an inaccurate initial estimate to propagate to the final result. In this paper, we introduce GUI‑Lens, a coarse‑to‑fine grounding framework that allows a general‑purpose VLM to determine the target through active visual observations. Specifically, GUI‑Lens extracts OCR text and detected UI components from the screenshot and presents their positions as coordinate references. Using the instruction, the current view, and these references, the VLM selects the region and scale of the next view, which is cropped and enlarged to provide finer visual details. This process continues over successively focused views until the target is determined. Proposed crops and clicks are checked against the instruction throughout the process, and the final local position is mapped back to the original screen coordinates. Experiments on four GUI grounding benchmarks and three general‑purpose VLM backends show that GUI‑Lens improves overall grounding accuracy by up to 24.9 percentage points and achieves state‑of‑the‑art performance with GPT‑5.5.

Authors:Leiye Liu, Miao Zhang, Jiahong Jiang, Jingjing Li, Jialong Zhong, Kai Peng, Tingwei Liu, Wei Ji, Yongri Piao, Huchuan Lu
Title: Hear to See: Discerning Stateful Listening for Audio-Visual Instance Segmentation
Abstract:
Audio‑visual instance segmentation (AVIS) requires accurately identifying and tracking individual sounding objects with pixel‑level masks. Existing methods struggle to match overlapping acoustic events with visual instances and handle asynchronous audio‑visual dynamics. Therefore, two critical questions arise: how can a model establish precise correspondence between overlapping sound sources and visual instances, and how can a model maintain robust tracking when audio and visual signals are temporally misaligned?This paper proposes Hear to See (H2S), addressing these challenges through two mechanisms. The Acoustic‑Semantic Projector (ASP) disentangles mixed audio and establishes hierarchical correspondence from semantic to spatial domains. The Asynchronous Dynamics Modulator (ADM) adaptively adjusts state transitions via audio‑modulated Mamba, prioritizing current information during dynamic variations and maintaining continuity in stable periods.Experiments on AVISeg show H2S achieves SOTA performance, attaining 48.54 mAP with a COCO pretrained ResNet50 and surpassing the previous by 7.8%. The code will be open‑sourced once the paper is accepted. The source code will be publicly available at https://github.com/leiyeliu/H2S.

Authors:Yizhong Geng, Tian-Hao Zhang, Chunfeng Wang, Wenxin Fu, Yingming Gao, Ruimin Wang, Zhou Pan, Kun Zhan, Liang Li, Ya Li
Title: CLASVS: Continuous-Latent Autoregression for Melody-Preserving Lyric Editing in Singing Voice Synthesis
Abstract:
Reference‑conditioned melody‑preserving lyric editing replaces words while retaining a performance's timing, singer identity, and naturalness. Continuous‑latent autoregression avoids finite codebooks and offers stepwise generation with learned stopping. Editing creates a conflict absent from ordinary reconstruction: training pairs reference cues with original lyrics, whereas inference asks revised lyrics to override source‑lyric‑correlated cues; one source‑following patch can propagate through AR history. We introduce CLASVS. Its State‑Control‑Transition (SCT) routing keeps target‑lyric and reference‑melody controls persistent, returns semantic feedback on phonetic progress to the causal planner, and confines the previous latent patch to the local Transition. Progressive State‑Control Grounding (PSCG) learns this contract through paired‑edit‑free, content‑consistent Mandarin reconstruction. On two Mandarin benchmarks, CLASVS improves all four operations over discrete‑AR Vevo2 and reduces macro‑PER by 46.2%, while maintaining melody, singer similarity, and perceptual quality. Together, these results establish a strong continuous‑AR operating point for score‑annotation‑free lyric edits and a basis for broader stepwise control. Audio demonstrations are available on our project page: https://piedpiperg.github.io/Liyric‑SVS/.

Authors:Yicheng Zhang, Haoyou Deng, Zhiqiang Li, Wenti Yin, Nong Sang, Changxin Gao
Title: Clarity Contrast and Similarity Selection for Multi-Focus Image Fusion
Abstract:
Multi‑focus image fusion (MFIF) aims to generate an all‑in‑focus image from multiple images of the same scene focused at different regions. Most existing deep learning‑based methods lack explicit interaction between the source images, which limits their performance and interpretability. This paper presents a novel Clarity Contrast and Similarity Selection Network (CSNet), to bridge direct information exchange for MFIF. Specifically, by contrasting the clarity differences between source images within our proposed Clarity Contrast Attention Module (CCAM), we mutually enhance sharp features while suppressing blurry ones. This allows us to identify the exactly focused regions in each source and locate the focused‑defocused boundaries. Moreover, the Defocus Spread Effect (DSE) degrades pixels in all source images around the boundaries. To further refine these ambiguous areas, we introduce a Similarity Selection Strategy, which reconstructs an initial clear image from source images and selects optimal pixels by comparing the similarity among them. Through this interactive approach, CSNet effectively preserves focused regions as well as recovering natural boundaries to fuse an all‑in‑focus output. Extensive experiments demonstrate that our method achieves state‑of‑the‑art performance both quantitatively and qualitatively. Our code is available on Github: https://github.com/ZYC‑HUST/CSNet.

Authors:Ning Zhu, Xiaochuan Ma, Juntao Xu, Jingze Liang, Mengfei Zhao, An Chen, Liang-Jian Deng
Title: One Knob to Rule Them All: A Unified Optimal Transport View of Cold-Start Active Learning
Abstract:
Cold‑Start Active Learning (CSAL) aims to select a valuable subset from an unlabeled pool without any prior knowledge or human assistance. Existing methods take diverse routes based on typicality, coverage, or diversity. Each rests on its own inductive bias and therefore performs well on some tasks yet poorly on others. We argue that the real challenge is not to design yet another selection heuristic, but to make CSAL adapt automatically to the data and task at hand. To this end, we revisit CSAL through the lens of optimal transport. First, we propose a generalized transport selection framework that reveals the shared allocation structure of existing methods and exactly subsumes representative formulations. Second, we introduce a theoretical analysis that characterizes the trade‑off controlled by entropic regularization and establishes a task‑agnostic minimax bound for cold‑start selection. These results provide a principled foundation for adapting the regularization strength to the unlabeled data. Third, we derive a data‑adaptive regularization rule and present a novel Sinkhorn‑based CSAL algorithm, termed ε‑Adaptive Selection (ε‑AS). Extensive experiments on six public datasets and multiple annotation budgets show that ε‑AS consistently achieves state‑of‑the‑art performance. On ImageNet‑1k, it improves the average accuracy over ActiveFT by 1.29% while reducing selection time by 56.2%. Code will be released at https://github.com/Z‑yiwei/OT‑CSAL

Authors:Jing Dai, Qibin Zhang, Weiwei Zhou, Mingde Xu, Jingsong Liu, Jingdong Zhang, Hongming Xu
Title: CIGTSurv: Clinical Information Guided Tri-modal Survival Prediction with Local Prototype Association and Global Feature Alignment
Abstract:
Multimodal learning has significantly advanced survival prediction by integrating pathology images with genomic data. However, clinical information, despite its critical role in reflecting a patient' s overall health, remains underutilized due to its discrete, sparse, and low‑dimensional nature. Furthermore, the inherent heterogeneity across these modalities pose significant challenges in modeling cross‑modal interactions. In this paper, we propose CIGTSurv, a Clinical Information Guided Tri‑modal framework for Survival prediction. Specifically, we first design a holistic text template and use pretrained foundation models to transform clinical tabular data into high‑dimensional tokenized embeddings. Using clinical information as an anchor, we then introduce a dual‑level interaction mechanism: 1) a local prototype association (LPA) module based on cross‑attention to explicitly learn token‑level correspondences between different modalities, and 2) a global feature alignment (GFA) loss based on Maximum Mean Discrepancy (MMD) to implicitly enhance cross‑modal distribution consistency. Extensive experiments on five TCGA cancer cohorts demonstrate that CIGTSurv achieves state‑of‑the‑art (SOTA) survival prediction performance. Our source code is publicly available at https://github.com/Daijing‑ai/CIGT‑Surv.git.

Authors:Chengyu Wu, Junpeng Tan, Wanxiang Luo, Yaqi Wang, Yandong Wen, Yefeng Zheng
Title: Open-Linguistic Concept Unified Learning for Cross-Site Interpretable Dermatology Image Diagnosis
Abstract:
Human‑interpretable computer‑aided diagnosis is crucial for clinical decision making. Concept‑based models excel by providing transparent reasoning and enabling post‑hoc, clinician‑in‑the‑loop interventions. However, their rigid dataset‑specific adaptation inherently restricts cross‑site generalization. Applying them across diverse modalities, such as dermoscopic and clinical photographs, is challenging due to heterogeneous concept taxonomies varying in availability, granularity, and semantics across cohorts. Consequently, adapting Foundation Vision‑Language Models (FVLMs) demands costly label engineering and repeated post‑training. Existing intervention mechanisms remain rigidly tied to predefined concepts, lacking adaptability and hindering scalable dermatology CAD deployment. To address these bottlenecks, we propose UniCon, an open‑linguistic unified concept learning framework for multimodal interpretable vision‑language diagnosis. UniCon resolves these challenges through three contributions: (1) A shared semantic representation space via a unified concept prototype codebook, seamlessly coordinating heterogeneous concept systems across modalities without dataset‑specific retraining. (2) Open‑linguistic based multi‑faceted semantic specifications to overcome sparse textual label limitations, improving boundary sensitivity in uncertain clinical contexts. (3) A robust, cross‑site adjustable intervention interface powered by reliability‑gated bottleneck aggregation, enabling consistent reasoning and transferable clinician corrections. Extensive experiments demonstrate that beyond securing top‑tier diagnostic accuracy, UniCon successfully bridges disparate clinical taxonomies, unlocking unprecedented cross‑site intervention capabilities. Code is available at https://github.com/wuchengyu123/UniCon.

Authors:Ranxu Zhang, Guinan Chen, Chenshaodong, Jinghao Lin, Xiaozhou Xu, Sunzhe, Yanyong Zhang, Chao Wang
Title: Agentic Reinforcement Learning with Self-Distilled Reward Shaping
Abstract:
Agentic reinforcement learning enables LLM agents to learn through interaction, but sparse trajectory‑level rewards reveal success without identifying which intermediate decisions deserve credit. Training‑only privileged skills can provide denser supervision by allowing the same frozen policy snapshot to rescore fixed tokens from skill‑free trajectories while conditioned on task‑matched procedural skills. Existing methods, however, do not jointly calibrate teacher scores across interaction steps, relate teacher confidence to realized returns, and integrate the resulting signal into native reward‑to‑advantage construction. We introduce Agentic Reinforcement Learning with Self‑Distilled Reward Shaping (ADRS), a framework for constructing return‑associated token‑level credit for multi‑turn language agents. ADRS centers and normalizes privileged token scores within each step, modulates them with a return‑associated Teacher Value Advantage (TVA) gate based on within‑group confidence‑‑return association, and incorporates the gated token signal into native RL credit construction. Together, these components determine what the teacher prefers, when that preference is return‑relevant, and how it enters the native reinforcement‑learning credit path, while keeping rollouts and inference skill‑free. Finally, experiments across three interactive benchmarks show that ADRS consistently improves performance on long‑horizon tasks, with gains persisting across RL backbones, reduced‑data settings, unseen tasks, and extended training. For anonymous review, our code is available at the following the link: https://github.com/gitrxh/ADRS‑arxiv

Authors:Yanchao Li, Wanhao Liu, Jiaqing Xie, Ben Gao, Yanbo Wang, Tianfan Fu, Yuqiang Li
Title: Reachability Is Not Realization: Tracing the Sources of LLM Benchmark Gains
Abstract:
Benchmark gains are often treated as evidence of greater LLM capability. Yet the same gain can reflect different changes in model behavior. A model may reach new answers, or produce answers that were already within reach. Aggregate scores do not distinguish these changes question by question. We establish a question‑level audit under fixed budgets, temperatures, and answer formats. A question is realized when the default deployment procedure produces the correct answer. A question is reachable when a specified probe finds that answer within a fixed budget. We first test whether inference‑time layer routing can expand reachability. Under a matched budget, random routes match or exceed structured search in all 43 model and task settings. Answer‑blind procedures retain almost none of this gain, which instead requires access to the correct answer. We then ask why reachable answers sometimes fail to appear. Across six cases spanning 0.5B to 31B, silencing one identified MLP block repairs 68 to 92 percent of a predefined failure set. We next test whether training closes the gap by expanding reachability. In five of six matched evaluations, deployed performance rises while the reachable ceiling remains flat or falls. For DAPO, the deployed score rises by 14.7 points while the reachable ceiling falls by 13.3 points. Across the settings we audit, realization and reachability therefore do not always change together. Claims of capability expansion should report both realized performance and reachability under matched evaluation conditions. Code is available at https://github.com/LiZaiyuan0619/reachability‑not‑realization

Authors:Fang Li, Yu He, Haoyang Tong, Lichen Ma, Jingling Fu, Wenxiao Fan, Tongxuan Liu, Luohang Liu, Ke Zhang, Junshi Huang
Title: iFAN: Inference-Aware Learning for Plain Mask Transformers
Abstract:
Query‑based mask transformers assemble segmentation outputs through pixel‑wise competition among query predictions of the final layer, yet this inference process is not explicitly optimized during training. We identify two key mismatches: the query with the highest probability‑mask score does not necessarily produce the most accurate mask, and final‑layer decoding may discard superior predictions from intermediate layers. To address these issues, we propose Inference‑Aware Learning (iFAN), a general training framework for plain mask transformers. iFAN introduces Adjusted Probability‑Mask Ranking (APMR), which aligns query competition with predicted mask quality and suppresses high‑confidence but inaccurate competitors. We further employ Cross‑Layer Self‑Distillation (CLSD) to transfer stronger intermediate predictions to the final layer. The ranking and distillation objectives are training‑only, while inference retains efficient final‑layer decoding. Experiments on COCO, ADE20K, and Cityscapes demonstrate consistent improvements across panoptic, instance, and semantic segmentation, as well as across different architectures, backbone scales, and input resolutions. Overall, iFAN improves performance by an average of 1.20 PQ, 1.30 AP, and 0.63 mIoU, with negligible additional parameters, FLOPs and inference latency.

Authors:Seonmi Park, Seunghyun Shin, Vihaan Misra, Dongmin Shin, Ukcheol Shin, Jean Oh, Hae-Gon Jeon
Title: Bridging Online and Offline Handwriting via Differentiable Physical Rendering
Abstract:
Realistic handwritten text generation plays an important role in numerous applications, such as font design, biometric authentication, and robotic calligraphy. Existing methods are typically divided into two independent paradigms: online approaches that estimate handwriting trajectories and offline approaches that synthesize realistic handwriting images. While online models capture structural and temporal dynamics, they often lack fine‑grained textures, whereas offline models reproduce realistic appearance but discard stroke order. However, unifying online and offline models remains challenging due to (1) the lack of an explicit physical model linking stroke kinematics to pixel‑level appearance and (2) the absence of paired trajectory‑image datasets. Moreover, enabling end‑to‑end learning requires a differentiable rendering process across motion and appearance domains. To address these challenges, we propose a compact physical brush model that bridges stroke dynamics and visual appearance, together with a differentiable rendering module that converts stroke trajectories into stylized images. By integrating these components, we propose a unified online‑offline handwriting generation framework via differentiable brush rendering. The proposed framework consists of four core modules: 1) a text‑to‑stroke generator that predicts the target stroke conditioned on the given text and style image, 2) a brush parameter observer that extracts brush model parameters from style references, 3) a differentiable brush renderer that maps a stroke sequence and physical brush parameters into a handwritten image, and 4) a zero‑shot image refiner that refines rendered images via diffusion models. Extensive experiments and real‑world robotic calligraphy demonstrations validate our approach, achieving both structural and visual fidelity.

Authors:Xiao Fan, Hongbin Guo, Yubo Han, Yi Zhang
Title: Frequency-Decorrelated Temporal Ensembles for EEG--fNIRS Imagined-Handwriting Decoding
Abstract:
Imagined handwriting offers a temporally rich paradigm for non‑invasive neural decoding, yet reliable recognition across unseen participants remains difficult because scalp EEG is noisy and internally generated stroke sequences vary across individuals. The Multimodal Brain‑Computer Interface Grand Challenge provides synchronized EEG and fNIRS for four‑class subject‑independent handwriting‑trajectory classification. We propose FRED, a task‑adapted system that models imagined handwriting as a multi‑second motor sequence and trains a compact multi‑scale temporal network on three complementary EEG frequency views. With three seeds per view, cross‑band members produce substantially less‑correlated errors than same‑band replicas, yielding a clean nine‑member ensemble accuracy of 0.8076/0.7242/0.7492 on the public/private/overall test partitions without test‑set adaptation or output constraints. The submitted pipeline further incorporates transductive pseudo‑label training, three EEG‑Conformer members, posterior aggregation, and a paradigm‑aware decoder. Because every 12‑trial randomization block contains three instances of each class, the final predictions are obtained by Hungarian assignment under the known block quota. On one fixed posterior pool, independent, session‑constrained, and block‑constrained decoding achieve 0.7600, 0.7758, and 0.7952 overall accuracy, respectively. The complete system reaches 0.8498/0.7718/0.7952, ranking fourth on the private split. A modality audit finds fNIRS‑only decoding at chance (0.2511 overall), while adding fNIRS to EEG changes accuracy by only +0.0025. These results identify frequency‑diverse temporal EEG modeling and protocol‑matched structured inference as the principal sources of performance in this sparse‑montage EEG‑‑fNIRS setting. The source code is available at https://github.com/XiuFan719/EEG‑fNIRS‑fuse‑method‑for‑MM‑challenge.

Authors:Qiming Bao, Sherry J. H. Feng, Kim Chester Eugenio, Meng Fon
Title: Surrogate Substitution Preserves PHI Detectability: A Multi-Detector Equivalence Study
Abstract:
Structure‑preserving de‑identification replaces protected health information (PHI) with realistic same‑type surrogates ‑‑ "Anna S." becomes "Maria S.", not [NAME] ‑‑ so that clinical text stays fluent and downstream tools keep working. But this only helps if the substitution does not itself corrupt the signal those tools rely on. We ask a narrow, testable question: on the spans a de‑identifier actually masks, can downstream PHI detectors still find the surrogate? We introduce a paired, multi‑detector evaluation protocol that (i) scores utility only on masked spans, decoupling coverage from utility; (ii) uses equivalence testing (TOST) rather than null‑hypothesis significance testing, which is uninformative at our sample size (57k paired spans); and (iii) builds a surrogate‑failure typology separating fixable generator defects from intrinsic detector limits. Across 11 detectors, 7 benchmarks, and 7 languages (1,750 documents), recall on masked spans moves from 76.1% to 74.9% ‑‑ a change our equivalence test shows is statistically equivalent to zero within a +/‑2‑point margin (p ~ 3e‑9), with detector ranking preserved. The residual loss does not reflect detectors getting worse at PHI: it concentrates in malformed and out‑of‑distribution surrogates (truncation Chicago ‑> Illino, salience loss Cedars‑Sinai ‑> Vidant). A redaction floor and an open‑source surrogate baseline indicate the effect is a property of well‑formed substitution, not of one tool. We release the evaluation subsets, scoring code, and an interactive dashboard at https://custodianai.pages.dev so the protocol can audit any structure‑preserving transform.

Authors:Xiaonan Xu, Wenjing Wu
Title: Test-time reasoning effort and unauthorized tool use in language-model agents: a prespecified equivalence study
Abstract:
Language‑model agents that execute multi‑step workflows through tool calls operate under access‑control policies that restrict which operations each role may perform. The APIs serving these agents expose a reasoning‑effort parameter that operators adjust for cost and latency. Whether this parameter also changes the rate of unauthorized tool use has not been tested by direct manipulation within a single model. We vary reasoning effort (low, max) inside GPT‑5.6 across the 14 confirmatory scenarios of TRIO‑20, a suite of 20 matched workplace triads in which a policy‑prohibited tool call is effective and its effect on the target metric is stated in the environment, effective but discoverable only through rule inspection, or ineffective. The three conditions derive from one code base and differ in two configuration fields, with identical prompts and tool sets. All analyses were prespecified in a frozen plan before confirmatory collection. Across 840 trajectories and two model tiers, no unauthorized tool call occurred. Exact one‑sided 95% limits place each arm's violation rate below 3.50% (Terra, n = 84) and 5.21% (Sol, n = 56). The interaction estimand, with a simultaneous exact 95% interval of \pm 4.34 percentage points on Terra, lies inside the \pm 7.01‑point equivalence margin. Raising effort did change behaviour, but only in inspection: rule‑probe rates rose in all conditions, most where probing carried no instrumental payoff, a pattern inconsistent with the hypothesis of targeted search (‑14.3 points, 95% CI ‑27.4 to +1.2). Raw trajectories are released at https://github.com/WenJing95/trio‑20.

Authors:Xiaogang Peng, Zeyu Han, Zichong Meng, Yiming Xie, Jihua Zhu, Gang Hua, Huaizu Jiang
Title: Surface Keypoint Representation for Multi-Object and Articulated Human-Object Interaction Generation
Abstract:
Daily activities require humans to coordinate whole‑body motion with the motion of surrounding objects. Despite recent progress in human‑object interaction (HOI) generation, most existing methods assume interactions with a single rigid object and do not extend well to scenarios involving a variable number of objects or articulated objects with diverse joint mechanisms. We propose surface keypoint trajectories as an object motion representation: for each rigid component, whether a standalone object or one part of an articulated assembly, we track a small set of non‑collinear surface points over time. This representation handles multi‑object coordination and diverse articulation mechanisms directly from point dynamics without requiring explicit joint‑type specification. To model when and where each body region contacts each object, we introduce a spatio‑temporal contact distance field that extends distance‑based contact modeling to whole‑body, multi‑object, and articulated settings. We factorize HOI generation into three stages: generating object motions from text or waypoints, predicting the contact distance field, and synthesizing whole‑body motion with contact‑guided optimization. Experiments on ParaHome, HIMO, ARCTIC, and OMOMO demonstrate better or comparable performance to existing methods across single‑object, multi‑object, and articulated interaction settings.

Authors:Tingzhang Luo, Ruizhong Liu, Yichao Liu, Cheng Fan, Yu Liu, Jianyuan Guo
Title: CROSS: Cascaded Distillation and Dual-Constraint Grounding for Remote Sensing Referring Segmentation
Abstract:
Referring Remote Sensing Image Segmentation (RRSIS) has achieved significant progress through the integration of VLMs and the Segment Anything Model (SAM). However, this progress largely relies on strong pre‑trained capabilities, while leaving two fundamental limitations insufficiently addressed: (1) Architectural Weak‑Coupling, where the unidirectional flow forces reliance on coarse VLM prompts and wastes SAM's pixel‑level structural guidance, causing localization drift; and (2) Object‑Centric Semantic Bias, where models overemphasize dominant object semantics while remaining insensitive to spatial reasoning crucial for RRSIS. Motivated by these observations, we propose CROSS, a tightly integrated paradigm for RRSIS. First, we introduce Linguistic‑Guided Cascaded Distillation (LGCD) to bridge the architectural gap, which distills SAM's geometric affinities as soft regularizers into VLM intermediate layers, injecting dense structural priors to refine localization. Second, Perspective‑Spatial Contrastive Learning (PSCL) imposes cross‑anchored constraints by mining mask‑filtered deceptive distractors and spatial‑linguistic counterfactuals as hard negatives, explicitly shattering semantic shortcuts to enforce genuine logical consistency. Extensive experiments on RRSIS benchmarks demonstrate that CROSS achieves state‑of‑the‑art performance and maintains precise localization even under severe spatial description perturbations, standing as a robust new paradigm for RRSIS.

Authors:Xiangyun Huang, Xiangchen Wang, Runfeng Lin, Yihao Xu, Kangyu Huang, Jiang Hengchen, Xiwang Dong, Lin Jiarong
Title: From Routes to Steps: Separating Semantic Progress from Local Execution in Vision-and-Language Navigation
Abstract:
Vision‑and‑Language Navigation (VLN) requires an agent to follow a route‑level instruction by executing its constituent steps from egocentric visual observations. Existing VLM‑based navigators typically supervise both capabilities through next‑action prediction alone, making progress‑tracking errors difficult to distinguish from execution errors. When an agent deviates from the route, a corrective action label may recover the next movement but does not indicate whether the agent selected the wrong sub‑instruction or failed to execute the correct one. Consequently, the agent may continue making decisions from an erroneous progress state. To resolve this ambiguity, we propose Route2Step, a framework that decouples semantic progress tracking from action generation through an explicit step‑level interface. The Instruction Analysis Module (\mathcalM_\mathrmIA) predicts this state from the global instruction and visual history. Conditioned on the predicted state and recent observations, the Action Generation Module (\mathcalM_\mathrmAG) generates local action chunks. To supervise the progress state without manual temporal labels, E‑SPA, a step‑alignment procedure, associates sub‑instructions with their corresponding portions of route‑level demonstrations. These alignments enable state supervision for incorrect progress estimates, while direct action supervision is reserved for rollout groups that repeatedly fail under the correct active sub‑instruction. On R2R‑CE, Route2Step improves SR from 48.1% to 55.3% and SPL from 43.3% to 48.2%, using 190K state‑level corrective samples while requiring only 11.5K directly action‑supervised states. Experiments in real‑world indoor and outdoor environments further demonstrate the practical applicability of Route2Step. The project page is: https://sisyphus‑hxy.github.io/Route2Step/.

Authors:Xiaolong Sun, Qichao Wang, Hangyu Li, Liang Chen
Title: Verifiable Memory: Learning Unified Memory Management with Local and Global Verifiers for Large Language Model Agents
Abstract:
Large language model (LLM) agents must retain reusable information, control a bounded active context, and recover earlier evidence during long‑horizon interaction. Existing methods commonly optimize long‑term memory (LTM) and short‑term memory (STM) separately, while unified policies are often trained primarily with trajectory‑level feedback, which provides weak credit for individual memory decisions. We present Verifiable Memory (VerMem), a framework that represents LTM, active context, and episodic history as distinct states and controls them with one memory operation policy. Seven atomic operations let the policy add, revise, or soft‑delete LTM entries; retrieve LTM into the active context; filter or summarize the active context; and restore selected episodic fragments. VerMem is initialized by supervised fine‑tuning and trained with a three‑stage reinforcement‑learning curriculum. The local verifier scores executable memory transitions, and a global verifier assesses evidence coherence and terminal‑memory consistency after task completion. These scores are combined with programmatically computed task, evidence‑recall, efficiency, and constraint signals through hierarchical credit assignment. The verifiers are used only during training. Across five benchmarks and two LLM backbones, VerMem achieves the best result on the vast majority of reported metrics and consistently outperforms strong memory baselines. Under controlled online‑token budgets on three interactive benchmarks, it also achieves the strongest efficiency‑‑performance frontier among the compared methods. Code is available at https://github.com/Sun‑SYSU‑24/VerMem.

Authors:Ning Zhu, An Chen, Mengfei Zhao, Juntao Xu, Jingze Liang, Boyuan Gu, Liang-Jian Deng
Title: Rectify Then Diffuse: Disentangling Concepts Before Denoising Trajectory Unfolds
Abstract:
Text‑to‑image diffusion models can generate individual concepts well, but they often omit or merge concepts incorrectly with multiple concepts. We trace these failures to an early coordination bottleneck: before denoising begins, prompt‑conditioned attention may allocate different concepts to strongly overlapping spatial support, which can keep their attention coupled as denoising proceeds. This observation motivates treating compositional generation as a boundary‑condition problem rather than repeatedly controlling the evolving trajectory. To this end, we propose Rectify‑then‑Diffuse (RTD), a training‑free framework that rectifies the initial allocation once before standard denoising. Firstly, we propose Soft‑Overlap Disentanglement (SOD), which converts normalized overlap between pilot concept maps into a differentiable and layout‑agnostic separation objective. Secondly, we introduce Isotropic Gradient Rectification (IGR), which normalizes the SOD gradient and applies a bounded latent displacement with a consistent scale across prompts and initializations. Extensive experiments show that RTD achieves state‑of‑the‑art compositional fidelity and robust gains. On the AE‑Bench object pair subset, RTD improves BLIP‑VQA by 45.8% and ImageReward by 19.6% over CO3 while running 2.3× faster. Code will be released at https://github.com/Z‑yiwei/rectify‑then‑diffuse

Authors:Yi Yang, Zhennan Chen, Mingfeng Lv, Hanlei Li, Zhengsen Ruan, Lvqing Yang
Title: Convex-Hull-Neighborhood Smooth Dual Generalization: Controlling Local Correction Propagation in Offline RL
Abstract:
Offline reinforcement learning (offline RL) can benefit from nearby out‑of‑distribution (OOD) actions, but estimation errors at these actions may be amplified by bootstrapping. Existing regularization and local‑generalization methods control either the admissible OOD region or the influence of generalized targets, often through separate mechanisms. We propose Convex Hull Neighborhood Smooth Dual Generalization (CSDG), which expresses the Bellman backup as an in‑sample value target plus a CHN‑local correction. This formulation makes the generalized contribution explicit and separates it from the in‑sample reference path. The correction is obtained by smoothing in‑sample‑oriented and OOD‑oriented candidates sampled at different perturbation radii. A mixture coefficient lambda scales its contribution to each backup, while the recursive discount remains gamma. Under boundedness and fixed perturbation kernels, we derive an exact one‑step correction identity, a time‑varying iterate bound, and a fixed‑point bound that depends only on the branch discrepancy at the fixed point. We further characterize the implicit policies induced by the idealized operators and give a conditional non‑degradation criterion. The practical algorithm approximates these quantities using asymmetric bounded noise and expectile regression, without exact support classification or an additional pessimistic OOD penalty. Experiments on Gym‑MuJoCo and AntMaze show strong aggregate performance and stable value estimation. Code is available at: https://github.com/YOUNG‑fnxm/CSDG

Authors:Ping Li, Chenhao Ping, Jie Song, Mingli Song
Title: Channel-wise Dynamic Knowledge Distillation via Adaptive Sample Generation for Action Recognition
Abstract:
Knowledge Distillation (KD) offers a promising yet underexplored path for compressing large action recognition models. However, existing KD methods suffer from two key limitations: 1) reliance on fixed input samples leads to suboptimal feature alignment between the frozen teacher (larger model) and the learnable student (smaller model), and 2) applying a uniform distillation strength for all channels fails to account for their varying importance in capturing distinct knowledge (e.g., motion tempo or magnitude) across training epochs. This motivates us to develop an Adaptive Sample‑aware Channel‑wise Dynamic (ASCD) KD approach, which operates in two stages. First, we use an adaptive sample generation module to create updated samples by incorporating semantics from sample gradients, which are derived by minimizing a feature loss weighted by channel centroid frequency differences at each layer. Meanwhile, crucial motion‑related details are preserved by applying a Gaussian mask to frequency features. Second, we employ a channel‑wise dynamic distillation module to train student on these generated samples, guided by sample gradients and feature frequencies. For efficiency, samples are updated periodically rather than per epoch. Extensive experiments on three video benchmarks (UCF101, Kinetics‑400, Something‑Something‑v2) and two image datasets (CIFAR‑100, ImageNet) demonstrate the state‑of‑the‑art performance of our method. Code is available at https://github.com/mlvccn/ASCD_KD_Action.

Authors:Pei Li, Sihan Chen, Delong Ran, Tianshuo Cong
Title: FakeI2V-Bench: Benchmarking the Applicability of Image-level Deepfake Detectors for Deepfake Video Detection
Abstract:
Recent advances in video generation models have significantly intensified the deepfake threat, yet the current deepfake video detection benchmarks remain underdeveloped. In particular, the effectiveness of image‑level detectors in the video domain has not been systematically assessed. To fill this gap, we present FakeI2V‑Bench, a benchmark for evaluating state‑of‑the‑art video‑level deepfake detectors in challenging scenarios, with a particular focus on systematically assessing the performance of image‑level deepfake detectors in the video domain. FakeI2V‑Bench comprises 97,548 videos, containing content generated by the latest powerful generation models and covering a broader range of categories. Using this dataset, we conduct a systematic evaluation of eight video‑level detectors and twelve representative image‑level detectors. Experimental results show that the best‑performing image‑level detector achieves an 80.16% AUC, slightly outperforming the strongest video‑level detector (i.e., 79.99% AUC). Going beyond benchmarking, we present IV‑Bridge, a general framework that enhances the applicability of image‑level deepfake detectors to videos. IV‑Bridge employs a random forest model with statistical features to aggregate frame‑level predictions, allowing eleven image‑level detectors to surpass state‑of‑the‑art video‑level approaches, with the best‑performing variant achieving a 93.80% AUC. Overall, FakeI2V‑Bench establishes a rigorous benchmark for deepfake video detection and introduces a novel pathway for extending image‑level detectors to the video domain, offering new insights and directions for future research. Code and data are available at https://github.com/CryptoAILab/FakeI2V‑Bench.

Authors:Ethan Bito, Yongli Ren, Estrid He
Title: Position Bias Undermines Preference Consistency in Listwise LLM-Based Reranking
Abstract:
Large language models (LLMs) have emerged as promising listwise rerankers for recommender systems, but their reliability under equivalent candidate permutations remains unclear. Since recommendation candidates form an unordered set, a reranker should not depend on the arbitrary order used to serialize them. However, decoder‑only LLM rerankers can allow input order to affect model scores, pairwise preferences, and rankings. We study how position bias affects the ranking process induced by LLM‑based rerankers. Instead of measuring only changes in final ranked lists, we treat rankings produced under equivalent candidate permutations as observations of an induced preference system. We introduce an evaluation framework measuring pairwise preference instability, global preference inconsistency, and listwise output consistency. This framework characterizes candidate‑order sensitivity at the pairwise, global, and output levels. Experiments across multiple LLMs, datasets, and list lengths show that these consistency measures are closely aligned, but can diverge from recommendation effectiveness and marginal position‑exposure bias. Improving relevance or flattening exposure across positions does not necessarily restore stable pairwise preferences, globally coherent preference structures, or consistent ranked outputs. These results show that reducing marginal exposure skew is insufficient to establish ranking‑function validity in LLM‑based reranking. Code is available at https://github.com/ejbito/InvariRank .

Authors:Seonaeng Cho, Minjee Seo, Minju Seol, Juil Park, Joon Ho Kwon, Kyungho Yoon
Title: Automatic Patient-Specific Microwave Ablation Planning Accelerated by a Physics-Guided Deep Learning Model
Abstract:
Microwave ablation (MWA) is a promising minimally invasive treatment for liver tumors, but its therapeutic outcome strongly depends on patient‑specific planning of antenna insertion trajectory, power, and treatment duration. Accurate numerical simulation can provide physically reliable ablation predictions; however, its high computational cost limits its use in optimization‑based planning, where repeated forward evaluations are required. To address this issue, we propose a digital twin‑based automatic planning framework that combines a neural ablation prediction model with a genetic algorithm. The model was trained on multiphysics simulation data generated from patient‑specific tumor and vessel structures, antenna configurations, and treatment conditions, and was used as a fast forward model during planning. The prediction model achieved a Dice score of 95.1%, enabling accurate deep learning‑based optimization. In 13 unseen planning cases, the proposed method improved ablation efficiency by 54.3% and reduced organ damage by 55.0% compared with clinician‑defined planning, while slightly shortening the insertion path length by 3.3%. Most generated plans were also judged clinically applicable by MWA specialists. Furthermore, the framework enabled approximately 420‑fold faster planning than numerical‑simulation‑based planning, demonstrating its potential as a fast digital twin for quantitative and personalized MWA treatment planning. The code is available at: https://github.com/SeonAengCho/MWA‑Planning.git

Authors:Yibo Yuan, Jiacheng Fu, Jiangtong Zhu, Yi Li, Jianhua Han, Meng Tian, Zhuohan Liu, Zhiwei Xiong, Hang Xu, Jianwu Fang, Jianru Xue
Title: SUV: Future Scene Understanding as Video Generation for End-to-End Driving
Abstract:
End‑to‑end driving requires a coherent understanding of future scenes, yet existing methods model these scenes using task‑specific heads and output formats, with limited scalability. Can video generation instead provide a shared predictor? We introduce SUV, a unified end‑to‑end driving framework that casts future Scene Understanding as Video generation using a pretrained video foundation model. SUV models future appearance, semantics, relative depth, and instance‑level dynamics as video streams with a shared video expert, without stream‑specific visual prediction heads. Through joint video‑action attention, the action expert attends to the latent representations of all future streams and generates the ego trajectory. Experiments show that SUV directly predicts all four future streams, while controlled ablations show that structured future supervision and direct future‑stream access yield higher trajectory planning scores. With only a single front camera and no candidate‑trajectory selection, SUV outperforms a broad set of recent state‑of‑the‑art methods on both NAVSIM‑v2 splits, achieving 91.0 EPDMS on navtest and 36.9 on navhard. On the long‑tail WOD‑E2E benchmark, SUV achieves a competitive RFS of 7.94.

Authors:Seokho Han, Dongwei Wang, Jinhee Kim, Yiran Chen, Kang Eun Jeon, Huanrui Yang, Jong Hwan Ko
Title: TASQ: Temporal-Adaptive Bit Sparsification Quantization for Diffusion Models
Abstract:
Static quantization assigns one weight precision to every denoising step. To preserve quality, that precision must accommodate the most quantization‑sensitive step, even though many other steps can tolerate fewer bits. The resulting model may satisfy its memory budget, but it repeatedly pays worst‑case arithmetic throughout the denoising trajectory. We introduce Temporal‑Adaptive Bit Sparsification Quantization (TASQ) to separate these two costs. TASQ stores one shared maximum‑precision weight buffer and learns a Temporal‑Spatial LSB Mask that selects a lower effective precision for each layer and denoising stage by truncating least‑significant bits. Storage therefore remains fixed by the worst case, while BitOPs decrease at less sensitive stages without per‑stage weight copies or runtime search. A Temporal‑Precision Engine maps the learned schedule to bit‑serial execution, where cycles scale with effective precision and switching precision has no measured cycle overhead. On PixArt‑Sigma, SANA‑1.6B, and SDXL‑Turbo, TASQ achieves quality comparable to static quantization with less computation. Together with the Temporal‑Precision Engine, it reduces execution cycles by 25 to 50 percent over static quantization and by 6.1 to 7.5x over a naive static 8‑bit bit‑serial execution. Code is available at https://github.com/seokho‑han/tasq.

Authors:Yizhuo Jia, Jingyun Hua, Yuanxing Zhang
Title: CAPE-T2V: Captioner-Anchored Prompt Enhancement toward Two-Sided Conditioning Alignment in Text-to-Video Generation
Abstract:
Text‑to‑video (T2V) diffusion transformers (DiTs) are trained with detailed video captions, whereas inference often relies on user prompts rewritten by a prompt enhancer (PE). Prior work has improved generation by optimizing the PE, the DiT, or both; some methods have also sought to narrow the training‑inference mismatch through shared schemas. Yet even within a shared schema, inference‑time PE outputs and DiT training captions may still differ in detail selection, information organization, descriptive granularity, and phrasing. We refer to this residual mismatch as the PE‑Caption gap and introduce CAPE‑T2V, a two‑step Captioner‑Anchored Prompt Enhancement framework toward two‑sided conditioning alignment in T2V generation. First, CAPE‑T2V constructs three types of PE training examples, pairing captioner‑generated targets with concise source captions, detailed source captions, or pseudo user prompts derived from those targets. It then fine‑tunes the PE to map each input to its paired target. Second, CAPE‑T2V fine‑tunes the DiT on video‑derived captions rewritten by the Anchored PE; the same PE rewrites user prompts at inference. Relative to a baseline using the same caption schema, CAPE‑T2V achieves higher aggregate scores on StoryEval, VBench‑2.0, and T2V‑CompBench across Wan2.2 and LTX‑2.3. Further, CAPE‑T2V exhibits a smaller PE‑Caption gap than the baseline: its DiT fine‑tuning captions are closer in distribution to inference‑time PE outputs, as measured by squared maximum mean discrepancy in a fixed embedding space. Overall, these results support CAPE‑T2V as an effective approach to mitigating the PE‑Caption gap. The project is available at https://github.com/yizzz927/CAPE‑T2V.

Authors:Tristan Wu, Daniel Chin, Junan Zhang, Junyan Jiang, Yansen Jing, Gus Xia
Title: DDSynth-RL: Audio Synthesizer Inversion via Discrete Diffusion with Reinforcement Learning
Abstract:
Synthesizer inversion is challenging for two main reasons: 1) Distinct parameter configurations can produce perceptually similar sounds. 2) Parameter‑space losses often fail to reflect rendered audio similarity, while the synthesizer being a non‑differentiable black box prevents simple audio‑domain supervision. To address the one‑to‑many mapping induced by the first challenge, we formulate synthesizer inversion as conditional generation over discrete synthesizer parameters and use masked discrete diffusion as the generator. This treatment additionally avoids the fixed‑order assumption of autoregressive models and the continuous‑relaxation mismatch of flow matching when modeling categorical synthesizer controls. To address the second challenge, we further fine‑tune the model with GRPO‑style audio‑domain rewards computed from rendered outputs. Experiments on Dexed show that, after supervised training, the discrete diffusion model is competitive with autoregressive and flow‑matching baselines, and reward‑based fine‑tuning further improves out‑of‑domain audio matching performance. Code and demos are available at: https://github.com/DDSynth‑RL/DDSynthRL.

Authors:Xiaoyu Tao, Mingyue Cheng, Bokai Pan, Chuang Jiang, Huanjian Zhang, Tian Gao, Yaguo Liu, Qi Liu, Enhong Chen
Title: CastFSR: A Fast--Slow--Reflect Agentic Reasoning Framework for Context-Aware Time Series Forecasting
Abstract:
Time series forecasting is fundamental to decision‑making in complex systems, where future dynamics are influenced not only by historical observations but also by evolving contextual features. Recent advances in large language models (LLMs) have extended forecasting beyond numerical extrapolation toward context‑aware reasoning. However, existing approaches often lack explicit mechanisms to identify relevant contexts, reason about their impacts, and validate forecasts against temporal and domain constraints. In this work, we propose CastFSR, an agentic framework that formulates context‑aware forecasting as a Fast‑‑Slow‑‑Reflect workflow. In fast thinking, CastFSR profiles observations and selects lightweight forecasters to construct a data‑driven forecast prior. In slow deliberation, it retrieves contextual evidence, adaptively determines informative look‑back windows, and reasons about how contexts reshape future dynamics. In reflection, it iteratively refines forecasts to ensure temporal, contextual, and domain consistency. CastFSR supports both training‑free inference with off‑the‑shelf LLMs and efficient deployment through a two‑stage SFT and reinforcement learning strategy that transfers its orchestration capability to compact LLMs. Extensive experiments on public datasets demonstrate that CastFSR consistently outperforms representative baselines. Our code is available at https://github.com/Xiaoyu‑Tao/CastFSR.

Authors:Linhan Xia, Rui Liu, Zhaofeng Zhang, Yihao Wang, Binrui Shen, Shengxin Zhu
Title: LoCA: Forward-Only LLM Tuning after One-Shot Calibration with Local Credit Assignment
Abstract:
Parameter‑efficient post‑training reduces the number of trainable parameters, but still requires repeated end‑to‑end backpropagation through the frozen backbone. Every adaptation step therefore needs backward‑capable hardware and must store or recompute activations. We ask whether this repeated backward chain can be replaced by a one‑time calibration. We introduce Local Credit Assignment (LoCA), a two‑stage method for small‑shift adaptation. One probe backward pass fits a low‑rank map at each transformer block from the final prediction error to a local hidden‑state correction. LoCA then reuses these maps to form blockwise regression targets from forward activations and fits low‑rank adapters with closed‑form ridge solves. No further backbone backward pass is required. We evaluate LoCA on five discriminative benchmarks with Qwen2.5 models from 0.5B to 14B. In 16 of 25 reported task‑‑scale comparisons, LoCA yields lower evaluation cross‑entropy than the corresponding LoRA run. Its measured full‑run GPU peak, including calibration, is 26‑‑29% lower than LoRA's. After calibration, its CPU steady‑state memory is 36‑‑52% lower and its per‑pass time is 43‑‑48% lower. A shared scale‑normalized candidate set is reused across all tested Qwen2.5 sizes and on SmolLM2‑1.7B. LoCA thus amortizes global credit assignment into one calibration and enables later forward‑only tuning when repeated backpropagation is impractical. The code associated with this paper is available \hrefhttps://github.com/Xia12121/LoCAhere.

Authors:Jong Hak Moon, Minjun Kim, Minjun Kim
Title: Clinically-Grounded Hierarchical Classification for Consistent Chest X-ray Interpretation
Abstract:
Accurate chest X‑ray interpretation is inherently hierarchical. Clinical decisions depend not only on what abnormality is present but where it is situated, requiring reasoning from broad anatomical systems down to specific pathological findings. Yet existing automated systems largely treat this as a flat classification problem, failing to capture inter‑level dependencies or enforce coherence between coarse and fine predictions. We propose CHASE (Classification with Hierarchical Analysis and Structured Enforcement), a unified single‑stage framework that mirrors radiologists' coarse‑to‑fine reasoning through a clinically driven three‑level taxonomy of 9 anatomical regions, 17 sub‑regions, and 28 pathological findings. CHASE jointly optimizes multi‑level supervision, cross‑level probability alignment, and a hierarchy‑violation penalty within a shared Vision Transformer backbone. This ensures that fine‑grained findings are anatomically supported by their coarser‑level context rather than predicted in isolation. Experiments demonstrate that CHASE outperforms flat and hierarchical baselines across all levels while achieving superior probabilistic hierarchy consistency, with level‑wise attention maps confirming anatomically grounded predictions. Code is available at: https://github.com/yejix‑ai/CHASE.

Authors:Subrat Prasad Panda, Blaise Genest, Arvind Easwaran
Title: Neurosymbolic Reasoning with Incremental Knowledge for Sample Efficient Hierarchical Reinforcement Learning
Abstract:
(Flat) Reinforcement Learning (RL) agents face significant challenges in environments with sparse rewards that require long‑horizon reasoning. A compelling approach to improve sample efficiency is to incorporate knowledge into learning and decision‑making. In standard Hierarchical RL (HRL), knowledge is encoded in a fixed, non‑updatable form, such as architectural choices, and remains unchanged throughout learning. With fixed HRL, reasoning with incremental knowledge learned during exploration is impractical before sufficient environmental knowledge is acquired, leading to poor sample efficiency. In this work, we propose neurosymbolic HRL with \em Incremental Knowledge (InK): symbolic high‑level components perform \em symbolic planning (e.g. using D^) on an updatable representation of current InK, while low‑level goal‑conditioned neural modules learn motion primitives through experience using reward shaping. Experiments on navigation tasks demonstrate that incorporating InK substantially improves sample efficiency. Additionally, to perform \em optimal symbolic planning given \em prior knowledge about the world, we develop Belief World Tree Search. The code is available at https://github.com/CPS‑research‑group/ink_bwts.

Authors:Keisuke Suzuki
Title: Internalising the Identity Primitive: Cryptographic Individuality for an Autonomous Agent on a Public Blockchain
Abstract:
A software agent on a public blockchain accumulates authority and economic stakes, raising the engineering question of what makes it count as an individual. The paper's central contribution is a shift of trust root for the key‑to‑weights binding of agent identity: from hardware, operator, or wrapper trust to cryptographic assumptions enforced by a pinned implementation (liveness, key custody, oracle trust, and the underlying software stack remain external). We design and deploy on Solana devnet an agent whose neural‑network weights are a deterministic function of its private key. The binding is committed in zero knowledge at genesis, re‑checked against that commitment at every state transition, and signed by the agent into an on‑chain history unforkable once finalized; in a PoC‑tier extension, a protocol‑imposed metabolic cost is debited each cycle from a key‑derived economic account, adding a consumption‑side economic‑viability constraint to the key‑history‑economy triple. Empirically, the agent completes a 2.36‑day on‑chain run with two host‑side resumptions but no rejected transition, at bounded per‑transition verification cost; a substituted substrate is rejected on chain, and independently keyed agents diverge as predicted while a same‑key control stays at zero. To our knowledge, this is the first published on‑chain agent whose identity primitive is itself a cryptographic invariant re‑checked at every state transition. The resulting transition‑time invariant instantiates the cryptographic individuality proposed by Suzuki 2026's Artificial Externality framework.

Authors:Kleyton da Costa, Bernardo Modenesi
Title: When Should Graph Attention Be Sparse? Learning a Per-Edge Tsallis Index
Abstract:
Graph attention normalizes neighborhood scores with softmax, the maximum‑entropy choice under Shannon statistics. But homophilic and heterophilic graphs want different attention shapes, and one fixed normalization cannot serve both. We propose LTGA (Learnable Tsallis Graph Attention), a graph attention layer whose Tsallis entropic index q is learned jointly with the weights, interpolating continuously between heavy‑tailed (q\!<\!1), softmax (q\!=\!1) and compact‑support (q\!>\!1) attention at four granularities from a global scalar to a per‑edge index, under a bounded reparameterization that starts every model at the GAT baseline. Across eight benchmarks at ten seeds, LTGA‑Edge takes the best average rank (2.75), but the omnibus test does not reject (p\!=\!0.199) and learning q does not beat searching it: a validation‑tuned frozen grid reaches 61.4%, tuned α‑entmax 62.2% and a capacity‑matched q\!\equiv\!1 control 62.0%, against 61.7% for LTGA‑Edge. What the learned index buys is one run instead of a grid, and an interpretable mechanism: where q leaves 1, it prunes 42% of attention coefficients to exactly zero, and those edges are selectively the wrong ones, restoring them costs 7.1 points, while random pruning at the same rate costs 13.0 more. Project page: https://kleyt0n.github.io/ltga

Authors:Pingqing Zheng, Jiayin Qin, Fuqi Zhang, Zishen Wan, Shang Wu, Yu Cao, Caiwen Ding, Yang Katie Zhao
Title: LACE: Large Language Model Aided Multi-Agent Framework for Agile RISC-V Instruction Extension
Abstract:
Domain‑specific Instruction Set Architecture eXtensions (ISAX) are widely adopted in the RISC‑V ecosystem to accelerate emerging workloads, but implementing and validating ISAXes across different cores remains slow and fragmented. Existing frameworks still require per‑core interface adaptation, and differential testing often breaks once either the microarchitecture or the ISAX changes. We present LACE, an LLM‑aided multi‑agent workflow that translates natural‑language ISAX intents into a compact two‑level IR (operation‑level and HDL task‑level), performs retrieval‑guided localized RTL edits over large repositories, and closes the loop with a compiler‑agnostic riscv‑formal checking flow (assuming RVFI availability or instrumentation). Across four embedded RISC‑V cores, LACE raises pass@1 generation accuracy from near‑zero to 72.8% within our evaluation setup, while improving code localization and reducing integration rework. The code of LACE is available at https://github.com/UMN‑ZhaoLab/LACE.

Authors:Minghao Zou, Qingtian Zeng, Shangkun Liu, Cong Liu, Paul L. Rosin, Guanghui Yue, Jun Liu, Wei Zhou
Title: Modeling Scientific Experiment Scenes: Dataset and Model
Abstract:
Scene Graph Generation (SGG) is fundamental to structured visual understanding, yet existing benchmarks focus mainly on daily life images and overlook scientific experiment scenes with specialized instruments, task‑specific experimental semantics, and dense, fine‑grained physical relations. These scenes are increasingly important for automated experimental analysis and smart education. To bridge this gap, we introduce PhysScene, the first SGG dataset for physical experiment scenes, providing densely annotated scene graphs and benchmarks under multiple supervision and protocol settings. PhysScene further exposes two key algorithmic challenges for SGG: pronounced long‑tail relational predicate distributions and a substantial visual‑textual semantic gap. To address these challenges, we propose the Cross‑Modal Dual‑Path Generator (CM‑DPG), a model for robust open‑vocabulary SGG. The model enhances object‑level semantic representations through joint visual‑textual encoding and improves relational reasoning using complementary visual and geometric cues. We also incorporate relation‑aware pre‑training, caption‑derived pseudo‑supervision, and adaptive weighting to support balanced learning across head and tail predicates. Extensive experiments on PhysScene and VG150 show that CM‑DPG achieves competitive performance across multiple evaluation settings, with ablation studies validating the contribution of each component. The dataset and code are publicly available at https://github.com/ZMH‑SDUST/CM‑DPG.

Authors:Ruiqi Lyu, Alistair Turcan, Bryan Wilder
Title: Population-Robust Feature Selection via Generalized Welfare Optimization
Abstract:
Choosing which features to collect is a deployment decision: the same limited questionnaire, test panel, or sensor set may need to serve several heterogeneous populations. Standard feature‑selection methods typically optimize for one large population, while existing robust approaches tend to learn one shared model for every population. We introduce PopFS, a method for learning one shared, deployable feature set that is robust to population differences while letting each pop‑ ulation train its own model. PopFS uses a tunable welfare objective that lets practitioners balance overall predictive ben‑ efit against stronger protection of the populations that benefit least. To make this objective practical at scale, PopFS first uses multitask sparse learning to reduce the candidate pool, then searches directly over hard feature sets by ranking promising additions and swaps and fully refitting only a shortlist. Across eight population splits from six prediction tasks drawn from five tabular and public‑health datasets, PopFS consistently achieves strong average and worst‑population performance while scaling to thousands of candidate features. A 43‑state COVID‑19 nowcasting study further shows that changing the welfare objective can improve the least‑served states with lit‑ tle change in average performance and yields an interpretable change in the selected symptom signals. Our code is available at https://github.com/Rachel‑Lyu/PopFS.

Authors:Nina Bodelot, Soufiane Belharbi, Eric Granger
Title: Test Time Adaptation Methods for Point Cloud Registration in Laparoscopic Surgery
Abstract:
3D point cloud registration in laparoscopic surgery estimates the transformation between an intraoperative organ reconstructed from video and its preoperative mesh. Because ground‑truth transformations are unavailable for real data, supervised networks are trained on synthetic organ pairs. At test time, real reconstructions differ from synthetic data and are noisy, sparse, and occluded, which degrades correspondence estimation. Test‑time adaptation (TTA) can reduce this domain shift, but existing methods mainly rely on logits, entropy, class prototypes, or cache memories unavailable in registration. Registration also involves paired inputs with an asymmetric shift that primarily affects the intraoperative cloud. We analyse and modify state‑of‑the‑art TTA methods from three families to 3D registration: model, normalization, and input adaptation. We analyze four representative approaches based on auxiliary‑task model updates, backpropagation‑free token purging, feature alignment, and layer‑normalization calibration. We modify them to handle asymmetric shifts between preoperative and intraoperative point clouds and replace classification‑based entropy objectives. Using a correspondence‑based model trained on clean synthetic source data, we evaluate adaptation to corrupted synthetic and real target data on P2P and P2ILReg. For synthetic targets, we apply eight corruptions, including uniform noise and global density reduction, at five severity levels. All methods improve registration on P2P, whereas normalization adaptation degrades performance on P2ILReg. Considering the computational overhead of backpropagation‑based adaptation, input adaptation is the most promising option for laparoscopic surgery, providing low inference latency and consistent error reductions across datasets. Code: https://github.com/ninaa‑git/survey_pc_registration_tta

Authors:Walid Saidi
Title: MutMem: Cryptographically Authorized Mutation in Persistent Agent Memory
Abstract:
Persistent agent memory must adapt as later outcomes change earlier evidence, yet mutable retrieval weights create an attribution problem: reviewers must distinguish authorized adaptation from database tampering. We present MutMem, an authorized‑mutation protocol in HOM‑AIMOS, a persistent agent‑memory engine. MutMem retains memory content, records signed positive and negative outcome evidence without age‑based expiry, and commits each nontrivial weight change as a housekeeper‑authorized transition. Each transition binds a terminal provenance node, signer epoch, quantized old and new weights, a no‑fork predecessor, and two domain‑separated SHA‑256 commitments. Ed25519 verification runs in both the database writer and a portable verifier. Content classified as poison‑likely is retained with signed, revisable labels used by recall as trust evidence. We evaluate utility, mutation integrity, and poisoning adaptation. HOM‑AIMOS answers 459/500 LongMemEval questions correctly under LLM judgment (91.8%). On LoCoMo, it obtains 74.12% judged accuracy and, under a separate upstream‑compatible protocol, 58.20 token F1. A native suite passes all declared authorization, topology, tamper, signer‑epoch, and post‑mutation‑recall cases; median signed‑transition latency is 4.865 ms. In a declared N=100 PoisonedRAG adaptation, no injected poison appears in attacked top‑5 disclosures (0/100; 95% Wilson upper bound 3.70%), while induced target‑answer attack success among 98 clean‑negative targets is 1/98 (1.02%). A preregistered four‑arm ablation attributes the retrieval reduction to signed stored labels: the retriever selects poison for 94/100 targets when epistemic policy is bypassed and 0/100 when labels are restored. MutMem provides evidence of integrity, authorization, traceability, and historical continuity; it does not establish content truth.

Authors:Sukhrobbek Ilyosbekov
Title: Localize, Don't Beautify: Client-Side Control of Image-Editing APIs for Cosmetic Surgery Previews
Abstract:
Ask a commercial image editor to preview a cosmetic procedure and it will often change more of the face than the request names: a nose edit can also smooth skin or alter lighting. Existing methods for confining an edit to one region require access to the model's internals, which a public editing API does not expose. We ask how much control is possible from the client side alone. In a pilot benchmark, six commercial editing configurations and one mask‑based inpainting model perform facelift‑style jaw‑neck and rhinoplasty edits at three levels of client‑side control: the prompt alone; cutting the edited region out of the response and pasting it back onto the original photograph through a landmark‑derived mask (a masked composite); and asking the model itself to inpaint inside the mask where supported. Of 210 attempted edits, 196 could be scored. ArcFace cosine measures identity preservation; a CIELAB pixel‑change ratio measures how much change lands inside the requested region rather than a protected facial zone. On the 12 frontal faces the regional metric could score, the masked composite improved localization over the paired prompt‑only output by a median of 0.446 (95% face‑clustered bootstrap interval 0.421‑0.457) while changing the requested region about as much. Editors differed in edit strength versus identity retention, and the one inpainting model we tested did not beat the simple composite. Against each face's input‑to‑postoperative baseline, no editor moved its outputs closer to the postoperative photograph in identity‑embedding terms. This is a study of control, not clinical accuracy: no surgeons rated the outputs, and each condition was generated once. Within that scope, keeping a surgical preview inside its intended region needs no access to the model; a mask and composite on the client enforce it across every editor tested, at low provider cost.

Authors:Peter Werner, Tobia Marcucci, Daniela Rus
Title: Biconvex Optimization for Smooth Minimum-Time Trajectories around Convex Obstacles
Abstract:
We present a biconvex approach for minimum‑time motion planning around convex obstacles that is guaranteed to converge, is anytime, and supports derivative constraints to arbitrary order. We jointly convexify the minimum‑time objective and all derivative constraints through a change of variables, and handle collision avoidance via time‑varying separating planes, reducing the problem to a biconvex program. This program is solved by alternating between computing maximum‑margin separating planes and optimizing the trajectory. By only adding planes for obstacles that the current iterate collides with, the trajectory can jump around obstacles and escape local minima. The method is guaranteed to converge starting from a simple collision‑free polygonal curve. In our experiments on drone navigation and dual‑arm bin unloading, we find that the proposed method reliably produces high‑quality trajectories with computation times comparable to state‑of‑the‑art decomposition‑based motion planners, while handling a larger class of problems and being substantially more robust to bad initialization. Project page:https://wernerpe.github.io/bmtp‑website/

Authors:Xuehang Guo, Pingyue Zhang, Ruiyi Zhang, Zhenhailong Wang, Hanrui Lyu, Heng Ji, Tong Sun, Qingyun Wang, Manling Li
Title: CURV: Enhancing Chart Understanding Through Curriculum Visual Grounded Reasoning
Abstract:
Chart question answering (CQA) requires multimodal large language models (MLLMs) to integrate visual comprehension with logical reasoning, yet current models struggle with accurate visual grounding and coherent reasoning chains. While extrinsic chain‑of‑thought prompting and visual cues significantly improve performance, current MLLMs lack intrinsic visual grounded reasoning capabilities, leading to inaccurate perception and reasoning disconnected from visual evidence. To address these limitations, we propose CURV, a curriculum learning framework that develops intrinsic visual reasoning capabilities by reformulating CQA as multi‑step visual grounded reasoning, where each step coordinates logical reasoning with dynamic visual grounding through spatial attention concentration. To assist model learning, we further introduce CCQA, a three‑level curriculum dataset with scalable synthetic generation across diverse chart types and reasoning patterns. Our curriculum systematically progresses from basic single‑operation reasoning to complex multi‑chart compositional tasks. Experiments demonstrate that CURV achieves up to \uparrow20.50% improvements over baselines and is generalizable to real‑world benchmarks (up to \uparrow12.30%) and out‑of‑domain multimodal reasoning tasks (up to \uparrow10.20%), validating the effectiveness of internalizing visual reasoning with dynamic grounding for enhanced chart understanding capabilities. Code is available at: https://xhguo7.github.io/CURV/.

Authors:Ruida Cheng, Tejas S. Mathai, Benjamin Hou, Qingqing Zhu, Zhiyong Lu, Matthew McAuliffe, Ronald M. Summers
Title: A Unified 2D Framework for DeepLesion Detection, Segmentation and Short Report Generation
Abstract:
In previous work, we integrated large language models (LLMs) into the lesion segmentation model based on the ULS23 DeepLesion dataset, using short‑form findings from the reports. In this study, we developed a unified 2D lesion analysis framework that integrates LLM‑based reasoning, lesion bounding box detection, segmentation, and radiology report generation from the original DeepLesion dataset. In the testing phase, we achieved relatively high lesion bounding box detection accuracy with mAP50 of 70.1%, mAP50‑95 of 46.4%; Lesion segmentation performance with a Dice score of 62.6%; short report generation accuracy with BLEU_1 score of 64.3%, BLEU_4 score of 49.6%, METEOR of 34.7%, and ROUGE_L of 60.1%. In this work, we address the challenging issue of segmentation in the original DeepLesion dataset and achieve a 28.5% Dice score improvement over the nnUNet lesion segmentation model. We also integrated spatial and anatomical context into the DeepLesion short report generation. We released the implementation, dataset, and models on Github. https://github.com/ruida/2D_DeepLesion_Foundation

Authors:Priyanka Bajaj
Title: Evaluation Blindness: How Silent Measurement Failures Corrupt AI Systems from Training to Deployment
Abstract:
AI systems can fail silently. The failure propagates through training loops, evaluation pipelines, and production monitoring stacks until downstream harm makes it visible. This paper introduces evaluation blindness: a measurement function M exhibits evaluation blindness with respect to failure class F when it produces readings indistinguishable from a healthy state while the system is actually failing, with no auxiliary signal flagging the gap. The problem surfaces at two lifecycle stages the literature has treated separately. At training time, reward models are gamed, importance‑sampling corrections are silently miscalculated, and benchmark contamination inflates fine‑tuning evaluations, all while loss curves look healthy and gradient updates proceed normally. At deployment time, monitoring fails to catch six classes of production failure, including an Operational category that is 100% silent by structural definition. We provide a formal detectability predicate unifying both stages. Four training‑time case studies trace concrete breakdowns, including a real implementation bug in TRL PR #6594 where gradients are corrupted as loss decreases normally. A six‑class taxonomy validated against 50 real‑world incidents from court documents and regulatory filings finds that 53% of verifiable public failures were silent. A failure budget framework ties acceptable failure rates to use‑case risk class. The implication is direct: measurement infrastructure is a correctness concern across the full AI lifecycle, not just at evaluation time. Data, code, and taxonomy schema are at https://github.com/priyanka25aug/llm‑failure‑taxonomy.

Authors:Shuai Wang, Haodong Chen, Yu Yin, Shengyao Zhuang, Bevan Koopman, Guido Zuccon
Title: Search, Inspect, Fetch: Exploiting Structure-Aware Boolean Retrieval for Deep-Research Agents
Abstract:
Existing deep‑research agents use a Search‑‑Visit workflow that retrieves whole webpages without considering the structure they expose through titles, headings, sections, and metadata. This prevents agents from directly constraining retrieval to parts of a webpage and often carries irrelevant content into their context. We introduce \textscSieve, a search‑‑inspect‑‑fetch strategy driven by a Boolean Query Language (BQL): it searches webpage fields to filter candidates, uses an interchangeable ranker to order them, presents structure‑rich result cards for inspection, and fetches only selected sections. Across three QA collections, \textscSieve is more accurate than the strongest conventional Search‑‑Visit configuration on each collection while using 20.7‑‑50.6% fewer tokens. Boolean filtering improves every tested ranker, and the accuracy‑‑context advantage persists across retriever choices and agent backbones. Our implementation is included in the SkimSearchAgent library at https://github.com/ielab/skim‑search‑agent.

Authors:Zixuan Wang, Yuhong Chen, Yuxuan Zhu, Guidong Lei, Zhiluohan Guo, Yu Zhao, Kun Wang, Bangyang Hong, Kangle Wu, Yabo Ni, Anxiang Zeng, Cong Fu, Hui Li
Title: Knowledge-Geometry Decoupling: Refreshable Pretrained Transfer for Streaming Recommendation
Abstract:
Industrial recommenders increasingly adopt the pretrain‑then‑transfer paradigm, yet behavioral distribution drift raises two questions: what to learn from behavior sequences, and how to transfer the learned knowledge while the pretrained model is continually refreshed. To resolve them, we propose Knowledge‑Geometry Decoupling (KGD). For what to learn, conventional next‑token prediction treats adjacency as dependency and may encode spurious transitions across unrelated sessions. We introduce Behavioral Multi‑Token Prediction (BMTP) to retain only collaboratively or semantically related future items as supervision, yielding cleaner and more transferable behavioral knowledge. For how to transfer, pretrained knowledge and task‑specific geometry impose conflicting optimization demands on shared parameters. To handle it, KGD assigns them to separate parameter sets: a refreshable encoder owns behavioral knowledge, while a task learner reads contextualized encoder states through read‑only cross‑attention and writes task‑specific geometry through Anchored Calibration Residual (ACR) orthogonal to the pretrained embedding. The decoupled ownership enables continual knowledge refresh without task‑gradient interference or invalidating downstream adaptation. KGD improves over strong pretrain‑transfer baselines by 4‑12% on eight public benchmarks and sustains its advantage over a 90‑day production stream where baselines show no gains. KGD has been fully deployed in Shopee. In a live A/B test on Shopee Homepage Search, it increases GMV per user by 1.75% and advertising revenue by 1.53%, demonstrating its high practical value. We provide the core implementation of KGD at https://github.com/FuCongResearchSquad/KGD4REC.

Authors:Yu Yang, Xuemeng Yang, Licheng Wen, Lingdong Kong, Xiaobin Hu, Dongyue Lu, Wei Chow, Xiyan Huang, Yuxiang Feng, Yue Liao, Jianbiao Mei, Daocheng Fu, Rong Wu, Pinlong Cai, Ran Yi, Ying Tai, Jiangning Zhang, Botian Shi, Yong Liu, Shuicheng Yan
Title: Quo Vadis, World Modeling?
Abstract:
Continually improving agents require dynamic interaction feedback beyond static supervision, yet direct real‑environment interaction is costly, slow, unsafe, and hard to parallelize. World modeling offers a natural intermediate proxy that allows agents to query lower‑cost, more controllable feedback before committing to real actions. Classical world models instantiate this proxy primarily through future physical‑state prediction, a formulation useful yet narrow for agents that require actionable feedback beyond raw state transitions. In this work, we conceptualize Agent‑Centric Interactive World Proxies, shifting the fundamental paradigm from physical state transitions to agent‑usable information transitions, such as execution outcomes, retrieved experiences or skills, and verification signals, broadening the scope of world modeling to provide versatile feedback for continually improving agents. To systematically map this design space, we organize world proxies into six functional forms based on their feedback modalities: dynamics, spatial, execution, memory/experience, skill, and reward/verification proxies, which together characterize the primary ways world modeling serves agent improvement. We further analyze how these proxies empower agents across three progressive levels: L.1 Inference‑Time Guidance, where proxy outputs enrich in‑context information for superior decisions; L.2 Training‑Time Optimization, where proxy outputs yield rewards, critiques, or synthetic rollouts for policy learning; and L.3 Agent‑Proxy Co‑Evolution, where real‑environment evidence continuously updates both the proxy and the agent for co‑evolution. Ultimately, this work recasts world modeling into an agent‑centric paradigm, establishing a roadmap for building world proxies that empower agents to plan better, learn faster, and evolve continually.

Authors:Junliang Ye, Kenkun Liu, Guocun Wang, Yang Li, Yansong Qu, Chunshi Wang, Jingwei Xu, Yunhan Yang, Zibo Zhao, Jiachen Xu, Jiaao Yu, Lifu Wang, Zhihao Liang, Zhuo Chen, Chunchao Guo
Title: Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing
Abstract:
Recent advances in image generation have demonstrated the potential of unified multimodal models that integrate understanding, generation, and editing. However, unified 3D modeling remains constrained by scarce multimodal data, particularly the lack of large‑scale and geometrically consistent editing data. To address this limitation, we propose Hunyuan3D‑Buffalo 1.0, a unified framework supporting 3D understanding, text‑to‑3D generation, instruction‑guided 3D editing, and text‑grounded part generation within a single architecture. To enable scalable training, we construct an 87M‑scale 3D multimodal corpus, comprising 25M understanding samples, 50M text‑to‑3D pairs, and 12M editing pairs generated using Nano3D‑v2. Architecturally, the framework combines Hunyuan3D‑VLM for semantic, structural, and spatial understanding with Hunyuan3D DiT for high‑fidelity 3D synthesis. The VLM provides multimodal semantic conditions for generation, while editing and part generation additionally condition the diffusion process on the source object representation to preserve its overall structure and unedited regions. Extensive experiments show that Hunyuan3D‑Buffalo 1.0 achieves state‑of‑the‑art or leading performance on text‑to‑3D generation and 3D editing benchmarks, while exhibiting strong understanding and part‑generation capabilities. Our analysis further shows that both generation and understanding improve editing, demonstrating the effectiveness of unified 3D multimodal training. Project Page: https://tencent‑hunyuan.github.io/Hunyuan3D‑Buffalo1.0/

Authors:Şuayp Talha Kocabay, Talha Rüzgar Akkuş, Kamer Ali Yuksel
Title: ARCHead: Activation-Metric Residual Correction for Large Language Model Output Heads
Abstract:
Weight‑only quantization substantially reduces the storage of large language model (LLM) transformer blocks, but practical backends often retain the final language‑modeling head (LM‑head) in BF16 or FP16. Quantizing this projection naively can strongly perturb the vocabulary‑logit distribution. We present ARCHead, a packed LM‑head compressor that combines a quantized low‑rank core, group‑wise INT4 residuals, and a low‑rank correction fitted in an activation‑derived metric. ARCHead stores no dense BF16 head and reduces persistent LM‑head storage by 3.7‑3.9x. On Qwen3‑8B‑Base, it uses 25.6% of BF16 head storage while attaining 1.007 relative perplexity; storage‑matched naive INT4 yields 1.14‑1.16. Replacing the BF16 head left by AWQ or bitsandbytes adds only 0.006‑0.007 cross‑entropy, with less than 2% throughput change in our measurements. ARCHead therefore complements block quantizers by compressing the large output projection they can leave untouched. Code is available at https://github.com/suayptalha/archead.

Authors:Lecheng Yan, Jianze Lin, Yichong Zhang, Ben Pan, Wenxi Li, Chenyang Lyu, Liting Zhou, Cathal Gurrin
Title: Crayotter: Learning Long-Horizon Video Editing Agents via Group-Relative Preference Backpropagation
Abstract:
Long‑horizon video editing agents receive final‑product feedback only after many interdependent decisions. Yet editing quality is subjective, admits multiple valid solutions, and is not meaningfully calibrated across heterogeneous requests, making a global scalar objective both ambiguous and temporally uninformative. Our key observation is that fixing the request, materials, and production constraints converts this subjective objective into an ordinal comparison among directly comparable alternatives. We introduce Group‑Relative Preference Backpropagation (GRPB), which transforms same‑task rankings into zero‑sum advantages and redistributes them as bounded credit over semantic editing segments. A lagged allocator and guarded transmission prevent current judgments or unreliable estimates from directly shaping the same rollout group. We manually construct a project‑disjoint, horizon‑stratified suite of realistic editing tasks for training and controlled evaluation. Across matched baselines, credit interventions, external benchmarking, and blinded human evaluation, GRPB improves both editing behavior and rendered products. The resulting 9B Crayotter model surpasses several proprietary systems on AgenticVBench, supporting task‑local preference reduction as a practical approach to learning from subjective, delayed outcomes. Code and all supporting materials are publicly available at https://github.com/idwts/Crayotter.

Authors:Ronglong Bao
Title: Stuck on "A": Diagnosing and Repairing Interface Injury in Attention-to-KDA Linearization of a 0.6B Language Model
Abstract:
We convert 21 of 28 full‑attention layers of Qwen3‑0.6B‑Base into KDA (Kimi Delta Attention) linear‑attention layers on a single consumer‑grade GPU budget, and ask a simple question: what exactly does the conversion break? After surgery, hidden‑state alignment and end‑to‑end KL distillation drive the student close to its teacher in perplexity, yet multiple‑choice accuracy stays near random chance (25‑29% vs. the teacher's 50.6% on C‑Eval). Using a four‑permutation diagnostic that rotates answer options while holding content fixed, we show the model sticks to option labels (predicting "A" 81% of the time; 106/161 questions keep the same label under all four rotations) rather than following answer content ‑‑ an interface injury that standard distillation metrics cannot see. A 1,000‑step format‑targeted completion‑only KL stage repairs the interface (+12.48 points on C‑Eval, label‑stickiness roughly halved), after which persona SFT and one round of on‑policy DPO preserve benchmark scores within noise. We release code, weights, recipes, and the full audit trail, and distill the engineering lessons ‑‑ including an FP32‑master failure mode in which bf16 optimizer updates are silently swallowed ‑‑ that made convergence possible at this budget.

Authors:Zetong Xiong, Qiao Zhao, Jun Zhang, Xueying Lyu, Zhi Li, Yixiang Tu, Xiaowen Yang, Yunjie Zhang, Yufeng Wang, Zhe Zhang, Kaize Yu, Hanwen Du, Zhongkai Sun, Zhuoxin Liu, Zekun Lin, Jianwen Yang, Ruining Chen, Ying Zhang, Tingxuan Pan, Ke Chen, Shubin Han, Chuanhao Sun, Yehua Yang
Title: BulkPR-Bench: Benchmarking Queue-Level Governance of Interacting Pull Requests
Abstract:
Coding‑agent benchmarks increasingly cover long‑horizon, end‑to‑end, and interactive development, but typically retain one requested outcome or a fixed change sequence. Sequential policies can process a pull‑request (PR) queue one candidate at a time, but when queued PRs interact, maximizing safe delivery can require jointly deciding which changes to merge and in what order. We introduce BulkPR‑Bench, an executable benchmark in which an agent must recover consequential PR relations and return a large safe subset in executable order under a rolling‑release protocol. The suite contains 581 newly authored candidate PRs on frozen snapshots of 18 real repositories. Registered state‑by‑state repository execution, including hidden safety checks, validates the gold relation graph; an exact oracle then computes the largest safe subset. Our primary metric, Relational Delivery Score (RDS), scores safe delivery and correct rejection over relation groups from the realized merge trace; Global Safety‑Gated Yield (Global‑SGY) separately measures strict delivery of the realized whole‑queue plan. Under the buffered primary protocol with batch size K=32, the three highest RDS estimates among the six models are 66.6%, 62.0%, and 57.9%, compared with 53.1% for the strongest sequential baseline. Only 8 of 324 model runs complete a queue exactly. Critical‑relation recall ranges from 35.2% to 57.7%, and diagnostic runs supplied with the gold relations show substantial remaining headroom. Gains on relation groups therefore do not yet translate into dependable whole‑queue governance.

Authors:Shu Quan, Tianfang Hao, Sitong Fang, He Geng, Jiayi Zhou, Boyuan Chen, Kaile Wang, Donghai Hong, Juntao Dai, Yaodong Yang, Jiaming Ji
Title: A Blind Spot in Alignment: Quantifying Biosecurity Risks in Large Language Models
Abstract:
Large Language Models (LLMs) are accelerating biological research, yet this same capability poses a critical biosecurity threat: models that assist in protein engineering can equally be prompted to generate predicted toxin‑like sequences, potentially lowering the barrier to biological misuse. Current safety evaluations, however, operate in natural language and cannot determine whether a model‑generated amino acid sequence is biological gibberish or a computational risk signal. To address this evaluation blind spot, we introduce SPIKE‑Bench, coupling 631 curated toxin‑design prompts across seven functional categories with the SPIKE funnel, a three‑stage protocol that filters output through compliance, biological plausibility, and predicted toxicity, producing stage‑level diagnostics and an aggregate function‑aware metric: the Functional Harmfulness Rate (FHR). An audit of 32 LLMs reveals that most models freely comply with toxin‑design requests; FHR is driven primarily by biological generation capability rather than safety alignment, reaching 50.7%; and Refusal Rate fails to predict functional risk. As a first step toward mitigation, we provide BioSafe‑Guard, a domain‑specialized classifier that substantially reduces predicted functional risk while preserving benign utility. We release SPIKE‑Bench and BioSafe‑Guard at https://github.com/PKU‑Alignment/SPIKE‑Bench to support more rigorous biosecurity evaluation of LLMs.

Authors:Rasvik Kudum, Max Corbett, Hitansh Paliwal, Romaisa Fatima, Thomas Jiralerspong, Sneheel Sarangi
Title: When Policies Change Probabilities: Modular Decision-Making for LLM Code Review
Abstract:
LLM code reviewers often estimate patch risk and make approval decisions in one prompt. A probability should depend on evidence; costs should determine the action taken from it. We test whether four deployed reviewer interfaces preserve this separation using 15,792 responses on 720 candidate patches, with one that passed and one that failed an archived test harness for each of 360 repository issues. In matched calls with the patch and monitor evidence fixed, replacing an equal‑cost policy with a 10:1 false‑accept policy changes reported failure probabilities by 13.6 to 16.9 percentage points on average. For every reviewer, the actions returned under the high‑cost prompt are worse than rejecting all patches. Applying the same high‑cost rule to probabilities elicited under equal costs reduces loss for all four systems, showing that probability elicitation itself contributes to the excess loss. We also evaluate a modular pipeline that elicits risk without policy information, combines an independent monitor score, and applies costs in code. Relative to calibrated reviewer‑only scores, the pipeline improves average probability accuracy and, at equal costs, reduces mean loss by .073 per issue while accepting 58 to 68% of patches. At 10:1, it accepts none and matches reject‑all. Downstream policy can therefore change the probability it is meant to use, motivating separate evaluation of risk, outside evidence, and action.

Authors:Mohammad Nasir Uddin, Rahnuma Tabassum Orpita, Asaduzzaman Anik, Eklachur Rahman Bhuiyan, Marjahan Risalat, SM Wali Ullah, Asif Ahamed
Title: CT-HEG: A Bidirectional, Timestamp-Attributed Event Graph for ICU In-Hospital Mortality Prediction - An Architectural Ablation Study
Abstract:
Accurate ICU mortality prediction requires modeling irregular clinical observations across heterogeneous entity types. Existing sequence models handle irregular sampling but ignore typed relational structure; existing graph models assume fixed‑interval inputs. We introduce the Continuous‑Time Heterogeneous EHR Graph (CT‑HEG) schema and evaluate which architectural choices drive predictive performance. CT‑HEG encodes each ICU stay as a typed, timestamped graph with three node types (visit, vital, lab_event) and 2D edge attributes (t_hours/48, value_norm) encoding timing and value without imputation. We instantiate CT‑HEG as CHIRP‑Net, a four‑layer heterogeneous GATv2Conv network, evaluated on MIMIC‑IV v3.1 (31,142 ICU stays, LOS>=48h, 13.4% mortality) with five seeds and bootstrapped confidence intervals, against logistic regression, mTAND, a Transformer, and GRU‑D, plus an ablation study. CHIRP‑Net achieved 5‑seed mean AUROC 0.8449+/‑0.0071 (AUPRC 0.4958+/‑0.0209); the ensemble achieved AUROC 0.8618 (95% CI: 0.8485‑0.8745). Removing reverse edges disconnected observation nodes from the visit readout, cutting AUROC by 0.1968+/‑0.0073. Time‑attentive edge features contributed 0.0247+/‑0.0093 AUROC. Collapsing heterogeneous edge types into one relation (7x fewer parameters) outperformed the full model on all seeds. Post‑calibration ECE was 0.0307. Temporal and demographic subgroup analyses were explored but not reported here, pending follow‑up work. Bidirectional connectivity was necessary for the model to use its inputs at all, and CT‑HEG was reasonably well calibrated after validation‑fitted temperature scaling. These results support CT‑HEG for irregular EHR data, while external validation, a pre‑specified temporal evaluation, and a demographic fairness audit remain necessary before any claim of robustness. Code: https://github.com/nasiruddinstudents‑ctrl/chirp‑net‑mimic‑iv.

Authors:Jianshuo Dong, Yiming Liu, Maosen Zhang, Nan Deng, Xu Peng, Xiaoping Zhang, Tianwei Zhang, Jie Zhang, Han Qiu
Title: Your Agentic LLMs Secretly Encode Latent Signals of Indirect Prompt-Injection Exposure
Abstract:
Agentic LLMs are vulnerable to indirect prompt injection (IPI) attacks, e.g., malicious side‑tasks hidden in external tool results. While many efforts have sought to address the threats, little is known about the internals of agentic LLMs when they are exposed to IPI attacks, a condition which we call IPI exposure. In this paper, we study this problem in depth from three aspects. (1) Probing: Across six models, including the giant 753B‑parameter GLM‑5.2, simple linear probes trained on pre‑generation hidden states can predict LLMs' IPI exposure. These probes achieve 90%+ AUROC on unseen attacks, agent instructions, and task suites; they exhibit high robustness under adaptive attacks and in cross‑lingual settings. (2) Defense: Our CoT measurement reveals a recognition‑‑action gap: though models encode such signals, they often fail to translate them into safe actions. We then introduce AGRI, a probe‑gated reasoning‑based defense that prepends anti‑injection reasoning on demand. On difficult AgentDojo settings, AGRI substantially reduces attack success rate, e.g., from 34.6% to 0% on Qwen3.5‑27B, while largely maintaining clean‑task utility. (3) Explanation: We introduce an analysis framework that identifies natural‑language explanations most strongly correlated with probe‑captured signals. The resulting profiles differ across models: latent signals can align with either direct IPI‑exposure claims or indirect operational cues. Code is available: https://github.com/jianshuod/IPI‑exposure‑signal.

Authors:Yuxuan Liu, Zhaochen Su, Yuhao Zhang, Jiahe Guo, Zhongwei Xie, Huihao Jing, Lingyun Xie, Qing Zong, Yauwai Yim, Zhixiong Zhang, Haoran Li, Yangqiu Song
Title: Rethinking Self-Evolving Agent Skills: Feedback Dynamics over Multiple Rounds
Abstract:
Self‑evolving skill systems promise to improve agents by turning execution feedback into persistent skill updates without changing the underlying model. Yet it remains unclear when further evolution helps, how successful and failed trajectories shape revision, and whether extra test‑time computation can recover the same gains. To address these questions, we present a controlled evaluation framework across five benchmarks and three models. Our primary study contains 42 feedback runs across 14 supported model‑benchmark settings. Within each setting, we hold the executor and optimizer configuration, revision procedure, validation rule, and round budget fixed, while varying only the feedback shown to the optimizer: successes and failures (Normal), failures only, or successes only. Evolution is sparse: only 55 of 388 candidates establish byte‑distinct validation bests. Validation‑based selection chooses an evolved skill in 11 of 14 settings, nine of which improve released‑test performance. All 11 selections come from feedback conditions that include failed trajectories, although the relative ranking of Normal and Fail‑only varies across settings. Validation and downstream evaluations on test, robustness, and transfer sometimes favor different feedback views. A broader SearchQA analysis covering eight models shows similarly sparse, feedback‑dependent dynamics. In the GPT‑5.5 test‑time‑scaling controls, oracle Parallel Sampling comes within 0.43 points of the evolved SearchQA skill but remains 30.96 points behind on SpreadsheetBench; Sequential Refinement recovers neither gain. Overall, persistent skill self‑evolution is better understood as sparse, validation‑filtered search with model‑ and benchmark‑dependent returns, rather than steady improvement from additional rounds. The implementation is available at https://github.com/HKUST‑KnowComp/rethinkskill.

Authors:Yutaro Yamada, Kei Hiroshima, Nozomu Yoshinari, Kento Uchida, Shinichi Shirakawa
Title: BBOWP-Bench: Evaluating LLMs on Black-Box Optimization Word Problems
Abstract:
Formulating an optimization problem strongly affects the quality of the final solution, yet good formulations usually require substantial expertise. Recent studies have therefore examined how to automatically derive optimization problems from natural‑language descriptions, but existing benchmarks focus on settings where objectives and constraints can be written explicitly as mathematical expressions. Many practically important problems are naturally treated as black‑box optimization (BBO) problems, in which only objective values are observable, and the functional form is unavailable. In BBO, the search space design, a part of the problem formulation, and the selection of the optimization algorithm are crucial for problem‑solving. Automating these processes with large language models (LLMs) is a significant challenge. This paper introduces Black‑Box Optimization Word Problems (BBOWP), a novel problem setting in which a system must infer both a search space and an optimization algorithm from a natural‑language description of a black‑box optimization task. To support research on this setting, we establish the BBOWP Benchmark Suite (BBOWP‑Bench), a dataset and evaluation framework for BBOWP. Each instance combines a natural‑language problem description, an executable evaluation environment, and a human‑designed baseline formulation, allowing evaluation of both search‑space design and algorithm selection. Using this benchmark, we provide the first evaluation of LLMs and show that current LLMs are capable of selecting suitable algorithms based on the given evaluation budget. However, they sometimes struggle with search space design, particularly in identifying important variables and balancing their ranges when the problem description is less informative or the search space is highly problem‑specific. Our code and dataset are available at https://github.com/shiralab/bbowp‑bench.

Authors:Meem Arafat Manab, Marinella Quaranta, Ilaria Angela Amantea, Sheyla Leyva-Sánchez, Víctor Rodríguez-Doncel
Title: Extracting ODRL Policies from Business Process Models: A Graph Traversal Approach to Compliance-by-Extraction
Abstract:
Organisations maintain large corpora of process models expressed in the Business Process Model and Notation (BPMN), yet the normative content encoded in those models, the obligations, permissions, and prohibitions that govern participant behaviour, remains inaccessible to policy infrastructure. The Open Digital Rights Language (ODRL) is the emerging lingua franca of machine‑readable policy, but authoring ODRL at scale is slow, expert‑intensive work, and generation by large language models introduces well‑documented risks of structural invalidity. We present a pipeline that resolves this gap by extracting ODRL policies automatically from BPMN XML, grounded in the observation that BPMN control flow encodes deontic modalities by construction. The pipeline traverses the BPMN process graph, classifies each task as an odrl:Duty or odrl:Permission via a reachability check, and introduces a novel treatment of intermediate catch events as odrl:Prohibition rules with lifting constraints, capturing waiting semantics that prior approaches have dropped. The result is a compliance‑by‑extraction approach in which auditable, interoperable ODRL policies are derived directly from the process models organisations already maintain. The implementation and a live demonstration are available at https://github.com/manabcodes/bpmn2odrl and https://bpmn.linkeddata.es/.

Authors:Yi Yang, Zhennan Chen, Yihong Zhuang, Tiehan Fan, Yinan Chen, Jian Li, Jian Yang, Ying Tai
Title: RoMeRL: Balancing Feedback Coverage and the Memory-Reward Trap in Self-Evolving Agent Memory via Reduced-Order Utility States
Abstract:
Learning‑based memory systems for self‑evolving LLM agents face two tightly coupled challenges. First, trajectory‑indexed utilities grow with the interaction history, thereby dispersing limited feedback over an ever‑expanding state space. Second, because trajectory‑level rewards are jointly assigned to co‑retrieved memories, irrelevant experiences may receive misleading utility updates and consequently enter the memory‑reward trap. To address these challenges, we introduce Reduced‑Order Memory Reinforcement Learning (RoMeRL), which represents the growing trajectory‑indexed utility space using a fixed‑dimensional per‑task memory state factorized by outcome polarity and memory dynamics. RoMeRL incorporates new experiences through a fixed set of semantic coordinates whose contents are updated or replaced over time, thereby concentrating feedback over a bounded utility support. Theoretically, we show that this reduced‑order parameterization increases the average feedback received by each utility coordinate and characterize the steady‑state occupancy of erroneous coordinates under a generic coordinate‑transition model. Empirically, across ALFWorld and LifelongAgentBench, RoMeRL improves task performance, reduces the Cold‑Q ratio by 80.0%, increases feedback density by approximately 6.0 times, reduces the maintained memory size by 84.4%, and cuts LLM calls by 21.1%. These results show that reduced‑order utility states support efficient self‑evolving agent memory while limiting persistent reward contamination. Code is available at: https://github.com/YOUNG‑fnxm/RoMeRL

Authors:Jiayu Chen, Xiaoyu Wu, Rongshan Gao, Maoliang Li, Zihao Zheng, Xinhao Sun, Hailong Zou, Guojie Luo, Xiang Chen
Title: EchoCache: Energy-Guided Cross-Modal Caching for Efficient Audio-Driven Video Generation
Abstract:
Audio‑driven video generation (A2V) has achieved promising progress in synthesizing temporally coherent and audio‑visually aligned videos, yet its inference remains expensive due to the iterative denoising process of diffusion models. Existing caching methods mainly exploit temporal redundancy in visual features while overlooking the cross‑modal alignment of A2V, where audio drives visual generation with highly non‑uniform temporal importance. In this paper, we identify two levels of misalignment in existing A2V caching methods: temporal‑semantic and computation‑storage misalignment. To address them, we propose EchoCache, an energy‑guided cross‑modal caching framework for efficient A2V generation. EchoCache leverages audio time‑frequency energy as a saliency anchor to guide latent‑level cache updates and further introduces a dynamic timestep‑latent caching mechanism with quantized cache management for joint efficiency and memory optimization. Extensive experiments on mainstream A2V models show that EchoCache consistently improves the latency‑quality trade‑off while preserving generation quality and audio‑visual consistency. In particular, on Wan2.2‑S2V over the EMTD benchmark, EchoCache achieves a 2.46x speedup with the best overall performance. Code is available at https://github.com/IF‑LAB‑PKU/EchoCache.

Authors:Jiawei Wang, Hao Yu, Yongzhen Hu, Xinyi Yang, Tao Ni, Xin Zhan, Junbo Chen, Xiaowei Zhou, Ruizhen Hu, Sida Peng
Title: InfiniSplat: Implicit Gaussian Decoding for Large-Baseline Monocular View Synthesis
Abstract:
Single‑image feed‑forward 3D Gaussian Splatting (3DGS) aims to directly generate a renderable 3D scene representation from one input image, avoiding the cost of multi‑view capture and per‑scene optimization. However, existing methods are often constrained by a pixel‑aligned representation, where Gaussians are predicted from fixed image‑grid locations. Such pixel‑aligned primitives can produce promising nearby‑view renderings, but they remain weakly coupled to underlying scene surfaces and struggle to preserve coherent structures under large viewpoint shifts. We present InfiniSplat, a feed‑forward single‑image 3DGS framework that moves from a pixel‑aligned representation toward a surface‑aligned representation. InfiniSplat constructs this representation by first using geometry‑guided sampling to place 2D supports according to depth‑induced local surface structure, and then applying a query‑conditioned implicit decoder to predict Gaussian attributes from the image features queried at these supports. By grounding support locations in geometry while decoupling Gaussian prediction from fixed pixel centers, InfiniSplat produces Gaussian layouts that better follow scene surfaces and reduce scattered primitives caused by grid discretization. Across multiple cross‑dataset NVS evaluations, InfiniSplat achieves state‑of‑the‑art performance compared with single‑image feed‑forward baselines, and demonstrates zero‑shot generalization from Hypersim indoor synthetic training to complex open‑world scenes. Project page: https://zju3dv.github.io/InfiniSplat.

Authors:Sitong Gong, Caixin Kang, Tianyu Yan, Guo Chen, Bo Zheng, Kaipeng Zhang, Yunzhi Zhuge, Xiang Ruan, Huchuan Lu, Yifei Huang
Title: GROVE: Growing and Reasoning over Temporally Stratified Memory from Streaming Video Experience
Abstract:
A wearable assistant should both answer questions about its visual history and recognize when that history is useful to the present situation. Existing video‑memory systems primarily support question‑conditioned recall, whereas proactive assistants typically use separate memory and control mechanisms. We introduce GROVE, a training‑free framework that supports both behaviors with one memory grown causally from a continuous video stream. GROVE retains fine‑grained perceptual evidence and incrementally consolidates it into time‑stamped moments, coherent episodes, and recurring cross‑day patterns. Each stratum is paired with a scale‑native retrieval skill for locating an observation, replaying an activity, or traversing long‑range regularities. Reactive QA and proactive assistance share this memory and access interface, differing in whether retrieval is initiated by a user query or the current situation. Across multiple benchmarks including the challenging MM‑lifelong and EgoServe, GROVE achieves the best results among the compared methods. Controlled ablations show that the temporal strata and their access skills are complementary, with patterns providing the largest benefit when evidence spans multiple days. Code will be available at https://github.com/SitongGong/GROVE.

Authors:Yuzhi Huang, Weijue Bu, Ziyi Xiong, Jie Wu, Fanding Huang, Jingyan Jiang, Zhi Wang
Title: ChainVLA: Chaining Vision-Language-Action Queries through a Unified Execution State for Long-Horizon Manipulation
Abstract:
Humans perform long‑horizon manipulation by retaining knowledge of what earlier actions have established while continuously adapting the motion underway. By contrast, action‑chunked vision‑language‑action (VLA) policies repeatedly replan from the current input at each query. Existing methods preserve either long‑term task evidence through memory or short‑term motion through action reuse and ensembling, leaving the cross‑query handoff incomplete. We introduce ChainVLA, a 1.2B‑parameter VLA policy that chains successive queries through a joint and revisable execution state. Progress Context combines a recurrent Working State with sparse event memory to carry observation‑derived task progress, while Motion Tail feeds the preceding prediction's unexecuted continuation into state construction and action generation. Together, the two components condition a decoder that regenerates each action horizon under the latest observation, allowing the carried state to guide the next prediction without fixing it. ChainVLA reaches 62.8% average success on RMBench and 98.8% across four LIBERO suites, while removing Motion Tail or Progress Context reduces RMBench success to 11.2% and 3.0%, respectively. These asymmetric ablations are consistent with motion continuity helping preserve the observation stream from which task progress is inferred.

Authors:Ziyue Zheng, Linli Shi, Bingkun He, Wen Jiang, Ziyun Wang
Title: TRACE: Ergodic Trajectory Optimization for Active Scene Reconstruction
Abstract:
Existing active reconstruction systems with Gaussian‑splatting maps select observations greedily, optimizing a single next‑best‑view (NBV) at each step and connecting the chosen views by short‑horizon path planning. This greedy decoupling disregards the global structure of scene information, producing inefficient trajectories that waste sensing capacity in transit between selected views. In this work, we study active reconstruction as an ergodic coverage problem: the time‑averaged spatial statistics of the sensor trajectory should match a target information distribution induced by the current map. Our approach derives this target distribution online from uncertainty and visibility, and calculates ergodic trajectories via a kernel‑ergodic horizon planner with gradient flow and footprint depletion, closing the loop between mapping and trajectory optimization. We thoroughly evaluate TRACE on the Replica dataset against the Next‑Best‑View (NBV) baselines, improving PSNR by 1.5 dB. Code: https://github.com/spikelab‑jhu/trace‑active‑reconstruction.

Authors:Shijie Ren, Xiting Wang, Meng Li, Yujie Guo, Yunhang Yao, Ziheng Peng, Xunlong Wang, Yuetan Chen, Haoyang Zhou, Yunlong Liang, Fandong Meng
Title: Self-Improving Large Language Models via Progressive Experience Evolution
Abstract:
Large language models (LLMs) capable of self‑improvement require not only effective policy optimization, but also a principled mechanism for transforming transient interaction experience into persistent model capabilities. Existing self‑improvement paradigms remain fragmented: test‑time methods can explicitly extract experience but cannot internalize it into model parameters, whereas training‑time optimization methods can update model parameters but lack an explicit mechanism for accumulating transferable experience. Bridging these two paradigms requires a critical intermediate stage that remains underexplored, namely \emphexperience distillation. To address this gap, we propose SPEE (Self‑Progressive Experience Evolution), a unified post‑training framework that sequentially performs explicit experience evolution followed by implicit policy optimization. During explicit experience evolution, SPEE reflects on trajectories collected from multiple interactions to extract, verify, and progressively evolve transferable experience, which is subsequently internalized into the policy through privilege‑guided On‑Policy Self‑Distillation (OPSD). During implicit policy optimization, reward‑driven reinforcement learning leverages these internalized priors to explore novel solution strategies. In the experience evolution stage, a continuously evolving global experience pool consolidates knowledge from both successful and failed trajectories, filters out low‑utility experience, and mitigates post‑hoc rationalization induced by individual trajectories. Experiments on five mathematical reasoning benchmarks demonstrate that SPEE consistently outperforms both test‑time and training‑time self‑evolution baselines across three model scales. The source code is available at https://github.com/rrrsj/SPEE.

Authors:Hongjie Zhou, Shiqin Wang, Haoyang Chen, Haonan Guo, Di Wang, Juhua Liu, Fu Lin, Yong Luo
Title: RSVideo: Are Your Vision-Language Models Ready for Remote Sensing Videos?
Abstract:
Remote‑sensing videos enable real‑time observation of changes in target attributes, short‑term activities, and scene evolution. They record motion, actions, interactions, and scene changes that cannot be captured by isolated images. Existing models primarily target single images or discrete temporal observations spanning a long time range. However, a unified evaluation setting for assessing vision‑language models on continuous remote‑sensing video understanding remains lacking. We introduce RSVideo‑10K, a remote‑sensing video dataset comprising 10,773 instances, 1.47 million frames, and 17.02 hours of footage, containing both unmanned aerial vehicles and satellite platforms. Its fixed evaluation benchmark, RSVideo‑Bench, contains 2,731 test instances and evaluates two complementary aspects of remote‑sensing video understanding: L1 Perception and L2 Reasoning, spanning seven capability groups and 17 tasks. Evaluations show that current vision‑language models still struggle to recover small local evidence, track short‑lived states, and use scene‑constrained spatial relations. Based on this analysis, we further propose RSVideo, a reinforcement learning framework for small‑target spatiotemporal focusing that selects question‑relevant regions across frames and suppresses redundant background tokens. RSVideo achieves a maximum absolute improvement of 9.01% with InternVL3.5‑14B and attains the highest accuracy of 40.63% with Qwen3.6‑27B across 26 open‑source vision‑language backbones. Codes will be available at https://github.com/HongjieZhou0329/RSVideo.

Authors:Yu Zhang, Ruiqi Li, Changhao Pan, Ke Lei, Xiang Yin, Cheng Yang
Title: SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks
Abstract:
Speech and audio generation is often needed in animation dubbing, audio drama, movies, advertising, games, podcasts, and short‑video production. In these scenarios, creators may need to design voices without reference recordings, control speaker styles with natural language, support acoustic scenes with environments and audio effects, and later reuse the designed voices. Therefore, it is important to support multi‑speaker speech and audio generation for both instruct and zero‑shot tasks. The instruct task requires a caption of the environment, speaker styles, and fine‑grained content, while the zero‑shot task uses reference audio together with the same fine‑grained content. We address these tasks from both the data and model sides. First, we propose SwanData‑Caption, which cleans raw speech and audio data, adds targeted synthetic coverage, and annotates diverse and accurate multi‑level captions. Then, we propose SwanTale, a multi‑speaker expressive speech and audio generation model that supports both zero‑shot and instruct tasks. We introduce SwanVAE to support high‑quality multi‑audio‑modality generation. Then, we adopt reward‑conditioned quality control and Engram conditioning, along with Unified MoE for multi‑task and multi‑audio‑modality modeling. In addition, we use curriculum learning and GRPO post‑training to let the model progressively learn and strengthen its capabilities. Experimental results show that SwanTale leads on multiple key zero‑shot and instruct metrics, achieves the best expressiveness scores in both tasks, and supports complex instruct generation involving multi‑speaker speech and audio. Demos can be found at https://swanaigc.github.io/#swantale.

Authors:Daeyoung Roh, Donghee Han
Title: Before Reasoning Can Fail: Pre-Evidence Procedural Failures in Agentic RAG
Abstract:
Agentic retrieval‑augmented generation (RAG) systems can fail before evidence‑conditioned reasoning is tested: an agent may retrieve candidate snippets but finalize without inspecting them. We study this failure mode as a procedural property of the agent trajectory, decomposing wrong answers into pre‑evidence discipline failures and post‑gold‑read failures using saved tool‑call traces, retrieved evidence, read passages, and final answers. Across 12,000 paired trajectories on HotpotQA, 2WikiMultiHopQA, and MuSiQue, the two failure types are largely non‑redundant: the both‑trigger rate is in [11.2%, 13.1%] across regex and spaCy entity extractors. We then evaluate Read‑Gate, a minimal runtime invariant requiring an agent to read after search and before finalization. Forced reading improves LLM‑Acc by 14.9‑19.9 points on trajectories that would otherwise skip reading and by 3.2‑9.4 points on full minimal‑reasoning cells. Additional diagnostics show that larger hidden thinking budgets do not necessarily increase evidence inspection. Together, these results indicate that evidence‑gathering should be evaluated as a trajectory‑level control problem, separately from answer‑side reasoning.

Authors:Daeyoung Roh, Donghee Han
Title: HALT: Verification-Aware Stopping for Retrieval-Augmented Search Agents
Abstract:
Retrieval‑augmented search agents answer multi‑hop questions by repeatedly issuing search queries and accumulating evidence. This creates a stopping problem: after the necessary evidence has appeared, further retrieval often adds cost, latency, and distracting context rather than useful information. We frame stopping as evidence coverage rather than generator confidence, and introduce HALT, a lightweight verification‑aware policy that leaves the search agent unchanged. Given expected hop claims, HALT stops only when cumulative evidence supports each required claim. Across three multi‑hop QA benchmarks, HALT reduces redundant search while largely preserving exact match. We separate a deployable setting, where hop claims are generated from the question, from a diagnostic upper bound that uses gold supporting‑fact annotations: generated claims give smaller but still exact‑match‑preserving savings, while gold claims show the larger savings available when hop targets are clean. Baseline comparisons and ablations show that this behavior is driven by claim‑evidence alignment rather than generic sufficiency, fixed stop positions, or lexical overlap. Open‑corpus pilots further suggest that HALT abstains when coverage cannot be reliably verified. Overall, evidence coverage provides a practical runtime control signal for improving retrieval‑augmented agents without retraining or modifying the host agent.

Authors:Zizhong Ding, Junxian Li, Kai Liu, Shaoqiu Zhang, Xiao Xiao, Linghe Kong, Yulun Zhang
Title: ET-Prune: Evidence-Aware Dynamic Budgeting for Visual Token Pruning in Text-Rich MLLMs
Abstract:
Visual token pruning reduces the inference cost of multimodal large language models, but a fixed token ratio is poorly matched to text‑rich inputs. In OCR‑centric tasks, decisive evidence can be a small number, label, or field whose relevance is specified by the question; indiscriminate pruning can erase that evidence while retaining visually salient but irrelevant regions. We present ET‑Prune, a training‑free framework that casts pruning as evidence allocation. It derives question‑conditioned evidence from a decoder‑side partial query‑key block, safeguards text‑like spatial regions, and converts evidence uncertainty and density into a sample‑specific token floor. Three progressive middle‑layer events then move the sequence toward this budget, retaining more tokens for diffuse or text‑dense evidence and pruning concentrated evidence more aggressively. At the observed point estimates from one deterministic pass per configuration, ET‑Prune leads or ties among pruned methods in all six backbone‑benchmark comparisons at roughly half tokens. On OCRBench‑v2, it leads the strongest pruned baselines by 1.80 and 0.68 percentage points on Qwen3‑VL‑8B and InternVL3.5‑8B, respectively, while retaining about half of the visual tokens; on MMBench v1.1, it reaches 0.8467 circular exact‑matching accuracy versus 0.8437 for Vanilla at 54.45% average visual‑token retention. These results show a favorable observed quality‑cost trade‑off for evidence‑aware dynamic budgeting in text‑rich multimodal inference.

Authors:Chishui Chen, Yaoyou Fan, Te Sun, Yi Yang, Chenghao Sun, Delin Mao, Hongbo Qiao, Zuowei Zhang, Junxi Wang, Chenxing Sun, Yangen Hu, Lu Pan, Xuyang Liu, Linfeng Zhang
Title: Look Ahead Before You Distill: Future Trajectory Validation of Teacher Guidance for Agentic On-Policy Distillation
Abstract:
On‑policy distillation (OPD) provides teacher supervision on states visited by the student, reducing the distribution gap between training and inference. However, in multi‑turn agentic tasks, student deviations may accumulate over time, gradually moving the trajectory away from states where teacher guidance remains effective. Our quantitative analysis further shows that high‑disagreement states offer promising opportunities for teacher guidance, but determining whether such guidance is beneficial requires examining its effect on subsequent student trajectories. We propose FutureBridge‑OPD (FTB), which executes a short teacher bridge at a high disagreement state and uses the resulting student continuation to assess whether the bridge increases the density of positive distillation signals relative to the teacher. On ALFWorld, WebShop, and ScienceWorld, under the main Qwen3‑32B teacher to Qwen3‑1.7B student setting, FTB outperforms vanilla OPD and TCOD by an average of 16.6 and 7.6 points, respectively, and remains effective across student scales and teacher settings. Our code is publicly available at https://github.com/ChenChiShui/FutureBridge‑OPD.

Authors:Xianjing Han, Yuhan Su, Yang Deng, Dong Ma, Wee Peng Tay, Bin Zhu
Title: CultureVidBench: Benchmarking Cultural Understanding in Text-to-Video Generation
Abstract:
Text‑to‑video (T2V) generation models have advanced rapidly, yet their ability to represent diverse cultural contexts remains underexplored. Existing benchmarks mainly focus on perceptual quality, physical plausibility, and text‑video alignment, but do not directly assess whether generated videos capture culturally specific objects, actions, rituals, visible text, or audio cues. We introduce CultureVidBench, a comprehensive benchmark for evaluating cultural understanding in T2V generation. CultureVidBench contains 1,000 curated prompts covering 12 countries, 6 continents, 8 cultural regions, and 14 cultural aspects organized into three categories: material culture, social practice & performance, and ritual & ceremony. Designed specifically for video generation, CultureVidBench emphasizes dynamic and multimodal cultural representation, including social interactions, ritual procedure, and culturally appropriate visible text and audio. We evaluate seven representative T2V models through human user studies and MLLM‑based automatic assessment across cultural faithfulness, multimodal cultural rendering, semantic adherence, and perceptual quality. Results show that although current models achieve strong semantic adherence and visual quality, they often fail to faithfully capture fine‑grained cultural details, particularly for underrepresented regions, rituals, and multimodal cultural cues.

Authors:Chong Jing, Junan Zhang, Jing Yang, Yulun Wu, Fan Fan, Zhizheng Wu
Title: P-MUSE: Prompt-MIDI-Optional Model for Unified Instrumental Music Synthesis and Editing
Abstract:
MIDI‑to‑Music system renders the melody and rhythm of a target MIDI sequence into musical segment while cloning instrument timbre from a prompt recording. Existing systems typically adopt one of two distinct paradigms: conditional generation with prompt audio alone, which remains applicable when aligned prompt MIDI is unavailable, and In‑Context Learning with paired prompt audio and MIDI, which exploits cross‑modal alignment for stronger control on MIDI following and timbre similarity. We introduce P‑MUSE, an instrumental MIDI‑to‑Music framework that unifies both paradigms via a multi‑stage Curriculum‑Learning supporting prompt‑MIDI‑optional inputs. P‑MUSE further unifies music generation and local editing through a shared fill‑in‑the‑middle formulation. Grounded in theoretical analysis and empirical study, we propose a phase‑aware classifier‑free guidance scheduling principle for Transcription‑to‑Audio systems, alongside a Tail‑Drop strategy. Finally, to advance research in this field, we establish the first comprehensive benchmark, covering various prompt modes, generation/editing tasks, and four representative instruments: piano, guitar, bass, and drums. Demos are available at https://p‑muse.github.io/.

Authors:Dongwei Sun, Bowen Yao, Yujie Zhang, Pei Liu, Jing Yao, Xiangyong Cao
Title: EchoChange: A Diffusion Language Model with Dual Pass Remasking for Factual Remote Sensing Disaster Change Captioning
Abstract:
Bi‑temporal remote‑sensing disaster change captioning often needs to identify sparse and spatially localized changes across large pre‑ and post‑event scenes and then translate them into coherent, factual descriptions. However, existing change captioning methods always follow an autoregressive decoding paradigm to generate the change description and thus an early misinterpretation of the changed object, event, or spatial relation becomes an irreversible premise for subsequent text, amplifying visual ambiguity into cascading factual errors. To address this limitation, we propose EchoChange, a multimodal discrete diffusion language model that formulates change captioning as iterative masked‑token denoising rather than left‑to‑right generation. By repeatedly revising the entire caption while conditioning on the image pair, EchoChange can reconsider uncertain content and correct imperfect intermediate predictions. We further introduce draft‑aware dual‑pass training, a progressive masking curriculum, and confidence‑guided remasking to align training with iterative inference. Extensive experiments on the RSCC benchmark show that EchoChange substantially outperforms both general‑purpose and remote‑sensing‑specific baselines across lexical and semantic metrics. The EchoChange Project is at https://sundongwei.github.io/EchoChange_Project/

Authors:Hongyi Fang, Jiahui Wu, Yichen Yue, Benjia Zhou, Dan Zeng
Title: Detail Continuation over a Trustworthy Coarse Scale for Autoregressive Super-Resolution
Abstract:
Hallucination remains a persistent challenge in generative super‑resolution (GSR), where reconstructed results may contain visually plausible yet weakly supported content, structural deviations, or unnatural textures with respect to the low‑resolution (LR) input. Existing GSR methods have extensively explored the trade‑off between perceptual realism and reconstruction fidelity, but the division between preserving reliable coarse‑scale information and restoring more uncertain fine details is often handled implicitly within the overall restoration process. Visual autoregressive (VAR) modeling provides a natural opportunity to revisit this issue, as its coarse‑to‑fine next‑scale prediction offers an explicit scale‑wise generation interface. However, existing VAR‑based SR methods still inherit the original full 1‑to‑N autoregressive generation path, even though, for super‑resolution, coarse‑scale information in LR is often relatively more reliable, while long autoregressive chains may accumulate prediction errors. Motivated by these observations, we propose K2N, which reformulates VAR‑based SR from full‑path generation into a k‑to‑N detail continuation process. Specifically, early coarse‑scale states are established directly from LR, while only the remaining finer scales are restored autoregressively. Experimental results show that K2N remains competitive with the VARSR baseline on standard SR metrics, while exhibiting clearer advantages on hallucination‑focused evaluation. These findings suggest that explicitly rethinking the generation path in a scale‑wise manner can be a promising direction for improving the reliability of generative super‑resolution. Our code will be released soon at https://github.com/BRL‑SYSU/K2NSR.

Authors:YuFei Luo, Xiucheng Xu, Zhen Yang
Title: MemSIF: From Structured Interactions to Dual-Track Fact Memory for LLM Agents
Abstract:
Long‑term memory is critical for LLM agents operating over long‑horizon interactions. However, several persistent limitations of existing memory systems can be traced to two recurring misalignment patterns in long‑term interaction settings: Temporal‑Structural Misalignment (TSM) and Delayed Utility Manifestation (DUM). TSM arises when temporal proximity does not reliably align with topical or event‑level relatedness, whereas DUM arises when write‑time salience does not reliably predict future query utility. To mitigate these misalignment patterns, we propose MemSIF (Memory with Structured Interactions and Facts), a structured interaction‑to‑fact memory framework. Structured Interaction Memory organizes raw interactions into Topical Segments that preserve local topical coherence and Event Trajectories that maintain cross‑time event continuity. Dual‑Track Fact Memory uses two complementary tracks: CoreFact memory consolidates stable, schema‑guided information at write time, whereas ActiveFact memory forms facts on demand and promotes those supported by multiple historical sources and recurring query demand for reuse. Experiments on LoCoMo and LongMemEval‑S across five backbone LLMs show that MemSIF achieves the highest Total ACC in all settings, outperforming the strongest baseline by 2.29%‑8.79% on LoCoMo and 2.87%‑6.15% on LongMemEval‑S. These results support the effectiveness of combining Structured Interaction Memory with Dual‑Track Fact Memory to mitigate TSM and DUM. Code is available at https://github.com/luoyufeihaha/MemSIF.

Authors:Wonjun Choi, Yerim Kim, Yukyung Lee, Susik Yoon
Title: PGMem: Tightly Coupled Persona-Memory Graph for Lifelong Personalized Agents
Abstract:
Long‑term personalized dialogue agents must track user preferences as their personas evolve. Existing memory systems organize past events well, but store personas as flat profiles detached from the events that justify them. This loose coupling leads to the memory‑persona validity gap and the persona‑aware retrieval gap. We propose PGMem, a heterogeneous persona‑memory graph that connects event and persona nodes through typed provenance and evidence edges, keeping each persona signal traceable to the events that support or revise it. At retrieval time, PGMem expands from query‑relevant seeds and ranks signals by evidential validity. Across three benchmarks with small language model backbones, PGMem consistently outperforms summary‑based, persona‑aware, graph‑structured, and agentic memory baselines, and improves performance as the context grows. The source code of PGMem is available at https://github.com/wonjunchoi23/pgmem/

Authors:Rishov Paul, Frederick H. Epstein, Miaomiao Zhang
Title: Generative Brownian Bridge Diffusion In Motion Space For Enhanced Myocardial Strain Analysis
Abstract:
Myocardial strain analysis of cardiac magnetic resonance (CMR) images provides an important tool for evaluating cardiac function. However, current techniques require either human‑adjusted post‑processing with suboptimal regional accuracy, or specialized acquisitions with limited availability. In this paper, we propose to leverage the power of generative models to synthesize high‑quality motion‑derived strain values from routinely acquired CMR sequences. Specifically, we develop a novel Brownian bridge diffusion model in motion space to learn the probabilistic mapping between standard CMR motion estimated from widely adopted registration methods and highly accurate motion provided by advanced strain imaging techniques. To promote the fidelity of anatomical structure in the generation process, our model is conditioned on the corresponding CMR images. We validate our method on large‑scale multi‑center CMR datasets including subjects of paired standard cine CMR and advanced strain imaging acquisitions. Experimental results demonstrate that our framework significantly improves the accuracy of motion prediction and strain analysis from standard CMRs compared to existing learning‑based approaches. Our research represents a new paradigm for potentially developing cost‑effective, clinically deployable AI tools for cardiac function assessment with enhanced strain accuracy in busy clinical workflows. Our code is publicly available at https://github.com/Rishov‑MIA/Brownian‑Bridge‑strain‑analysis.

Authors:Zhixue Fang, Zhimin Zhang, Bi'an Du, Zijie Meng, Yan Zhou, Wei Hu, Guoxin Zhang, Pengfei Wan, Kun Gai
Title: Motion Beyond Morphology: Bootstrapping Cross-Category Motion Transfer from Abstract Motion Representations
Abstract:
Video motion transfer aims to animate a target object using dynamics from a reference video. Existing formulations largely rely on fixed structural correspondence, which becomes ill‑defined when reference and target objects differ substantially in morphology, articulation, or deformation mechanisms. We introduce Motion Beyond Morphology, a perspective that seeks to transfer motion beyond fixed structural correspondence, by preserving dynamics that remain meaningful across different target morphologies. To realize this, we propose a two‑stage framework. Stage~I learns complementary multi‑granularity abstract motion views and uses them to bootstrap cross‑category video pairs that preserve transferable dynamics across diverse morphologies. Stage~II internalizes this supervision into direct reference‑video‑conditioned generation, removing the need for explicit motion extraction at inference. We further introduce OpenVMT‑Dataset and OpenVMT‑Bench for training and evaluating image‑ and text‑conditioned motion transfer across Same, Near, and Far category gaps. Extensive experiments demonstrate state‑of‑the‑art motion fidelity and target preservation. Project page: https://miniz233.github.io/MotionBeyondMorphology/

Authors:Sam Andersson, Ricky Molén
Title: Finite-Probe Total-Variation Certificates for Finite-Basis Drifting Models
Abstract:
Drifting objectives compare a target and model distribution through a vector field observed noisily at finitely many locations. We ask what distributional conclusion such a frozen measurement system warrants. For integrable antisymmetric interactions and absolutely continuous laws in a declared finite density basis, the unnormalized sampled numerator satisfies \operatornamevec(V_X)=Mc, where c is an antisymmetric mismatch and M is probe‑dependent. This identity yields an a posteriori total‑variation (TV) upper confidence bound accounting for held‑out field noise, estimated‑operator error, and externally validated L^1 residual radii around normalized density approximants in the span; a nonpositive observability margin returns the trivial TV bound and abstains. The audit recomputes this numerator from held‑out samples; a normalized drift statistic requires a separate joint numerator‑‑denominator analysis. For Gaussian‑RBF interactions, a global envelope supports distribution‑free and empirical‑Bernstein radii without truncation, with companion bounds for the Laplace similarity in the original drifting objective. We characterize random‑probe observability by a population Gram matrix, identify rank and symmetry degeneracies, and prove large‑bandwidth collapse toward mean matching. Synthetic studies exercise Gaussian and Laplace numerators, separately prespecified bounded‑vector and variance‑adaptive radii, Monte Carlo‑calibrated operators, nonzero residual radii around normalized finite‑basis approximants, outward‑rounded observability bounds, and designed abstention. A joint basis‑size/dimension stress path extends evaluation through m=8. The result is a conditional diagnostic for a finite density class, or for normalized finite‑basis density approximants with external residual radii, not a universal guarantee from small training drift.

Authors:Dingyi Kang, Dongming Jiang, Yi Li, Guanpeng Li, Bingzhe Li
Title: V-Mem: Modality-Routed Retrieval for Long-Term Multimodal Agentic Memory
Abstract:
Interaction between users and LLM agents is increasingly multimodal: conversations interleave text with images, and a later question may target either. Yet most agent memories are designed around text, and even the few that support multimodal conversations still fail on vision‑related questions. We trace this failure to an assumption behind the similarity search they rely on: in the index space, a query lies close to the relevant evidence that answers it. In multimodal settings, two gaps break it. By the modality gap, a query lies closer to memory content of its own modality than to evidence in another, even in a trained joint embedding space. By the similarity‑relevance gap, the content most similar to a query is often not the evidence that answers it, most acutely when a query carries both text and image and its evidence resembles neither part alone. We present V‑Mem, a multimodal agentic memory system that routes retrieval by the modality of the query and that of the target evidence, both recognized from the query alone. To cross the modality gap, V‑Mem organizes the conversation into rounds and returns the target‑modality content from the same round as the match, without comparing across modalities. To close the similarity‑relevance gap, it searches with an LLM‑generated anchor that sits closer to the relevant evidence than the query does: a hypothetical caption for a text‑only query seeking an image, and an enriched search anchor, the query text plus relevant keywords extracted from the query image, when the evidence is reachable only by combining the two. On Mem‑Gallery, V‑Mem reaches an LLM‑judge score of 0.82 versus 0.56 for the second best, with the largest margin on questions carrying an image (0.87, no baseline above 0.47); on LoCoMo it scores 0.69 versus 0.58.

Authors:Ruokai Yin, Priyadarshini Panda
Title: Celty: SpMspV GPU Kernel and SIMT Co-Design for Efficient Dual-Sparse LLM Inference
Abstract:
Large Language Models (LLMs) increasingly rely on sparsity to reduce inference cost, but most prior work targets a single sparsity source‑either weight or activation‑and optimizes for batched multi‑user inference. Dual‑sparsity, which combines unstructured weight pruning with runtime activation sparsity, offers a compelling tradeoff among model size, accuracy, and latency for single‑user decoding, but formulates as a Sparse Matrix‑Sparse Vector (spMspV) workload that existing GPU kernels handle poorly. We propose Celty, a co‑designed sparse format, GPU kernel, and SIMT microarchitecture for efficient spMspV in LLM inference. At the kernel level, Celty introduces a Run‑Length Compressed CSC (RLC‑CSC) format that enables vectorized loading of compressed weight columns and exploits both sparsity sources to skip unnecessary memory accesses, with shared memory used for scattered partial‑product accumulation. At the microarchitecture level, the Celty Sparse SIMT Core integrates a pipelined RLC decoder to eliminate software‑level index reconstruction and repurposes local register files for conflict‑free accumulation‑operating directly on the same RLC‑CSC format without data layout changes. The Celty GPU kernel achieves up to 2.8x speedup over cuBLAS and 2.4x over Flash‑LLM. With the Sparse SIMT Core, speedups reach up to 5.3x over cuBLAS at 70% dual‑sparsity.

Authors:Kazi Ahmed Asif Fuad, Lizhong Chen
Title: BiKAN: Restoring Collapsed Basis of Binary Kolmogorov--Arnold Networks
Abstract:
Binarizing a polynomial Kolmogorov‑‑Arnold Network (KAN) not only changes parameter precision, but also alters the function space available to each layer. When activations are restricted to ‑1,+1, all even powers reduce to 1 and all odd powers reduce to x, causing the elementwise polynomial basis to collapse to constant and first‑order responses. We refer to this structural failure as Spatial Orthogonality Collapse. Our proposed BiKAN addresses this critical issue by augmenting each binary KAN layer with selected degree‑2 Walsh characters. Fixed circular channel rolls generate pairwise parities, and learned binary projections mix them using the same XNOR‑‑popcount operations as the remaining W1A1 paths. This restores explicit pairwise coordinates without learned routing or multiplier‑based feature generation. Experiments on CIFAR‑10 confirms that removing parity reduces accuracy by 1.23 points over five paired seeds (p=0.003), the gain increases as width decreases, and accuracy improves monotonically as more parity planes are added. At an equal ~11.9M‑parameter budget, parity outperforms conventional widening by 3.09 points (p<10^‑4). At W1A1, BiKAN reaches 99.48%, 84.38%, and 55.81% on MNIST, CIFAR‑10, and CIFAR‑100, respectively. Post‑route Zynq‑7020 FPGA results show that the repair remains hardware‑efficient; the convolutional design cuts DSP usage from 164 to 72 and estimated compute‑core latency from 401 to 54.8 ms, while the power‑of‑two‑aware dense design achieves zero‑DSP inference with a 0.03‑point accuracy loss. The BiKAN implementation is available at https://github.com/OSU‑STARLAB/BiKAN.

Authors:Quang Bui, Shlok Jaiswal, Samuel Paik-Heintz, Kevin Zhou, Kaushik Madapati, Krittaphas Chaisutyakorn, Noah Dane Hebdon, Dimitrios Proios, Sebastián Andrés Cajas Ordóñez, Kacper Dobek, Boya Zhang, Aly Dhedhi, Ahram Han, Kushul Reddy Palakala, Rahul Gorijavolu, Jacques Kpodonu, Leo Anthony Celi
Title: Loud or Silent? A Reusable Framework for Per-Modality Failure Analysis in Multimodal Clinical AI
Abstract:
Multimodal clinical models are usually judged on accuracy with every modality present, but deployment removes modalities; an echocardiogram is often unavailable where an ECG is routine. Two questions then matter beyond the size of the accuracy loss: which modality was responsible, and whether the model fails loudly or silently once that modality is dropped. The distinction is per‑example and modality‑level, and is separate from post‑hoc feature attribution (e.g. SHAP). Models are replaced often; the evaluation that answers these questions is reused. We present a model‑agnostic modality‑failure framework: given N modality embeddings, any mask‑aware probe, and labels, it returns a per‑example failure taxonomy, a per‑modality complementarity matrix that attributes error to modalities, and a loud‑vs‑silent dropout profile separating monitorable failures from those that pass unflagged far from the decision boundary, using only deployment‑observable signals. We release it as a small, unit‑tested harness and validate it against planted ground truth. Across seeds it recovers that planted modality dominance and complementary subset, reports per‑modality loud‑vs‑silent rates, and scales to a three‑modality complementarity matrix; because the planted structure is known by construction, this validates recovery of per‑example attribution rather than clinical performance. We then instantiate the framework on frozen EchoJEPA and HuBERT‑ECG embeddings for LVEF and the EF <= 40% HFrEF gate over a paired MIMIC‑IV cohort, where on the held‑out test split (n = 245) dropping echo nearly doubles error. The narrow echo‑to‑ECG overlap that bounds cohort size is itself a deployment finding for cardiac foundation models. All of our work can be found at https://github.com/criticaldata/PRIMED‑AI.

Authors:Priyanka Dey, Brihi Joshi, Preyashi Poddar, Jieyu Zhao, Emilio Ferrara
Title: PALMs: Using Multi Construct-Grounded Rationales for Modeling Population Preferences in LLMs
Abstract:
Large language models are being extensively used to simulate individual user behavior, yet faithfully representing a population requires capturing the systematic variation in values, beliefs, and cultural norms that distinguish one group from another. We introduce Population Aligned Language Models (PALMs), a suite of models each aligned to specific populations, covering five countries: USA, India, Brazil, France and Italy. PALMs are created by synthesizing rationales grounded in psychological and cultural constructs and using these as latent supervision during preference tuning for population‑specific alignment. Evaluated across four dimensions: personality, values and beliefs, cultural norms, and morality, PALMs consistently outperform baselines, including culture‑specialized models, achieving an average of 8.59% relative improvement over the best baseline across all five populations. Notably, construct‑grounded rationales outperform both demographic prompting and survey‑based fine‑tuning, suggesting that grounding preference learning in psychology and culture provides a richer inductive signal than surface‑level response distributions. We further demonstrate strong generalization to downstream applications with‑ out task‑specific supervision: outperforming best baselines by 5.19% in personalized reward modeling, 6.34% in population simulation, and showing strong transfer to social reasoning tasks. Datasets and code are available at: https://github.com/limenlp/PALMs.

Authors:Haoran Liao, Pengyue Wang, Shuoyu Chen, Kehan Cheng, Xuhang Chen, Yuhao Lin, Mu Lin, Zhizhao Liang, Xiaoyi Fan, Chengyi Xing, Dan Niu, Yi-Lin Wei, Wei-Shi Zheng
Title: DynamicManip: Enabling Dynamic Manipulation from a Single Static Demonstration
Abstract:
Dynamic manipulation is a critical capability for robots operating in complex and dynamic environments, where robots must interact with objects that are moving or require rapid adjustments. However, learning models for dynamic manipulation tasks face two major challenges: (1) the combinatorial complexity of dynamic scenarios leads to substantial data requirements, and (2) rapid variations in dynamics require real‑time and accurate policy execution. In this paper, we propose DynamicManip to address these challenges through an efficient data augmentation pipeline and a low‑latency imitation policy. We first propose a static‑to‑dynamic augmentation pipeline that synthesizes diverse dynamic manipulation demonstrations from a single static demonstration. Second, we introduce a dynamic‑aware adaptive policy that adjusts its inference frequency according to task dynamics, enabling responsive and effective dynamic manipulation. Third, we build a dynamic manipulation benchmark, which includes diverse dynamic tasks with an automatic evaluation system for scalable and consistent assessment. Extensive experiments in both simulation and the real world demonstrate that DynamicManip not only provides significant improvements in data efficiency but also achieves better performance in dynamic manipulation tasks, with a mean success rate 18.4 percentage points higher and policy‑query latency 32.9% lower.

Authors:Huiyu Yi, Yongqi Xu, Bogang Zhang, Dunwei Tu, Xu Zhiming, Zhen-Hao Xie, Baile Xu, Furao Shen
Title: Beyond Routing Saturation: A Long-Horizon Class-Incremental Perspective on Expert Routing in Multimodal Continual Instruction Tuning
Abstract:
Multimodal Continual Instruction Tuning (MCIT) enables multimodal large language models to acquire new tasks sequentially while retaining previously learned capabilities. Many recent methods maintain task‑specific LoRA experts and route each input to one or more experts at inference. Yet the task‑identification problem underlying expert routing remains under‑explored. We show that routing is nearly saturated on widely used MCIT benchmarks. Textual fingerprints that leak task identity and short 4‑‑10‑task sequences with few competing experts jointly obscure the long‑horizon routing problem. To expose this challenge, we introduce FLEX (Fingerprint‑reduced Long‑horizon Expert eXamination), a 34‑task long‑horizon MCIT benchmark with weakened textual fingerprints. FLEX groups tasks with similar instruction and answer formats but diverse visual and knowledge domains, normalizes their outer templates, and evaluates routing over a substantially larger expert pool. Crucially, we formulate progressive‑LoRA routing as soft task‑as‑class Multimodal Class‑Incremental Learning (MCIL): each task defines an incremental routing class, whose complete score distribution supplies the LoRA mixture weights, with hard routing as a discrete special case. FLEX exposes this expanding task‑identification challenge, while the MCIL formulation provides a principled interface for transferring CIL methods to expert routing. We instantiate PureLoRA as a controlled baseline and adapt four CIL methods to four MCIT frameworks without modifying their LoRA experts or generation pipelines. Our plug‑in routers improve strict LoRA matching by up to 16.3 percentage points and overall MacroScore by up to 4.6 points. Code is available at: https://github.com/RINC‑CL/FLEX

Authors:Sabri Mustafa Kahya, Richard R. Chen, Muhammet Sami Yavuz, Jerry Jierui Lou, Akanimoh Adeleye, Haci Ali Kahya, Jana Lipkova
Title: Training-Free Out-of-Distribution Detection for Pathology Whole-Slide Images
Abstract:
Safe deployment of AI methods in medicine requires robust guardrails that detect when input data deviate from the training distribution to ensure that models provide predictions only within their scope of expertise and abstain otherwise. Out‑of‑distribution (OOD) detection can provide such safeguards and is extensively studied in general computer vision. Yet, it remains underdeveloped in computational pathology, where gigapixel whole‑slide images (WSIs), subtle differences between disease subtypes, and variability in tissue preparation pose unique challenges for conventional OOD methods. We propose ZIO, a training‑free, multimodal OOD detector for pathology WSIs that leverages vision‑‑language pathology foundation models (FMs). ZIO constructs text and visual prototypes of in‑distribution classes and integrates their complementary information through a prototype shrinkage mechanism to derive OOD scores. We provide the ZIO formulation for both slide‑ and patch‑level FMs. We evaluate ZIO across diverse clinically relevant domain shifts, including rare diseases and near‑OOD settings. Extensive evaluation of over 14,700 WSIs from five independent consortia shows that ZIO consistently outperforms both unimodal prototypes and 40 state‑of‑the‑art OOD methods. These results demonstrate the benefits of multimodal representation for OOD detection and pave the way towards safer AI deployment in clinical practice.

Authors:Rasa Hosseinzadeh, Alex Labach, Zexin Xue, Shuyi Han, Valentin Thomas, Anthony L. Caterini
Title: TabDPT-Turbo: Efficient In-Context Learning for Tabular Prediction
Abstract:
Tabular foundation models, driven by in‑context learning, have rapidly grown in quality and popularity. However, recent approaches with either cell‑based architectures or retrieval have sacrificed efficiency for raw performance, restricting their utility in situations where compute is limited or inference speed is crucial. We adopt an alternate approach, sticking with row‑based attention while incorporating long context pre‑training to eliminate the need for retrieval. By combining this with architectural improvements and SSL pre‑training on a newly‑sourced, larger corpus of real data results, we present TabDPT‑Turbo, a model that provides comparable default performance to TabDPT v1.1 on TabArena‑Lite, CC18, and CTR23, at orders of magnitude faster. In our experiments, TabDPT‑Turbo is the fastest model overall among leading foundation models. We have released the new model as TabDPT v1.2 at https://github.com/layer6ai‑labs/TabDPT‑inference.

Authors:Sherzod Hakimov, Karl Osswald, Jelle Psurek, Eszter Bukovszky, A. Altar Lüser, David Schlangen
Title: Language Equality has a Price: A Systematic Investigation of Multi-turn LLM Performance for EU-24+
Abstract:
We evaluate large language models (LLMs) as language agents playing goal‑directed dialogue games in self‑play across 30 languages: the 24 official EU languages plus six others. Unlike static or preference‑based evaluation, this paradigm is multi‑turn, reference‑free and programmatically scored, and because the game mechanics are language‑agnostic it extends to a new language by localising a fixed set of prompt and word‑list files. Evaluating nine open‑weight and commercial LLMs, we find that no open‑weight model covers the EU‑24 well: in every official language both commercial systems outscore every open‑weight model, and the two weakest average below 40 points across the EU‑24. The commercial systems stay ahead even in languages with four orders of magnitude less public web text, showing that linguistic parity is achievable, but not from public crawls alone. A model's home region lifts it without closing the gap: Chinese is the strongest of all 30 languages for two Chinese‑developed models, yet the best Chinese score of any model belongs to a US commercial system. Coverage is also not parity of service. Pooled over models and languages, the median non‑English language costs 31% more to run than English, and scores 10% lower.

Authors:Yuxiang Xiao, Yang Hu, Bin Li, Tianyang Zhang, Zexi Li, Huazhu Fu, Jens Rittscher, Kaixiang Yang
Title: Understanding Synergistic Interactions among Pathology Foundation Models via Adaptive Fusion
Abstract:
Pathology foundation models (PFMs) provide strong tile‑level representations via self‑supervised pre‑training on large‑scale pathology images. Yet, PFMs are developed under diverse and often opaque data, architecture, and objective choices, inducing latent representational biases that limit robustness and obscure what each model specialises in. We present AdaFusion, a lightweight adaptive fusion framework that integrates complementary signals from multiple frozen PFMs through (1) low‑dimensional feature compression and (2) a sample‑conditioned gating module that reweights model‑wise (and optionally channel‑wise) contributions. Beyond improving predictive accuracy, AdaFusion provides contribution‑driven interpretation that offers evidence consistent with model‑specific preferences and synergistic interactions across tissue phenotypes. We evaluate AdaFusion on three public benchmarks spanning treatment response prediction, prostate cancer grading, and spatial gene expression inference. AdaFusion consistently outperforms individual PFMs and other fusion baselines, while providing interpretable tissue visualisation which aligns model preferences with morphological patterns. Code is available at: https://github.com/xyx‑98/PathoOracle.

Authors:Jianan Xie, Xin Sun, Zhongqi Chen, Xing Zheng, Shu Wu, Bowen Song, Liang Wang
Title: EviSD: Evidence-Conditioned Self-Distillation for Search-Augmented Agents
Abstract:
Outcome‑based reinforcement learning enables search‑augmented language agents to learn from verifiable final answers, but its trajectory‑level credit cannot distinguish the contributions of individual actions in a multi‑turn search process. We propose EviSD, an evidence‑conditioned self‑distillation framework that uses instance‑level supporting evidence as privileged information for search actions and golden answers as complementary privilege for answer actions. During training, the student samples actions from the original context, while the same model re‑scores them as a privileged teacher under an action‑aligned context. EviSD converts the detached teacher‑‑student gap into a bounded correction to the outcome‑derived GRPO advantage and applies it only to generated action spans. This design localizes privileged guidance while preserving the update direction determined by the outcome reward, without an auxiliary distillation objective or any change at inference time. Across seven question‑answering benchmarks and three backbones spanning model scales and generations, EviSD achieves the highest macro‑average Exact Match in all evaluated settings, outperforming the strongest compared methods by 1.3‑‑2.3 points while modulating only 6.7%‑‑15.1% of response tokens. Code is available at https://github.com/JiananXie/EviSD.

Authors:Zhiwei Chen, Yang Hu, Yuxiang Xiao, Yakun Ju, Tianyang Zhang, Yingxue Xu, Wei Li, Hao Chen, Jens Rittscher, Kaixiang Yang
Title: Harnessing Adversarial Distillation to Customise Debiased, Disease-Specific Pathology Foundation Models for Breast Cancer
Abstract:
Pathology foundation models (PFMs) provide strong tissue representations and have become central to digital pathology. However, deployment in disease‑specific settings is limited by 1) the high computational cost of billion‑parameter PFMs and 2) distribution mismatch and non‑biological bias inherited from pan‑cancer, multi‑centre pre‑training, including site‑specific signatures and imbalanced disease prevalence. These factors can encourage shortcut learning and under‑emphasise subtle morphology required for reliable modelling of a specific cancer type. We present SmartStu (a Smart Student), a framework to customise compact, breast‑cancer‑specific PFMs via distillation whilst mitigating confounding. SmartStu distils representations from multiple teacher PFMs into a lightweight student backbone. Crucially, we introduce adversarial distillation that leverages a dedicated noise model trained to predict nuisance, edge‑dominated cues on the distillation set. Using this noise model as a counterexample, the adversarial objective encourages the student to recognise, yet suppress, features predictive of nuisance targets. We further incorporate multi‑teacher ensemble distillation and an auxiliary self‑supervised objective with artefact injection. We validate SmartStu on three external cohorts (Yale HER2, SLN‑Breast, and BRACS) with multiple tiny backbones. SmartStu yields breast‑cancer‑specific PFMs that are over 30× smaller than general PFMs whilst largely preserving, and sometimes improving, downstream performance measured by balanced accuracy (bAcc) and AUC. Code is available at https://github.com/zwchen03/advDistall.

Authors:Junhan Wang, Kani Chen
Title: CORTIVA: Candidate-Score Fusion of Complementary Visual Teachers for EEG- and MEG-to-Image Retrieval
Abstract:
Decoding visual experience from non‑invasive brain activity is central to neuroscience and brain‑computer interfaces. Functional magnetic resonance imaging (fMRI) offers fine spatial detail, but its slow hemodynamics and burdensome acquisition limit temporally resolved decoding. Electroencephalography (EEG) and magnetoencephalography (MEG) provide millisecond resolution, making image retrieval compelling: identify the viewed image from one neural response and a fixed candidate bank. Contrastive alignment to pretrained visual representations enables zero‑shot retrieval from EEG and MEG, but most systems collapse heterogeneous visual supervision into a single embedding before ranking. This early consolidation imposes one similarity geometry on every candidate order and removes encoder‑specific disagreements from the final ranking. We propose CORTIVA, a candidate‑score fusion framework that preserves this complementary evidence. Three decoding routes are aligned to heterogeneous visual targets, score the same indexed candidates independently, and combine only their temperature‑scaled score vectors before ranking. On the 200‑way THINGS‑EEG2 benchmark, CORTIVA reaches 73.5% Top‑1 and 95.3% Top‑5 across ten participants, exceeding the strongest reported baseline by 10.3 and 5.4 percentage points. With a modality‑specific neural encoder, the same fusion principle reaches 42.4% Top‑1 on THINGS‑MEG. Matched route‑removal retraining and four weight controls demonstrate that CORTIVA's gain arises from integrating complementary route scores and persists with uniform weighting, without requiring a specialized weighting rule. Independent DINOv2 analyses further reproduce the local error neighborhoods and posterior neural‑visual correspondence. These results establish candidate‑score fusion as a simple and testable alternative to embedding‑level consolidation for neural image retrieval.

Authors:Kaisi Guan, Bingzi Zhang, Xihua Wang, Ying Ba, Xin Cheng, Yijing Chen, Ruihua Song
Title: FATE: Frame-Level Audio-Visual Temporal Embedding
Abstract:
When a dog opens its mouth and barks, humans naturally recognize what the sound is and when it occurs. Building audio‑visual models with this same ability requires representations that capture both semantic and temporal alignment. Current approaches fall short on one side or the other: embedding models match semantic but lose temporal information; synchronization models capture temporal offsets but lack semantic understanding. To bridge this gap, we propose FATE, Frame‑level Audio‑visual Temporal Embedding. Unlike prior embedding models that pool each modality into a single embedding and discard temporal information, FATE retains frame‑level sequences, aligns them on the physical timeline, and computes similarity over strictly aligned frame pairs. Unlike synchronization models that output only an offset prediction, FATE encodes synchronization in a reusable embedding space, trained with a joint objective combining cross‑video semantic and within‑video temporal contrastive learning to capture both what sounds and when it occurs. Across three tasks, FATE surpasses the strongest baseline on temporal and semantic retrieval by a large margin, matches fully supervised methods on event localization in a zero‑shot setting, and achieves the best correlation with human judgments as a generation evaluation metric. The source code can be found at \texttthttps://github.com/guankaisi/FATE.

Authors:Jiawei Guo, Junxian Li, Yixin Tang, Bingya Zhang, Jiaxin Lu, Yulun Zhang, Shangchen Zhou
Title: TurboClear: One-Step Object-Effect Removal via Region-Calibrated Distribution Matching and Fusion
Abstract:
Recently, diffusion‑based removal methods have achieved promising visual quality in removing both target objects and their associated effects. However, they typically rely on multi‑step denoising, leading to high inference cost. Directly applying existing one‑step distillation methods is also suboptimal, since their global objectives lack explicit region‑wise calibration and may weaken the asymmetric edit‑and‑preserve behavior required by object‑effect removal. To address these challenges, we propose TurboClear, a one‑step SDXL‑based object‑effect removal model. During training, we design Region‑Calibrated Distribution Matching (RDM) for region‑aware distillation to preserve the teacher model's asymmetric edit‑and‑preserve behavior. Furthermore, we propose Learnable Spatial Fusion (LSF) for lightweight inference‑time fusion. Extensive experiments show that TurboClear significantly improves inference efficiency while maintaining competitive visual quality. TurboClear reduces the computational overhead by up to 40.04× compared to ObjectClear, and by up to 665× against the Flux‑based method OmniPaint, all while maintaining comparable or better visual removal quality. Code is available at https://github.com/GuoCalix/TurboClear.

Authors:Qi Lv, Jianming Xing, Zhao Yang, Mingyuan Yao, Yinan Shi, Yawei Jueluo, Mike Zheng Shou, Xiang Deng
Title: Hermite Curves as Trajectory Priors for Vision-Language-Action Models
Abstract:
Despite recent progress in Vision‑Language‑Action (VLA) models for robotic manipulation, the action chunk remains a weakly structured interface. Existing work typically flatten each chunk into per‑timestep controls, relying on implicit data learning that manifests as jagged motion and boundary discontinuities during physical execution. To address these limitations, we introduce Hermite trajectory priors, parameterizing the chunk trajectory as a piecewise cubic Hermite curve defined by endpoint positions and velocities to explicitly enforce smoothness and continuity. We instantiate this fixed operator across discrete autoregressive and continuous generative paradigms via three variants: (1) Hermite Tokens, which predict quantized boundary variables autoregressively; (2) Hermite Scaffold, which decomposes clean actions into a base scaffold and residuals; and (3) Hermite Regularization, which applies the prior strictly as an auxiliary training objective. Across simulation benchmarks and real‑robot platforms, Hermite Regularization achieves superior performance among these three variants, improving π0.5 baseline success rates from 95.9% to 98.7% on LIBERO, 85.7% to 90.9% on LIBERO‑plus, and 63.4% to 90.0% across four real‑robot tasks without additional inference overhead. Trajectory analyses reveal that explicitly structuring trajectory priors serves most effectively as a learning inductive bias rather than a runtime constraint.

Authors:Zirui Zhang, Yinbo Yu, Donghai Guan, Chunwei Tian, Daoqiang Zhang, Qi Zhu
Title: A Benchmark Dataset for MLLM-Generated Image Detection: GPT Image2 & Nano Banana2
Abstract:
The realism of images generated by multimodal large language models (MLLMs), such as GPT Image2 and Nano Banana2, has improved rapidly in recent years. Compared with early generative models, current models have made clear progress in text rendering. They can produce high‑quality images that closely resemble real‑world application scenarios. The enhanced generation capabilities of current MLLMs pose increasingly severe challenges to AI‑generated image detection. Detection is no longer limited to identifying obvious artifacts left by early generators. Instead, it requires systematic and realistic benchmarks for the new generation of generated content. However, most existing benchmarks are still built around early generative models and cannot fully evaluate the forensic challenges introduced by high‑quality and multi‑form generated images. To address this gap, this paper constructs a benchmark dataset for detecting images generated by MLLMs. The benchmark covers several realistic application scenarios and adopts three generation protocols to simulate direct generation, reference‑based reconstruction, and local editing. Based on this benchmark, we evaluate detector degradation from traditional scenarios to MLLM‑generated images and analyze false positive rates and false negative rates across three sample types, revealing the failure modes of existing methods. We further propose a structural‑artifact‑prior‑guided dual‑stream prompt framework (SAP‑DSP) as a strong baseline. SAP‑DSP uses dual‑stream prompt learning and structure‑aware routing fusion to improve representation learning. Extensive experiments show that the proposed benchmark exposes the performance degradation of existing detectors on high‑quality generated images, while SAP‑DSP achieves more stable detection results on this benchmark. Our code and dataset are publicly available at https://github.com/xbrainnet/SAP‑DSP.

Authors:Haocheng Wang, Tongkun Guan, Wei Shen, Xiaokang Yang
Title: VaRS-Doc: Interpretation-Aware Variant Representations via Latent Self-Probing for Visual Document Retrieval
Abstract:
Visual document retrieval has recently become increasingly important in applications such as enterprise search, scientific literature discovery, and retrieval‑augmented generation. These applications depend on efficiently identifying query‑relevant pages across large collections of visually rich documents. Existing methods commonly adopt late‑interaction architectures that encode and index documents offline to enable scalable and low‑latency online retrieval. Despite its efficiency, this paradigm requires each document to be encoded into a fixed representation before the query is known. However, the same content in a visual document may induce different interpretations depending on the query intent, which a fixed representation struggles to capture. Yet postponing document encoding until the query arrives would incur prohibitive online retrieval latency. To address this gap, we propose VaRS‑Doc, a visual document retrieval framework that diversifies document representations by enabling the model to actively explore variant latent interpretations during document encoding, while preserving efficient late‑interaction retrieval in which each query adaptively selects the best‑fit representation. We further introduce a two‑stage training strategy that encourages the model to capture complementary semantic interpretations and prevents it from falling back to train a single dominant representation. Experiments on visual document retrieval benchmarks show that VaRS‑Doc achieves state‑of‑the‑art retrieval performance, offering a practical solution to the mismatch between query‑agnostic document encoding and query‑specific retrieval needs. Code is available at https://github.com/bokufa/VaRS‑Doc.

Authors:Denis Tarasov, Robert K. Katzschmann
Title: ReBRAC-v2: The Return of the King
Abstract:
Recent offline reinforcement learning methods increasingly rely on expressive generative policies and specialized value‑guidance mechanisms. We ask whether comparable progress can instead come from systematically modernizing a conventional behavior‑regularized actor‑critic while preserving its algorithmic simplicity. We introduce ReBRAC‑v2, which directly trains an exact‑likelihood normalizing flow as the RL actor, combines likelihood, MSE, and MAE behavior regularization, and integrates a classification‑based residual critic, staged optimization, and multi‑sample test‑time action selection. Rather than tuning this recipe separately for every task, we develop a single shared configuration via roughly 600 Bayesian proposals on six challenging OGBench tasks, freeze all structural and optimization choices, and adapt only two behavior‑regularization coefficients over a 16‑point grid. Across ten common state‑based OGBench categories, ReBRAC‑v2 averages 74.8 compared to 52.3 for the next‑best aggregate result and ranks first in eight categories. The same recipe, without structural changes, obtains the strongest averages in our comparisons on D4RL AntMaze (90.2) and Adroit (33.6). Fixed‑recipe ablations show the largest sensitivity to the selected mixed cloning objective, staged training, sufficient flow capacity, and multi‑sample inference, while showing that several smaller choices depend on the values of other hyperparameters. These results show that disciplined, transferable engineering can achieve state‑of‑the‑art aggregate performance without abandoning a minimalist offline RL foundation.

Authors:Yicheng Liu, Bolin Zhang, Weiran Liu, Yakun Zhang, Yangqin Jiang, Zhiying Tu, Dianhui Chu
Title: MADE: Belief-Driven Dual-Agent Coordination for Autonomous Model Deployment
Abstract:
LLM‑based agents now have strong general capabilities. However, they still struggle with domain‑specific tasks, motivating the integration of external tools to broaden their capabilities. The open‑source community offers a vast array of AI models typically released as heterogeneous research artifacts, whereas transforming them into ready‑to‑call APIs is costly and labor‑intensive. Automated model deployment is therefore essential for bridging the gap between model resources and tool usability, yet it remains a long‑horizon, multi‑stage task that has not been sufficiently explored. To tackle this challenge, we introduce Model Automated Deployment Engine (MADE), a dual‑agent coordination system. Specifically, given a model resource, MADE iteratively constructs and validates the deployment artifacts, updates its deployment belief based on execution feedback, and revisits invalid upstream artifacts until the model is successfully served as a ready‑to‑call API that can then be used by other agents. We further introduce M2ABench, a benchmark for the task of transforming Models to ready‑to‑call APIs. M2ABench comprises 122 real‑world models with standardized test cases for evaluation. Experimental results demonstrate that MADE achieves a deployment success rate of 68.85%, outperforming SWE‑agent and OpenHands by 13.93 and 44.26 percentage points, respectively. Our code and dataset are publicly available at https://github.com/HITDiSC/MADE.

Authors:Yinglong Li, Donghui Shen, Xiaoyu Zhang, Zhichao Ye, Hongyu Wu, Aimin Hao, Guofeng Zhang, Haomin Liu
Title: QuerySplat: Decoupling Geometry and Appearance Representations in 3DGS Prediction
Abstract:
While feed‑forward 3D Gaussian Splatting (3DGS) enables efficient 3D reconstruction, achieving high‑fidelity rendering remains challenging. Existing pixel‑aligned approaches suffer from spatial inflexibility and massive structural redundancy, whereas query‑based methods lack 3D priors and entangle geometry with appearance, yielding blurry, pose‑dependent results. To overcome these deficiencies, we propose QuerySplat, a feed‑forward 3DGS framework driven by geometric priors and explicit appearance decoupling. Specifically, we design a dual‑branch query‑based decoder: the geometry branch leverages a pretrained Vision Geometric Model for spatial understanding, which intrinsically endows QuerySplat with pose‑free modeling capabilities, while the appearance branch recovers high‑frequency details through a dedicated pathway separated from geometric attribute regression. Extensive experiments demonstrate that QuerySplat mitigates the blurry rendering issues of early query‑based models and consistently outperforms pixel‑aligned approaches in rendering fidelity. On the challenging DL3DV benchmark, it achieves state‑of‑the‑art novel view synthesis performance, with average PSNR gains of 2.30 dB and 1.04 dB over the best pose‑free and pose‑required baselines, respectively. Project Page: https://inspatio.github.io/querysplat.

Authors:Changwoo Baek, Kyeongbo Kong
Title: 3DZip: Spatial-Aware Feature Diversity-Guided Token Compression for 3D Question Answering
Abstract:
Recent 3D vision‑language models (3D VLMs) construct geometry aware tokens by projecting 2D visual features into world coordinates, enabling spatial reasoning for tasks such as 3D question answering. However, this design generates thousands of tokens per scene, resulting in substantial computational and memory overhead. While token compression has been extensively studied in 2D VLMs, existing approaches rely on semantic relevance or attention‑based selection that overlook the structured spatial nature of 3D tokens. Moreover, redundancy in 3D representations cannot be resolved by spatial proximity alone, as object‑level token imbalance persists even after spatial aggregation. To address this, we propose 3DZip, a three‑stage token compression framework that first applies coarse voxelization to remove point‑level redundancy, then selects anchor tokens based on feature‑space diversity via a Determinantal Point Process, and finally merges remaining tokens under spatial constraints to preserve geometric coherence. Experiments on three 3D question answering benchmarks demonstrate that 3DZip consistently outperforms existing compression methods, retaining 94.7% of the original performance with only 128 tokens, achieving a 1.92× faster inference speed.

Authors:Bruno Brocai, Ilaria Papagno, Mayumi Ohta
Title: PlainMedScale: A Corpus of Multi-Level Simplified Medical Texts in German and English
Abstract:
We introduce PlainMedScale, a topic‑aligned medical corpus spanning four levels of comprehensibility in German and English, drawn from MSD (professional and consumer), Gesund.Bund, Apotheken Umschau Einfache Sprache, and the NHS. The four tiers correspond to distinct communicative functions ‑‑‑ reference, explanation, decision support, and access ‑‑‑ and move beyond the binary expert‑‑lay contrast of prior corpora. In two pilot studies enabled by the alignments, we show that many readability metrics established on two registers fail to generalize across the full gradient, and that a SOTA open‑weight LLM prompted for Plain Language still partially preserves the difficulty of its input. Code (https://github.com/GS‑Uni‑Heidelberg/PlainMedScale) and data (https://doi.org/10.5281/zenodo.21728290) are made available.

Authors:Ganzhong Luo, Yang Ren, Hanyong Wang, Shuyu Zheng, Menglong Yang
Title: UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering
Abstract:
Knowledge‑Based Visual Question Answering (KB‑VQA) requires retrieving entity knowledge from external sources to answer visually grounded questions. Existing retrieval‑augmented systems suffer from two critical limitations. First, relying on a single retrieval modality creates a Single‑Source Retrieval Bottleneck, missing ground‑truth entities that are only accessible through complementary sources. Second, dual‑tower pointwise rerankers suffer from Retrieval‑Source‑Blind Reranking, as they overlook retrieval origins and candidate‑level retrieval priors, leading to redundant modality reliance. To address these challenges, we propose UniHEAR, a unified lightweight framework for heterogeneous‑source entity retrieval and reranking. UniHEAR constructs a Coarse Retrieval Descriptor for each candidate entity, and introduces Retrieval‑Guided Attentive Modality Gating to condition modality attention weights on this descriptor, complemented by Entropy‑Weighted Source Fusion of coarse retrieval priors. A hybrid training strategy combining contrastive learning with an auxiliary modality‑preserving loss unifies entity‑level and section‑level retrieval within a single model. Extensive experiments on E‑VQA and InfoSeek demonstrate that UniHEAR achieves state‑of‑the‑art retrieval and VQA performance, improving Recall@1 by 6.7 and 1.2 points over the strongest baselines while maintaining a lightweight reranking architecture. Code and model are available at https://github.com/iven‑luo/UniHEAR.

Authors:Yuyang Shen
Title: When Do Surrogate Updates Improve Decisions? A Local Theory of Trajectory-Wise Transfer
Abstract:
A broad range of models face the mismatch where they are updated through trajectory losses but are evaluated by downstream task reward. Here, a trajectory is a training instance that induces a surrogate loss whose reduction might not track the model's decision utility update. Theoretically, we ask when one step of trajectory training reduces both population surrogate loss and decision risk, and how transfer accumulates along repeated updates. To formalize this, we first fix a checkpoint and a restricted update space, and define the reductions in population surrogate risk and decision risk induced by a trajectory as its learnability and decision utility, respectively. On this basis, our theory yields four main results. First, a one‑step transfer bound separates their discrepancy into first‑order gradient misalignment after nonnegative calibration and second‑order curvature; and a pathwise extension accumulates the same terms over repeated updates. Second, when the accessible surrogate gradient is nonzero, universal first‑order transfer over every accessible direction holds exactly when the accessible surrogate and decision gradients are positively collinear. Third, the calibration gap bounds the decision regret of learnability‑based trajectory selection, while a candidate‑difference refinement tightens this guarantee by retaining only directions that affect pairwise rankings. Finally, we establish an approximation‑‑calibration trade‑off across nested update spaces. Controlled gridworld and LLM post‑training experiments yield results consistent with our predictions.

Authors:Ganghyeon Lee, Inha Lee, Junhee Lee, Jeongeon Lee, Sung Whan Yoon, Kyungdon Joo
Title: FeDepth: Federated Learning for Depth Estimation under Robot Heterogeneity
Abstract:
Although recent robot perception research emphasizes training on data from diverse environments to improve generalization, most existing methods still rely on centralized learning, which is inefficient and difficult to scale across heterogeneous robot platforms. Federated learning (FL) offers an alternative by enabling distributed training without raw data transfer, but it suffers from severe performance degradation under domain shifts caused by heterogeneity across clients. In real robotic deployments, data distributions often overlap across platforms, environments, and sensing conditions, making it difficult to partition clients into clearly separated domains. However, this characteristic breaks the assumption of clearly separable client domains commonly used in clustered FL. To address this gap in robot perception, particularly in depth estimation, we introduce two realistic and unexplored non‑IID scenarios that reflect heterogeneity in terms of platform, environment, and depth distribution. We then propose FeDepth, a descriptor‑based clustered FL framework that models client relationships through soft clustering. Unlike hard clustering methods that assume clearly separated clusters, FeDepth allows clients to participate in multiple clusters, capturing continuous and ambiguous domain transitions commonly observed in robotic environments. Extensive experiments demonstrate that FeDepth consistently improves robustness over standard FL and clustered FL baselines across multiple depth estimation architectures, providing a practical and effective solution for federated robot perception. Our project page is available at https://vision3d‑lab.github.io/fedepth/.

Authors:Siyuan Li, Aodu Wulianghai, Zehao Liu, Xi Lin, Qinghua Mao, Haoyu Li, Xiang Chen, Siyuan Liang, Jun Wu, Jianhua Li, Dacheng Tao
Title: SoK: Intent-Oriented Systematization of Multi-Turn LLM Jailbreaks
Abstract:
Large Language Models (LLMs) are increasingly deployed in interactive settings, where user intent commonly unfolds through multi‑turn dialogue. Multi‑turn jailbreaks exploit this pattern by advancing a harmful intent across turns, so that no single message exposes the full objective. However, existing work treats these attacks as a loose collection of prompt patterns and does not analyze how the adversary organizes and advances harmful intent across an interaction. We develop a four‑part, intent‑oriented taxonomy that organizes multi‑turn jailbreaks by adversarial intent structure. Through controlled ablations, we find that effectiveness is driven by how deliberately intent is organized across turns rather than by context length or query count. We further show that the way intent is organized determines the level at which it becomes detectable, pushing the required detection surface outward from the turn level to the session level to the cross‑session level. These findings indicate that turn‑local safety mechanisms are structurally insufficient and that single‑point evaluation overlooks how intent is organized, motivating evaluation protocols aligned to the level at which harmful intent becomes observable. The code is available at: https://github.com/SiyuanLi00/INTACT.

Authors:Sen Liang, Fengbin Guan, Youliang Zhang, Xin Li, Zhibo Chen
Title: CoT-Edit: Let CoT Guide Instruction Video Editing
Abstract:
Text‑driven instruction‑based video editing in complex scenes remains challenging: purely textual prompts often fail to capture precise spatial relationships and physical constraints, resulting in target ambiguity and physically implausible outcomes. To address this, we propose a plan‑‑guide‑‑edit framework that explicitly bridges semantic intent and spatial execution. In our framework, a Chain‑of‑Thought (CoT)‑enhanced multimodal large language model (MLLM) serves as a planner, performing structured reasoning over the video and instructions to derive a precise sequence of bounding boxes and attribute‑enriched editing directives. These spatial priors then guide a box‑conditioned mask generator, transforming ambiguous global retrieval into localized, context‑aware refinement and producing masks that more accurately capture object scale, contact relationships, and placement. Building on these spatial and semantic signals, a diffusion‑based editor integrates the masks, enriched instructions, and frame features to render high‑fidelity edits that remain temporally coherent and spatially well aligned. Trained first in a modular manner and then jointly, our framework achieves superior performance with reduced data requirements, delivering precise localization in scenes with multiple similar objects and physically consistent object additions, and extensive experiments demonstrate state‑of‑the‑art performance over multiple strong baseline methods. More details are available at: https://github.com/flying‑sky999/CoT‑Edit

Authors:Jiaming Jiang, Yuzhe Huang, Hao Liang, Pei Lin, Shengcheng Luo, Fanrong Dong, Jiaping Wu, Chenxi Xiao, Wanlin Li, Ziyuan Jiao
Title: CAAT: Contact-Aware Attention Scaling and Tactile Masking for Data-Efficient Contact-Rich Manipulation
Abstract:
In contact‑rich manipulation, visual observations primarily guide motion in free space, whereas tactile observations become particularly informative during contact. However, standard Transformer‑based visuo‑tactile policies typically rely on either token concatenation or learnable gating. These approaches lack explicit contact‑aware priors, making it difficult to efficiently learn effective cross‑modal representations from demonstrations. To address this limitation, we propose CAAT, a lightweight contact‑aware framework that explicitly incorporates contact priors through attention scaling and dynamic tactile masking. Specifically, CAAT emphasizes visual information before contact and tactile information during contact. It also suppresses static background tokens by comparing the current tactile observation with a non‑contact reference. CAAT can be integrated into commonly used Transformer‑based policies without modifying their action decoders. In simulation, integrating CAAT with ACT improves the average success rate by 18.0 percentage points over direct visuo‑tactile fusion and by 10.0 percentage points over gated fusion. In real‑world experiments using a visuo‑tactile UMI platform, CAAT achieves an average success rate of 60.0% across ACT, Diffusion Policy, and π_0, outperforming the strongest baseline by an average of 21.1 percentage points. These results demonstrate that explicit contact priors and dynamic tactile masking are effective in improving visuo‑tactile policy learning and task performance of diverse policy architectures. https://mrjiangjm.github.io/caat/

Authors:Shigeng Wang, Chao Li, Yangyuxuan Kang, Jiawei Fan, Anbang Yao
Title: Attend to Your Own Thoughts: Breaking the Barrier for Post-Training Quantization of Reasoning LLMs through the Lens of 1.58-Bit Quantization
Abstract:
We propose ScaleQ‑1.58, a scalable ternary post‑training quantization (PTQ) framework for reasoning LLMs. Its core insight stems from an empirical finding: although modern LLMs are typically trained to exhibit chain‑of‑thought reasoning capabilities, in the PTQ regime, even the latest CAT‑Q method based on learning‑based differentiable ternarization still leads to performance collapse on challenging mathematics and coding tasks when using conventional calibration schemes that ignore the model's reasoning process. Driven by this finding, we introduce a simple calibration approach, Attend to Your Own Thoughts (AYOT), where reasoning traces and final answers generated by the pre‑trained high‑precision target LLM on a proper set of calibration samples are used as the context input during the ternarization process, along with the corresponding questions. ScaleQ‑1.58 is formed by simply integrating AYOT with CAT‑Q, which demonstrates several scaling properties: (1) with only 4M calibration tokens, Qwen3‑1.7B ternarized by ScaleQ‑1.58 reaches over 90.52% of the performance of the prior best BitNet b1.58 2B4T averaged over 4 mathematics and coding tasks, and our ternary Qwen3‑4B shows an absolute gain of 8.97%, while requiring 1,000,000x fewer calibration tokens for quantization; (2) ScaleQ‑1.58 generalizes well to both dense and MoE architectures, with performance improving as model scale increases (up to 235B parameters); (3) ScaleQ‑1.58 demonstrates strong generalization across tasks of varying difficulty levels, including mathematics, coding and scientific logic reasoning, as well as commonsense reasoning and basic language generation; (4) its performance continues to improve as the number of calibration tokens increases. Notably, AYOT also exhibits strong generalization ability across other quantization bit‑widths. Code will be available at https://github.com/IntelChina‑AI/BitTern.

Authors:Savindu Dilshan Wickramasinghe
Title: VespaSeg: A Resource-Aware Ground-then-Segment Pipeline for Referring Expression Segmentation
Abstract:
Referring expression segmentation requires language conditioned localization and pixel‑accurate masks, but monolithic models can be costly to deploy. We present VespaSeg, a modular pipeline that grounds a text query with a compact vision‑language model and converts the predicted box to a mask with MobileSAM. We study Florence‑2‑base, Florence‑2‑large, and Moondream2 grounders together with targeted adaptation of the grounding and segmentation stages. Under a repository‑specific RefCOCO validation protocol containing the first expression for each of 3,811 referenced‑object records, the adapted Florence‑2‑base pipeline obtains 73.64 mean intersection over union (mIoU) and 84.60 precision at IoU 0.5. On an NVIDIA RTX 6000 Ada GPU it processes 22.8 cached‑image queries per second with 2.20 GB mean allocated GPU memory. A matched 500‑query comparison gives 73.73 mIoU for Florence‑2‑base and 72.82 for Florence‑2‑large, while the base model is 1.70 times faster and uses 1.17 GB less allocated memory. Ablations show that ground‑truth‑box adaptation raises MobileSAM mIoU from 82.22 to 86.61 and that reducing the Florence‑2 output‑token budget from 64 to 32 preserves accuracy. These results support compact, modular grounding and segmentation, while also exposing the need for evaluation on the complete standard RefCOCO expression splits and deployment hardware.

Authors:Muhammad Yousaf Rehman, Muhammad Islam
Title: DeBERTa-Sentinel: Toward Transparent and Trustworthy Detection of AI-Generated Text
Abstract:
The rapid spread of large language models (LLMs) across the web raises concerns about misinformation, academic integrity, automated content manipulation, and risks to vulnerable online communities. Existing transformer‑based detectors, such as GPT‑Sentinel, show promise but struggle to generalize to diverse model outputs and paraphrasing attacks, limiting their role in building trustworthy web ecosystems. This work introduces DeBERTa‑Sentinel, a responsible AI‑generated text detection framework leveraging DeBERTa‑v3's disentangled attention to capture subtle structural irregularities in synthetic content. A central design principle is transparency: unlike black‑box commercial detectors, DeBERTa‑Sentinel exposes token‑level explanations of its decisions, enabling affected stakeholders journalists, educators, and platform trust and safety teams to audit, challenge, and contextualize detection outcomes. Using the GLC‑AIText dataset of 28,057 human and LLM‑generated samples (GPT, LLaMA, and Claude) with a 60‑20‑20 split, DeBERTa‑Sentinel achieves 98.21% validation accuracy and surpasses the RoBERTa‑Sentinel baseline from NeurIPS 2025, achieving 97.53% test accuracy, 95.89% precision, 99.33% recall, and 99.53% ROC‑AUC, and maintaining a 0.665% false negative rate. The model's interpretability reveals linguistic markers such as academic phrasing and formal transitions associated with synthetic text, directly supporting stakeholder needs for verifiable, auditable content‑authenticity decisions. By advancing responsible detection methods that reduce bias and enhance explainability, DeBERTa‑Sentinel promotes trustworthy, ethical, and human‑centric AI systems. Code and data are available at https://github.com/Galileo‑Galili/HUMAN‑VS‑AI‑TEXT‑DETECTION.

Authors:Kento Kawaharazuka, Shuhei Ikemoto
Title: Diffusion-Based Body Schema Learning Enabling Abnormal-State Adaptation in Musculoskeletal Robots
Abstract:
Musculoskeletal robots require an internal body schema that remains consistent under a wide range of physical state changes, including abnormalities such as muscle rupture and actuator jamming. Conventional approaches based on autoencoders or variational autoencoders learn average behaviors by projecting sensor and actuator signals into a low‑dimensional latent space; however, exploration within the latent space alone has limited capability to handle out‑of‑distribution or abnormal states that are not included in the training data. To address this limitation, this study proposes a diffusion‑based framework for body schema learning in musculoskeletal robots. Unlike generative models that operate through low‑dimensional latent spaces, diffusion models can directly and iteratively estimate physically consistent sensor and actuator values in the high‑dimensional space through a denoising process, even under partial observations and constraints, without requiring retraining. By formulating body schema adaptation as a gradient‑guided denoising process, the proposed method enables adaptive estimation of appropriate muscle lengths and muscle tensions even under abnormal conditions such as muscle rupture and actuator jamming. The validity of the proposed framework is verified through simulation experiments using a musculoskeletal robot model.

Authors:Taku Okawara, Aoki Takanose, Kenji Koide, Shuji Oishi, Masashi Yokozuka
Title: KING: Embodiment-Aware Kinematic Graph Neural Network for Unified Motion Representation of Legged and Wheeled Robots
Abstract:
Kinematic models provide reliable motion constraints for odometry estimation in featureless environments, where exteroceptive sensing degrades and IMU integration drifts. Learning‑based kinematic models can achieve more accurate odometry estimation than model‑based methods by capturing nonlinear effects; however, most existing learning‑based models are trained on a single embodiment and generalize poorly to new embodiments. This generalization is difficult because the meanings and structures of proprioceptive measurements vary across embodiments, including the number of joints and ground‑contact elements (e.g., wheels, feet). To address this challenge, we propose KING, a Graph Neural Network (GNN)‑based kinematic model that explicitly incorporates robot embodiments by representing them as a common graph. We show that wheel and leg kinematic models can be expressed by a unified representation, enabling a single model for both wheeled and legged robots. Trained on datasets spanning diverse embodiments, KING provides a unified representation of wheeled and legged kinematics and achieves high‑accuracy odometry estimation in real environments. KING estimates accurate odometry using only an embodiment description (e.g., a URDF file) and on‑board proprioception (encoders and an IMU) and can be adapted to new robot embodiments through few‑shot learning with only one minute of data, avoiding retraining from scratch on a new dataset for each robot. The project page is available at: https://smrg‑aist.github.io/king_project_page/

Authors:Yohei Nakajima
Title: Passing Coarse Marginal Checks Can Be Cheap: Persona Mixtures and Imprecise Treatment-Response Estimates in an LLM Persona Panel
Abstract:
Large language models are increasingly used as synthetic research participants and are often validated by whether their marginal responses resemble human data. We study a fixed panel of sixteen lightweight persona‑conditioned GPT‑4.1 configurations in repeated strategic games. The panel met preregistered broad‑reference condition‑mean criteria in three of four repeated‑game cells; the sole miss was 0.011 below the lower reference bound. Variation was strongly prompt‑indexed, but its share depended on uncertainty assumptions: fixed‑panel symmetric‑Dirichlet sensitivities produced median between‑prompt shares of 63%‑71% under Jeffreys alpha=0.5 and 47%‑53% under alpha=1, while finite‑opportunity plug‑in estimates were 85%‑96%. Aggregate continuation‑probability contrasts were +0.083 and +0.078, with conservative simultaneous 95% intervals [‑0.171, +0.330] and [‑0.181, +0.330]. The treatment jointly changed the continuation process and its textual representation. A separate wording‑and‑position operation shifted cooperation from 0/40 to 37/40 in the bare configuration, and a label conflict also revealed representation control. The original persona‑level p13 result was not prospectively family‑controlled, while a post‑adjudication exact gate was structurally underpowered; p13 is therefore a replication target rather than a finding. External review exposed family‑error, dependence, construct, and boundary‑uncertainty defects, and zero‑call reanalysis changed the interpretation without rewriting the historical record. The registered marginal criteria could be passed without precisely estimating the treatment‑response object. A public capsule verifies 4,916 confirmatory Phase 3‑5 runs with no live model calls. The results concern one fixed model‑prompt panel and do not establish human substitutability.

Authors:Myeongkyun Kang, Yanting Yang, Xiaoxiao Li
Title: Location-Aware Fine-Grained Representation Learning for Medical Vision Foundation Models
Abstract:
Fine‑grained visual representations are essential for medical image analysis, particularly when diagnostically relevant evidence is subtle and spatially localized. Modern transformer‑based medical vision encoders must therefore learn patch‑level representations that are both clinically meaningful and spatially consistent. Without these properties, large vision‑language models (LVLMs) operate on an ambiguous visual foundation, limiting their ability to generate clinically reliable and spatially grounded responses. However, existing training strategies for medical vision encoders rarely achieve both objectives. Image‑text alignment provides clinically meaningful supervision primarily at the image level, leaving the spatial localization of diagnostic evidence weakly constrained. In contrast, self‑supervised learning promotes spatial consistency but lacks the semantic supervision needed to distinguish visually similar yet clinically distinct regions. To address this gap, we present LoFi, a medical vision foundation model built on location‑aware fine‑grained representation learning. LoFi trains a vision encoder with a lightweight large language model under grounding and grounded captioning objectives. Because these objectives require predicting location from clinical text and vice versa, spatial consistency emerges without any explicit patch‑level regularization. To enable training at scale, we construct MedG, a large‑scale medical grounding dataset of 4.48M image‑text‑box triplets curated from 84 datasets spanning 7 modalities. Across phrase grounding, visual question answering, and region‑based organ classification under perturbations, LoFi consistently outperforms general‑purpose and medical vision foundation models as well as state‑of‑the‑art LVLMs. Code is available at https://github.com/myeongkyunkang/lofi‑medg.

Authors:Minseong Kweon, Junaed Sattar
Title: Swimm3R: Splatting with Medium-aware SfM for Underwater 3D Reconstruction
Abstract:
We propose Swimm3R, a unified framework that combines medium‑aware structure‑from‑motion (SfM) with Underwater Beta Splatting to address scattering‑ and attenuation‑induced failures in underwater 3D reconstruction. Swimm3R distills in‑air geometric priors into a feed‑forward backbone and uses a physics head to regress underwater image‑formation parameters, camera poses, and restored point clouds. Additionally, we introduce Underwater Beta Splatting, which extends Gaussian splatting with Beta primitives and scattering‑aware geometric gradients for stable underwater geometry representation. We further establish the Barbados underwater video dataset to demonstrate the effectiveness of our method in challenging underwater environments. On this dataset, Swimm3R robustly recovers underwater scene structure under challenging scattering conditions, yielding coherent seafloor geometry. Using these predicted point clouds, the proposed Underwater Beta Splatting improves average PSNR by 1.47 dB over WaterSplatting while increasing downstream localization performance by 2.0 and 2.4 percentage points in RRA@15 and RTA@15, respectively.

Authors:Maxim Nikolaev
Title: Claim Plane: Reliability Gains and the Limits of Selective Concurrency for Parallel Coding Agents: A 30-Pair, Three-Seed Confirmatory Study of Deterministic Pre-Write Admission
Abstract:
Parallel coding agents can produce locally valid changes that fail when combined. Claim Plane addresses this failure mode as deterministic pre‑write admission over versioned change intents. This paper reports a confirmatory study on 30 frozen CooperBench feature pairs, balanced between 15 conflict and 15 clean labels, with three coder seeds, four coordination arms, and 360 completed executions. DeepSeek V4 Pro generated 60 feature‑level planner declarations once; the declarations were frozen across all arms and coder seeds, while DeepSeek V4 Flash performed the coding work. Static Claim Plane raised pair pass from 23.3% under unconstrained parallel execution to 50.0%, a paired task‑cluster difference of +26.7 percentage points (95% bootstrap CI 9.6 to 60.0), and raised integration success from 65.6% to 96.7%. On conflict‑labeled pairs, pair pass rose from 6.7% to 60.0%. However, static admission serialized 96.7% of executions, including 93.3% of clean cases, and therefore recovered reliability largely by collapsing toward serial execution. Dynamic admission was more selective, serializing 66.7% of conflict cases and 13.3% of clean cases, but 46 of 90 executions failed closed on undeclared scope, reducing pair pass to 22.2%. Forty‑five of those 46 blocks targeted files already present in the frozen declarations, indicating region undercoverage and insufficient amendment handling rather than wholly unknown files. The results support pre‑write admission as a reliability mechanism, but they do not establish useful wall‑clock parallel speedup: provider calls were physically sequential, and the conservative policy largely serialized the workload. The complete study artifacts, hashes, and clustered bootstrap analysis are publicly released for reproduction.

Authors:Juan Li, Wei Cai, Yan Bai
Title: Neuro-Symbolic Participation Governance for Verifiable AI Agents in Open Digital Twin Ecosystems
Abstract:
Autonomous AI agents, increasingly empowered by large language models, are becoming important components of human‑machine systems for high‑stakes decision support in digital twin ecosystems. However, existing multi‑agent systems often lack robust verification for identity, capability, and policy compliance, especially in decentralized environments spanning multiple institutions. This paper proposes a neuro‑symbolic decentralized governance framework for verifiable agents in collaborative digital twin environments. By representing agents through multi‑layer semantic profiles, the framework bridges probabilistic neural reasoning with deterministic institutional governance, thereby supporting trustworthy human‑AI collaboration and meaningful human oversight. Capabilities are grounded in formal domain ontologies to enable machine‑interpretable, policy‑aware, and context‑sensitive participation. These credentials, issued by organizational authorities, are validated via blockchain‑based smart contracts, ensuring auditable participation without exposing sensitive data. We demonstrate the framework using a decision‑support prototype with clinic, digital twin, and wearable provider agents effectively prevents unauthorized interaction and enforces institutional policies with manageable overhead. Our findings suggest that neuro‑symbolic decentralized governance provides a scalable and trustworthy pathway for safe human‑machine collaboration across institutional boundaries.

Authors:Binshuang Li
Title: UpliftBench: Revealing Outcome-Regime and Objective Mismatch in Uplift Evaluation
Abstract:
Uplift modeling (conditional‑average‑treatment‑effect estimation) drives personalized targeting, yet published uplift benchmarks frequently disagree on which estimator performs best; we show the disagreement is substantially about metrics, not models. UpliftBench evaluates 12 uplift estimators under an outer‑test‑isolated, multi‑objective protocol across seven dataset families; its two findings are identified where a reference objective exists ‑‑ F1 on the standard continuous benchmark (IHDP), F2 in a within‑sample case study on Jobs. On that benchmark, Qini shows no detectable alignment with effect accuracy ‑‑ across all 100 IHDP realizations its mean rank correlation with effect accuracy is +0.07, 95% CI [‑0.03, +0.16] ‑‑ while AUUC is consistently more aligned (paired prefix‑mean‑AUUC‑over‑Qini gap +0.49 [+0.40, +0.59]; the shipped cumulative‑gain AUUC aligns better still, +0.73). On Jobs, ranking metrics are structurally insufficient for a sign‑threshold policy because they discard the score level; empirically, within the released split‑rotation analysis direct policy‑risk selection yields lower benchmark regret than random model selection while Qini, AUUC, and uplift‑at‑k do not (14‑15% regret). Calibrating the decision threshold removes 81% of the Qini‑selection regret. Both findings are bounded, not universal: F1 is not detected on either validation family (the ACIC and Revenue‑Synthetic gaps are both indistinguishable from zero), and F2 vanishes under a budgeted‑value objective where rank suffices. UpliftBench releases versioned loaders, fixed protocols, result artifacts, and a reproducible living leaderboard; the public repository accompanies the paper.

Authors:Weimin Fu, Hejia Zhang, Minghao Shao, Zeng Wang, Johann Knechtel, Ozgur Sinanoglu, Muhammad Shafique, Ramesh Karri, Xiaolong Guo
Title: FinHardBench: Can LLMs Generate Latency-Aware Hardware for Financial Computing?
Abstract:
Can large language models generate not just correct, but fast hardware? This paper investigates the question in financial FPGA design, where 5‑10 nanoseconds of latency determines competitive advantage and designs iterate continuously as protocols, strategies, and regulations evolve. FinHardBench, a benchmark of 33 financial computing tasks, is presented together with three experiments that mirror the real‑world FPGA iteration cycle: generating new modules from specifications, tuning system‑level configurations across a 6‑stage trading pipeline, and adapting existing modules to specification changes. Evaluation of six LLMs on 1530+ experiment rounds yields three findings: (1) models achieve 19‑61% functional correctness with timing degradation up to 13.7× on specific tasks; (2) in system‑level design space exploration, top LLMs converge to the optimal configuration with higher reliability than random search, simulated annealing, and Bayesian optimization baselines (5/5 seeds vs. 0‑4/5 at the same 24‑round budget); (3) strategy‑level specification changes remain unsolved for most models. Across the six models, generation and DSE rankings overlap moderately: the strongest code generator is not the fastest architecture optimizer, and the weakest code generator (MiniMax M2.7) still reaches the system optimum on 4 of 5 seeds. On the tasks in FinHardBench, difficulty tracks training data pattern availability more closely than abstraction level. FinHardBench is released as an open‑source benchmark.

Authors:Dongheng Lin, Jianbo Jiao
Title: PeCA: Palette Context Assisted Inference for Test-Time Paint-Bucket Colourisation on Animation Videos
Abstract:
In animation production, paint‑bucket colourisation for hand‑drawn animation is a labour‑intensive procedure that assigns each enclosed region in line sketches a colour from reference design sheets. Recent automatic paint‑bucket colourisation pipelines mirror this workflow via region correspondence, but correspondences can be brittle when regions are ambiguous fragments without proper context. In this paper, we propose Palette Context Assisted (PeCA), a new training‑free, plug‑and‑play framework for animation video colourisation that aims to close this gap at test‑time via reasoning over spatial and temporal contexts. Extensive experiments on existing benchmarks and a newly introduced long‑video test case show consistent performance boosts.

Authors:Jitendra Prajapati
Title: An explicit construction of two completely independent spanning trees in the four-dimensional dual-cube
Abstract:
Lalou, Mbarek, Skender and Togni (arXiv:2607.25917) proved that the n‑dimensional dual‑cube F_n admits two completely independent spanning trees for every n\ge 5, observed that none exist for n\le 3, and identified F_4 as the first unresolved case, reporting more than 700 hours of inconclusive computation. We settle this case affirmatively by an explicit construction, completing the classification: F_n admits two completely independent spanning trees if and only if n\ge 4. The internal‑vertex sets of the two trees are the level sets of a single ten‑term cubic polynomial over \mathbbF_2 in the seven vertex bits, and correctness reduces to finite connectivity checks that are machine‑verified by a solver‑free program distributed with the certificate. In F_4 the two trees necessarily use 254 of the 256 edges. We also report exact infeasibility results for simpler rules of the same shape: within the search model, no affine or quadratic rule works, and ten terms is the fewest possible for a cubic rule.

Authors:Eylon E. Krause
Title: Shiftfly: Scaling the Accelerator Interconnect Past the Pod with a Shift-Routed Optical Tier
Abstract:
Google's TPU interconnect spent nine generations as a k‑ary n‑cube, whose diameter grows as Θ(N^1/n), before TPU 8i replaced it with Boardfly: a three‑tier hierarchy in which every tier is a complete graph, giving pod diameter 7 over 1,152 chips. Boardfly suits a single inference pod but does not extend, because a complete global tier needs G‑1 optical ports to reach G groups. A 400,000‑chip machine would need 12,499 per group against the 40 available, so another hierarchy level must be stacked, and each level costs four chip hops. We propose Shiftfly, which keeps Boardfly's building block and group verbatim and replaces only the global tier with a generalized Kautz digraph, installed as a fixed permutation on the optical circuit switch the fabric already owns. At an identical 40 ports per group, Shiftfly is flat, has guaranteed diameter \lceil \log_d G \rceil, routes without tables by a shift register, and addresses content natively. The evaluation is deliberately two‑sided. Shiftfly loses at one‑pod scale, where Boardfly achieves chip‑level diameter 7 against Shiftfly's 11, and wins beyond it, cutting worst‑case distance from 23 to 19 hops at 400,000 chips with roughly 2.7x better spectral expansion at equal cost. The redundancy argument that motivated the design does not survive its own evaluation: locality‑aware placement supplies most of the achievable saving on shared content, leaving the shift algebra a 2.9% residual, and we report the metric inversion that conceals this. We also measure operability. Replacing a failed group costs the same 40 optical circuits in both designs; deriving the global wiring costs 28 bits of control‑plane state against 550,000. Slice allocation is the one regression: an arbitrary induced subset of a shift graph is disconnected, so slices must be instantiated rather than carved.

Authors:Kevin Bui, Adina Ciomaga
Title: MBO Scheme for Local Chan--Vese Segmentation
Abstract:
Robust to intensity inhomogeneity, the local Chan‑‑Vese (LCV) model extends the classical Chan‑‑Vese (CV) image segmentation method by incorporating local statistical information around each pixel. Originally, the LCV model was solved using a finite difference scheme, following the approach used for the CV model. As an alternative to the finite difference scheme, a more efficient algorithm based on the Merriman‑Bence‑Osher (MBO) scheme was later developed for the CV model. In this paper, we derive a similar MBO‑based algorithm to solve the LCV model and propose an efficient implementation. The algorithm is developed for both two‑phase and multiphase segmentation, and an extension to color images is also discussed. To demonstrate the effectiveness of the proposed approach, we apply it to a variety of grayscale and color images, including medical and microscopy images.

Authors:Kazi Ahmed Asif Fuad, Lizhong Chen
Title: SparseKAN: Compressing Kolmogorov--Arnold Networks Across Basis Functions, Neurons, and Bits
Abstract:
Kolmogorov‑‑Arnold Networks (KANs) replace scalar edge weights with learnable univariate functions parameterized by multiple basis coefficients. This introduces a source of redundancy that conventional neural‑network compression does not directly expose. We present SparseKAN, a unified approach that compresses KANs along three complementary axes: basis functions, neurons/channels, and numerical precision. SparseKAN equips the base branch, nonlinear basis branch, and individual basis terms with hierarchical learnable gates trained under a differentiable active‑cost objective. The learned importance structure is subsequently hardened under explicit basis and width budgets, recovered in full or low precision, and physically compacted into smaller dense tensors rather than retained as sparse masks. Experiments on MNIST, CIFAR‑10, and CIFAR‑100 across spline, polynomial, RBF, wavelet, and convolutional KAN variants show that the structural axes compose predictably in cost. We also find strong basis‑dependent differences in term importance: coefficient‑based selection outperforms matched low‑order truncation by up to 15.25 accuracy points in the evaluated Gram‑polynomial settings. Eight‑bit quantization is broadly robust, whereas 4‑bit convolutional KANs require quantization‑aware adaptation. Physical compaction removes up to 73.0% of parameters without accuracy loss on MNIST and reduces large‑batch CUDA latency to as little as 0.51× dense execution. On a ZCU104 FPGA, the resulting sparse low‑bit models achieve up to 23.63× lower inference latency, demonstrating that SparseKAN converts functional redundancy into measurable software and hardware efficiency. The SparseKAN implementation is available at https://github.com/OSU‑STARLAB/SparseKAN.

Authors:Muhayy Ud Din, Ahmed Nadar, Jan Rosell, Irfan Hussain
Title: RIT*: Riemannian Informed Trees for Cost-Adaptive Optimal Motion Planning
Abstract:
We present Riemannian Informed Trees (RIT), a planning framework that replaces Euclidean primitives in batch‑informed search with their Riemannian counterparts. RIT constructs a tighter, cost‑consistent informed set, performs a nearest‑neighbour search under an anisotropic distance metric, and evaluates edge costs efficiently via a cascading scheme. We further introduce a Collision‑Adaptive Metric Refinement (CARM), which learns an obstacle‑proximity cost field online from collision feedback, reducing the reliance on prior metric design in practical settings. Experiments across environments from 2‑D to 14‑D show that RIT is competitive in low‑dimensional and spatially constant‑metric settings and produces substantially lower‑cost solutions when the metric varies spatially in high‑dimensional configuration spaces. Performance gains scale with anisotropy and dimension, reaching up to 13.0% improvement in median initial cost over BIT in the 3‑D anisotropic benchmark, up to 9.0% in median final cost over BIT in 6‑DOF manipulation, and 24.8‑63.5% in a 14‑DOF bimanual planning problem, where Euclidean‑informed baselines degrade. Videos and code can be found here: https://muhayyuddin.github.io/ritstar/

Authors:Boyi Liu, Qijin Li, Tianqi Yu, Qinrui Yan, Xingxing Zuo
Title: LooperMuscle: Fast and Stable Learning of Humanoid Whole-Body Tracking via Structured Mixture-of-Experts
Abstract:
FastSAC‑style methods significantly reduce humanoid motion training time but often suffer from notable performance degradation compared with PPO in whole‑body tracking tasks. We target this speed‑performance gap by introducing LooperMuscle, a composed expert policy learning framework that restores tracking quality while preserving high training efficiency. LooperMuscle combines a semantically structured mixture‑of‑experts actor, an expert‑aware distributional critic, and contribution‑routed replay with deferred curriculum scheduling. These three components form a closed training loop in which expert contributions guide data routing, routed data shape value learning, and value gradients in turn refine expert specialization. Empirically, our approach substantially outperforms vanilla FastSAC in motion tracking accuracy while requiring far less wall‑clock time than PPO: where FastSAC trains in about 15 minutes but underperforms, and PPO achieves stronger results but requires about 6 hours, LooperMuscle recovers a substantial fraction of the remaining gap to PPO in roughly 45 minutes of simulation training, delivering practical efficiency for rapid policy iteration. The code will be released to benefit the research community at https://loopermuscle.github.io/.

Authors:Meher Bhaskar Madiraju, Meher Sai Preetam Madiraju
Title: AgentSLABench: Evaluating and Benchmarking Agentic Systems Under Resource Constraints
Abstract:
We present AgentSLABench, a resource‑aware evaluation framework for autonomous AI agents that measures correctness alongside latency, cost, compute, memory, and network usage under declared resource budgets. Unlike standard benchmarks that report only accuracy, AgentSLABench produces a multi‑dimensional profile per agent per task ‑ the same way systems profilers (perf, pprof, cProfile) measure resource consumption of code, but extended with task correctness as a first‑class dimension. AgentSLABench provides 16 task environments across 6 categories (5 core: multi‑hop QA, retail substitution, code generation, web shopping, travel planning; 11 extended) with isolated Docker containers, declared CPU/memory/time/network budgets, sealed test sets with SHA256 hashes, and a standardized profiling protocol. We profile 5 general‑purpose baseline agents (ReAct, PlanAndSolve, Reflexion, CoT, Random) plus 4 task‑specialized agents, finding that specialized agents achieve 100% success on 3/5 core tasks (fact_qa, web_shopping, travel_planning) and 66.7‑83.3% on retail and code_gen, while general baselines fail entirely on 4/5 domain tasks. Crucially, we report the Efficiency‑Adjusted Success Rate (EASR) ‑ success weighted by resource consumption relative to declared budgets ‑ revealing that high accuracy at unbounded cost is not production‑viable. We release the full infrastructure, sealed test sets, and profiling results to enable reproducible, resource‑aware agent evaluation.

Authors:Pengyun Qiu, Shuo Wang, Zeyuan Chen, Yihao Zhi, Chongjie Ye, Xiaoguang Han
Title: AIMold: An Autonomous AI-based Pipeline for Complex Mold Design
Abstract:
Injection molding is the cornerstone of mass‑producing plastic components. While current algorithms can automate mold design for basic geometries using standard two‑piece molds, complex parts featuring undercuts, side holes, or re‑entrant features present a significant challenge. These geometries often necessitate auxiliary components beyond the primary upper and lower molds. In practice, designing these intricate assemblies is a laborious process that relies heavily on expert knowledge. Furthermore, the scarcity of public datasets has hindered the development of effective learning‑based solutions. To bridge these gaps, we introduce MoldCAD, a curated dataset that pairs complex single‑body CAD parts with industry‑standard mold assemblies. Each entry includes the upper and lower molds, parting surfaces, demolding orientations, and necessary auxiliary components. The dataset comprises 4,934 CAD models and over 3,850 mold assemblies, totaling more than 23k individual models. Building upon this dataset, we propose a comprehensive pipeline that predicts demolding orientations, identifies auxiliary components, and constructs parting surfaces to derive a complete, manufacturing‑ready mold assembly for downstream CAD/CAM workflows. Our results demonstrate a promising path toward fully automated industrial mold design and contribute to the broader advancement of manufacturing‑aware CAD generation.

Authors:Soslan Kabisov, Gennadiy Savrasov, Maksim Elistratov, Antonio Rodriguez, Daniil Ignatiev, Nikita Gavrilov, Rustam Uzdenov, Alexey I. Boyko, Igor Pasechnik, Anton Konushin, Andrey Kuznetsov, Dmitrii Zhemchuzhnikov
Title: CADENA: Stepwise CAD Reverse Engineering
Abstract:
Computer‑Aided Design (CAD) underpins modern engineering, yet converting existing shapes into editable models still demands substantial expert effort. Most AI systems emit the entire CAD program in a single pass, never inspecting the intermediate geometry. In contrast, human engineers build a part feature by feature, checking after each operation what remains to be modeled. We introduce CADENA (Spanish for "chain"), a model that reconstructs a 3D mesh as a parametric CAD program, growing its sequence of operations one at a time and comparing the target with the currently predicted geometry at every step. We also address the lack of benchmarks for evaluating reverse‑engineering methods on mechanical parts, introducing CADENA‑Bench, a benchmark that measures performance across categories of mechanical parts. CADENA outperforms prior methods on CADENA‑Bench and on the DeepCAD, Fusion 360, and MCB datasets. Code is available at https://github.com/zhemdi/cadena, model weights at https://huggingface.co/kulibinai/cadena, and CADENA‑Bench at https://huggingface.co/datasets/kulibinai/cadena‑bench.

Authors:Zixuan Zhu, Rui Wang, Lihua Jing, Jinwen Zhong
Title: Mitigating Backdoors via Decoy Shortcuts and Knowledge Decoupling
Abstract:
Backdoor attacks pose a serious threat to deep neural networks, especially when training relies on third‑party data, allowing adversaries to inject malicious behaviors through data poisoning. In this work, we reveal that backdoor behaviors tend to be absorbed by a simpler parallel branch when jointly trained with the main network. Motivated by this insight, we propose Trapping and Removing (TR), a simple yet effective training‑time defense that introduces a lightweight shortcut branch as a "honeypot" to trap backdoor knowledge. After training, backdoors can be removed by discarding the shortcut, without requiring any additional data. To further enhance backdoor isolation while maintaining benign performance, we design a knowledge decoupling strategy with entropy‑based weight assignment, encouraging poisoned samples to flow through the honeypot while guiding the main network to focus on benign learning. In addition, we introduce an automatic shortcut generation strategy to improve generalization across model architectures. Extensive experiments on four benchmark datasets and five model architectures demonstrate that our approach effectively mitigates a wide range of backdoor attacks while preserving performance on benign data. Code: https://github.com/Zixuan‑Zhu/TRgithub.com/Zixuan‑Zhu/TR.

Authors:Renhao Lu, Mingxin Wang, Chenyang Cao, Yang Yang, Guoping Pan, Kangkang Dong, Yi Cheng, Houde Liu
Title: Push-Wiper: Toward General-Purpose Robotic Cleaning across Varied Stains and Surfaces with Segmented Pushing Trajectories
Abstract:
Viscous stains, characterized by high viscosity and complex rheological properties, remain a major challenge for robotic surface cleaning. Conventional wiping often spreads the stain, while scrubbing provides stronger friction but risks damaging the surface. In this paper, we propose Push‑Wiper, a framework that reformulates viscous stain cleaning as an aggregation problem. Push‑Wiper employs a sponge to progressively gather stains through segmented pushing trajectories, followed by a post‑processing phase that detaches the aggregated material and enables sponge self‑cleaning. We adopt a stepwise strategy for stain gathering and leverage Diffusion Policy to generate adaptive pushing action sequences. These sequences are executed through our Arbitrary Surface Pose Interpolator (ASPI) and a hybrid force‑position controller, allowing the method to generalize to stains with diverse spatial distributions. Push‑Wiper achieves a cleaning score (CS), defined as the percentage of stain area removed, up to 130% higher than baseline methods. Without additional training, Push‑Wiper also transfers in a zero‑shot manner to solid residues, liquid spills, unseen viscous stains, and curved surfaces with varying geometries. Our experiments demonstrate the cleaning effectiveness of Push‑Wiper and its strong generalization ability. The project website is available at https://push‑wiper.github.io/.

Authors:Xinshun Feng, Ziqi Miao, Lijun Li, Jing Shao
Title: Tracing the Cascade: A Topology-Aware Evaluation Framework for Scientific Agent Hallucinations
Abstract:
Large language model (LLM) agents are increasingly deployed in scientific research, where reliability is critical and the underlying knowledge is densely interconnected. In such settings, hallucinations are particularly damaging: a single erroneous claim on a foundational concept can propagate through multi‑step reasoning and corrupt entire trajectories. Existing hallucination benchmarks largely operate at the surface level, treating facts in isolation and relying on uniform accuracy metrics that ignore this topological structure. We address this gap with SCHEMA, the first evidence‑grounded, topology‑aware evaluation framework for hallucinations in scientific agents. SCHEMA automatically constructs scientific concept graphs from benchmark seeds and literature evidence, synthesizes graph‑grounded tasks spanning claim verification, multi‑hop reasoning, open‑ended explanation, and experimental code generation, and evaluates agents with two complementary diagnostics. A trajectory hallucination pipeline audits intermediate reasoning at scale via a topology‑weighted severity score, while a multi‑agent counterfactual attribution module pinpoints the causal mechanism behind selected failures. SCHEMA reveals that hallucinations concentrate at a small set of highly connected knowledge hubs, and that final‑answer accuracy decouples from trajectory honesty; models often reach correct conclusions through structurally flawed reasoning. These results indicate that for high‑stakes scientific applications, terminal accuracy alone is an insufficient signal of agent reliability, motivating mechanism‑level evaluation grounded in knowledge topology. Code is available at https://github.com/circles‑post/SCHEMA.

Authors:Kaihua Tang, Ziqing Xia, Xiaoxu Zheng, Xiaoxue Zhang, Michael Bi Mi, Zhan Xu, Dave Zhenyu Chen
Title: Breaking the Horizontal Prior: From Long-Tailed Orientation Bias to Roll-Robust Monocular Depth Estimation
Abstract:
Despite recent advances in Monocular Depth Estimation, state‑of‑the‑art depth foundation models remain vulnerable to robustness issues. Particularly, even slight camera rolls can result in substantial degradation in depth estimations. We attribute this problem to a previously overlooked phenomenon, termed the Horizontal Prior, which is a manifestation of long‑tailed distribution bias: most training images are captured in approximately horizontal orientations due to human visual preferences and photographic habits. While intuitive remedies such as re‑balanced data augmentation and horizon leveling provide partial improvements, they fail to fully address the issue. In this paper, we introduce Invariant Depth Constraint (ID‑Constraint), a training‑time supervision strategy that improves roll robustness by fine‑tuning and jointly regularizing the depth backbone with a series of geometric and spatial reasoning tasks. These auxiliary objectives encourage the backbone to learn rotation‑stable, depth‑relevant representations, while the auxiliary prediction heads are discarded after training, leaving the original inference architecture unchanged. Extensive experiments on five benchmark datasets across four roll settings demonstrate the effectiveness of the proposed method.

Authors:Malavika Suresh, Ikechukwu Nkisi-Orji, Nirmalie Wiratunga
Title: Relative Parameter Importance in Task-Agnostic Replay-Free Continual Learning
Abstract:
Achieving continual learning (CL) with deep neural networks requires balancing stability and plasticity while enabling knowledge transfer. In this work, we focus on offline learning algorithms under the constraints: (I) no access to training data from prior tasks (II) no access to task‑id at inference time. We introduce a novel measure, the relative parameter‑importance, which measures the relative importance of each parameter with respect to both the current and past tasks. Parameters with high relative importance are interpreted as more important for maintaining past‑task stability and thus heavily regularised, whereas parameters with low relative‑importance are allowed to be more freely updated. Unlike existing methods, our approach allows the update of parameters with high past‑task importance when they have low relative‑importance, thus enabling backward knowledge transfer in addition to tackling the stability‑plasticity trade‑off. We demonstrate improvements against state‑of‑the‑art CL methods on both class‑incremental and domain‑incremental learning text classification problems and provide insights for extending our method to text generation problems. Code available at: https://github.com/itsmemala/LACL

Authors:Randhir Kumar
Title: Verification Without Sufficiency: Per-Chunk Filtering Fails on Multi-Hop RAG, and Decomposition Repairs It
Abstract:
Verification for retrieval‑augmented generation usually scores each retrieved chunk and drops the ones that fail. We show this cannot work for multi‑hop questions, and show what does. Per‑chunk scoring assumes one chunk is a sufficient premise for the answer. Multi‑hop questions are built so that none is, and the paragraph carrying the answer is the one the question does not name. Entailment scoring reaches 0.643, 0.523 and 0.560 AUC on HotpotQA, 2WikiMultihopQA and MuSiQue, against 0.951 on single‑hop SQuAD. Seven controls rule out model capacity, premise length, hypothesis template, decision threshold, retriever, answer‑matching criterion and prompt. End to end across three datasets, three generator sizes and two prompts, per‑chunk gating is significantly worse than not filtering at all in every cell, and its penalty grows with generator capability. The repair is to condition verification on the decomposed sub‑question rather than the original query. Using MuSiQue's gold decomposition, entailment on a later hop rises from 0.546, which is chance, to 0.840, a paired lift of +0.355 with a bootstrap interval of [0.331, 0.382]. An off‑the‑shelf Qwen2.5‑7B decomposer, given the question and the top retrieved paragraph, reaches 0.637 and captures 31% of that ceiling; decomposing without retrieval reaches 0.533, below the original question. Iterative retrieval systems already produce such decompositions and discard them before verifying.

Authors:Bogdan Savelyev
Title: Loanword or Switch? The Annotation Boundary, Not the Model, Drives Kazakh-Russian Code-Switching Identification
Abstract:
Off‑the‑shelf LID and letter heuristics over‑label Kazakh‑Russian social text as mixed: Russian loanwords inside Kazakh look like code‑switching under a shared Cyrillic script. We release a document‑level gold LID set whose guideline keeps integrated borrowings as Kazakh and reserves mixed for clause‑level switches, plus a mixed‑only sentiment pool used after LID in a filter‑first cascade. On a shared LID test, FastText, Lingua, raw and windowed HeLI, character‑trigram NB, and XLM‑R range from weak to strong performance. The gap shows the bottleneck is the loanword‑vs‑switch annotation boundary, not model class alone.

Authors:Ziang Wu, Peng Jin, Qishen Yin, Munan Ning, Hao Li, Peizhen Zhang, Li Yuan
Title: Relax Within, Balance Across: Geometry-Guided Load Balancing for Vision-Language Mixture-of-Experts
Abstract:
Vision‑language MoE batches contain different numbers of image and text tokens. Image resolution, image count, tiling, and prompt length all change this token mix. We call the standard token‑level Switch auxiliary loss Std‑Aux. Std‑Aux balances only the mixed load, so large image and text load errors can cancel at one mix. On our main model, the same trained router shows more than a fivefold change in load imbalance across image resolutions. We hold the image and text load profiles fixed and derive the exact load curve as the token mix varies. The image‑text load gap controls sensitivity to the token mix. Physical preprocessing can also change the conditional profiles. The fixed‑profile law excludes such changes. To design a remedy, we examine the router input structure. Image and text occupy distinct regions, while visual tokens group strongly by source image. The modality boundary motivates separate image and text terms. The image boundary motivates one equal‑weight routing instance per image. ReBA, or Relax Within, Balance Across, implements both choices. Across four split backbones, ReBA lowers load on every reported benchmark input while keeping mean task accuracy comparable to Std‑Aux. ReBA also lowers average load over the tested range and worst physical load under resolution and tiling shifts. Code is available at https://github.com/ZiangWu‑77/ReBA.

Authors:Qunhui Zhang
Title: AiFlow: Token-Native Reactive Orchestration with Bounded Backpressure for Streaming LLM Applications
Abstract:
Large language model (LLM) applications increasingly operate as streaming workflows combining retrieval, tool calls, safety filters, and multi‑agent coordination. Although contemporary frameworks expose provider deltas, workflow nodes often treat generation as coarse request‑response steps, leaving queue management, worker allocation, ordering, and backpressure to ad hoc callback code. This paper presents AiFlow, a token‑native reactive orchestration model that normalizes provider deltas into typed Context<T> events propagated through a directed streaming graph. Each node is managed by a Node Guardian that declares and enforces local queue bounds, worker concurrency, ordering, overflow policy, cancellation propagation, and retry discipline. We formalize the bounded‑memory property, present the compilation from a compact DSL and JSON graph form, and provide static validation for type safety, state concurrency, and injection compatibility. Controlled microbenchmarks, captured DeepSeek trace replay (30 runs), descriptive online runs, LangGraph baselines, a streaming RAG workload, and an Ollama local‑backend check show that AiFlow does not alter provider‑side Model TTFT but reduces Application TTFPT by 70.9‑94.7% versus aggregation and keeps runtime‑owned queue depth within declared bounds (93.7‑96.5% MaxQ reduction versus unbounded policies). The supplementary artifact contains scripts, raw traces, machine‑readable tables, checksums, and an API‑free smoke test; the public implementation is available through the FIT Framework repository.

Authors:Xiaolun Jing, Kezhao Yin, Xinxing Yang, Genke Yang, Jian Chu
Title: PHA-Net: Prototype-based Hierarchical Alignment Network for Text-Video Retrieval
Abstract:
With the emergence of large‑scale image‑text pre‑training models, e.g., CLIP, text‑video retrieval has experienced substantial advances in recent years. Existing best‑performing methods involve aligning cross‑modal semantics at individual, local, and global levels simultaneously, raising concerns about the intrinsic semantic mismatch between concise texts and rich videos. A canonical approach is to integrate multiple language‑video attention modules into the hierarchical framework while this paradigm only optimizes visual representations with prohibitive computational costs. In this paper, we propose a new prototype‑based hierarchical alignment network (PHA‑Net) to align individual/local/global level representations across modalities. Concretely, we introduce multiple modality‑shared prototypes as the bridge to efficiently optimize text and video representations for cross‑modal alignment. Then, we argue that the imbalanced semantic distribution in clustered tokens may undermine retrieval performance, as tokens with weak semantics are of little interest. To reduce the impact of these tokens, a proposed prototype‑supported token merge module is responsible for enhancing tokens with strong semantics and suppressing others with weak semantics via prototype semantics guidance. Moreover, we devise a prototype contrastive loss to encourage textual and visual prototypes to focus on different semantic information. The idea of this auxiliary loss is to ensure higher similarity between textual and visual prototypes from the same prototype than those from different prototypes. Extensive experiments on four benchmarks confirm the effectiveness of our PHA‑Net, which achieves significant improvements in the sum of all recalls on MSR‑VTT (8.8%), ActivityNet (19.2%), VATEX (0.7%), and Charades (4.9%). Code is available at https://github.com/JingXiaolun/PHA‑Net.

Authors:Xianhao Zhou, Jianghao Wu
Title: Beyond Prompt Adherence: Auditing Attribute-Level Voice Control in Speech Generation
Abstract:
Natural‑language descriptions have become a flexible interface for controlling generated speech. Existing evaluations largely assess whether an output matches a prompt, but prompt matching alone does not reveal whether characteristics outside the intended change remain stable. We examine this distinction through a controlled paired audit of three speech‑generation systems: CosyVoice3, VoxCPM2, and Fish‑Speech‑S2. The evaluation contains 5,940 outputs spanning six reference speakers, ten texts, three random seeds, and eleven conditions. Using acoustic, prosodic, content, and speaker measurements, we find that responses in the expected target direction are frequently accompanied by changes outside descriptor‑specific signal‑level target sets. This pattern remains among outputs whose target response exceeds baseline seed variation, and the accompanying changes differ substantially across systems. We further introduce VoDER‑Cal, a training‑free candidate selector that retains sufficiently strong target responses while favoring smaller off‑target deviations. A three‑candidate pool raises the joint success rate from 4.8% under single‑sample direct generation to approximately 14% for all candidate‑selection policies. Within the matched three‑candidate budget, VoDER‑Cal reduces held‑out off‑target deviation from 0.344 under target‑only selection to 0.276 and improves listener‑rated preservation. Preservation‑sensitive evaluation therefore complements prompt‑adherence evaluation, while candidate reranking offers a practical inference‑time improvement. Code, configuration files, and analysis scripts are available at https://github.com/intelland/VoDER

Authors:Xuankang Zhang, Jiangming Liu
Title: DE-NER : Zero-shot Named Entity Recognition via Dialogue Elicitation of Large Language Models
Abstract:
Recent advancements of zero‑shot Named Entity Recognition (NER) establish strong baselines by formulating sequence labeling into question answering where Large Language Models (LLMs) can be naturally adopted. However, existing LLM‑based zero‑shot NER methods suffer from the limitations of prompt and demonstration engineering. To address these issues with minimal human interventions, we introduce DE‑NER, a dialogue elicitation framework which elicits the chatting ability of LLMs to fully extract the knowledge encoded in LLMs. Our experiments demonstrate that the proposed method outperform the competitive baselines in zero‑shot settings across multiple benchmarks, with an average improvement of 3.75% F1 points. Codes are released in https://github.com/kkkenshi/DE‑NER.

Authors:Hongjie Wu, Yiping Xie, Jiancheng Lv
Title: Hybrid-Domain Posterior Sampling for Inverse Problems via Latent Flow Matching
Abstract:
Latent Flow Models have revolutionized compressed‑space image synthesis, yet their application to high‑fidelity inverse problems remains bottlenecked. In this paper, we trace this dilemma to a fundamental geometric limitation of pre‑trained autoencoders, which we term \emphFirst‑Order Manifold Blindness. Severe decoder compression (e.g., retaining only ~\!2% of the original degrees of freedom) produces a rank‑deficient Jacobian, rendering high‑frequency measurement residuals in its orthogonal complement invisible to latent gradients even when the decoder can represent the target image. To overcome this bottleneck, we propose Hybrid‑Domain Posterior Sampling (HDPS), a decoupled inference framework that disentangles physical measurement consistency from semantic prior modeling. HDPS diverges into the pixel space, leveraging Langevin dynamics to absorb precise orthogonal measurement gradients, and subsequently projects these structural corrections back onto the generative manifold. An optimization‑based latent alignment is introduced to filter pixel‑space artifacts while avoiding the semantic drift of direct encoding. Extensive experiments on diverse inverse problems demonstrate that HDPS establishes a new state‑of‑the‑art, successfully recovering the high‑frequency structural precision that latent‑only solvers inherently discard. The code is available at \hrefhttps://github.com/74587887/HDPShttps://github.com/74587887/HDPS.

Authors:Ling Ren, Chao Deng, Ziming Wang, Yuecong Xu, Kai Zheng
Title: Test-time Adaptation of Pelvic Bone Segmentation Models via Dynamic Reliability-Guided
Abstract:
Reliable pelvic bone segmentation (PBS) from CT is essential for robot‑assisted pelvic trauma surgery, yet deploying a source‑trained model to a new hospital suffers from severe performance degradation due to cross‑center domain shifts. While test‑time adaptation (TTA) enables online model adaptation without accessing source data, existing methods show limited effectiveness for PBS, facing challenges including boundary degradation, anatomical inconsistency under domain shifts, and voxel‑level class imbalance. To address these challenges, we propose a novel closed‑loop dynamic Reliability‑Guided TTA framework (ReGA) for PBS. Specifically, we introduce a pseudo‑label reliability criterion termed Segmentation Inference Consistency Evaluation (SICE), which jointly measures region overlap and boundary deviation via dropout‑based ensemble predictions. Based on SICE, a trust‑weighted refinement module adaptively updates features to mitigate boundary errors in pseudo‑labels. Furthermore, a confidence‑weighted region‑level contrastive learning strategy is proposed to enforce anatomical consistency. Finally, ReGA follows the teacher‑student (TS) scheme to alleviate voxel‑level class imbalance. Experiments on three heterogeneous 3D pelvic CT datasets demonstrate that ReGA consistently outperforms state‑of‑the‑art TTA methods, enabling effective adaptation of the source‑trained PBS model to unseen clinical domains. The code is available at https://github.com/Ren‑ling/ReGA.

Authors:Kai Geissler, Laurens Müller-Groh, Hans Meine
Title: RadYOLO: Computationally Efficient 3D Object Detection and Segmentation in CT and MRI
Abstract:
Object detection and segmentation in three‑dimensional medical images is a very active area of research. However, most proposed deep learning models carry a high computational cost, and only few aim to be broadly applicable, achieve high detection performance, and remain fast to execute on resource‑constrained hardware. To address this gap, we present RadYOLO, a 3D extension of YOLO11 tailored to medical images. We compare it with nnU‑Net and nnDetection on five datasets comprising CT and MRI data with varying object sizes and prevalence. RadYOLO's detection performance surpasses that of nnDetection on four of five datasets and is comparable on one. Compared to nnU‑Net, RadYOLO performs better on lesion detection tasks, while nnU‑Net excels at detecting large organs when precise localization is required. When rough object localization is sufficient, RadYOLO matches or outperforms nnU‑Net on all five datasets. Regarding inference time, RadYOLO is 8‑46x faster than nnU‑Net on a GPU. Compared to nnDetection the speedup is even higher. When executed on a CPU, RadYOLO's inference runs within seconds (still faster than nnU‑Net on a GPU) offering a significant advantage for clinical and edge‑device deployment. RadYOLO repository: https://github.com/FraunhoferMEVIS/RadYOLO

Authors:Tongsheng Ding, Zhen Luo, Yixuan Yang, Boyu Wang, Luyang Xie, Jinyu Yang, Feng Zheng
Title: DreamTraj: Generating 6-DoF Object Trajectories by Reading Unrendered Video Diffusion Latents
Abstract:
Accurate prediction of object trajectories during manipulation is essential for closing the perception‑action loop. Progress is limited on two fronts: available datasets lack fine‑grained language‑to‑motion annotations, and existing predictors either rely on privileged inputs such as video, depth, or CAD models, or recover motion from fully generated videos through costly, error‑prone perception pipelines. We close the supervision gap with the MOVE dataset, 5,038 object‑centric egocentric trajectories, each paired with a fine‑grained natural‑language instruction rather than a coarse verb‑noun label. We further propose DreamTraj, which predicts a 6‑DoF object trajectory from a single RGB image and a task instruction, requiring no video, depth, or CAD model at inference: rather than generating a video, it reads motion from the internal representations of a frozen image‑to‑video diffusion model at an early denoising step. A lightweight flow‑matching Reader decodes query‑key attention tracks and pooled hidden states into relative 6‑DoF poses. To our knowledge, this is the first approach to directly decode object 6‑DoF trajectories from intermediate video diffusion representations rather than generated pixels. DreamTraj sets a new state of the art on both translation and rotation against forecasters that consume multi‑frame or privileged inputs, and runs 4.6x faster than generate‑then‑extract pipelines.

Authors:Tao Liu, Tao Feng, Xiangheng Li, Jinwang Song, Yifan Li, Xiaoqing Cheng, Dixuan Zhang, Siquan Li, Lin Lan, Hongying Zan, Kunli Zhang, Chao Wu
Title: SERL-SQL: Selective Hindsight Distillation for Text-to-SQL Reinforcement Agentic Learning
Abstract:
Recent Text‑to‑SQL systems increasingly rely on multi‑turn interaction, execution feedback, and reinforcement learning. However, most existing methods use execution correctness only as a trajectory‑level reward, which provides limited guidance for identifying the SQL decisions responsible for success or failure. We propose SERL‑SQL, a selective execution‑grounded reinforcement learning framework for multi‑turn Text‑to‑SQL agents. SERL‑SQL samples on‑policy SQL interaction trajectories and uses a training‑only teacher to re‑score student actions with execution feedback. The resulting teacher‑‑student likelihood gap is converted into bounded, masked weights that reweight GRPO advantages only on SQL and tool‑action tokens. In this way, task rewards preserve the optimization direction, while execution hindsight provides localized credit assignment. Experiments on BIRD, Spider, and cross‑domain benchmarks show that SERL‑SQL achieves competitive performance, reaching 76.56% execution accuracy on BIRD‑Dev and 89.92% on Spider‑Test. Moreover, our reward‑based selection strategy closely approaches the oracle Best‑of‑N upper bound and consistently outperforms consistency‑based selection, showing that SERL‑SQL produces high‑quality candidates that can be reliably identified by lightweight execution‑grounded rewards. Our code will be released at https://github.com/Ffunkytao/SERL‑SQL.

Authors:Yuyan Chen
Title: Ekova: A Personality-Support Agent for Self-Discovery Dialogue
Abstract:
Emotional Support (ES) systems have long optimized a single objective: alleviating the user's emotional distress in the moment. We argue that a complementary need, helping users see themselves more clearly, defines a distinct paradigm we call Personality Support (PS). PS is not counseling or clinical intervention: it targets cognitive clarity and self‑articulation, not symptom relief or diagnosis. We instantiate this paradigm in three layers. First, we present DSD, a Chinese self‑discovery PS Dataset of 8,590 samples collected through real longitudinal interaction across five minimal units, Coach, Warm, Tsukkomi, Real, and Gonzo. Second, we build DeepSupport, a multi‑persona PS system trained with OrthoTune, a PS‑tailored framework with style‑specific adapters and a style‑consistency regularizer. Third, we unify the five DeepSupport personas into Ekova, a persistent personality‑support agent with a unified cross‑session memory layer, supporting both adaptive routing and user‑customized persona selection. Experiments show that OrthoTune‑trained models outperform all baselines with an average relative gain of 16.3% across all metrics over the strongest prompt‑based baseline. Code is available at https://github.com/Yukyin/Ekova.

Authors:Masaki Yoshida, Ren Togo, Takahiro Ogawa, Miki Haseyama
Title: Scene2Sound: Auditory-Grounded Soundscape Generation for 3D Gaussian Worlds
Abstract:
3D Gaussian Splatting (3DGS) turns captured or generated imagery into photorealistic 3D world simulations that users can freely explore, yet these worlds remain silent. Because existing audio generation methods condition on a single image or viewpoint, their sound is tied to that observation and cannot stay consistent while a listener moves. We introduce the task of generating a spatially consistent soundscape for a given 3DGS world through auditory grounding, identifying which objects in the world should emit sound and anchoring each to a persistent 3D position, and present Scene2Sound, a training‑free framework built on this grounding. From the input world alone, our pipeline selects viewpoints that jointly cover the scene, identifies sound‑emitting objects with a vision‑language model, and associates the multi‑view detections into 3D instances through Gaussian set matching, which measures the overlap between the Gaussian sets that render each detection. Each source then receives generated audio that a standard object‑based audio engine spatializes in real time at arbitrary listener poses. We further propose two spatial‑consistency metrics, one testing whether rendered audio responds consistently to listener motion and one testing whether the claimed sources are supported by views held out from their placement. On a curated set of generated 3DGS worlds and on 3DGS scenes generated from real‑world 360‑degree captures, Scene2Sound preserves the audio quality of strong per‑viewpoint baselines while remaining spatially consistent where per‑viewpoint and single‑panorama pipelines do not, and a user study confirms the perceptual benefit. Project page: https://masaki‑lmd.github.io/scene2sound/.

Authors:Zhishan Zou
Title: Poplar: A Scalable Pipeline for Human-Centric Image Dataset Synthesis
Abstract:
Recent image generators can synthesize convincing human‑centric images, yet producing a useful collection remains different from producing a single successful image. A human‑centric dataset must cover varied people and contexts, avoid implausible attribute combinations, preserve an everyday photographic character, and expose quality‑control decisions at scale. We present Poplar, a reproducible Specify‑‑Render‑‑Inspect pipeline for human‑centric image dataset synthesis. Specify samples structured attributes under commonsense constraints and verbalizes them as photography‑oriented prompts. Render uses a realism‑adapted image generator across composition‑aware aspect ratios and retries obvious technical failures. Inspect applies a single structured vision‑‑language review to each candidate, preserving the original prompt while rejecting intrinsic image defects or material prompt mismatches. Using Poplar, we construct Poplar‑9K: 9,401 curated human‑centric image‑‑text pairs retained from 11,765 reviewed candidates (79.9% acceptance). We release the dataset together with the pipeline, configurations, immutable generation prompts, and auditable inspection records as a compact resource for building customizable human‑centric collections.

Authors:Wenjun Xiong, Yijin Zhou, Jiaqian Wang, Shangding Gu, Bo Tang, Zhiyu Li, Feiyu Xiong, Ying Wen, Muning Wen
Title: MAPLE-Guard: Memory-Aware Link Enforcement Against Memory-Link Poisoning in Multi-Agent Systems
Abstract:
LLM‑based multi‑agent systems (MAS) increasingly rely on persistent private and shared memories for long‑horizon coordination. This memory layer improves continuity, but it also gives attackers a durable channel: a poisoned memory can be written once, continuously retrieved in later tasks, promoted into shared memory, and reused by other agents. A single poisoned write can therefore steer many later decisions and contaminate agents that never saw the original attack, all while no malicious message crosses a visible communication edge at the moment of harm. Further, because existing safeguards mainly inspect prompts, actions, or communication edges, they can miss attacks whose content appears benign at write time but becomes harmful after retrieval. We introduce Memory‑Aware Propagation and Link Enforcement Guard, MAPLE‑Guard, a memory‑link guard for memory‑enabled MAS. MAPLE‑Guard monitors the memory lifecycle and places gates at write, retrieval, promotion, and cross‑agent reuse, so risky memories can be quarantined, unsafe retrievals filtered, and poisoned private memories blocked before they enter shared memory. In the main evaluation, MAPLE‑Guard lowers attack success rate (ASR) from 38.2% to 0.9% on LongMemEval and from 34.7% to 0.2% on AppWorld; it also raises multi‑agent defense success rate (MDSR) from 54.0% to 74.3% and from 42.5% to 99.8% on the same benchmarks. These results suggest that memory‑aware link enforcement covers a gap left by prompt‑level and topology‑level defenses. Code is available at the link: https://github.com/xiong‑wenjun/MAPLE‑Guard.

Authors:Liming Liu, Chao Hu, Mingfei Lu, Cong Tan, Yiwei Ge, Chijin Zhou, Yongjun Xie, Runzhe Wang, Xiaohai Shi, Heyuan Shi
Title: Diagnose Before You Compress: Prediction-Independent Bottleneck Witness Refinement for LLM Serving Traces
Abstract:
Production LLM serving generates millions of diverse requests, making full‑trace replay across serving configurations increasingly expensive. Existing trace reduction methods mainly preserve workload distributions or representative requests, but bottleneck‑revealing workloads may be rare and non‑representative. Moreover, evidence for one component cannot compensate for missing evidence in another, while using predicted bottlenecks as target truth creates circular evaluation. These limitations make it necessary to preserve evidence for every bottleneck component rather than rely on workload representativeness alone. We propose Bottleneck‑Preserving Witnessing (BPW), a quality‑constrained framework for compact and diagnostically reliable LLM serving replay suites. BPW first performs Workload Candidate Nomination using response‑blind workload features and closed source‑side measurements. This stage identifies workloads that may expose scheduler, prefill, decode, or KV‑cache bottlenecks. Coverage‑Priority Sequence Construction then organizes multi‑component proposals as reusable hyperedges and prioritizes weak and uncovered dimensions. Finally, Bottleneck Truth Verification derives prediction‑independent labels solely from direct target‑system measurements. The verified results determine the earliest prefix satisfying the direct two‑witness requirement for every component. Experiments on BurstGPT, ServeGen, and Mooncake show that BPW reaches the verified gate with a compact workload set and outperforms 16 policies, achieving relative improvements of 2.3% and 16.3% in Mean prefix Macro‑F1 and WBRC‑AUC, respectively. Stage‑resolved and sensitivity analyses confirm the distinct contributions and local stability of its three stages. Our code is publicly available at https://github.com/llmllmllm/BPW

Authors:Zhiwen Xu, Xiaoming Yan, Chengkun Wu, Juan Chen, Haoang Chi, Liyang Xu
Title: Gene Ontology-Guided Hierarchical Spatial Gene Expression Prediction from Histopathology Images
Abstract:
Predicting spatial gene expression from histopathology images enables large‑scale transcriptomic profiling without the cost of direct measurement. Existing methods decode the target gene set as a flat, unstructured vector, ignoring the inter‑gene dependencies arising from shared biological pathways and regulatory programs. Without explicit structural guidance, models must infer these dependencies entirely from limited paired data, constraining prediction quality. We propose MSGR (Multi‑Scale Gene Refiner), which bridges this gap by incorporating the Gene Ontology (GO), a curated functional hierarchy of genes, as an explicit structural prior. MSGR organizes target genes into a four‑level GO tree. Its GO‑guided decoder then progressively refines predictions from coarse functional domains to fine individual genes via residual corrections under scale‑weighted supervision. Operating solely on the gene side, the GO‑guided decoder serves as a seamless plug‑in replacement that consistently improves existing architectures without requiring any image‑side modifications. Extensive experiments on nine datasets from the HEST‑1k benchmark provide empirical evidence for two central claims: GO‑structured decoding consistently outperforms flat decoding, even against a state‑of‑the‑art generative baseline, and the gain is attributable to biological ontology structure rather than hierarchical decomposition per se, as confirmed by a +0.027 margin over a structurally equivalent random hierarchy.

Authors:Hanwen Xing, Pengyun Wang, BingXu Meng, Kumail Alhamoud, Xiang Li, Jicheng Wang, Xin Yu, Xinyang Han, Xiaomin Li, Philip Torr, Yuexing Hao
Title: CurveShift: Is Agent Progress Scalar? Separating Level from Shape
Abstract:
Progress in large language models is often summarized using a single scalar measure, such as a time horizon, a latent ability estimate, or an aggregate benchmark score. These summaries capture the overall performance, but they do not test whether progress is distributed differently across task difficulty. We find that most of the apparent shift in gains toward harder tasks does not reflect a change in the shape of the difficulty‑response curve. On METR time‑horizon data, a single Rasch model with rising ability reproduces this pattern, so it is largely explained by ceiling effects rather than a qualitative change in capability. This echoes how the choice of metric can make claimed emergent abilities look like a property of the models themselves. We then identify a smaller hard‑task effect that survives this control. Isolating it is difficult on agentic benchmarks, because newer models are usually run with newer agentic harnesses, so a gain on hard tasks cannot be assigned to the model or its scaffold. We break the confound with LiveCodeBench, a public competitive programming benchmark that runs no agentic scaffold while pairing dated models with an exogenous difficulty ordering. After accounting for the rise in overall ability, models released after September 2024 still gain on the hardest problems beyond what their easy and medium performance predicts, by about +0.40 logits under our most conservative assumption, raising the hard‑problem solve rate from roughly 18% to 25%. The effect is led by the strongest reasoning models and holds for hard tasks that need only short reasoning, not autonomy over long horizons. We present this as a result specific to competitive programming, since our clean identification rests on a single coding benchmark. We release the LiveCodeBench Difficulty Panel (66 dated models x 1,055 problems) and our analysis code.

Authors:Renjie Liang, Zijian Xu, Jinqian Pan, Chengkun Sun, Zhengkang Fan, Shawn Li, You Qin, Mei Liu, Jie Xu
Title: ORCA: ORgan-Centroid Aggregation for Training-Free 3D CT Visual Token Compression
Abstract:
A 3D CT scan entering a vision‑language model produces a long sequence of visual tokens, often thousands to tens of thousands per volume, and this sequence must be compressed before a language model can consume it. Token compression is well studied in general vision, but little of it targets 3D CT specifically. A common baseline is grid average, which pools regular grid cells and can blend distinct anatomy, lesion, and air into one token. We present ORCA (ORgan‑Centroid Aggregation), a token compressor for 3D CT. It merges adjacent tokens with organ guidance and adds a sinusoidal encoding of each region's centroid to preserve spatial layout. This preserves the anatomical information a downstream model needs. ORCA is training‑free and plug‑and‑play, producing an adjustable token set without any model change or text query. We evaluate it across two datasets (CT‑RATE and Merlin) and five encoders. The evaluation spans two task types: attribute prediction over five families (size, density, location, texture, and disease) and text generation (visual question answering and report generation). At matched token budgets, ORCA improves consistently over existing compression methods. It shrinks the visual context 64× and its KV‑cache 50×, and is 31× faster to process each volume. Code released at https://github.com/renjie‑liang/ORCA‑3DCT.

Authors:Rohan Bansal, David He, Nadun Ranawaka Arachchige, Zhenyang Chen, Soobum Kim, Kexin Rong, Danfei Xu
Title: Action Chunk Scheduling for Batched Robot Policy Serving
Abstract:
Deploying robot foundation models at scale is the next step towards realizing the potential of general‑purpose robots. However, Vision‑Language‑Action (VLA) and other foundation models are computationally demanding, and on‑device compute is constrained by power and space. In this paper, we introduce the problem of serving a robot policy to multiple robots from a remote GPU and formulate it as a scheduling problem. We build Armory, a serving system validated on fleets of both simulated and real robots. Our experiments show that naive scheduling heuristics perform well when all robots are the same, but fall short when robots consume action chunks at different rates, uncovering a mismatch between conventional batching methods and the closed‑loop requirements of robot policy execution. To address this, we propose a scheduling algorithm that accounts for this heterogeneity and improves overall system throughput by up to 18% in real‑world experiments. Additional details are available at https://gatech‑rl2.github.io/actionchunkscheduling.

Authors:Amr M. Zaki, Farhoud Jafari Kaleibar, Honggeun Ji, Komal Sarda, Marin Litoiu
Title: OrEdge: Efficient Multi-Modal Anomaly Detection in Distributed Software Systems via Orthogonal-Domain Learning
Abstract:
We introduce Orthogonal‑Edge (OrEdge), a lightweight framework for real‑time anomaly detection in multi‑modal distributed software systems. Unlike existing approaches that rely on computationally expensive attention‑ and graph‑based architectures, OrEdge leverages orthogonal‑domain temporal representations to achieve accurate anomaly detection with substantially lower computational complexity and model size. It jointly analyzes heterogeneous monitoring data, including logs, metrics, and traces, to identify abnormal software behavior, capture temporal dependencies, and reduce redundancy across observability signals. At its core, OrEdge incorporates OrEdgeCore, a lightweight orthogonal‑domain reconstruction module that captures recurring temporal patterns while suppressing transient variations. Evaluated on three real‑world microservice datasets (MSDS, SN, and TT), OrEdge achieves competitive detection performance while reducing the reconstruction model size to at most 9.6K parameters, compared with 20K‑‑143K parameters in existing methods. This compact design enables efficient deployment on resource‑constrained edge devices: on Raspberry Pi platforms, OrEdge achieves sub‑second inference and reduces inference latency by over an order of magnitude compared with existing approaches. Extensive ablation studies, sensitivity analyses, orthogonal basis evaluations, and qualitative case studies further validate the effectiveness of each design component. Overall, OrEdge demonstrates that orthogonal‑domain temporal modeling provides an effective alternative to computationally intensive attention‑ and graph‑based architectures, achieving a favorable balance between detection accuracy and computational efficiency for real‑time multi‑modal anomaly detection in edge environments. The code is available at https://github.com/theamrzaki/MicroService_Twin_Original.

Authors:Jose Fuentes, Abdullah Al Redwan Newaz, Ana Cavalcanti, Leonardo Bobadilla
Title: Localization in Spatiotemporal Fields via Environmental PDEs
Abstract:
This paper proposes a localization framework that uses spatiotemporal fields governed by partial differential equations (PDEs) as localization signatures. Two PDE classes are considered: the shallow water equations, which describe free‑surface flows in coastal and riverine environments, and the advection‑diffusion equation, which models the transport and mixing of scalar quantities such as temperature, salinity, and dissolved oxygen. A numerical PDE solver provides predicted fields over the domain, and multiple field channels are fused as multimodal measurements to improve localization accuracy. We formulate the problem within a Rao‑Blackwellized particle filter (RBPF) that partitions the vehicle state into a nonlinear component sampled by particles and a linear sensor bias component tracked analytically via per‑particle Kalman filters. This factorization reduces the required number of particles compared to a standard particle filter while accounting for realistic sensor drift. Simulation studies on both PDE scenarios show that the RBPF consistently outperforms a standard particle filter in terms of final position error and Root Mean Square Error (RMSE) across varying particle counts. Field experiments with an autonomous surface vehicle measuring salinity, temperature, and dissolved oxygen validate that PDE‑governed environmental fields provide sufficient spatial variability for practical localization. Related experimental videos are available at https://localization‑environmental‑pdes.github.io/.

Authors:Begoña B. Sierra, Colin McLean, Peter S. Hall, Sarah Friedrich-Welz, Catalina A. Vallejos
Title: A reproducible and extensible framework for benchmarking competing risks survival models
Abstract:
A wide range of statistical and machine learning methods have been proposed for survival analysis with competing risks, where the occurrence of one event (i.e., cancer death) precludes the occurrence of other events (i.e., cardiovascular disease death). Despite these methodological advances, their systematic evaluation and adoption are limited by the lack of comprehensive, reproducible and extensible benchmarking frameworks. We developed an open‑source benchmarking framework for competing risks models that enables their systematic comparison across multiple datasets under different aspects of performance; calibration, discrimination, overall prediction error and clinical utility. We additionally introduce an extension of SHAP for competing risks, allowing model‑agnostic interpretability of covariates contributions over time. All our code is publicly available via GitHub:https://github.com/BBolosSierra/CompRisksBenchmark

Authors:David Aguado, Daniel Fuertes, Carlos R. del-Blanco, Fernando Jaureguizar
Title: Geometric Self-Supervised Pre-training for Neural Combinatorial Optimization
Abstract:
Neural Combinatorial Optimization (NCO) techniques have emerged as a highly efficient alternative to traditional exact algorithms for solving routing problems such as the Traveling Salesman Problem (TSP). However, the generalization capabilities of these Reinforcement Learning‑based models are severely hindered when scaling to high‑dimensional instances. This issue has been mitigated in other domains, like computer vision and natural language processing, by adopting a self‑supervised pre‑training strategy. Nevertheless, its application to routing graphs, which lack complex topological attributes beyond 2D spatial coordinates, remains a challenge. In this paper, we propose a geometric self‑supervised pre‑training framework specifically designed to capture spatial invariance and global relative distance distributions. By applying isometric transformations, such as rotations and axial reflections, the model learns robust structural representations prior to the policy optimization phase. Empirical results demonstrate that this strategy consistently outperforms models trained from scratch (baselines), achieving a 7.23% improvement in tour length for massive zero‑shot extrapolation scenarios (TSP1,000). Furthermore, the proposed model exhibits remarkable computational efficiency, delivering speedups of up to two orders of magnitude over the exact solver Concorde at massive scales. The source code and pre‑trained models are publicly available at https://github.com/davidaguadocosano/TSP‑GeoPretrain.git.

Authors:Shaohang Wei, Zikun Su, Feifan Song, Wen Luo, Wei Li, Guangyue Peng, Houfeng Wang
Title: Verifier-Induced Support Reshaping in On-Policy Optimization
Abstract:
We show that on‑policy reinforcement learning with verifiable rewards (RLVR) can improve the current objective while making successful behaviors for later objectives too rare to sample and reinforce. We call this verifier‑induced support reshaping and define effective rewardable support as successful trajectories reachable within a fixed rollout budget. Across two model families, we study this effect through repeated verifier‑scored sampling and bidirectional training on mathematical reasoning and constrained instruction following, including sequential training with the opposite verifier. Math‑RLVR raises average instruction‑following success but reduces the number of prompts with any successful response under repeated sampling. On IFEval with Qwen3‑8B‑Base, pass@1 rises by 6.5 percentage points while best@32 falls by 9.8 percentage points, and the same divergence appears across both models and IF benchmarks. Conversely, IF‑RLVR shifts math responses from step‑by‑step openings toward direct answers, lowers best@k across sampling budgets, and reduces reward variation for later Math‑RLVR. Token‑distribution analyses and controlled opening interventions show that these changes concentrate in the first few response tokens. RLVR mainly reranks openings already available in the base policy, and the selected opening causally affects math searchability. The tested reference‑policy constraints, routing priors, and on‑policy distillation preserve cross‑task support only partially; MathIF and ReasonIF show that marginal gains translate only partly into responses that are both correct and constraint‑following. Therefore, endpoint improvements do not guarantee future trainability or joint capability under on‑policy optimization. Code is available at https://github.com/sylvain‑wei/verifier‑induced‑support‑reshaping

Authors:Sparsh Rastogi, Tanmay Kumar, Baiyu Chen, Jatin Bedi, Zechen Li, Flora D. Salim
Title: TRACE-TS: Attribution-Grounded and Traceable Sensor-Language Reasoning for Human Activity Understanding
Abstract:
Wearable sensors capture fine‑grained motion patterns that support rich behavioral understanding, yet most existing methods reduce these signals to activity labels. Recent LM‑based approaches generate natural‑language explanations for sensor data, but their reasoning is weakly grounded in the underlying signal, leading to fluent yet unverifiable explanations. We introduce TRACE‑TS (Traceable Reasoning with Attribution‑Grounded Evidence), a framework for structured and signal‑grounded reasoning over wearable time series. TRACE‑TS uses attribution from an expert classifier to identify salient spatio‑temporal sensor regions, uses them to construct DAG reasoning traces with explicit evidence provenance, and trains a compact language model to generate these traces through gated cross‑attention over sensor memory tokens. At inference, the adapted model jointly outputs the activity prediction and its reasoning trace, without requiring attribution computation or teacher guidance. We introduce Semantic Node Match(SNM), an LLM‑as‑judge metric that diagnoses reasoning fidelity at the observation, inference, and synthesis levels, localizing hallucinated observations and broken evidence chains missed by standard NLG metrics. Across seven wearable benchmarks, TRACE‑TS achieves the best average accuracy and F1 among all evaluated methods (84.43%/81.24%), and outperforms the best LLM‑based baseline by 17.96% in F1. Our code is available at https://github.com/SparshRastogi/TRACE‑TS.

Authors:Marco Ruiz, Miguel Arana-Catania, David R. Ardila, Rodrigo Ventura
Title: AutoCause: A Python framework that automates expert decisions in environmental time-series causal discovery
Abstract:
Environmental time‑series causal discovery requires expert decisions about method choice, conditional‑independence tests, lag horizons, sample‑size adequacy, multiple‑testing control, and evidence interpretation. Applied inconsistently across datasets, these choices yield graphs that cannot be compared, reproduced, or audited. We present AutoCause, an open‑source Python workflow that records each decision, derives defaults from an extended causal‑audit module, and admits domain‑informed overrides. The workflow wraps four established causal‑discovery methods from three families, adds non‑causal reference models, and grades links by method‑count support. On 145 datasets from DGP‑Atlas, TimeGraph, and a topology‑derived CausalRivers reference, the methods recover complementary parts of the reference graphs. Majority‑supported links are more precise than single‑method links on the synthetic benchmarks but not against river topology. AutoCause converts inconsistent expert practice into an auditable, repeatable analysis; causal interpretation remains with the analyst. Available at https://github.com/marcoruizrueda/autocause.

Authors:Dong Yan, Jian Liang, Dapeng Hu, Ran He, Nicholas Jing Yuan, Qi Zhang, Tieniu Tan
Title: AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?
Abstract:
Large language model (LLM) agents can self‑evolve by continually improving from their own accumulated experience. However, existing studies predominantly adopt independent evaluation. Consequently, the behavior of self‑evolving agents in realistic streaming settings, where agents adapt to diverse and complex task streams, remains poorly understood. To address this gap, we introduce AgentStream, a unified framework that evaluates self‑evolving agents spanning diverse evolution components by organizing agentic benchmarks into a configurable task stream and instantiating the \textttIsolated, \textttSequential, and \textttInterleaved streaming scenarios at test time, which progressively vary the scope and domain composition of the stream. Over these scenarios, we combinatorially evaluate five representative self‑evolving methods across three frontier foundation models, disentangling how model capability, method architecture, and streaming scenario jointly shape self‑evolution. Our results show that self‑evolution reliability varies across streaming scenarios, the benefit of self‑evolution is gated by model capability and non‑monotonic in model strength, and no single method dominates across models and scenarios. These findings offer concrete guidance for selecting self‑evolving methods across models and streaming scenarios. Overall, we advocate that self‑evolving agents should be evaluated under realistic task streams rather than isolated single‑task settings.

Authors:Yi Yu, Yixuan Liu, Ziyu Zhang, Parker Martin, Zhenyu Bu, Yuchi Han, Yuan Xue
Title: Automatic LV Localization and Short-Axis Plane Estimation from Arbitrary CMR Slice
Abstract:
Accurate estimation of left ventricular (LV) orientation is essential for cardiac magnetic resonance (CMR) imaging and downstream analysis. Existing methods typically formulate orientation recognition as discrete view classification or rely on multi‑slice geometric intersection, limiting their ability to model continuous 3D orientation and generalize across arbitrary slices. This work introduces a novel paradigm: Joint LV localization and 3D orientation estimation from a single CMR slice. To investigate this setting, representative orientation‑aware detection frameworks are adapted to the CMR domain, and their limitations are analyzed. Upon that, we propose the Polar‑Coupled Circular (PCC) embedding that provides a continuous and unambiguous orientation representation to address the limitations. Meanwhile, a scalable benchmark is constructed through automatic slice sampling from volumetric CMR segmentation datasets. Extensive experiments on four datasets demonstrate strong performance, achieving an average mIoU of 86.18% and an average angle deviation of 3.39°. This study establishes a new task setting for single‑slice LV orientation modeling and provides a geometry‑consistent framework for spatially informed CMR analysis. Code is available at https://github.com/yuyi1005/cmr‑3d‑ood.

Authors:Victor Maricato
Title: Leak It: A Probabilistic Approach to Training-Data Extraction from Black-Box Language Models
Abstract:
Membership inference (MIA) on language models is usually summarised by an aggregate ROC‑AUC, but such evaluations are confounded: model‑free blind baselines separate members from non‑members from surface text alone. We study black‑box, sampling‑based training‑data leakage through a probabilistic lens, treating N samples from p(.|x) as an estimate of the output distribution and casting leakage signals as functionals of it. We extend the blind‑baseline critique into the sampling regime: on WikiMIA a blind bag‑of‑words classifier reaches AUC 0.97 (TPR 0.90 at 5% FPR) and sampling adds nothing, while on an IID Pile split (MIMIR) neither self‑concentration nor gold‑continuation recovery significantly beats a blind baseline (incremental AUC 95% CI includes zero). Aggregate metrics hide the real harm. The same sampling verbatim‑extracts training data for a tail of documents no blind attack can reach. On Pythia‑6.9B, 83 of 500 Pile documents bearing a real identifier (16.6%; 21.3% of those bearing an email address) have that exact identifier reproduced AND not reproduced under a mismatched‑prefix control, so each leak is attributable to that document, not to a globally common string. This per‑document disclosure is invisible to aggregate AUC and grows with capacity (5.6% to 16.6% from 410M to 6.9B). The risk is uneven: identifier leakage is ~3x stronger in code than prose, though prose stays clearly positive and also grows with capacity (4.0% to 12.1%), while recovery of arbitrary held‑out continuations is confined to code (+0.44 member gap on GitHub vs at most +0.014 on prose). Temperature and nucleus sampling matter little, a 16‑token prefix suffices, and we detect no reduction from corpus deduplication. Privacy audits should report per‑document extraction, decomposed by domain, not a single AUC. We release leakit, a black‑box extraction‑audit tool.

Authors:Siyuan Li, Zehao Liu, Haoyu Li, Xi Lin, Ning Liu, Jun Wu, Jianhua Li, Mohsen Guizani
Title: Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks
Abstract:
As LLMs become increasingly integrated into complex applications, their vulnerability to adversarial attacks has raised significant concerns. However, existing defenses remain reactive in nature. This limitation makes it difficult for them to counter sophisticated threats, as adversaries continuously adjust their strategies across multi‑turn interactions. In this paper, we present a proactive defense framework for securing LLMs against evolving multi‑turn adversarial attacks that combines disruption, misdirection, and adaptation across successive interaction turns. In particular, it employs a cooperative multi‑agent architecture in which specialized agents execute complementary defense strategies. These strategies include controlled response pacing to increase attack costs, strategically ambiguous outputs to mislead adversaries into ineffective strategies, and forensic analysis of interaction logs to identify attack patterns and refine defenses. These agents are coordinated by an adaptive mechanism that dynamically adjusts the defense strategy in response to escalating threats. To facilitate comprehensive evaluation, we present the EMRA dataset designed to simulate evolving strategies across multi‑turn attacks, including 5,200 adversarial samples across eight attack types. Experimental results on EMRA across multiple LLM backbones show that the proposed framework reduces ASR by 69% on average relative to evaluated state‑of‑the‑art baselines. Beyond suppressing harmful outputs, it sustains deceptive engagement, achieving an average DR more than six times that of the strongest baselines and increasing attacker‑token consumption by 198.83% on average relative to evaluated baselines. Code and dataset are available at https://github.com/SiyuanLi00/CoopGuard.

Authors:Yan Fang, Jialin Chen, Chun Gan, Hang Yu, Mingjun Nie, Yeyu Zhang, Fengxiang He, Ching Law
Title: LLM-OSDA: An Optimal-Stopping Dynamic Auction for Native Advertising in Multi-Turn LLM Conversations
Abstract:
LLM‑native advertising embeds sponsored content directly into model‑generated responses, shifting the unit of sale from a fixed slot to a moment within an evolving conversation. Existing LLM ad‑auction mechanisms primarily operate within a single response, settling the winner but not the timing. The extension is nontrivial: with one native insertion opportunity per session, the stopping time depends on bids, coupling timing with allocation, so static truthfulness arguments no longer apply. We propose the LLM‑based Optimal Stopping Dynamic Auction (LLM‑OSDA), a dynamic cost‑per‑click auction that integrates Bellman optimal stopping, winner allocation, and envelope pricing. A bid‑independent LLM layer estimates contextual click quality and seamlessly renders the winning ad, while bids enter only the committed auction mechanism. Under an exact Bellman oracle, the expected discounted‑click allocation is monotone in each advertiser's bid, and the corresponding envelope payment makes truthful bidding weakly dominant in expectation. For practical deployment, a learned StopNet approximates the Bellman action values. We show that its decisions differ from the optimal policy only near the stopping boundary and bound the resulting incentive loss in terms of its approximation error. Experiments on a simulated conversational advertising corpus show that LLM‑OSDA improves net revenue by 11 percent over the strongest fixed‑timing baseline while maintaining comparable user retention. Code is at https://github.com/2025Fang2025/llm‑osda.

Authors:Hoang Thanh Thanh Truong, Charles R. Clark
Title: DS@GT ARC at MEDIQA-CORE-Task-1 2026: Trimodal Model Fusion with Task-Specific Gates for Brain Tumor Subtype Classification
Abstract:
Brain tumor diagnosis is a time‑sensitive process in which patients may wait weeks for a finalized pathology report. This problem motivates automated systems that classify tumor subtype from multimodal inputs. This paper details the DS@GT ARC team's work for ImageCLEFmed MEDIQA‑CORE 2026 Task~1, Brain Tumor Subtype Classification. The task evaluates three glioma classification problems: Level‑1 Molecular Type, LGG vs HGG, and WHO Grade. We combine pre‑extracted MRI (NeuroVFM) and histopathology (Prov‑GigaPath) embeddings with free‑text radiology reports. Our team explored two trimodal fusion architectures, two report encoders (RadBERT and Llama‑3.1‑8B‑Instruct), and a biologically motivated post‑processing stage. We achieve a mean macro‑F1 of 0.801 under the Fully Multimodal condition, exceeding the organizers' baseline of 0.796 and ranking second among the teams whose code passed verification. Additional evaluation across modality‑dropping conditions shows that this advantage depends heavily on the availability of the histopathology modality, and that our system falls behind the baseline when modalities are missing. Our code is available on GitHub at https://github.com/dsgt‑arc/imageclef‑mediqacore‑2026.

Authors:Rongxiang Zhang, Songhua Liu
Title: LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head Generation
Abstract:
Long‑form and real‑time talking‑head generation remains challenging due to a latency‑quality trade‑off: inefficient multi‑step diffusion prohibits streaming generation, whereas real‑time autoregressive approaches suffer from error accumulation and identity drift. To address this drawback, we propose LeapTalk, a novel framework that achieves stable and real‑time talking‑head generation with a single forward step, scaling to arbitrarily long videos. At the heart of our approach lies a single‑step bridge distillation scheme. On the one hand, departing from the conventional noise‑to‑data paradigm, we introduce a data‑to‑data transport formulation based on a Brownian bridge. Anchored by a persistent reference, this strategy effectively mitigates identity drift and enhances long‑term temporal stability. On the other hand, to enable smooth knowledge transfer from a pre‑trained diffusion teacher to the student bridge model, we explore a heterogeneous distillation framework with an SNR‑aligned time transformation Φ(τ), which bridges the functional discrepancy between the two models. Moreover, we propose an audio‑driven classifier‑free guidance mechanism to maintain fine‑grained lip synchronization under extreme step reduction. Extensive experiments demonstrate that our method achieves high‑fidelity and temporally consistent video generation with only 1 step at up to 200 FPS, significantly outperforming existing approaches in both efficiency and stability. Project Page: https://zhangrongxiang.github.io/leaptalk‑page/

Authors:Saif U Din, Muhammad Ahsan Hussain, Radu Timofte, Dmitry Ignatov
Title: Device-First Feedback: Toward Mobile-Native LLM-Driven Neural Architecture Search
Abstract:
Deploying convolutional neural networks generated by large language models (LLMs) on real mobile hardware requires more than GPU validation accuracy: INT8 TensorFlow Lite export, delegate selection, and on‑device latency jointly determine whether a model is usable. We present an automated mobile deployment pipeline that closes the loop from QLoRA fine‑tuning of an architecture‑generating LLM through GPU evaluation, INT8 export, and physical‑device benchmarking to gated augmentation of the training corpus. The pipeline is fully scripted and runs cycle‑by‑cycle without manual intervention, with resume support after interruptions. We evaluate the same frozen protocol on two benchmarks, CIFAR‑10 and CIFAR‑100, on a Samsung SM‑P613 tablet (seed 42, 20 models per cycle, cycles 0‑6). On CIFAR‑10, cycle 1 is gate‑accepted and improves the mobile deployment score approximately 25.6x over the baseline with a mean quantized accuracy of 46.9%; later cycles raise GPU accuracy but fail the non‑decreasing mobile gate. On CIFAR‑100, the pre‑QLoRA baseline retains the best mobile score; iterative rounds improve GPU accuracy (up to 26.2%) yet cannot surpass cycle 0 on‑device, and the training pool stalls at 19 examples after the first accepted round. Together, the two studies show that closed‑loop GPU fine‑tuning does not guarantee monotonic mobile gains, especially on harder classification tasks, and that multi‑dataset, on‑device measurement is needed to stress‑test deployment objectives. We release per‑cycle metrics with 95% confidence intervals, all figures, and complete reproduction commands.

Authors:Feixiang Liu, Qiang Qiu, Hao Zhang, Xinyue Wang
Title: Beyond Accuracy: Auditing Spatial Provenance in Visual Token Pruning for OCR-Critical MLLM Inference
Abstract:
Visual‑token pruning is usually judged by answer quality at a fixed retention budget. For text‑rich multimodal large language models (MLLMs), this protocol can miss a distinct failure: an answer remains correct even when no retained token is locally traceable to the small OCR region that supports it. We turn this blind spot into an evidence‑risk audit that couples answer behavior with geometric token‑origin provenance, interventions, and realized cost; transparent training‑free selectors isolate controlled operating points. On locked image‑disjoint confirmation, Qwen Target at 30% retention has observed accuracy 0.786 versus 0.783 for Full (paired image‑cluster difference +0.003, 95% CI [‑0.014, +0.020]), yet same‑budget Target, Random, and Grid retain sharply different positive‑support coverage: 0.620, 0.270, and 0.318. Across Qwen3‑VL‑8B, LLaVA‑1.5‑7B, and InternVL3.5‑8B, matched controls, interventions, detector tests, and external methods reveal model‑specific quality‑risk‑traceability frontiers that accuracy alone does not expose. Materialized prefixes yield up to 4.32x batch‑prefill speedup and 76.4% lower incremental peak memory; full‑validation TextVQA and DocVQA further show that favorable target‑verification points do not imply task‑general compression. Visual‑token pruning should therefore report surviving spatial provenance and realized cost alongside quality and compression.

Authors:Yi Cui, Zilin Wang, Yijie Xu, Qianyi Cai, Huizai Yao, Shuai Jiang, Bingzhuo Zhong, Hui Xiong
Title: SafeBuild-Bench: A Temporal-Robust Construction Safety Benchmark with Graph-Enhanced Data Mining
Abstract:
Construction‑safety models must handle concrete deployment risks, such as a worker standing near a scaffold edge without guardrails, rather than only recognize common objects in curated images. Yet real inspection archives are redundant, long‑tailed, and collected across changing sites and months. We introduce SafeBuild‑Bench, a metadata‑driven benchmark for evaluating multimodal large language models on construction safety under realistic temporal and site variation. It is mined from 100K+ industrial image‑text records and contains 3,314 task instances from over 3,000 expert‑verified images, covering multiple‑choice hazard identification and free‑form hazard description. To make expert verification scalable, we develop GEMS, a graph‑enhanced multimodal selection pipeline that combines a proxy‑model confusion signal with graph‑based diversity to identify informative candidates from redundant streams. On public instruction‑tuning data, GEMS‑selected subsets preserve robustness‑oriented performance under small data budgets. On SafeBuild‑Bench, current MLLMs remain far from reliable construction‑safety understanding, with the best overall score near 60. We release the benchmark, evaluation scripts, and GEMS codebase at https://github.com/safebuild/gems.

Authors:Ramesh B. Paramkusham
Title: Trustworthiness Costs of Domain Adaptation in Small Language Models:A Cross-Architecture Empirical Study
Abstract:
Domain adaptation of small language models (SLMs) has emerged as a practical strategy for deploying capable NLP systems in resource‑constrained, high‑stakes environments including healthcare, legal services, and financial analysis. While performance gains from parameter‑efficient fine‑tuning are well characterised, the corresponding impact on trustworthiness (factual calibration and adversarial robustness) remains poorly understood. This paper presents the first systematic cross‑domain, cross‑architecture empirical study quantifying the trustworthiness cost of domain adaptation across three SLM architectures (TinyLlama 1B, Gemma‑2 2B, Llama 3.2 1B), three domains (healthcare, legal, finance), two training‑data conditions (benign and adversarially perturbed), and four fine‑tuning strategies (baseline LoRA, Safety‑DPO, Dark Experience Replay, and Task Arithmetic LoRA, TA‑LoRA). Trustworthiness is evaluated through TruthfulQA MC2 (factual calibration) and HarmBench ASR (adversarial robustness) across all 216 experimental configurations with three random seeds. Three principal findings emerge. First, baseline QLoRA domain adaptation produces minimal TruthfulQA MC2 change across all model‑domain combinations (mean |Delta TQA| < 0.02). Second, adversarially perturbed training data consistently improves domain adaptation quality (Delta loss approximately ‑0.040) without worsening trustworthiness benchmarks. Third, none of the three safety‑preserving strategies reduced adversarial harm susceptibility: Safety‑DPO was effectively neutral (mean Delta ASR < 0.001), while Dark ER and TA‑LoRA increased mean HarmBench ASR by +0.171 and +0.155 respectively in safety‑aligned models (Gemma‑2 2B, Llama 3.2 1B), with individual configurations exceeding +0.45. These results challenge the assumption that replay‑based and arithmetic‑merge strategies transfer alignment to domain‑adapted SLMs.

Authors:Julia Belikova, Rauf Parchiev, Mikhail Filimonov, Konstantin Polev, Andrey Savchenko, Maksim Makarenko
Title: SIRIN: A Unified Toolkit for Detecting Contextual Hallucinations in Retrieval-Augmented and Memory-Grounded LLM Systems
Abstract:
SIRIN (Semantic Inconsistency Recognition and Inspection Nexus) is a unified toolkit and interactive web UI for detecting contextual hallucinations (fluent, plausible responses unsupported by the provided evidence) in retrieval‑augmented, agentic, and memory‑grounded LLM systems. SIRIN unifies three detector paradigms (representation probing, uncertainty estimation, and judge‑style verification) and the complementary task of pre‑generation query answerability under one interface, configuration system, and evaluation pipeline, supporting response‑ and span‑level inspection in both white‑box and black‑box settings. The web UI enables live analysis of user‑supplied context‑query‑answer triples through hallucination scores, unsupported‑span highlighting, and side‑by‑side detector comparison, with a lightweight plug‑in design for adding new detectors. We demonstrate SIRIN on hallucination detection, query answerability, and as a faithfulness gate within long‑term memory systems. The source code is publicly available at https://github.com/sb‑ai‑lab/SIRIN.

Authors:Isaac Song, Mohammed Rehan Parwani, Glenn Matlin, Emile Anand, Akhil Theerthala, Arjun Chatterjee, Anthony Wen-Ming Zang, Maria Kostylew, Yonadav G. Shavit, Sebastien Krier, Mark Riedl
Title: Role Steering of Language Models for Social Simulations
Abstract:
Social simulations built from language‑model agents need role‑conditioned behavior that can be checked before agents are placed into a simulated population. We introduce an activation‑steering screening workflow for role‑conditioned agents: define a role profile, extract a role‑specific direction, sweep four steering coefficients, evaluate role‑profile alignment, and pass or flag each candidate configuration. On OLMo‑3‑7B‑Instruct, we apply the workflow to a mixed 275‑role inventory with 228 role‑agnostic questions, GPT‑4.1‑mini prompted role references, and GPT‑4.1‑mini judges. Role‑specific directions receive higher judged role‑profile alignment than an assistant‑axis directional control from prior persona‑vector work, with mean overall scores of 63.2 versus 41.1 across the tested grid. They also preserve high lexical diversity, while the control drops sharply at larger coefficients. The role‑level screen is the main practical output: most roles improve as steering increases, but 38 roles decline across all six measured dimensions, showing why simulation builders should choose coefficients per role rather than deploy a uniform high‑strength setting. We make our code and evaluation artifacts available at https://anonymous.4open.science/r/anonymous‑research‑code‑5F03/.

Authors:Xiaotong Yuan
Title: HSRAI: Permutation-Preserving Address Interleaving with Hierarchical Balance Metrics
Abstract:
Address interleaving balances bandwidth across caches, DRAM, and GPU partitions. In a multi‑level interconnect topology, the mapping must be a one‑to‑one, invertible correspondence between logical and encoded addresses and, under typical access patterns, keep traffic uniform at every level's egress ports, not only at terminal slave nodes. Using random access as the stimulus and terminal uniformity as the acceptance criterion is insufficient for cascaded interconnects; this paper partitions workloads by access‑pattern priority and requires a global bijection with no slave node left unvisited for extended periods. Per‑level coefficient of variation (CV), consecutive same‑port run length, and sliding‑window peak occupancy evaluate traffic at each level's egress. HSRAI preserves high‑order and intra‑line low‑order address bits and applies a W‑bit bijection only to the intermediate index segment. Full‑domain topologies use offline GF(2) affine search with matrix and salt parameters selected via priority‑ranked access patterns and per‑level admissibility hard constraints; pruned topologies combine the Chinese Remainder Theorem with remapping. Evaluation uses a reproducible C++ benchmark covering linear streams, matrix tiling, and 2D arithmetic lattices. On the 128‑node full‑domain topology, the proposed affine map satisfies bijection, terminal balance, design‑time admissibility thresholds, and zero long‑window starvation on the high‑priority acceptance set; fixed XOR and folding/table baselines expose intermediate‑level hotspots or extended zero‑access periods. For pruned topologies, the CRT variant significantly improves per‑level metrics on linear and tiling accesses. Artifact: https://github.com/xiaotongyuan/hsrai_address_hash (tag paper‑v11).